Skip to content

HTML to PDF Python: a requests guide

This HTML to PDF Python guide shows how to send HTML or a public web address to the AIIPWorld API with requests and save the PDF. Every example was run before it was published.

In one sentence

A Python guide to turning HTML or a public web address into a PDF with the requests library and the render API, with tested code and a comparison with pdfkit, WeasyPrint and Playwright.

Key facts

  • POST /v1/render answers 202 with a job id; poll GET /v1/jobs/:id until the status is succeeded, failed or timeout, then fetch download_url (a short-lived signed link, no API key needed).
  • Polls count toward the per-minute rate limit of the plan (30 requests per minute on the free plan); a 429 answer carries a Retry-After header.
  • Every example was run against the gateway and render engine on 2026-10-11 with Python 3.11 and requests 2.34 (the first example also with Python 3.9).
  • There is no batch endpoint, no template_id support and no webhooks yet; each document is one request.

HTML to PDF Python: how the API works

The AIIPWorld API turns HTML or a public web address into a PDF. From Python that takes three steps with the requests library: send the HTML to POST /v1/render, wait for the job to finish, and download the file.

The API answers the first request with status 202 and a job id, not with the PDF. The PDF is made by a job that runs headless Chromium inside a resource-limited worker container, with a hard time limit. You poll GET /v1/jobs/:id until the status is succeeded, failed or timeout. A succeeded job carries a download_url, a short-lived signed link that you fetch without your API key.

There is no mode that returns the PDF in the first response, and webhooks are not available yet. The small client below hides the waiting, so the rest of your code can call one function and get the file back.

  • Python 3.9 or newer and the requests library (pip install requests). We ran the examples with Python 3.11 and requests 2.34, and the first example also with Python 3.9.
  • An API key from your dashboard. Keep it in an environment variable named AIIPWORLD_API_KEY, never in your source code.
  • Jinja2 (pip install jinja2) for the template example only.

We ran every example on this page against our own gateway and render engine on 2026-10-11 before publishing it. The base address in the code is https://api.aiipworld.com; the output shown is what the scripts printed.

Install requests and set your API key

Install the library with pip install requests. Create a key in the dashboard and export it, for example export AIIPWORLD_API_KEY=your-key on macOS and Linux. The code reads the key from that variable, so it never appears in your repository.

The API accepts the key as Authorization: Bearer followed by the key, or in an x-api-key header. The examples use the Authorization header. A free account includes 100 credits a month, and a simple render costs 1 credit.

A small Python client for the render API

Save this file as aiipworld_pdf.py. The other examples import it. It has four small functions and one that joins them.

aiipworld_pdf.py
import os
import time

import requests

API = "https://api.aiipworld.com"
HEADERS = {"Authorization": f"Bearer {os.environ['AIIPWORLD_API_KEY']}"}


class RenderError(Exception):
    """A request or a job failed. `code` is the HTTP status or the job status."""

    def __init__(self, code, detail):
        super().__init__(f"{code}: {detail}")
        self.code = code
        self.detail = detail


def call(method, path, retries=3, **kwargs):
    """Send one API request. On 429 (rate limited) wait for Retry-After and try again."""
    for attempt in range(retries + 1):
        r = requests.request(method, f"{API}{path}", headers=HEADERS, timeout=30, **kwargs)
        if r.status_code == 429 and attempt < retries:
            time.sleep(int(r.headers.get("retry-after", "1")))
            continue
        return r


def submit(payload):
    """POST /v1/render. The API answers 202 with a job id; return that id."""
    r = call("POST", "/v1/render", json=payload)
    if r.status_code == 202:
        return r.json()["job_id"]
    try:
        body = r.json()
    except ValueError:
        body = {"error": r.text[:200]}
    raise RenderError(r.status_code, body.get("message") or body.get("error"))


def wait(job_id, timeout=120):
    """Poll GET /v1/jobs/:id until the job has ended, then return the job."""
    deadline = time.monotonic() + timeout
    pause = 0.5
    while time.monotonic() < deadline:
        time.sleep(pause)
        pause = min(pause * 1.5, 5)  # every poll counts toward your rate limit
        r = call("GET", f"/v1/jobs/{job_id}")
        r.raise_for_status()
        job = r.json()
        if job["status"] not in ("queued", "running"):
            return job
    raise RenderError("client_timeout", f"job {job_id} did not finish in {timeout} s")


def download(job, path):
    """Save the file of a succeeded job. download_url is signed, so no API key is sent."""
    if job["status"] != "succeeded":
        raise RenderError(job["status"], job["error"])
    r = requests.get(job["download_url"], timeout=60)
    r.raise_for_status()
    with open(path, "wb") as f:
        f.write(r.content)
    return len(r.content)


def render(payload, path, timeout=120):
    """Submit, wait and save in one call. Returns the finished job."""
    job = wait(submit(payload), timeout)
    download(job, path)
    return job

How it works:

  • submit() sends POST /v1/render and returns the job id. Any status other than 202 becomes a RenderError with the API's message.
  • wait() polls GET /v1/jobs/:id until the job has ended. It starts at half a second and waits longer between polls, up to five seconds. Every request counts toward your plan's rate limit, polls included, so a tight loop wastes requests.
  • call() sends one request and, on 429 (rate limited), waits for the number of seconds in the Retry-After header and tries again.
  • download() fetches the signed link. It sends no API key, because the link already carries its own signature. If the job did not succeed it raises a RenderError with the job status and the error text.
  • render() runs the three steps in order and returns the finished job, which holds the page count and the size of the file.

Python convert HTML to PDF from a string or a file

Most scripts that convert HTML to PDF Python code runs in memory: the HTML comes from a template or an f-string and the PDF goes to a file. This example sends a small HTML document and saves report.pdf.

html_to_pdf.py
from aiipworld_pdf import render

html = """<!doctype html>
<html>
<head>
<meta charset="utf-8">
<style>
  body { font-family: "Liberation Sans", Arial, sans-serif; color: #0f1222; }
  h1 { color: #4338ca; }
  table { border-collapse: collapse; width: 100%; }
  th, td { border-bottom: 1px solid #d9dbe8; padding: 6px; text-align: left; }
</style>
</head>
<body>
  <h1>Monthly report</h1>
  <p>Made with Python and the AIIPWorld API.</p>
  <table>
    <tr><th>Item</th><th>Amount</th></tr>
    <tr><td>Subscriptions</td><td>$1,200.00</td></tr>
    <tr><td>Support</td><td>$300.00</td></tr>
  </table>
</body>
</html>"""

job = render({"html": html, "output": "pdf", "options": {"format": "A4"}}, "report.pdf")
print(job["status"], job["result"]["pages"], "page,", job["result"]["bytes"], "bytes")

Run it with python html_to_pdf.py. When we ran it, it printed succeeded 1 page, 17375 bytes. The options object holds the PDF settings. format sets the paper size and is A4 by default. The default margin is 10 mm on all four sides, and backgrounds are printed unless you set print_background to false.

To convert a file on your disk, read it with open() and pass the text as html. The API cannot see your disk, so relative links to images and stylesheets inside that file will not load. Inline them with data URLs, use absolute https addresses, or set base_url to a public site so that relative paths resolve against it.

Fonts on your own machine are not on the render server. It has the Noto and Liberation families and colour emoji; Arial and Helvetica map to Liberation Sans. Load any other font with @font-face from a public URL or a data URL.

Python HTML to PDF from a web address

When the page already has a public http or https address, send url instead of html. Send exactly one of the two. This script prints the pricing page of this site on Letter paper with narrower margins.

url_to_pdf.py
from aiipworld_pdf import render

payload = {
    "url": "https://aiipworld.com/pricing/",
    "output": "pdf",
    "options": {
        "format": "Letter",
        "margin": {"top": "15mm", "bottom": "15mm", "left": "12mm", "right": "12mm"},
        "wait_until": "networkidle",
        "media": "print",
    },
}

job = render(payload, "pricing.pdf")
print(job["status"], job["result"]["pages"], "pages,", job["result"]["bytes"], "bytes")

When we ran it, it printed succeeded followed by the page count and file size. wait_until decides when the page counts as loaded: load (the default), domcontentloaded, networkidle or commit. For pages that draw their content late, add wait_for_selector with a CSS selector, or delay_ms for an extra pause of up to 10000 ms. JavaScript runs by default.

The PDF uses print styles by default. Set media to screen if you want the page as it looks on a monitor. Pages behind a login show the login screen, because no cookies are sent. You can send up to 20 custom headers. Requests to private networks, localhost and cloud metadata addresses are blocked, including redirects.

Page numbers, headers and footers

Set display_header_footer to true and pass header_template and footer_template as HTML strings. Inside them, an element with the class pageNumber shows the current page and one with the class totalPages shows the page count. Leave room for both in the top and bottom margin.

header_footer.py
from aiipworld_pdf import render

rows = "".join(
    f"<tr><td>Order {1000 + i}</td><td>${(i * 37) % 500 + 20}.00</td></tr>" for i in range(1, 121)
)
html = f"""<!doctype html>
<html><head><meta charset="utf-8">
<style>
  body {{ font-family: "Liberation Sans", Arial, sans-serif; }}
  table {{ border-collapse: collapse; width: 100%; }}
  th, td {{ border-bottom: 1px solid #ddd; padding: 4px 6px; text-align: left; }}
</style></head>
<body><h1>Orders</h1><table><thead><tr><th>Order</th><th>Total</th></tr></thead>
<tbody>{rows}</tbody></table></body></html>"""

header = '<div style="font-size:9px; width:100%; text-align:center;">Orders report</div>'
footer = (
    '<div style="font-size:9px; width:100%; text-align:center;">'
    'Page <span class="pageNumber"></span> of <span class="totalPages"></span></div>'
)

payload = {
    "html": html,
    "output": "pdf",
    "options": {
        "display_header_footer": True,
        "header_template": header,
        "footer_template": footer,
        "margin": {"top": "20mm", "bottom": "20mm", "left": "15mm", "right": "15mm"},
    },
}

job = render(payload, "orders.pdf")
print(job["status"], job["result"]["pages"], "pages")

We checked the PDF this script made with pdftotext: every page had Orders report at the top and Page 1 of 4 to Page 4 of 4 at the bottom, and the table header repeated on each page. The sample below is that file.

Style the templates inline, as the example does. They are rendered apart from your page, so your stylesheet does not reach them. In a check we made, a header template without an inline font-size came out about one point tall, even though the page text was set to 40 px.

Example output

Four-page orders report with a header and Page 1 of 4 footer (page 1)
Output of the header and footer example in this guide, run from Python. Generated by this site's own engine. Page 1 shown. Open the PDF (or select the picture) for the full document.

Python generate PDF from a template: many documents

To python generate pdf from template data, fill an HTML template with your values and send the result as html. Jinja2 is a common choice. The example below renders four invoices. With autoescape=True, a customer name that contains a character such as & or < is shown as text and cannot break the layout.

Submit every job first and collect the results afterwards. Each submit() returns at once with a job id, and the jobs wait in the queue until a worker is free. Queue priority follows the plan, with Business first and Free last.

batch.py
from jinja2 import Template

from aiipworld_pdf import RenderError, download, submit, wait

# autoescape=True turns characters such as < and & in your data into plain text
TEMPLATE = Template("""<!doctype html>
<html><head><meta charset="utf-8"></head>
<body style="font-family: 'Liberation Sans', Arial, sans-serif; margin: 0">
  <h1>Invoice {{ number }}</h1>
  <p>Bill to: {{ customer }}</p>
  <p>Total: ${{ "%.2f"|format(total) }}</p>
</body></html>""", autoescape=True)

invoices = [
    {"number": "INV-1001", "customer": "Harbor Lane Studio", "total": 1250.00},
    {"number": "INV-1002", "customer": "Bluebird Supplies", "total": 89.50},
    {"number": "INV-1003", "customer": "Maple & Pine Co.", "total": 4300.00},
    {"number": "INV-1004", "customer": "Northgate Design", "total": 610.25},
]

# 1. Submit every job first. Each call returns at once with a job id.
job_ids = {}
for invoice in invoices:
    html = TEMPLATE.render(**invoice)
    job_ids[invoice["number"]] = submit({"html": html, "output": "pdf"})

# 2. The jobs run in parallel on the server. Now collect each result.
for number, job_id in job_ids.items():
    try:
        job = wait(job_id)
        size = download(job, f"{number}.pdf")
        print(number, "saved,", size, "bytes")
    except RenderError as e:
        print(number, "failed:", e)

It printed one saved line per invoice. The API has no batch endpoint and does not take a template id yet: the template_id field is reserved and returns an error. Each document is its own request, and results are collected by polling.

When you size a batch against your plan's rate limit, count at least two requests per document: one POST and at least one GET.

Handle errors and rate limits

Errors come in two kinds. A request can be refused straight away with an HTTP status, or a job can be accepted and then end as failed or timeout. RenderError carries both: code is the HTTP status or the job status, and detail is the message.

errors.py
from aiipworld_pdf import RenderError, render

attempts = {
    "unknown option": {"html": "<h1>Hi</h1>", "output": "pdf", "options": {"colour": True}},
    "private address": {"url": "http://127.0.0.1/", "output": "pdf"},
    "html and url together": {"html": "<h1>Hi</h1>", "url": "https://aiipworld.com/", "output": "pdf"},
}

for name, payload in attempts.items():
    try:
        job = render(payload, "out.pdf")
        print(name, "->", job["status"])
    except RenderError as e:
        # e.code is an HTTP status (401, 402, 429 ...) or a job status (failed, timeout)
        print(f"{name} -> {e.code}: {e.detail}")
WhereCodeMeaningWhat to do
Request400The request is not valid, for example html and url sent togetherRead the message and fix the payload
Request401 invalid_api_keyThe key is missing or not recognisedCheck AIIPWORLD_API_KEY
Request402 insufficient_creditsThe account has too few credits for this renderWait for next month's credits or add credits
Request429 rate_limitedToo many requests this minuteWait for Retry-After seconds; the client does
Jobfailed, invalid_inputAn option name or value was rejected by the engineFix the option
Jobfailed, ssrf_blockedThe address is a private network, localhost or metadata addressUse a public address
Jobfailed, navigation_failedThe page could not be loadedCheck the address and whether the site is up
Jobfailed, output_limit or memory_limitThe result or the page was too largeSplit or simplify the document
JobtimeoutThe job ran out of timeSimplify the page or split the work

Credits for failed and timed-out jobs are returned. With a valid key, errors.py printed unknown option -> failed: invalid_input: params: Unrecognized key(s) in object: 'colour', then private address -> failed: ssrf_blocked: blocked address 127.0.0.1, then html and url together -> 400: provide exactly one of template_id, html, url. With a wrong key every line ended in 401: invalid_api_key, and with an account that had no credits the first two ended in 402: insufficient_credits.

Retry a request only when it makes sense. A 429 is safe to retry after the pause. A 400, a failed job or a 401 will fail the same way until you change something.

HTML to PDF converter Python options compared

There are four common ways to turn HTML into a PDF from Python. They differ in the rendering engine, in what you install and in who runs it. The facts below come from each project's own pages, checked on 2026-10-11.

OptionHow it worksLicenceWorth knowing
pdfkit with wkhtmltopdfpdfkit is a Python wrapper that calls the wkhtmltopdf programpdfkit: MIT. wkhtmltopdf: LGPL-3.0You install wkhtmltopdf yourself. The pdfkit README warns that the Debian and Ubuntu package has reduced functionality, such as headers, footers and outlines. The wkhtmltopdf repository is archived and its status page advises against untrusted HTML.
WeasyPrint 70.0A Python library that lays out HTML and CSS itself, without a browserBSD-3-ClauseNeeds Python 3.10 or newer and the Pango library on the system. Its documentation says it has no JavaScript.
Playwright for PythonYou run headless Chromium and call page.pdf()Apache-2.0You install a browser with playwright install and look after timeouts, fonts and clean-up yourself.
The AIIPWorld APIYou send HTTP requests and we run headless ChromiumA hosted serviceNothing to install. Your HTML is sent to our servers and deleted after the retention time of your plan.

Sources: pypi.org/project/pdfkit and its README; github.com/wkhtmltopdf/wkhtmltopdf and wkhtmltopdf.org/status.html; doc.courtbouillon.org/weasyprint (First Steps and Going Further); playwright.dev/python/docs/api/class-page; github.com/microsoft/playwright-python; all checked on 2026-10-11.

pdfkit python scripts are short, and pdfkit html to pdf calls work as long as wkhtmltopdf is installed. The catch is the engine behind it. wkhtmltopdf uses Qt WebKit, which its status page says was deprecated in 2015 and removed in 2016. If you are moving away from it, read the wkhtmltopdf alternative guide for an option-by-option mapping, including a Python HTML to PDF route without wkhtmltopdf.

WeasyPrint suits documents written for print, such as invoices and reports with simple layouts, and it runs entirely on your machine. Because it does not run JavaScript, a page that builds itself with scripts will not look the same as in a browser.

Running Playwright yourself gives you the same kind of engine as this API, on your own machine. It is the right choice when your documents must not leave your network. This is the equivalent of the first example; we ran it with Playwright 1.63.0 and its Chromium headless shell, and it wrote a one-page A4 PDF.

playwright_self_hosted.py
from playwright.sync_api import sync_playwright

html = "<h1>Monthly report</h1><p>Made with Playwright on this machine.</p>"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.set_content(html, wait_until="load")
    page.pdf(
        path="self-hosted.pdf",
        format="A4",
        print_background=True,
        margin={"top": "10mm", "right": "10mm", "bottom": "10mm", "left": "10mm"},
    )
    browser.close()

Playwright's page.pdf() prints with print CSS by default, as the API does. What you add is the work around it: installing and updating the browser, setting a timeout for each job, cleaning up temporary files and installing the fonts you need. The API takes that work off your hands in exchange for credits.

Python create PDF from HTML: common problems

  • The PDF has the wrong fonts. The render server has a fixed set of fonts. Load a web font with @font-face from a public https address or a data URL.
  • Images or styles are missing. Relative paths cannot be read from your disk. Use absolute https addresses, data URLs or base_url.
  • Content that a script draws is missing. Add wait_for_selector with a selector that only exists once the content is there, or use delay_ms.
  • A table row or a card is split between two pages. Add break-inside: avoid to it. Our print CSS guide covers page breaks, @page sizes and margins.
  • The download fails after a while. The download_url is short-lived. Download the file as soon as the job succeeds, or read the job again to get a fresh link. Results are deleted when their retention time ends.
  • Your key appears in a repository. Keep it in an environment variable or a secrets manager and send it only from your server code.

Limits, credits and how long files are kept

A simple render costs 1 credit. PDFs longer than 10 pages cost one more credit for each further 10 pages, and a result over 5 MB costs one more credit for each further 5 MB. Credits for failed and timed-out jobs are returned. These costs are provisional until launch.

Each plan has its own limit on requests per minute: 30 on the free plan, 60 on Starter, 180 on Pro and 600 on Business.

Results are kept for 1 day on the free plan, 7 days on Starter and Pro, and 30 days on Business. See pricing for the plans.

Questions

How do I convert HTML to PDF in Python?

Install requests, send your HTML to POST /v1/render with output set to pdf, poll GET /v1/jobs/:id until the status is succeeded, then download the file from download_url. The client on this page does all three in one call.

Do I need wkhtmltopdf or pdfkit to use the API?

No. The API renders in headless Chromium on our servers, so there is nothing to install except an HTTP library such as requests. pdfkit is a separate Python wrapper that needs the wkhtmltopdf program on your machine.

How does pdfkit work in Python?

pdfkit calls the wkhtmltopdf program for you, so calls such as pdfkit.from_string are short, but wkhtmltopdf must be installed. Its README warns that the Debian and Ubuntu package has reduced functionality, and the wkhtmltopdf project is archived (pypi.org/project/pdfkit and github.com/wkhtmltopdf/wkhtmltopdf, checked 2026-10-11).

Why does the API return a job id instead of the PDF?

Rendering takes a moment, so the first response is 202 with a job id and the work happens in a queue. Poll GET /v1/jobs/:id until the job has ended. Webhooks are not available yet.

How do I make a python pdf from html without installing a browser?

Send the HTML to the API with requests. The PDF is made on our servers in headless Chromium, so your machine needs only Python and an HTTP library. No browser, wkhtmltopdf or system libraries are installed on your side.

Can I use a template and JSON data?

Not through the API yet: the template_id field is reserved and returns an error. Fill a template yourself, for example with Jinja2, and send the finished HTML, as the batch example does.

How do I add page numbers to a PDF made in Python?

Set display_header_footer to true and pass a footer_template that contains an element with the class pageNumber and one with the class totalPages. Style the template inline and leave room in the bottom margin.

Can I convert a URL to PDF in Python?

Yes. Send url instead of html. The address must be public http or https. Pages behind a login show the login screen, and private network addresses are blocked.

Does JavaScript run before the PDF is made?

Yes, by default. If the page draws its content late, use wait_for_selector or delay_ms. You can turn scripts off with javascript set to false.

How do I handle rate limits?

A 429 answer has a Retry-After header with the seconds to wait. The client on this page waits and retries. Polls count toward the limit, so poll with growing pauses instead of in a tight loop.

Should I use WeasyPrint, Playwright or the API?

It depends on where your HTML may go and what it needs. WeasyPrint and Playwright run on your own machine, which suits documents that must not leave your network. WeasyPrint has no JavaScript. Playwright needs a browser you maintain. The API needs nothing installed but sends your HTML to our servers.

Can I use httpx or aiohttp instead of requests?

Yes. The API is plain HTTP and JSON, so any client works. The examples use requests because it is the most familiar, and we only ran the code shown.

How long are the PDFs kept?

Results are kept for 1 day on the free plan, 7 days on Starter and Pro, and 30 days on Business. Download links are short-lived, so save the file soon after the job succeeds.

Which Python versions did you test?

The client and examples ran on Python 3.11 with requests 2.34. The first example also ran on Python 3.9 with requests 2.32. Playwright 1.63.0, used in the comparison, needs Python 3.10 or newer.

Related tools