PuffinParse docs
/ llms.txt GitHub

OpenAI (Responses API)#

Status: docs-only. Implemented from OpenAI's published API documentation and tested against fixture payloads built from it. It has not yet been run against the live API, so expect wire-format differences. Help verify it: issue #10.

1. Summary#

Provider name openai
Base URL https://api.openai.com (override: base_url on the request, or OPENAI_BASE_URL)
API key OPENAI_API_KEY (or api_key on the request) — sent as Authorization: Bearer sk-…
Docs https://developers.openai.com/api/docs (the old platform.openai.com/docs/* links 301 here)
Endpoint POST /v1/responses — one synchronous call per document, no job id, no polling
Modes parse, ocr (derived from parse), extract
Checked against 2026-09-11 against the published docs only — no OpenAI key was available, so the fixtures are built from the documented response shape and the live tests are #[ignore]
Implementation crates/puffinparse-core/src/providers/openai.rs

This is not a document-parsing product: it is a general vision LLM asked, with a strict JSON schema, to transcribe a document page by page. That buys layout-aware markdown of figures, handwriting and messy scans, and costs you geometry — there are no bounding boxes, no per-block types and no confidences. Every page comes back as a single text block whose content is the page markdown.

2. Models exposed by PuffinParse#

Model Provider parameters PuffinParse sets Estimated price (pricing.json)
openai/gpt-5.6-luna (default) model=gpt-5.6-luna, reasoning.effort=low ~$0.00114 / page
openai/gpt-5.6-terra model=gpt-5.6-terra, reasoning.effort=low ~$0.0114 / page
openai/gpt-5.6-sol model=gpt-5.6-sol, reasoning.effort=low ~$0.0200 / page
openai/gpt-6-astra model=gpt-6-astra, reasoning.effort=low ~$0.0500 / page

Pricing is per token, not per page, so pricing.json holds an estimate: 1,500 input + 700 output tokens per page (OpenAI bills a PDF page as extracted text plus a page image; 1.5k input tokens is the low end of the published 1,500–3,000 range, so a dense page costs more). Rates per 1M tokens (https://developers.openai.com/api/docs/pricing, 2026-09-11): luna $0.20/$1.20, terra $2/$12, sol $4/$20, astra $10/$50.

The authoritative number for a call is response.usage.provider_cost_usd, which PuffinParse computes from the actual usage.input_tokens/usage.output_tokens and the per-token table embedded in openai.rs (PRICES, dated 2026-09-11). It always wins over the per-page estimate, so cost_usd is exact even when the estimate is not. metadata.openai_input_tokens / openai_output_tokens carry the raw counts.

Any other model can be reached without a registry change via provider_options={"model": "gpt-4.1-mini"}; the price table covers the gpt-5.x, gpt-4.1 and gpt-4o families as well, and cost falls back to None for a model it does not know.

3. Request flow PuffinParse uses#

One call. POST {base}/v1/responses, Authorization: Bearer …, JSON body:

{
  "model": "gpt-5.6-luna",
  "input": [{
    "role": "user",
    "content": [
      // PDFs — the filename matters, the model sees it:
      {"type": "input_file", "filename": "invoice.pdf", "file_data": "data:application/pdf;base64,JVBER…"},
      // …or, for images: {"type": "input_image", "image_url": "data:image/png;base64,iVBOR…", "detail": "high"}
      {"type": "input_text", "text": "Transcribe this document. …"}
    ]
  }],
  "text": {"format": {"type": "json_schema", "name": "puffinparse_pages", "schema": { … }, "strict": true}},
  "max_output_tokens": 32000,
  "reasoning": {"effort": "low"},
  "store": false
}
  • Input handling. Path and bytes inputs are read locally; a URL input is downloaded by PuffinParse and inlined as a data: URL (the API would only fetch public URLs, and the response is identical either way). MIME type comes from the URL's Content-Type when it is specific, else from the filename extension. application/pdf → input_file; image/* → input_image; anything else is an input_error before any network call.
  • parse schema (name: puffinparse_pages): {"pages": [{"page_number": <int>, "markdown": <string>}]}.
  • extract schema (name: puffinparse_extraction): the caller's JSON Schema, passed through the strict sanitiser.
  • pages and language are prompt-level, not API parameters: the prompt says "transcribe only pages 1-3,7, keeping the original page numbers" and "the document is mainly in de". Page selection is therefore best-effort — the whole document is still uploaded and billed.
  • output="text" adds a rule telling the model to emit plain text instead of Markdown.
  • store: false by default so the document is not retained for the response store; set provider_options={"store": true} if you want to fetch the response later.
  • temperature: 0 is only sent to gpt-4.x models. gpt-5.x and gpt-6 are reasoning models and return 400 Unsupported value: 'temperature' does not support 0 with this model — PuffinParse sends reasoning: {"effort": "low"} to those instead (transcription is perception, not reasoning).
  • provider_options is merged into the body verbatim (deep merge, so {"reasoning": {"effort": "medium"}} replaces only the effort). The key strict is consumed by PuffinParse and never forwarded.

4. Response mapping#

Responses API field PuffinParse unified field Notes
id (resp_…) ParseResponse.provider_job_id
output[].content[].text where type == "output_text" parsed as JSON → the structured answer Parts are concatenated in order; reasoning items are skipped. A refusal part becomes a provider error. If nothing is found, the SDK-style output_text aggregate is used as a fallback.
pages[].page_number Page.page_number Missing or 0 falls back to the array index + 1.
pages[].markdown Page.markdown (trimmed) With output="text" this is already plain text.
— Page.text markdown_to_text(markdown), or the string itself in text mode.
— Page.blocks Exactly one text block per page, content = page markdown, bbox: None, confidence: None.
— Page.width / height Always None — the API reports no page geometry.
number of returned pages Usage.pages The API never reports a page count; extract mode reports 0.
usage.input_tokens / output_tokens Usage.provider_cost_usd (and metadata.openai_*_tokens) Cost = tokens × the embedded per-token prices, so it is exact.
— Usage.credits Always None — OpenAI has no credit unit.
status: "incomplete" + incomplete_details.reason provider error Typically max_output_tokens.
status: "failed" + error.message provider error
everything else (reasoning, annotations, output_tokens_details, …) — Visible with include_raw=True.

extract mode returns the tool-free JSON object as ExtractResponse.data. ExtractResponse.fields is always empty: the Responses API has no per-field grounding, so citations=True is recorded as metadata.openai_citations_unsupported = true rather than silently pretending to support it.

Trimmed response (crates/puffinparse-core/tests/fixtures/openai_responses_parse.json, built from the documented shape):

{
  "id": "resp_68f0a1b2c3d4e5f60123456789abcdef",
  "object": "response",
  "status": "completed",
  "model": "gpt-5.6-luna",
  "output": [
    { "id": "rs_…", "type": "reasoning", "summary": [] },
    { "id": "msg_…", "type": "message", "status": "completed", "role": "assistant",
      "content": [{ "type": "output_text", "annotations": [],
                    "text": "{\"pages\":[{\"page_number\":1,\"markdown\":\"# Hello PuffinParse\\n\\nInvoice #1234…\"}]}" }] }
  ],
  "text": { "format": { "type": "json_schema", "name": "puffinparse_pages", "strict": true } },
  "usage": { "input_tokens": 3120, "output_tokens": 412,
             "output_tokens_details": { "reasoning_tokens": 64 }, "total_tokens": 3532 },
  "store": false
}

5. Errors, status codes, rate limits, timeouts#

The error envelope is {"error": {"message", "type", "code", "param"}}; Error::from_http picks up the nested error.message and classifies by status:

Status Trigger PuffinParse ErrorKind
400 invalid schema for text.format, unsupported parameter (temperature on a reasoning model), file too large / unreadable bad_request
401 missing or revoked key authentication
403 key/region or model not permitted for the org authentication
404 unknown model id bad_request
429 rate limit or insufficient quota (code: "insufficient_quota" — not actually retryable) rate_limit (retried)
5xx server_error, 502, 503 provider (retried)

Retries. max_retries (default 2), exponential backoff with full jitter, on rate-limit, network and 500/502/503/504 only. 4xx is never retried.

Rate limits. Per-model RPM/TPM quotas by usage tier; responses carry x-ratelimit-remaining-* and retry-after headers, which PuffinParse does not read yet — backoff is blind.

Limits. A file must stay under 50 MB, and all files in one request under 50 MB total; a PDF page is billed as extracted text plus a page image, so context, not page count, is the practical limit. max_output_tokens defaults to 32,000 (~20 dense pages); a longer document truncates with status: "incomplete", reason: "max_output_tokens".

Timeouts. timeout_secs (default 300) is the whole-call deadline and also caps the single HTTP request. Big documents on a reasoning model are slow: budget minutes, not seconds.

6. Gotchas#

  • Strict structured outputs restrict JSON Schema. With strict: true, every object must set "additionalProperties": false and list every property in required; there are no optional fields — express "may be absent" as a nullable type ("type": ["string", "null"]). The root must be an object. PuffinParse applies sanitize_strict_schema() to the caller's extract schema recursively (through properties, items, prefixItems, $defs/definitions, anyOf/oneOf/ allOf, if/then/else, not, contains), rewriting required to the full property list and forcing additionalProperties: false. This widens required: fields you marked optional come back as null rather than missing.
  • Some keywords are still rejected in strict mode (minimum/maximum, minLength/maxLength, pattern, recursive $ref beyond #). If the API answers 400 Invalid schema, either drop those keywords or turn the feature off with provider_options={"strict": false} — PuffinParse then sends your schema untouched and asks for JSON by prompt instead.
  • temperature is not accepted by gpt-5.x / gpt-6. See §3. If you force it through provider_options, expect a 400.
  • Reasoning tokens are billed as output tokens and are invisible in the text. effort: "low" keeps them small; raise it with provider_options={"reasoning": {"effort": "medium"}} if a dense or handwritten document transcribes badly.
  • No geometry, ever. Block.bbox, Block.confidence and Page.width/height are None, and every page holds exactly one text block. Anything that draws overlays needs a layout provider (Reducto, Extend, Azure, …). ocr mode is derived from parse by the default trait method, so its words[] carry no boxes either, and metadata.puffinparse_derived_from = "parse" says so.
  • pages is a prompt instruction, not an API parameter. The whole document is uploaded and billed even when you ask for one page, and the model may ignore the restriction on a bad day. Split the PDF client-side if this matters.
  • Page numbers come from the model. They are usually the document's own, but a model can repeat or skip one; PuffinParse falls back to the array position when the number is missing or 0, and sorts pages by number when assembling the document markdown.
  • Hallucination is the failure mode. A layout parser drops text it cannot read; an LLM can invent plausible text instead. The prompt forbids it ("transcribe verbatim, never invent"), but for high-stakes extraction prefer a provider that returns citations.
  • store: false is PuffinParse's default — responses are not kept in OpenAI's response store, which also means previous_response_id chaining is unavailable unless you opt back in.
  • detail: "high" is set on image inputs for legibility of small print; {"detail": "low"} through provider_options on the input block is not possible (the block is built by PuffinParse) — use the openai/gpt-5.6-luna model and a downsampled image instead if you need to save tokens.

7. Useful provider_options passthrough#

# 1. Longer documents: raise the output ceiling (the default 32k covers ~20 dense pages).
puffinparse.parse("contract.pdf", model="openai/gpt-5.6-luna",
              provider_options={"max_output_tokens": 64000})

# 2. Harder documents: more reasoning (billed as output tokens).
puffinparse.parse("handwritten.pdf", model="openai/gpt-5.6-terra",
              provider_options={"reasoning": {"effort": "medium"}})

# 3. A model that is not in the registry (pricing falls back to the embedded token table).
puffinparse.parse("scan.png", model="openai/gpt-5.6-luna",
              provider_options={"model": "gpt-4.1-mini", "temperature": 0})

# 4. Extraction with a schema that uses keywords strict mode rejects.
puffinparse.extract("invoice.pdf", schema=schema_with_patterns,
                model="openai/gpt-5.6-luna", provider_options={"strict": False})

# 5. Keep the response in OpenAI's response store, and tag it.
puffinparse.parse("report.pdf", model="openai/gpt-5.6-sol",
              provider_options={"store": True, "metadata": {"run": "bench-2026-09"}})