PuffinParse docs
/ llms.txt GitHub

Google Gemini#

Status: docs-only. Implemented from Google's published API documentation and tested against fixture payloads built from it. It has not yet been run against the live API, so expect wire-format differences. Help verify it: issue #10.

1. Summary#

Provider name gemini
Base URL https://generativelanguage.googleapis.com (override: base_url on the request, or GEMINI_BASE_URL)
API key GEMINI_API_KEY (or api_key on the request) — sent as the header x-goog-api-key: AIza…
Docs https://ai.google.dev/gemini-api/docs
API version Path-versioned; PuffinParse uses /v1beta (the Files API and the newest models live there). /v1 serves the same generateContent method.
Modes parse, ocr (derived from parse), extract
Checked against 2026-09-11 against the documented wire format; the live call was not exercised — see §6, "Unverified against a live key"
Implementation crates/puffinparse-core/src/providers/gemini.rs

Gemini is not a document-AI product but a general vision LLM: PuffinParse sends the file plus a transcription prompt and pins the answer's shape with generationConfig.response_schema (structured output), so the model must return one entry per page instead of a single markdown blob. There is exactly one HTTP call per parse — no upload step, no polling — unless the document is larger than ~14 MB, in which case the resumable Files API is used first.

The consequence of using an LLM is that nothing geometric comes back: no bounding boxes, no page dimensions, no per-block confidence. Anything that needs overlays or coordinates should use a layout provider (Reducto, Extend, LlamaParse, Azure, Textract) instead.

2. Models exposed by PuffinParse#

The PuffinParse model name is the Gemini model id minus the gemini- prefix.

Model API model id Token price (in / out, per 1M) pricing.json per-page estimate
gemini/2.5-flash (default) gemini-2.5-flash $0.30 / $2.50 $0.00245
gemini/2.5-pro gemini-2.5-pro $1.25 / $10.00 (>200k prompt tokens: $2.50 / $15.00) $0.009875
gemini/2.5-flash-lite gemini-2.5-flash-lite $0.10 / $0.40 $0.00047
gemini/3.5-flash gemini-3.5-flash $1.50 / $9.00 $0.00945
gemini/3.5-flash-lite gemini-3.5-flash-lite $0.30 / $2.50 $0.00245
gemini/3.8-flash gemini-3.8-flash $0.75 / $3.75 (introductory, through 2026-12-31; $1.50 / $7.50 after) $0.004125

All six support parse, ocr and extract — the mode is a prompt + schema, not an endpoint, so every model serves every mode. The 2.5 trio is the safe default; the 3.x entries come from the models page and the changelog (3.5 Flash GA 2026-05-19, 3.5 Flash-Lite GA 2026-07-21, 3.8 Flash GA 2026-09-02) and have not been exercised against a live key here.

Cost is computed from tokens, not from pages. Each response's usageMetadata is priced exactly (promptTokenCount × input + (candidatesTokenCount + thoughtsTokenCount) × output) and lands in usage.provider_cost_usd, which puffinparse-core prefers over the price table. The per-page numbers in pricing.json are only a fallback for the case where a response carries no usageMetadata (and for puffinparse providers-style estimates). They were derived, not measured:

per_page_usd = (1500 × input_price_per_1M + N × output_price_per_1M) / 1e6
N = 800 output tokens for parse/ocr, 300 for extract

so gemini/2.5-flash parse = (1500 × 0.30 + 800 × 2.50) / 1e6 = $0.00245/page, and the extract column of pricing.json is the same formula with N = 300 ($0.0012/page).

1 500 input tokens/page is a deliberately conservative blend: a PDF page bills at a flat 258 tokens plus the rendered-image tokens, while a full-page scan image is ~1 000–1 600 tokens. Thinking tokens are not in the estimate; with a thinking budget left at its default they can double the output side.

A model that is not in this table can be used without a registry change: provider_options={"model": "gemini-3.1-pro-preview"} replaces the API model id verbatim (PuffinParse then has no token prices for it and falls back to the registry model's per-page price).

3. Request flow PuffinParse uses#

One call: POST {base}/v1beta/models/{api_model}:generateContent, headers x-goog-api-key and content-type: application/json.

{
  "contents": [{"role": "user", "parts": [
    {"inline_data": {"mime_type": "application/pdf", "data": "JVBERi0xLjQK…"}},
    {"text": "Transcribe the attached document to GitHub-flavoured Markdown, page by page. …"}
  ]}],
  "generationConfig": {
    "temperature": 0,
    "response_mime_type": "application/json",
    "response_schema": {
      "type": "object",
      "properties": {"pages": {"type": "array", "items": {
        "type": "object",
        "properties": {"page_number": {"type": "integer"}, "markdown": {"type": "string"}},
        "required": ["page_number", "markdown"],
        "propertyOrdering": ["page_number", "markdown"]}}},
      "required": ["pages"]
    }
  }
}

Input handling.

Input What PuffinParse does
Path / bytes ≤ 14 MB base64 into inline_data
Path / bytes > 14 MB resumable Files API upload, then file_data: {file_uri, mime_type}
URL downloaded first (Gemini cannot fetch URLs), then treated as bytes

The MIME type is sniffed from the magic bytes (%PDF-, PNG, JPEG, WebP, GIF) and only falls back to the extension guess, because Gemini rejects a mismatched mime_type outright. PDFs are native input (up to 1 000 pages / 50 MB); the accepted image types are image/png, image/jpeg, image/webp, image/heic and image/heif — not GIF, which PuffinParse still labels correctly so that Gemini's rejection names the real reason.

The 14 MB inline cut-off exists because a generateContent request is capped at ~20 MB and base64 inflates the payload by a third. The Files API path is POST {base}/upload/v1beta/files with X-Goog-Upload-Protocol: resumable and X-Goog-Upload-Command: start → the x-goog-upload-url response header → a second request with X-Goog-Upload-Command: upload, finalize carrying the bytes → {"file": {"uri", "name", "state"}}; if state is not ACTIVE PuffinParse polls GET {base}/v1beta/files/{id} (0.5 s, ×1.5, max 5 s) until it is. Uploaded files expire after 48 hours.

Per mode.

  • parse — the prompt above plus the pages schema. The model returns {"pages": [{"page_number": 1, "markdown": "…"}, …]}, so the page split is exact rather than guessed from a separator.
  • ocr — not implemented natively; the trait default derives a TextResponse from parse (metadata.puffinparse_derived_from = "parse"). Lines come from the page text, words carry no boxes. Passing output="text" additionally tells the model to skip Markdown syntax.
  • extract — the request schema is rewritten into Gemini's schema subset (see §4) and used as response_schema; the schema and instructions are also restated in the prompt, which measurably improves adherence. ExtractResponse.data is the parsed JSON. citations=True is accepted but cannot be honoured (metadata.gemini_citations_unsupported = true).

Page selection. Gemini always reads the whole file, so pages is a prompt instruction ("Transcribe ONLY these pages of the PDF: 2-3 … keep the original page numbers"), applied only when the input is a PDF. It is best effort: the full document is still uploaded and still billed as input tokens, and a model may ignore the restriction. For a non-PDF input pages is dropped and metadata.gemini_pages_ignored = true is set.

Where provider_options are merged: the whole object is deep-merged into the request body after PuffinParse builds it, so generationConfig, safetySettings, systemInstruction, tools, … can all be set or overridden. Three keys are consumed by PuffinParse and never reach Gemini: model (API model id), prompt (replaces the built-in instruction entirely) and prompt_suffix (appended to it).

4. Response mapping#

{
  "candidates": [{
    "content": {"parts": [{"text": "{\"pages\": [{\"page_number\": 1, \"markdown\": \"# …\"}]}"}],
                 "role": "model"},
    "finishReason": "STOP", "index": 0
  }],
  "usageMetadata": {
    "promptTokenCount": 1809, "candidatesTokenCount": 1060, "totalTokenCount": 2869,
    "promptTokensDetails": [{"modality": "DOCUMENT", "tokenCount": 1548},
                            {"modality": "TEXT", "tokenCount": 261}]
  },
  "modelVersion": "gemini-2.5-flash",
  "responseId": "HfPJaNXKO7OXm9IPk7P1kQc"
}

(The full payloads are crates/puffinparse-core/tests/fixtures/gemini_parse_multipage.json and gemini_extract_invoice.json, which drive the normalisation tests.)

Gemini field PuffinParse unified field Notes
responseId ParseResponse.provider_job_id Gemini has no job concept; this is the only per-call id.
candidates[0].content.parts[].text — All text parts are concatenated, then parsed as JSON (a stray ```json fence is stripped defensively).
pages[].page_number Page.page_number 1-based; falls back to the array index + 1 if the model omits it.
pages[].markdown Page.markdown Trimmed. With output="text" the markdown-to-text conversion is stored instead.
derived Page.text Always markdown_to_text(markdown).
derived Page.blocks Exactly one text block per non-empty page, content = the page, bbox: None, confidence: None.
— Page.width / height Always None — Gemini reports no geometry.
number of returned pages Usage.pages For an inline PDF the /Type /Page count is compared against it (see below).
usageMetadata Usage.provider_cost_usd prompt × input_price + (candidates + thoughts) × output_price.
— Usage.credits Always None; Gemini bills tokens, not credits.
usageMetadata.* metadata.gemini_tokens {prompt, candidates, thoughts, cached, total}.
modelVersion metadata.gemini_model_version The exact served version behind an alias.
the API model id sent metadata.gemini_api_model Useful when provider_options.model overrode it.
finishReason ≠ STOP metadata.gemini_finish_reason
promptFeedback.blockReason → error bad_request, message gemini blocked the prompt: <reason>.

For extract, the parsed JSON becomes ExtractResponse.data verbatim, fields stays empty (no citations, no per-field confidence) and usage.pages is the PDF page count, or 1 for an image.

PDF page-count sanity check. For an inline PDF PuffinParse counts /Type /Page objects (ignoring /Type /Pages) in the raw bytes and records it as metadata.gemini_pdf_page_count. If it differs from the number of pages the model returned — a dropped or hallucinated page, or simply a page-selection request — metadata.gemini_page_count_mismatch = true is set and a warning is logged. Usage.pages still follows what the model returned. The count is best effort and inline-only: PDFs whose page objects live in compressed object streams return no count, and neither does a document that went through the Files API — in extract mode that means usage.pages falls back to 1.

sanitize_schema (public, unit-tested) rewrites a draft-2020-12 JSON Schema into the subset Gemini's response_schema accepts:

  • keeps only type, format, title, description, nullable, enum, items, prefixItems, properties, required, propertyOrdering, minItems, maxItems, minimum, maximum, minLength, maxLength, pattern, anyOf — everything else ($schema, $id, additionalProperties, unevaluatedProperties, default, examples, allOf, not, if/then, patternProperties, …) is dropped;
  • type: ["string", "null"] → type: "string" + nullable: true;
  • oneOf → anyOf; const: X → enum: [X] with the type inferred;
  • local $refs into $defs / definitions are inlined (depth-limited to 12; a recursive schema degrades gracefully instead of expanding forever, and an unresolvable $ref becomes {"type": "object"});
  • a node with properties but no type gets type: "object", with items gets type: "array".

5. Errors, status codes, rate limits, timeouts#

Errors are {"error": {"code": 400, "message": "…", "status": "INVALID_ARGUMENT", "details": [...]}}; Error::from_http lifts error.message and classifies by status:

Status status field Trigger PuffinParse ErrorKind
400 INVALID_ARGUMENT malformed body, unsupported mime_type, invalid API key, schema Gemini rejects bad_request
400 FAILED_PRECONDITION free tier not available in the caller's country / billing required bad_request
403 PERMISSION_DENIED key lacks access to the model, or key restrictions (IP/referrer) authentication
404 NOT_FOUND unknown model id, or an expired Files API file_uri bad_request
429 RESOURCE_EXHAUSTED per-minute request/token quota rate_limit (retried)
500 / 503 INTERNAL / UNAVAILABLE server error, model overloaded provider (retried)
504 DEADLINE_EXCEEDED prompt too large to finish in time provider (retried)

Note the odd one: an invalid API key is a 400, not a 401, so it surfaces as bad_request rather than authentication. A missing key is caught before the call and is authentication.

Failures that arrive with HTTP 200 are turned into errors too:

  • promptFeedback.blockReason set, or a candidate with finishReason: "SAFETY" and no text → bad_request / provider with the reason in the message.
  • finishReason: "MAX_TOKENS" → provider, "gemini hit the output token limit, so the JSON is truncated — raise generationConfig.maxOutputTokens … or parse fewer pages per call".
  • Text that is not the promised JSON → provider, with the first 200 characters in the message.

Retries. max_retries (default 2) with exponential backoff and full jitter, on 429, 5xx and network errors only (the shared http::with_retry). Gemini's 429 body sometimes carries a RetryInfo detail with retryDelay; PuffinParse does not read it yet and backs off blind.

Rate limits. Counted per project and per model in three dimensions — requests/minute, input tokens/minute and requests/day — and they depend on the usage tier, so the authoritative numbers are the ones shown in Google AI Studio rather than any number quoted here. A single large document can exhaust the token dimension in very few requests.

Timeouts. timeout_secs (default 300) is the whole-call deadline and caps every individual request, including the Files API upload and the URL download.

6. Gotchas#

  • Unverified against a live key. Everything here follows the published wire format, and both the normalisation path (fixture tests) and the transport (mod transport — a loopback HTTP server that asserts the exact URL, headers, base64 inline_data, schema, 429 retry and 400 mapping) are covered by tests, but the GEMINI_API_KEY present in the development environment is rejected by Google with 400 API_KEY_INVALID on every endpoint, so no call has ever reached the real service. The fixtures under tests/fixtures/gemini_*.json were therefore constructed from the documented response schema, not captured. Run cargo test -p puffinparse-core gemini_live -- --ignored --nocapture with a working key, then replace the fixtures with real payloads and fill in measured latency and cost here.
  • No geometry, ever. Block.bbox, Page.width/height and Block.confidence are always None. ocr mode returns words without boxes. This is a property of the model, not of PuffinParse.
  • The page split comes from the model, not from the file. Structured output makes it reliable in practice, but a model can still merge or drop a page. metadata.gemini_page_count_mismatch is the tripwire for PDFs; there is no equivalent check for multi-page TIFFs or images.
  • pages is a suggestion. Whole-file upload, prompt-level selection, full input billing.
  • Thinking tokens are billed as output and are invisible in candidatesTokenCount. On 2.5 models a page of dense tables can spend more on reasoning than on the transcript. Set provider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}} to turn thinking off on Flash / Flash-Lite (Pro cannot go below 128; Gemini 3 models use thinkingLevel instead).
  • Long documents hit MAX_TOKENS before they hit the page limit. The 1 000-page ceiling is theoretical: the transcript of ~40 dense pages already approaches the 64k output budget. Split long PDFs with pages, or raise generationConfig.maxOutputTokens.
  • temperature: 0 does not make the output deterministic. Repeated runs differ slightly, which matters for benchmark reproducibility.
  • Gemini cannot fetch a URL. PuffinParse downloads it first; a URL that needs authentication has to be fetched by the caller and passed as bytes.
  • additionalProperties is the classic 400. Most schema generators (Pydantic, zod) emit it and Gemini rejects it; sanitize_schema strips it, along with $schema, default and examples. The current docs page lists additionalProperties as supported, but rejecting it has been the observed behaviour for long enough that stripping is the safe default.
  • The API is versioned in the path and the newest surface is moving. Google's Interactions API (POST /v1beta/interactions, GA June 2026) is now the documented default and uses a different shape (input, steps, response_format). generateContent remains supported and is still the recommended path for stable deployments — PuffinParse targets it deliberately; expect the docs links below to show the Interactions shape.
  • Files API objects expire after 48 hours and are scoped to the project, so a file_uri cannot be reused across keys. PuffinParse uploads per call and never reuses or deletes (files are free and capped at 20 GB per project).
  • Image vs. PDF tokenisation differ. A PDF page is a flat 258 tokens plus image tokens; a standalone image is tiled at 768×768 (≈258 tokens per tile). A scan sent as PNG can therefore cost several times what the same page costs inside a PDF.
  • x-goog-api-key and ?key= are equivalent; PuffinParse uses the header so keys never land in request logs or URLs.

7. Useful provider_options passthrough#

# 1. Turn thinking off for cheap, fast transcription (Flash / Flash-Lite only).
puffinparse.parse("scan.png", model="gemini/2.5-flash",
              provider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}})

# 2. Raise the output budget for a long PDF.
puffinparse.parse("report.pdf", model="gemini/2.5-pro",
              provider_options={"generationConfig": {"maxOutputTokens": 65536}})

# 3. Reach a model that is not in the registry.
puffinparse.parse("doc.pdf", model="gemini/2.5-flash",
              provider_options={"model": "gemini-3.1-pro-preview"})

# 4. Steer the transcription without rewriting the whole prompt.
puffinparse.parse("statement.pdf", model="gemini/2.5-flash",
              provider_options={"prompt_suffix": "Keep every stamp and handwritten note, "
                                                 "and transcribe struck-through text as ~~text~~."})

# 5. Replace the instruction entirely (the pages schema still applies).
puffinparse.parse("form.pdf", model="gemini/2.5-flash-lite",
              provider_options={"prompt": "Return each page of this form as a Markdown table of "
                                          "field name and value, one row per field."})

# 6. Loosen safety filters for documents that trip them (medical, legal, incident reports).
puffinparse.parse("report.pdf", model="gemini/2.5-flash",
              provider_options={"safetySettings": [
                  {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_NONE"},
                  {"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"}]})

# 7. A system instruction, e.g. to pin the output language.
puffinparse.parse("brief.pdf", model="gemini/2.5-flash",
              provider_options={"systemInstruction": {"parts": [{"text": "Always answer in German."}]}})