Google Gemini#
Status: docs-only. Implemented from Google's published API documentation and tested against fixture payloads built from it. It has not yet been run against the live API, so expect wire-format differences. Help verify it: issue #10.
1. Summary#
| Provider name | gemini |
| Base URL | https://generativelanguage.googleapis.com (override: base_url on the request, or GEMINI_BASE_URL) |
| API key | GEMINI_API_KEY (or api_key on the request) — sent as the header x-goog-api-key: AIza… |
| Docs | https://ai.google.dev/gemini-api/docs |
| API version | Path-versioned; PuffinParse uses /v1beta (the Files API and the newest models live there). /v1 serves the same generateContent method. |
| Modes | parse, ocr (derived from parse), extract |
| Checked against | 2026-09-11 against the documented wire format; the live call was not exercised — see §6, "Unverified against a live key" |
| Implementation | crates/puffinparse-core/src/providers/gemini.rs |
Gemini is not a document-AI product but a general vision LLM: PuffinParse sends the file plus a
transcription prompt and pins the answer's shape with generationConfig.response_schema
(structured output), so the model must return one entry per page instead of a single markdown blob.
There is exactly one HTTP call per parse — no upload step, no polling — unless the document is
larger than ~14 MB, in which case the resumable Files API is used first.
The consequence of using an LLM is that nothing geometric comes back: no bounding boxes, no page dimensions, no per-block confidence. Anything that needs overlays or coordinates should use a layout provider (Reducto, Extend, LlamaParse, Azure, Textract) instead.
2. Models exposed by PuffinParse#
The PuffinParse model name is the Gemini model id minus the gemini- prefix.
| Model | API model id | Token price (in / out, per 1M) | pricing.json per-page estimate |
|---|---|---|---|
gemini/2.5-flash (default) |
gemini-2.5-flash |
$0.30 / $2.50 | $0.00245 |
gemini/2.5-pro |
gemini-2.5-pro |
$1.25 / $10.00 (>200k prompt tokens: $2.50 / $15.00) | $0.009875 |
gemini/2.5-flash-lite |
gemini-2.5-flash-lite |
$0.10 / $0.40 | $0.00047 |
gemini/3.5-flash |
gemini-3.5-flash |
$1.50 / $9.00 | $0.00945 |
gemini/3.5-flash-lite |
gemini-3.5-flash-lite |
$0.30 / $2.50 | $0.00245 |
gemini/3.8-flash |
gemini-3.8-flash |
$0.75 / $3.75 (introductory, through 2026-12-31; $1.50 / $7.50 after) | $0.004125 |
All six support parse, ocr and extract — the mode is a prompt + schema, not an endpoint, so
every model serves every mode. The 2.5 trio is the safe default; the 3.x entries come from the
models page and the changelog (3.5 Flash GA 2026-05-19, 3.5 Flash-Lite GA 2026-07-21, 3.8 Flash GA
2026-09-02) and have not been exercised against a live key here.
Cost is computed from tokens, not from pages. Each response's usageMetadata is priced exactly
(promptTokenCount × input + (candidatesTokenCount + thoughtsTokenCount) × output) and lands in
usage.provider_cost_usd, which puffinparse-core prefers over the price table. The per-page numbers in
pricing.json are only a fallback for the case where a response carries no usageMetadata (and for
puffinparse providers-style estimates). They were derived, not measured:
per_page_usd = (1500 × input_price_per_1M + N × output_price_per_1M) / 1e6
N = 800 output tokens for parse/ocr, 300 for extract
so gemini/2.5-flash parse = (1500 × 0.30 + 800 × 2.50) / 1e6 = $0.00245/page, and the extract
column of pricing.json is the same formula with N = 300 ($0.0012/page).
1 500 input tokens/page is a deliberately conservative blend: a PDF page bills at a flat 258 tokens plus the rendered-image tokens, while a full-page scan image is ~1 000–1 600 tokens. Thinking tokens are not in the estimate; with a thinking budget left at its default they can double the output side.
A model that is not in this table can be used without a registry change:
provider_options={"model": "gemini-3.1-pro-preview"} replaces the API model id verbatim (PuffinParse
then has no token prices for it and falls back to the registry model's per-page price).
3. Request flow PuffinParse uses#
One call: POST {base}/v1beta/models/{api_model}:generateContent, headers x-goog-api-key and
content-type: application/json.
{
"contents": [{"role": "user", "parts": [
{"inline_data": {"mime_type": "application/pdf", "data": "JVBERi0xLjQK…"}},
{"text": "Transcribe the attached document to GitHub-flavoured Markdown, page by page. …"}
]}],
"generationConfig": {
"temperature": 0,
"response_mime_type": "application/json",
"response_schema": {
"type": "object",
"properties": {"pages": {"type": "array", "items": {
"type": "object",
"properties": {"page_number": {"type": "integer"}, "markdown": {"type": "string"}},
"required": ["page_number", "markdown"],
"propertyOrdering": ["page_number", "markdown"]}}},
"required": ["pages"]
}
}
}
Input handling.
| Input | What PuffinParse does |
|---|---|
| Path / bytes ≤ 14 MB | base64 into inline_data |
| Path / bytes > 14 MB | resumable Files API upload, then file_data: {file_uri, mime_type} |
| URL | downloaded first (Gemini cannot fetch URLs), then treated as bytes |
The MIME type is sniffed from the magic bytes (%PDF-, PNG, JPEG, WebP, GIF) and only falls back to
the extension guess, because Gemini rejects a mismatched mime_type outright. PDFs are native input
(up to 1 000 pages / 50 MB); the accepted image types are image/png, image/jpeg, image/webp,
image/heic and image/heif — not GIF, which PuffinParse still labels correctly so that Gemini's
rejection names the real reason.
The 14 MB inline cut-off exists because a generateContent request is capped at ~20 MB and base64
inflates the payload by a third. The Files API path is
POST {base}/upload/v1beta/files with X-Goog-Upload-Protocol: resumable and
X-Goog-Upload-Command: start → the x-goog-upload-url response header → a second request with
X-Goog-Upload-Command: upload, finalize carrying the bytes → {"file": {"uri", "name", "state"}};
if state is not ACTIVE PuffinParse polls GET {base}/v1beta/files/{id} (0.5 s, ×1.5, max 5 s) until
it is. Uploaded files expire after 48 hours.
Per mode.
parse— the prompt above plus thepagesschema. The model returns{"pages": [{"page_number": 1, "markdown": "…"}, …]}, so the page split is exact rather than guessed from a separator.ocr— not implemented natively; the trait default derives aTextResponsefromparse(metadata.puffinparse_derived_from = "parse"). Lines come from the page text, words carry no boxes. Passingoutput="text"additionally tells the model to skip Markdown syntax.extract— the request schema is rewritten into Gemini's schema subset (see §4) and used asresponse_schema; the schema andinstructionsare also restated in the prompt, which measurably improves adherence.ExtractResponse.datais the parsed JSON.citations=Trueis accepted but cannot be honoured (metadata.gemini_citations_unsupported = true).
Page selection. Gemini always reads the whole file, so pages is a prompt instruction
("Transcribe ONLY these pages of the PDF: 2-3 … keep the original page numbers"), applied only when
the input is a PDF. It is best effort: the full document is still uploaded and still billed as input
tokens, and a model may ignore the restriction. For a non-PDF input pages is dropped and
metadata.gemini_pages_ignored = true is set.
Where provider_options are merged: the whole object is deep-merged into the request body after
PuffinParse builds it, so generationConfig, safetySettings, systemInstruction, tools, … can all be
set or overridden. Three keys are consumed by PuffinParse and never reach Gemini: model (API model id),
prompt (replaces the built-in instruction entirely) and prompt_suffix (appended to it).
4. Response mapping#
{
"candidates": [{
"content": {"parts": [{"text": "{\"pages\": [{\"page_number\": 1, \"markdown\": \"# …\"}]}"}],
"role": "model"},
"finishReason": "STOP", "index": 0
}],
"usageMetadata": {
"promptTokenCount": 1809, "candidatesTokenCount": 1060, "totalTokenCount": 2869,
"promptTokensDetails": [{"modality": "DOCUMENT", "tokenCount": 1548},
{"modality": "TEXT", "tokenCount": 261}]
},
"modelVersion": "gemini-2.5-flash",
"responseId": "HfPJaNXKO7OXm9IPk7P1kQc"
}
(The full payloads are crates/puffinparse-core/tests/fixtures/gemini_parse_multipage.json and
gemini_extract_invoice.json, which drive the normalisation tests.)
| Gemini field | PuffinParse unified field | Notes |
|---|---|---|
responseId |
ParseResponse.provider_job_id |
Gemini has no job concept; this is the only per-call id. |
candidates[0].content.parts[].text |
— | All text parts are concatenated, then parsed as JSON (a stray ```json fence is stripped defensively). |
pages[].page_number |
Page.page_number |
1-based; falls back to the array index + 1 if the model omits it. |
pages[].markdown |
Page.markdown |
Trimmed. With output="text" the markdown-to-text conversion is stored instead. |
| derived | Page.text |
Always markdown_to_text(markdown). |
| derived | Page.blocks |
Exactly one text block per non-empty page, content = the page, bbox: None, confidence: None. |
| — | Page.width / height |
Always None — Gemini reports no geometry. |
| number of returned pages | Usage.pages |
For an inline PDF the /Type /Page count is compared against it (see below). |
usageMetadata |
Usage.provider_cost_usd |
prompt × input_price + (candidates + thoughts) × output_price. |
| — | Usage.credits |
Always None; Gemini bills tokens, not credits. |
usageMetadata.* |
metadata.gemini_tokens |
{prompt, candidates, thoughts, cached, total}. |
modelVersion |
metadata.gemini_model_version |
The exact served version behind an alias. |
| the API model id sent | metadata.gemini_api_model |
Useful when provider_options.model overrode it. |
finishReason ≠ STOP |
metadata.gemini_finish_reason |
|
promptFeedback.blockReason |
→ error | bad_request, message gemini blocked the prompt: <reason>. |
For extract, the parsed JSON becomes ExtractResponse.data verbatim, fields stays empty (no
citations, no per-field confidence) and usage.pages is the PDF page count, or 1 for an image.
PDF page-count sanity check. For an inline PDF PuffinParse counts /Type /Page objects (ignoring
/Type /Pages) in the raw bytes and records it as metadata.gemini_pdf_page_count. If it differs
from the number of pages the model returned — a dropped or hallucinated page, or simply a
page-selection request — metadata.gemini_page_count_mismatch = true is set and a warning is logged.
Usage.pages still follows what the model returned. The count is best effort and inline-only:
PDFs whose page objects live in compressed object streams return no count, and neither does a
document that went through the Files API — in extract mode that means usage.pages falls back
to 1.
sanitize_schema (public, unit-tested) rewrites a draft-2020-12 JSON Schema into the subset
Gemini's response_schema accepts:
- keeps only
type,format,title,description,nullable,enum,items,prefixItems,properties,required,propertyOrdering,minItems,maxItems,minimum,maximum,minLength,maxLength,pattern,anyOf— everything else ($schema,$id,additionalProperties,unevaluatedProperties,default,examples,allOf,not,if/then,patternProperties, …) is dropped; type: ["string", "null"]→type: "string"+nullable: true;oneOf→anyOf;const: X→enum: [X]with the type inferred;- local
$refs into$defs/definitionsare inlined (depth-limited to 12; a recursive schema degrades gracefully instead of expanding forever, and an unresolvable$refbecomes{"type": "object"}); - a node with
propertiesbut notypegetstype: "object", withitemsgetstype: "array".
5. Errors, status codes, rate limits, timeouts#
Errors are {"error": {"code": 400, "message": "…", "status": "INVALID_ARGUMENT", "details": [...]}};
Error::from_http lifts error.message and classifies by status:
| Status | status field |
Trigger | PuffinParse ErrorKind |
|---|---|---|---|
| 400 | INVALID_ARGUMENT |
malformed body, unsupported mime_type, invalid API key, schema Gemini rejects |
bad_request |
| 400 | FAILED_PRECONDITION |
free tier not available in the caller's country / billing required | bad_request |
| 403 | PERMISSION_DENIED |
key lacks access to the model, or key restrictions (IP/referrer) | authentication |
| 404 | NOT_FOUND |
unknown model id, or an expired Files API file_uri |
bad_request |
| 429 | RESOURCE_EXHAUSTED |
per-minute request/token quota | rate_limit (retried) |
| 500 / 503 | INTERNAL / UNAVAILABLE |
server error, model overloaded | provider (retried) |
| 504 | DEADLINE_EXCEEDED |
prompt too large to finish in time | provider (retried) |
Note the odd one: an invalid API key is a 400, not a 401, so it surfaces as bad_request rather
than authentication. A missing key is caught before the call and is authentication.
Failures that arrive with HTTP 200 are turned into errors too:
promptFeedback.blockReasonset, or a candidate withfinishReason: "SAFETY"and no text →bad_request/providerwith the reason in the message.finishReason: "MAX_TOKENS"→provider, "gemini hit the output token limit, so the JSON is truncated — raisegenerationConfig.maxOutputTokens… or parse fewer pages per call".- Text that is not the promised JSON →
provider, with the first 200 characters in the message.
Retries. max_retries (default 2) with exponential backoff and full jitter, on 429, 5xx and
network errors only (the shared http::with_retry). Gemini's 429 body sometimes carries a
RetryInfo detail with retryDelay; PuffinParse does not read it yet and backs off blind.
Rate limits. Counted per project and per model in three dimensions — requests/minute, input tokens/minute and requests/day — and they depend on the usage tier, so the authoritative numbers are the ones shown in Google AI Studio rather than any number quoted here. A single large document can exhaust the token dimension in very few requests.
Timeouts. timeout_secs (default 300) is the whole-call deadline and caps every individual
request, including the Files API upload and the URL download.
6. Gotchas#
- Unverified against a live key. Everything here follows the published wire format, and both the
normalisation path (fixture tests) and the transport (
mod transport— a loopback HTTP server that asserts the exact URL, headers, base64inline_data, schema, 429 retry and 400 mapping) are covered by tests, but theGEMINI_API_KEYpresent in the development environment is rejected by Google with400 API_KEY_INVALIDon every endpoint, so no call has ever reached the real service. The fixtures undertests/fixtures/gemini_*.jsonwere therefore constructed from the documented response schema, not captured. Runcargo test -p puffinparse-core gemini_live -- --ignored --nocapturewith a working key, then replace the fixtures with real payloads and fill in measured latency and cost here. - No geometry, ever.
Block.bbox,Page.width/heightandBlock.confidenceare alwaysNone.ocrmode returns words without boxes. This is a property of the model, not of PuffinParse. - The page split comes from the model, not from the file. Structured output makes it reliable in
practice, but a model can still merge or drop a page.
metadata.gemini_page_count_mismatchis the tripwire for PDFs; there is no equivalent check for multi-page TIFFs or images. pagesis a suggestion. Whole-file upload, prompt-level selection, full input billing.- Thinking tokens are billed as output and are invisible in
candidatesTokenCount. On 2.5 models a page of dense tables can spend more on reasoning than on the transcript. Setprovider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}}to turn thinking off on Flash / Flash-Lite (Pro cannot go below 128; Gemini 3 models usethinkingLevelinstead). - Long documents hit
MAX_TOKENSbefore they hit the page limit. The 1 000-page ceiling is theoretical: the transcript of ~40 dense pages already approaches the 64k output budget. Split long PDFs withpages, or raisegenerationConfig.maxOutputTokens. temperature: 0does not make the output deterministic. Repeated runs differ slightly, which matters for benchmark reproducibility.- Gemini cannot fetch a URL. PuffinParse downloads it first; a URL that needs authentication has to be fetched by the caller and passed as bytes.
additionalPropertiesis the classic 400. Most schema generators (Pydantic, zod) emit it and Gemini rejects it;sanitize_schemastrips it, along with$schema,defaultandexamples. The current docs page listsadditionalPropertiesas supported, but rejecting it has been the observed behaviour for long enough that stripping is the safe default.- The API is versioned in the path and the newest surface is moving. Google's Interactions API
(
POST /v1beta/interactions, GA June 2026) is now the documented default and uses a different shape (input,steps,response_format).generateContentremains supported and is still the recommended path for stable deployments — PuffinParse targets it deliberately; expect the docs links below to show the Interactions shape. - Files API objects expire after 48 hours and are scoped to the project, so a
file_uricannot be reused across keys. PuffinParse uploads per call and never reuses or deletes (files are free and capped at 20 GB per project). - Image vs. PDF tokenisation differ. A PDF page is a flat 258 tokens plus image tokens; a standalone image is tiled at 768×768 (≈258 tokens per tile). A scan sent as PNG can therefore cost several times what the same page costs inside a PDF.
x-goog-api-keyand?key=are equivalent; PuffinParse uses the header so keys never land in request logs or URLs.
7. Useful provider_options passthrough#
# 1. Turn thinking off for cheap, fast transcription (Flash / Flash-Lite only).
puffinparse.parse("scan.png", model="gemini/2.5-flash",
provider_options={"generationConfig": {"thinkingConfig": {"thinkingBudget": 0}}})
# 2. Raise the output budget for a long PDF.
puffinparse.parse("report.pdf", model="gemini/2.5-pro",
provider_options={"generationConfig": {"maxOutputTokens": 65536}})
# 3. Reach a model that is not in the registry.
puffinparse.parse("doc.pdf", model="gemini/2.5-flash",
provider_options={"model": "gemini-3.1-pro-preview"})
# 4. Steer the transcription without rewriting the whole prompt.
puffinparse.parse("statement.pdf", model="gemini/2.5-flash",
provider_options={"prompt_suffix": "Keep every stamp and handwritten note, "
"and transcribe struck-through text as ~~text~~."})
# 5. Replace the instruction entirely (the pages schema still applies).
puffinparse.parse("form.pdf", model="gemini/2.5-flash-lite",
provider_options={"prompt": "Return each page of this form as a Markdown table of "
"field name and value, one row per field."})
# 6. Loosen safety filters for documents that trip them (medical, legal, incident reports).
puffinparse.parse("report.pdf", model="gemini/2.5-flash",
provider_options={"safetySettings": [
{"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_NONE"},
{"category": "HARM_CATEGORY_HARASSMENT", "threshold": "BLOCK_NONE"}]})
# 7. A system instruction, e.g. to pin the output language.
puffinparse.parse("brief.pdf", model="gemini/2.5-flash",
provider_options={"systemInstruction": {"parts": [{"text": "Always answer in German."}]}})
8. Links#
- Docs home: https://ai.google.dev/gemini-api/docs
- Document (PDF) understanding: https://ai.google.dev/gemini-api/docs/document-processing
- Image understanding: https://ai.google.dev/gemini-api/docs/image-understanding
- Structured output: https://ai.google.dev/gemini-api/docs/structured-output
- Models: https://ai.google.dev/gemini-api/docs/models · live list:
GET /v1beta/models - Pricing: https://ai.google.dev/gemini-api/docs/pricing
- Rate limits: https://ai.google.dev/gemini-api/docs/rate-limits
- Files API: https://ai.google.dev/gemini-api/docs/files
generateContentreference: https://ai.google.dev/api/generate-content- API versions: https://ai.google.dev/gemini-api/docs/api-versions
- Interactions API (the newer surface PuffinParse does not use): https://ai.google.dev/gemini-api/docs/migrate-to-interactions