PuffinParse gateway (puffinparse serve)#
A small HTTP server in front of every provider PuffinParse supports, in the spirit of the LiteLLM proxy: clients call one endpoint with one API shape and a gateway-issued key; the gateway holds the provider keys, routes aliases to provider models with fallbacks, enforces per-key model lists, monthly budgets and rate limits, and emits JSON-lines request logs and Prometheus metrics.
It is a thin layer over puffinparse-core (crate crates/puffinparse-server, axum + tower-http). It adds
no provider logic: every call is puffinparse_core::parse / ocr / extract with the request fields
below, so responses are exactly the unified types of SPEC §5, or
a vendor shape via output_format (COMPAT.md). Long documents can go through the
async jobs API instead (POST /v1/jobs + GET /v1/jobs/{id}, core submit_parse /
retrieve_parse, SPEC §15), so no connection is held
open while the provider works.
Run it#
export PUFFINPARSE_MASTER_KEY=sk-master-change-me
export PUFFINPARSE_KEY_BILLING=sk-billing-change-me PUFFINPARSE_KEY_RESEARCH=sk-research-change-me
export REDUCTO_API_KEY=... EXTEND_API_KEY=... LLAMA_API_KEY=...
puffinparse serve --config examples/server/puffinparse.toml # --host / --port override the file
Without --config it reads ./puffinparse.toml if present (or $PUFFINPARSE_CONFIG), else runs with
defaults: 127.0.0.1:4000, no aliases, no auth (a warning is printed). Anything reachable
beyond localhost should have a master_key or [[keys]].
Docker (multi-stage build, distroless runtime, runs as non-root). Releases publish a linux/amd64
image to ghcr.io/ajinkyashejul/puffinparse (tags latest, 0.1.0, 0.1); or build it yourself
with docker build -t puffinparse .:
docker run --rm -p 4000:4000 -v $PWD/examples/server/puffinparse.toml:/etc/puffinparse/puffinparse.toml:ro \
-e PUFFINPARSE_MASTER_KEY -e PUFFINPARSE_KEY_BILLING -e PUFFINPARSE_KEY_RESEARCH \
-e REDUCTO_API_KEY -e EXTEND_API_KEY -e LLAMA_API_KEY ghcr.io/ajinkyashejul/puffinparse
The image's default command is serve --host 0.0.0.0 --config /etc/puffinparse/puffinparse.toml; it exits
if no config is mounted rather than serving an open gateway.
Configuration (puffinparse.toml)#
Every secret may be written as "env:VAR", resolved once at startup. A virtual key or master key
whose reference resolves to nothing is a startup error (fail closed); a provider key that resolves
to nothing prints a warning and the core falls back to the provider's default env var.
master_key = "env:PUFFINPARSE_MASTER_KEY" # full access: all models, no budget, no rate limit
[server]
host = "127.0.0.1"
port = 4000
log_stdout = true # JSON-lines request log on stdout
log_file = "requests.jsonl" # optional append-only copy
state_file = "usage.json" # optional: per-key monthly spend survives restarts
max_body_mb = 50 # multipart upload / base64 JSON body cap
max_timeout_secs = 300 # cap (and default) for a request's `timeout`
max_retries = 2 # per provider call, when the client sends none
allow_direct_models = true # false: only the aliases below may be requested
job_retention_hours = 168 # how long POST /v1/jobs handles stay readable
[webhooks] # optional provider webhook receiver (off by default)
enabled = false
secret = "env:PUFFINPARSE_WEBHOOK_SECRET" # required when enabled; sent as ?token= or a header
[providers.reducto] # one table per provider (name as in `puffinparse providers`)
api_key = "env:REDUCTO_API_KEY"
# base_url = "https://eu.platform.reducto.ai"
[[models]] # an alias clients can send as "model"
name = "invoices"
targets = ["reducto/standard", { model = "extend/parse_performance", api_key = "env:EXTEND_KEY_2" }]
strategy = "ordered" # or "round_robin"
fallback_on = ["provider", "rate_limit", "timeout", "network"] # the default
[[keys]]
id = "billing-team" # appears in logs, metrics, usage; never the secret
key = "env:PUFFINPARSE_KEY_BILLING" # the bearer token the client sends
models = ["invoices", "llamaparse/*"] # aliases, provider/model, provider/*, or "*"; empty = all
monthly_budget_usd = 50.0 # calendar month, UTC
rpm = 60 # requests per minute, sliding 60 s window
A complete sample is examples/server/puffinparse.toml.
Routing. model is looked up as an alias first, then (if allow_direct_models) as a registry
model (reducto/standard, or a bare provider for its default model in that mode). An alias expands
to its targets: ordered always starts at the first, round_robin rotates the start per request,
and both fall back in order when the error kind is in fallback_on. A request's own fallbacks
list is appended after that (aliases expand in order). Every target is checked against the
endpoint's mode before any provider is called. Target-level api_key / base_url override the
[providers.*] ones, so two deployments of the same provider (two accounts, two regions) can sit
behind one alias. When a fallback served the call, the response metadata carries
puffinparse_fallback_index and puffinparse_fallback_from_error, as with the SDK Router.
Budgets. Spend is the response's cost_usd (provider-reported cost when available, otherwise
the list-price estimate from pricing.json), summed per key per UTC calendar month. The check runs
before the call: a key whose spend has reached its budget gets 402. One in-flight request can take
a key past its budget; nothing is charged for failed attempts (a provider may still bill a failed
job). An async job is charged to the key that submitted it, once, when it is first observed
succeeded (by GET /v1/jobs/{id} or a webhook); submitting and polling are free. State is in
memory; with state_file it (spend, counts and submitted jobs) is rewritten (temp file + rename)
after every change and reloaded on startup. Rate-limit windows are not persisted.
API#
Authenticate with Authorization: Bearer <key> or x-api-key: <key>. Send x-request-id to
choose the request id (echoed in the x-request-id response header and the log), otherwise a
UUID is generated.
POST /v1/parse, POST /v1/ocr, POST /v1/extract#
Body is JSON or multipart/form-data with the same field names (multipart fields are text; the
JSON-valued ones — provider_options, schema, metadata, fallbacks — are JSON strings, and
fallbacks may also be comma-separated).
| Field | Type | Notes |
|---|---|---|
model |
string | Required. Alias or provider/model. |
document_url |
string | Public http(s) URL, passed to the provider. |
document |
string | Base64 bytes (a data:…;base64, prefix is accepted). Needs filename. |
file |
multipart file part | The upload; its filename sets the type (or send filename). |
filename |
string | Sets the MIME type for document / overrides the part's filename. |
pages, language |
string | As in SPEC §4.1. |
output |
"markdown" | "text" |
Block content format (parse). |
output_format |
"reducto" | "extend" | "llamaparse" | "puffinparse" |
Vendor-native response shape (parse, extract). |
provider_options |
object | Merged into the provider request verbatim. |
include_raw |
bool | Attach the provider payload as raw. |
timeout |
number (s) | Capped at server.max_timeout_secs. |
max_retries |
int | Per provider call (max 10). |
fallbacks |
string[] | Extra aliases/models tried after model's own targets. |
metadata |
object | Echoed back in the response. |
schema, instructions, citations |
/v1/extract only; schema required there. |
Exactly one of document_url, document, file. api_key and base_url are rejected: a
client must never be able to point the gateway's provider credentials at another host. Local file
paths are not accepted either.
The response is the unified ParseResponse / TextResponse / ExtractResponse JSON (or the
vendor shape), with headers x-puffinparse-model (served model), x-puffinparse-cost-usd, x-request-id.
POST /v1/jobs, GET /v1/jobs/{id} (async parse)#
POST /v1/jobs takes the /v1/parse body (JSON or multipart, same fields) plus an optional
webhook_url, uploads the document, starts the provider job and answers 202 at once:
{"id": "job_4f0c…", "object": "job", "status": "pending",
"job": {"provider": "reducto", "model": "reducto/standard", "job_id": "c1e2…",
"submitted_at": "2026-09-24T20:23:52.140Z", "output": "markdown", "include_raw": false}}
job is the core JobHandle (SPEC §15) minus the operator's base_url. Authentication, the
key's model allow-list, aliases, the budget pre-check and rpm apply exactly as for
/v1/parse. Only providers with a job queue can take jobs (reducto, extend, llamaparse;
others give 400 unsupported_model_error). Fallback does not apply to jobs: an alias submits
to its first target (in round_robin, the next one in rotation) and a request fallbacks list is
rejected, because a provider failure only shows up later, on retrieve. webhook_url is
forwarded to the provider's per-job webhook (Reducto async.webhook, LlamaParse webhook_url;
Extend rejects it) and is refused by the synchronous endpoints.
GET /v1/jobs/{id} asks the provider once and returns
{"id": "job_4f0c…", "object": "job", "status": "pending" | "succeeded" | "failed",
"model": "reducto/standard", "provider": "reducto", "provider_job_id": "c1e2…",
"submitted_at": "…", "result": {…ParseResponse…}, "error": {…error object…}}
with result only when succeeded (the unified ParseResponse, or the vendor shape for the
output_format sent at submit time; ?output_format=reducto overrides it per call) and error
only when failed (the same object as an error body's error, with the provider's message and
job_id). A failed job is still HTTP 200; a failed status check (provider credentials,
network) is an error response as usual. Job ids are opaque and bound to the key that created the
job: any other key gets 404 not_found, exactly as for an unknown id (the master key can read all
jobs). Polls count toward the key's rpm but are not budget-gated. The job's cost is charged to
its owner once, the first time it is seen succeeded. Credentials are resolved from the config on
every poll (the alias target's own key, else [providers.*]), never stored; handles are kept for
job_retention_hours (and in state_file when set).
POST /v1/webhooks/{provider} (optional)#
Off by default (404). With [webhooks] enabled = true and a secret, the gateway accepts the
body a provider POSTs when a job changes state — Reducto direct webhooks, Extend parse_run.*
events, LlamaCloud parse.* events — authenticated by ?token=<secret> or the
x-puffinparse-webhook-secret header (compared in constant time). The body is read with core
parse_webhook; when it only says the job finished, the gateway makes one status check. The
provider's job id must match a job submitted through this gateway (else 404); a success is
charged to the job's owner (once, shared with GET), and the answer is an acknowledgement
{"ids": [...], "status"} — clients still collect the result with GET /v1/jobs/{id}. ids can
hold several gateway jobs: LlamaParse returns the same (cached) job id for an identical upload, so
two submissions of the same file share one provider job, and each is settled and charged to its
own key. Point a provider's
webhook at https://<gateway>/v1/webhooks/reducto?token=<secret> (per job via webhook_url, or a
workspace-level endpoint for Extend / LlamaCloud). This is a shared secret, not the vendors' HMAC
signatures: keep the URL private and serve the gateway over TLS. A LlamaParse webhook_url
result push that names no job id cannot be attributed and is rejected with 400.
Errors#
Every error has the same body:
{"error": {"type": "provider_error", "message": "…the provider's own message…", "provider": "reducto",
"provider_status": 503, "job_id": null, "request_id": "5c0f…"}}
| HTTP | type |
Cause |
|---|---|---|
| 400 | input_error, bad_request_error, unsupported_model_error |
Malformed request, provider 4xx, unknown model or model without this mode |
| 401 | unauthorized |
Missing or unknown gateway key |
| 402 | budget_exceeded |
Key's monthly budget spent |
| 403 | model_not_allowed |
Model (or a fallback) not in the key's models |
| 404 | not_found |
Unknown job id, a job another key owns, or webhooks disabled |
| 413 | payload_too_large |
Body over max_body_mb |
| 429 | key_rate_limited |
Key's rpm reached (Retry-After set) |
| 429 | rate_limit_error |
Provider rate limit after retries and fallbacks |
| 502 | provider_error, network_error, authentication_error |
Provider 5xx / failed job, network failure, provider rejected the gateway's credentials |
| 504 | timeout_error |
timeout exceeded |
Provider credential failures are 502, not 401: the caller's key was fine, the operator's was not.
GET /v1/models#
Aliases (targets, strategy, the modes every target supports) and, with allow_direct_models,
every registry model with its modes, per-page list price per mode and whether a provider key is
configured. Filtered to what the caller's key may use.
GET /v1/usage#
This month's spend, requests, pages, remaining budget and limits: the caller's own key, or all keys for the master key.
GET /health, GET /metrics#
/health → {"status": "ok", "version": …}. /metrics is Prometheus text, unauthenticated
(it carries model names, counts and costs, no secrets or content; firewall it if that matters):
| Metric | Labels |
|---|---|
puffinparse_requests_total |
mode, model (served model, or requested alias on failure, - if it never resolved), status |
puffinparse_errors_total |
type (the error type above) |
puffinparse_request_duration_seconds (histogram, 0.25 s – 300 s buckets) |
mode |
puffinparse_pages_total |
model |
puffinparse_cost_usd_total |
model |
puffinparse_fallbacks_total |
— |
puffinparse_jobs_total |
event: submitted, and succeeded / failed the first time a job is seen terminal |
mode is parse / ocr / extract for the synchronous endpoints and job_submit,
job_retrieve, webhook for the jobs API, so job latencies do not mix with blocking calls. A
job's pages and cost are counted once, when it is first seen succeeded.
Request log#
One JSON object per request, on stdout and/or log_file:
{"ts":"2026-09-24T20:23:52.140Z","request_id":"bf397168-…","key_id":"demo","method":"POST",
"path":"/v1/parse","mode":"parse","model":"cheap","served_model":"llamaparse/cost_effective",
"provider":"llamaparse","fallback_index":0,"pages":1,"cost_usd":0.00375,"latency_ms":10384,
"status":200,"error_type":null,"provider_status":null}
Jobs API lines add job_id (the gateway id) and job_status (pending / succeeded /
failed, as observed by that request); method is GET for status checks, and key_id is
webhook for provider webhooks. pages and cost_usd are set only on the request that charged
the job.
The record has no field for document bytes, URLs, extracted content, provider error text (some providers echo document text in errors), provider keys or gateway key secrets.
curl examples#
GW=http://127.0.0.1:4000; KEY=sk-billing-change-me
# Upload a file (multipart) to an alias
curl -s $GW/v1/parse -H "Authorization: Bearer $KEY" -F model=invoices -F file=@invoice.pdf | jq .markdown
# URL input, vendor-native shape, a request-level fallback
curl -s $GW/v1/parse -H "Authorization: Bearer $KEY" -H 'content-type: application/json' -d '{
"model": "invoices", "document_url": "https://example.com/invoice.pdf",
"output_format": "reducto", "fallbacks": ["llamaparse/cost_effective"]}'
# Base64 input, OCR mode
curl -s $GW/v1/ocr -H "x-api-key: $KEY" -H 'content-type: application/json' \
-d "{\"model\": \"reducto\", \"filename\": \"scan.png\", \"document\": \"$(base64 -w0 scan.png)\"}" | jq .text
# Extract with a schema
curl -s $GW/v1/extract -H "Authorization: Bearer $KEY" -F model=invoice-fields -F file=@invoice.pdf \
-F 'schema={"type":"object","properties":{"total":{"type":"number"}}}'
# Async job: submit, then poll until status is succeeded or failed
JOB=$(curl -s $GW/v1/jobs -H "Authorization: Bearer $KEY" -F model=invoices -F file=@big.pdf | jq -r .id)
curl -s $GW/v1/jobs/$JOB -H "Authorization: Bearer $KEY" | jq .status
curl -s "$GW/v1/jobs/$JOB?output_format=reducto" -H "Authorization: Bearer $KEY" | jq .result
curl -s $GW/v1/models -H "Authorization: Bearer $KEY" | jq '.data[].id'
curl -s $GW/v1/usage -H "Authorization: Bearer $KEY"
curl -s $GW/metrics
Not in scope (yet)#
Streaming, jobs for ocr / extract (jobs are parse-only, like the core), fallback for jobs,
the gateway registering its own webhook URL with providers automatically, vendor HMAC signature
checks on webhooks, key management over HTTP (keys live in the config file; restart to change them), a database, response
caching, per-key budgets by model, TLS termination (put it behind a reverse proxy), and metrics
auth.