AWS Textract#
Status: docs-only. Implemented from AWS's published API documentation and tested against fixture payloads built from it. It has not yet been run against the live API, so expect wire-format differences. Help verify it: issue #10.
1. Summary#
| Provider name | textract |
| Base URL | https://textract.{region}.amazonaws.com (override: base_url on the request, or TEXTRACT_BASE_URL) |
| API key | A key pair, not a single token. AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY, optional AWS_SESSION_TOKEN, region from AWS_REGION (then AWS_DEFAULT_REGION, then us-east-1). api_key on the request overrides the access key id only. |
| Auth header | Authorization: AWS4-HMAC-SHA256 … — SigV4, computed in providers/textract.rs; no AWS SDK crate is vendored |
| Protocol | AWS JSON 1.1 RPC: always POST /, operation chosen by X-Amz-Target, Content-Type: application/x-amz-json-1.1 |
| API version | textract-2018-06-27 (implicit in the target names; there is no version header) |
| Checked against | 2026-09-11 — request/response shapes and quotas from the AWS API reference; pricing from the AWS Price List API (AmazonTextract, us-east-1, publication 2026-08-31). Response fixtures are built from the AWS documentation's own examples, not from a live capture (see §6). |
| Implementation | crates/puffinparse-core/src/providers/textract.rs |
Textract is not a "parse a document to markdown" product: it returns a flat array of Block objects
linked by ids. PuffinParse reassembles those into pages, blocks, markdown tables and extraction results.
There is no upload endpoint and no remote-URL input — synchronous calls carry the document inline as
base64, asynchronous ones read it from S3.
2. Models exposed by PuffinParse#
| Model | Modes | Textract operation + FeatureTypes |
List price (pricing.json) |
|---|---|---|---|
textract/detect-text (default for ocr) |
ocr |
DetectDocumentText (no features) |
$0.0015 / page |
textract/layout |
parse, ocr |
AnalyzeDocument, ["LAYOUT", "TABLES"] |
$0.015 / page |
textract/queries (default for extract) |
extract |
AnalyzeDocument, ["QUERIES"] + QueriesConfig |
$0.015 / page |
textract/forms |
extract |
AnalyzeDocument, ["FORMS"] |
$0.050 / page |
Prices are the us-east-1 pay-as-you-go list prices for the first 1M pages/month, taken from the
AWS Price List API rather than the marketing page, and used only to fill
cost_usd = per_page_usd × usage.pages:
| Usage type (us-east-1) | $ / page (0–1M) | $ / page (1M+) |
|---|---|---|
USE1-SyncTextPagesProcessed (DetectDocumentText) |
0.0015 | 0.0006 |
USE1-SyncTablesPagesProcessed (AnalyzeDocument TABLES) |
0.015 | 0.010 |
USE1-SyncQueriesPagesProcessed (AnalyzeDocument QUERIES) |
0.015 | 0.010 |
USE1-SyncFormsPagesProcessed (AnalyzeDocument FORMS) |
0.050 | 0.040 |
USE1-SyncLayoutPagesProcessed (AnalyzeDocument LAYOUT alone) |
0.004 | 0.003 |
Why textract/layout is priced at the TABLES rate and not TABLES + LAYOUT. AWS bills a combined
AnalyzeDocument call at a single combination rate, and there is no Layout…Tables usage type in the
price list: "Layout is available for free when used with the Tables feature." So ["LAYOUT","TABLES"]
bills exactly like ["TABLES"] — $0.015 / page. A LAYOUT-only call would be $0.004 / page, which you
can get with provider_options={"FeatureTypes": ["LAYOUT"]} (the passthrough replaces the array), at
the cost of losing markdown tables. Prices are region-dependent and PuffinParse's table is us-east-1 only,
so cost_usd is an estimate for any other region.
Async (Start*/Get*) pages cost the same as sync pages; the USE1-Async… usage types carry
identical rates.
3. Request flow PuffinParse uses#
Every call is POST {base}/ with these headers, all of them signed:
Content-Type: application/x-amz-json-1.1
X-Amz-Target: Textract.<Operation>
X-Amz-Date: 20260911T123456Z
X-Amz-Content-Sha256: <sha256 hex of the body>
X-Amz-Security-Token: <AWS_SESSION_TOKEN> # only when set
Authorization: AWS4-HMAC-SHA256 Credential=<id>/<date>/<region>/textract/aws4_request,
SignedHeaders=content-type;host;x-amz-content-sha256;x-amz-date;
[x-amz-security-token;]x-amz-target,
Signature=<hex>
SigV4 is implemented in-file with hmac + sha2 (~60 lines: canonical request → string to sign →
four chained HMACs for the signing key). It is re-signed on every retry, because the signature
expires with X-Amz-Date. Unit tests assert AWS's own published iam/ListUsers example vectors
(signing key c4afb1cc…a4b9, signature 5d672d79…b5d7), so canonicalisation drift is caught
without a network call.
3.1 Synchronous (default)#
- Load bytes. Path and bytes inputs are read locally. A URL input is downloaded by PuffinParse and sent inline — Textract cannot fetch URLs itself.
- Multi-page guard. If the bytes are a PDF,
pdf_page_count()counts/Type /Pageobjects (falling back to the page tree's/Count). More than one page ⇒ abad_requesterror before any billable call (§5). - Call.
Textract.DetectDocumentTextorTextract.AnalyzeDocumentwith
json
{
"Document": { "Bytes": "<base64 of the file>" },
"FeatureTypes": ["LAYOUT", "TABLES"],
"QueriesConfig": { "Queries": [{ "Text": "What is the total?", "Alias": "total" }] }
}
FeatureTypes is omitted for detect-text; QueriesConfig only for textract/queries.
3.2 Asynchronous (multi-page, provider_options.s3_object)#
Textract's async API reads only from S3 — there is no way to hand it bytes. PuffinParse has no S3 client and deliberately does not grow one, so you upload the object yourself and name it:
provider_options={"s3_object": {"bucket": "my-bucket", "name": "invoices/2026-q3.pdf"}}
# optional: "version": "<S3 object version id>"
Then PuffinParse:
Textract.StartDocumentTextDetection/Textract.StartDocumentAnalysiswith{"DocumentLocation": {"S3Object": {"Bucket": …, "Name": …}}, "FeatureTypes": […]}→{"JobId": …}.- Polls
Textract.GetDocumentTextDetection/Textract.GetDocumentAnalysiswith{"JobId": …, "MaxResults": 1000}starting at 2 s, backing off ×1.5 to 10 s, untilJobStatusleavesIN_PROGRESS.SUCCEEDEDandPARTIAL_SUCCESScontinue; anything else becomes aprovidererror carryingStatusMessageand the job id. - Follows
NextTokenuntil it is absent, concatenating everyBlockspage (a 3 000-page document is many round trips — budgettimeout_secsaccordingly).
NotificationChannel / SNS is not used: PuffinParse polls. You can still set it (and OutputConfig,
KMSKeyId, JobTag, ClientRequestToken, AdaptersConfig) through provider_options, which is
deep-merged into whichever body is actually sent.
Where provider_options are merged: the keys region and s3_object are consumed by PuffinParse;
everything else is deep-merged verbatim into the Textract request body, so the remaining keys are
top-level Textract request members in PascalCase (FeatureTypes, QueriesConfig, AdaptersConfig,
HumanLoopConfig, OutputConfig, KMSKeyId, JobTag, ClientRequestToken, NotificationChannel).
3.3 Page selection#
Textract has no page-range parameter. pages="1-3,7" is therefore applied client-side: blocks
whose Page falls outside the selection are dropped after the response arrives. It reduces the
output, never the bill. The one exception is textract/queries in async mode, where the ranges are
also forwarded as Query.Pages (["1-3", "7"]), which Textract does honour.
language is ignored — Textract auto-detects (English, French, German, Italian, Portuguese, Spanish)
and never reports which language it found.
4. Response mapping#
Every operation returns the same envelope: {"Blocks": [...], "DocumentMetadata": {"Pages": n},
"AnalyzeDocumentModelVersion": "1.0"} (plus JobStatus / NextToken / Warnings for Get*).
| Textract field | PuffinParse unified field | Notes |
|---|---|---|
DocumentMetadata.Pages |
Usage.pages |
Falls back to the highest Block.Page, minimum 1. |
JobId (async only) |
provider_job_id |
None for synchronous calls — Textract returns no id there. |
Geometry.BoundingBox.{Left,Top,Width,Height} |
Block.bbox / Line.bbox / Word.bbox |
Already normalised 0–1, origin top-left; from_normalized_ltwh only clamps and converts to {x0,y0,x1,y1}. Geometry.Polygon and RotationAngle are ignored. |
Confidence (0–100) |
confidence (0–1) |
Divided by 100 and clamped. |
Page |
page_number |
1-based. Always 1 for JPEG/PNG, even multi-page scans. |
| — | Page.width / height |
Always None: Textract never reports page dimensions. |
AnalyzeDocumentModelVersion |
metadata.textract_model_version |
|
| resolved region | metadata.textract_region |
|
FeatureTypes sent |
metadata.textract_feature_types |
|
Warnings[] |
metadata.textract_warnings |
Present only when non-empty (e.g. INVALID_REQUEST_PARAMETERS for an over-quota query page). |
| — | Usage.credits, Usage.provider_cost_usd |
Always None; cost_usd comes from pricing.json. |
4.1 ocr mode — LINE / WORD#
LINE blocks become TextPage.lines, WORD blocks become TextPage.words, both with box and
confidence; TextPage.text is the lines joined by \n in response order. This is the native path
for textract/detect-text, and textract/layout uses it too — AnalyzeDocument returns all lines
and words regardless of FeatureTypes, so no second call is needed (and no extra charge).
4.2 parse mode — LAYOUT_* + TABLE#
Textract returns LAYOUT_* blocks in implied reading order (left to right, top to bottom;
column by column on multi-column pages). PuffinParse walks them in array order, so Page.markdown is in
reading order. A layout block's text is its descendant LINE blocks via Relationships[CHILD].
BlockType |
Block.type |
Block.content |
|---|---|---|
LAYOUT_TITLE |
title |
# <text> |
LAYOUT_SECTION_HEADER |
section_header |
## <text> |
LAYOUT_HEADER |
header |
plain text |
LAYOUT_FOOTER |
footer |
plain text |
LAYOUT_TEXT, LAYOUT_KEY_VALUE |
text |
plain text, one line per child LINE |
LAYOUT_LIST |
list |
- per child LAYOUT_TEXT |
LAYOUT_TABLE |
table |
markdown table (below) |
LAYOUT_FIGURE |
figure |
the caption lines, if any |
LAYOUT_PAGE_NUMBER |
other |
plain text |
LAYOUT_LIST points at LAYOUT_TEXT children, and those children also appear at the top level of
Blocks; PuffinParse suppresses any layout block that is another layout block's child so list items are
not emitted twice.
Tables. A LAYOUT_TABLE is matched to its TABLE block by a direct CHILD reference, or — when
it only points at lines — by the highest-overlap unclaimed TABLE on the same page (>10 % of the
table's area). The TABLE's CHILD CELL blocks are placed on a RowIndex × ColumnIndex grid
(MERGED_CELL blocks are skipped: they repeat content already present in the individual cells) and
rendered as markdown. Cell text is the CHILD WORD texts joined by spaces; a SELECTION_ELEMENT
child becomes [x] / [ ]; | is escaped. The header row is the lowest RowIndex among cells with
EntityTypes: ["COLUMN_HEADER"], else row 1; rows above the header row and TABLE_TITLE /
TABLE_FOOTER blocks are emitted as plain lines around the table.
Fallback. If a response carries no LAYOUT_* blocks at all (LAYOUT disabled through
provider_options, or nothing detected), PuffinParse emits every TABLE as a markdown table plus one
text block per LINE, skipping lines whose words are all inside table cells so table content is
not duplicated. output="text" renders every block through markdown_to_text.
Trimmed fixture (crates/puffinparse-core/tests/fixtures/textract_layout.json, most blocks elided):
{
"DocumentMetadata": { "Pages": 1 },
"AnalyzeDocumentModelVersion": "1.0",
"Blocks": [
{ "BlockType": "LAYOUT_TITLE", "Confidence": 98.12, "Id": "lay-title", "Page": 1,
"Geometry": { "BoundingBox": { "Left": 0.12, "Top": 0.06, "Width": 0.26, "Height": 0.02 } },
"Relationships": [ { "Type": "CHILD", "Ids": ["ll1"] } ] },
{ "BlockType": "LAYOUT_TABLE", "Confidence": 97.4, "Id": "lay-table", "Page": 1,
"Relationships": [ { "Type": "CHILD", "Ids": ["tbl-1"] } ] },
{ "BlockType": "TABLE", "Confidence": 99.21, "Id": "tbl-1", "Page": 1,
"EntityTypes": ["STRUCTURED_TABLE"],
"Relationships": [ { "Type": "CHILD", "Ids": ["c11","c12","c21","c22","c31","c32"] } ] },
{ "BlockType": "CELL", "RowIndex": 1, "ColumnIndex": 1, "RowSpan": 1, "ColumnSpan": 1,
"EntityTypes": ["COLUMN_HEADER"], "Id": "c11", "Page": 1,
"Relationships": [ { "Type": "CHILD", "Ids": ["tw1"] } ] },
{ "BlockType": "LINE", "Confidence": 98.9, "Text": "Quarterly Report", "Id": "ll1", "Page": 1,
"Relationships": [ { "Type": "CHILD", "Ids": ["lw1","lw2"] } ] },
{ "BlockType": "WORD", "Confidence": 99.0, "Text": "Region", "TextType": "PRINTED", "Id": "tw1", "Page": 1 }
]
}
4.3 extract mode — textract/queries#
Each flat property of the request schema becomes one Textract query:
{"Text": <description> ?? <title> ?? <key with _ and - as spaces>, "Alias": <sanitised key>}
The alias keeps [A-Za-z0-9_.:-] (everything else becomes _, duplicates get a numeric suffix) and
maps back to the original property name, so "invoice number" in the schema is queried as
invoice_number and returned under "invoice number".
Responses are QUERY blocks carrying Query.Alias and a Relationships[{"Type":"ANSWER"}] list of
QUERY_RESULT ids. PuffinParse takes the highest-confidence non-empty QUERY_RESULT, coerces its Text
to the schema's type (number/integer strip currency and separators; boolean understands
yes/no/true/false/selected/[x]), and writes it to data[<key>]. Confidence lands in
fields["/<key>"].confidence; with citations=True the answer's page, box and text land in
fields["/<key>"].citations. A query with no answer yields data[<key>] = null and no entry in
fields — the key is always present, so the shape of data matches the schema.
Only flat string-like fields are supported. Properties typed
objectorarray(or carryingproperties/items) cannot be expressed as a Textract query; they are skipped, reported inmetadata.textract_unsupported_fields, and set tonull. Textract Queries answers one question with one span of text — there is no nesting and no repeated-row extraction. Flatten the schema (line_item_1_total, …) or use a provider with native structured extraction.
instructions on the request is ignored (Textract has no free-text guidance parameter); the fact is
recorded in metadata.textract_instructions_ignored.
Textract allows 15 queries per page synchronously and 30 asynchronously. A schema with more flat
properties than the applicable limit is rejected with an input error before any call is made.
4.4 extract mode — textract/forms (best effort)#
FORMS returns KEY_VALUE_SET blocks: a block with EntityTypes: ["KEY"] holds the label (its
CHILD WORDs) and points at its value block through Relationships[{"Type":"VALUE"}]; the value
block's CHILDren are WORDs or a SELECTION_ELEMENT.
PuffinParse matches each schema property to a detected key by case- and punctuation-insensitive
comparison (both sides reduced to lowercase alphanumerics), trying the property name and its title,
first for an exact match and then for "one name contains the other" (≥3 characters). Each detected
key is consumed at most once. Values are coerced by the schema's type exactly as in §4.3, and a
selected checkbox becomes true.
This is a heuristic, not schema-driven extraction. Textract decides what the keys are; PuffinParse only tries to line them up with your field names. Properties with no match are set to
nulland listed inmetadata.textract_unmatched_fields;metadata.textract_form_keys_foundreports how many key-value pairs Textract actually detected, which is the first thing to look at when fields come back empty. Nested (object/array) properties are never matched. At $0.050/pagetextract/formsis also the most expensive model here — prefertextract/querieswhen you know what you want.
5. Errors, status codes, rate limits, timeouts#
Textract reports failures as {"__type": "<Exception>", "message": "…"}, almost always with
HTTP 400 — including throttling and server-side failures, so the shared Error::from_http
mapping (4xx → bad_request) is wrong for Textract. textract::map_error keys off __type instead
(a fully-qualified com.amazonaws.textract#ThrottlingException is accepted too) and prefixes the
exception name onto the message:
__type |
HTTP | PuffinParse ErrorKind |
Retried? |
|---|---|---|---|
AccessDeniedException |
400 | authentication |
no |
UnrecognizedClientException, InvalidClientTokenId, ExpiredTokenException |
400/403 | authentication |
no |
IncompleteSignature, InvalidSignatureException, MissingAuthenticationTokenException |
400/403 | authentication |
no |
ThrottlingException |
400/500 | rate_limit |
yes |
ProvisionedThroughputExceededException |
400 | rate_limit |
yes |
LimitExceededException (too many concurrent async jobs) |
400 | rate_limit |
yes |
InternalServerError, ServiceUnavailable, InternalFailure |
500 | provider |
yes |
BadDocumentException, UnsupportedDocumentException, DocumentTooLargeException |
400 | bad_request |
no |
InvalidParameterException, InvalidS3ObjectException, InvalidKMSKeyException, InvalidJobIdException |
400 | bad_request |
no |
IdempotentParameterMismatchException, HumanLoopQuotaExceededException |
400 | bad_request |
no |
| anything unrecognised | — | falls back to Error::from_http |
per status |
An async job that ends FAILED becomes a provider error carrying StatusMessage and the job id.
PARTIAL_SUCCESS is accepted, logged, and its Warnings surfaced in metadata.
Multi-page PDFs. Synchronous Textract accepts PDF and TIFF at one page only. PuffinParse raises a
bad_request before the call:
textract: synchronous operations accept single-page PDF/TIFF only (this document has 2 pages). Multi-page documents require Textract's asynchronous API, which reads the file from S3. Upload the file yourself and pass
provider_options={"s3_object": {"bucket": "my-bucket", "name": "path/doc.pdf"}}, or split the PDF into single pages first.
The page count is best effort — it reads uncompressed page objects and the /Count in the page tree,
which covers most producers but not PDFs that hide their structure in object streams. When the count
cannot be determined, the call goes out and Textract's own 4xx is rewritten into the same message, so
the advice is identical either way.
Retries. max_retries (default 2), exponential backoff with full jitter, only on the
rate_limit, network and 5xx provider kinds above. Textract returns no Retry-After and no
rate-limit headers, so backoff is blind. Default us-east-1 quotas worth knowing (all adjustable in
Service Quotas, and lower in most other regions): synchronous AnalyzeDocument 10 TPS,
DetectDocumentText 25 TPS; StartDocumentAnalysis 10 TPS, StartDocumentTextDetection 15 TPS;
GetDocumentAnalysis 10 TPS, GetDocumentTextDetection 25 TPS; at most 600 asynchronous jobs
existing simultaneously per account (exceeding that is LimitExceededException, which PuffinParse
retries).
Timeouts. timeout_secs (default 300) is the whole-call deadline — download, signing, the call,
job polling and NextToken pagination — and also caps each individual HTTP request, shrinking as the
budget is spent.
6. Gotchas#
- The "API key" is a key pair.
api_keyon a PuffinParse request can only stand in forAWS_ACCESS_KEY_ID;AWS_SECRET_ACCESS_KEYmust be in the environment, otherwise the request is refused with anauthenticationerror that says so. SetAWS_SESSION_TOKENas well for STS / assumed-role credentials — it is signed asx-amz-security-tokenand included inSignedHeaders. PuffinParse reads only those environment variables: it does not parse~/.aws/credentials, does not honourAWS_PROFILE, and does not call IMDS or the ECS credential endpoint. - The region is part of the signature, not just the URL. Signing with the wrong region gives a
400 that talks about the credential scope, not about the host. If you override
base_urlto a VPC endpoint or a mock, setprovider_options={"region": …}to match. - No fixtures were captured live. The credentials available in this repository's build
environment are proxy placeholders; a real
DetectDocumentTextcall againsttextract.us-east-1.amazonaws.comreturned{"__type":"UnrecognizedClientException","message":"The security token included in the request is invalid."}. That round trip does confirm the transport and the signature format (a malformed canonical request yieldsIncompleteSignature/InvalidSignatureException, notUnrecognizedClientException) and the error mapping (ErrorKind::Authentication), but everytextract_*.jsonfixture is assembled from the shapes in the AWS API reference, with confidences, ids and boxes filled in to be realistic. Re-capture them from a real account before trusting the numbers, and run the#[ignore]d live tests at the bottom oftextract.rs. - Confidence is 0–100, not 0–1. Every
Confidenceis a percentage; PuffinParse divides by 100. TheQUERY_RESULTsample in the AWS docs shows"Confidence": 1.0, which is 1 %, not certainty. - Boxes are already normalised. Unlike most providers,
BoundingBoxis 0–1 relative to the page, so no page dimensions are needed — which is just as well, because Textract never reports them andPage.width/heightare alwaysNone. Pageis always 1 for JPEG/PNG, even for a scanned image that visually contains several pages. Only PDF and TIFF producePage > 1, and only through the async API.- Synchronous PDFs are single-page, full stop (10 MB in memory). Async accepts 500 MB / 3 000
pages but only from S3 — there is no bytes variant of
StartDocument*. This is the single biggest limitation of this provider and the reasonprovider_options.s3_objectexists. - Password-protected PDFs and XFA PDFs are rejected; images must be ≤10 000 px per side.
AnalyzeDocumentalways returns everyLINEandWORD, whateverFeatureTypessays. That is whytextract/layoutcan serveocrmode from the same response, and why aFORMS-only call still gives you the full text inraw.LAYOUT_LISTchildren areLAYOUT_TEXT, notLINE— one level of indirection that also means thoseLAYOUT_TEXTblocks appear twice inBlocks(once nested, once at the top level). Naive iteration duplicates every list item.MERGED_CELLduplicates content. A merged cell'sCHILDids are the individualCELLs, which are also children of theTABLE. Rendering both repeats the text; PuffinParse renders only the plain cells, so a row/column span shows its text in the first cell and blanks beside it.- Queries are English-only and capped at 15 per page synchronously / 30 asynchronously. Answers
are capped at 128 characters. A query aimed at a page that does not exist comes back as an
INVALID_REQUEST_PARAMETERSentry inWarnings, not as an error. - Throttling arrives as HTTP 400. Treating Textract's 400s as non-retryable (the usual rule)
means giving up on
ThrottlingExceptionandProvisionedThroughputExceededException; the mapping in §5 exists entirely for this. Warningsis silent data loss. APARTIAL_SUCCESSjob returns blocks for the pages that worked and lists the failures inWarnings— checkmetadata.textract_warningsbefore trusting the page count.- No
provider_job_idfor sync calls. Textract returns the request id only in thex-amzn-RequestIdresponse header, which PuffinParse does not surface today. - Pagination is per 1 000 blocks, not per page. A dense 50-page document can need dozens of
Get*round trips; they all come out oftimeout_secs.
7. Useful provider_options passthrough#
# 1. Multi-page PDF: upload to S3 yourself, then let PuffinParse drive Start*/Get* + NextToken.
puffinparse.parse("ignored-when-s3.pdf", model="textract/layout",
provider_options={"s3_object": {"bucket": "my-bucket", "name": "reports/q3.pdf"}},
timeout=1200)
# 2. Cheapest layout: drop TABLES to bill at the LAYOUT rate ($4 vs $15 per 1k pages).
# Tables then come back as plain lines instead of markdown grids.
puffinparse.parse("memo.png", model="textract/layout",
provider_options={"FeatureTypes": ["LAYOUT"]})
# 3. A different region (also changes the signing scope, not just the host).
puffinparse.ocr("scan.png", model="textract/detect-text", provider_options={"region": "eu-west-1"})
# 4. Signatures alongside layout, and the raw Block array for anything PuffinParse does not map.
puffinparse.parse("contract.png", model="textract/layout", include_raw=True,
provider_options={"FeatureTypes": ["LAYOUT", "TABLES", "SIGNATURES"]})
# 5. A trained Custom Queries adapter (adapters are Queries-only).
puffinparse.extract("claim.png", model="textract/queries", schema=schema,
provider_options={"AdaptersConfig": {"Adapters": [
{"AdapterId": "abc123", "Version": "1"}]}})
# 6. Async with your own output bucket, KMS key and an idempotency token.
puffinparse.parse("big.pdf", model="textract/layout", timeout=1800,
provider_options={"s3_object": {"bucket": "in", "name": "big.pdf"},
"OutputConfig": {"S3Bucket": "out", "S3Prefix": "textract/"},
"KMSKeyId": "alias/textract",
"ClientRequestToken": "big-pdf-2026-09-11"})
8. Links#
- What is Amazon Textract: https://docs.aws.amazon.com/textract/latest/dg/what-is.html
AnalyzeDocument: https://docs.aws.amazon.com/textract/latest/dg/API_AnalyzeDocument.html ·DetectDocumentText: https://docs.aws.amazon.com/textract/latest/dg/API_DetectDocumentText.htmlStartDocumentAnalysis: https://docs.aws.amazon.com/textract/latest/dg/API_StartDocumentAnalysis.html ·GetDocumentAnalysis: https://docs.aws.amazon.com/textract/latest/dg/API_GetDocumentAnalysis.htmlBlockreference: https://docs.aws.amazon.com/textract/latest/dg/API_Block.html- Layout response objects: https://docs.aws.amazon.com/textract/latest/dg/layoutresponse.html
- Tables: https://docs.aws.amazon.com/textract/latest/dg/how-it-works-tables.html · Form data: https://docs.aws.amazon.com/textract/latest/dg/how-it-works-kvp.html · Queries: https://docs.aws.amazon.com/textract/latest/dg/queryresponse.html
- Quotas: https://docs.aws.amazon.com/textract/latest/dg/limits-document.html · https://docs.aws.amazon.com/general/latest/gr/textract.html
- Pricing: https://aws.amazon.com/textract/pricing/ · machine-readable price list: https://pricing.us-east-1.amazonaws.com/offers/v1.0/aws/AmazonTextract/current/index.json
- Signature Version 4: https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_sigv4-signing-examples.html