Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion api.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,7 +100,7 @@ from landingai_ade.types.v2 import (
```

- <code><a href="./src/landingai_ade/types/v2/job.py">Job</a></code> -- unified job shape: `job_id`, `status` (<code><a href="./src/landingai_ade/types/v2/job.py">JobStatus</a></code>: `pending` / `processing` / `completed` / `failed` / `cancelled`), `created_at`, `completed_at`, `progress`, `result` (a `V2ParseResponse` for parse jobs, a `V2ExtractResult` for extract jobs, a `V2BuildSchemaResponse` for build-schema jobs, or `None` until completion), `error` (<code><a href="./src/landingai_ade/types/v2/job.py">JobError</a></code>), `metadata` (the result's metadata receipt as a `dict`, populated top-level only when `output_save_url` was set and the result was delivered to `output_url` instead of inline; `None` otherwise, since inline jobs carry it on `result.metadata`), `raw` (the full original envelope as a `dict`), and the `.is_terminal` property.
- <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseResponse</a></code> -- `markdown`, `structure`, `grounding`, `metadata` (<code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseMetadata</a></code>, which nests <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseBilling</a></code> and carries `output_markdown_chars`, `range_units`, and `openapi_spec`). `structure` is a typed <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseStructure</a></code> tree (`document` → <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParsePage</a></code> → <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseElement</a></code>); each node below the root carries its spatial data inline in a <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseNodeGrounding</a></code> (`page`, <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseRange</a></code>, <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseBox</a></code>, normalized page coordinates, and an optional `confidence` in `[0, 1]` that is present only on word-granularity `atomic_grounding` segments (`dpt-3-verity`), where it is the lowest per-character OCR confidence in the word, and that is `None` on node-level grounding and on line-granularity models (`dpt-3-pro`)), and leaf elements additionally carry an `atomic_grounding` list. With `options.inline_markdown`, each node also carries its `markdown` slice. The legacy top-level `grounding` tree (<code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseGrounding</a></code> → `V2ParseGroundingPage` → `V2ParseGroundingElement` → `V2ParseGroundingEntry`) is retained for older gateway responses. Element `type`/page `status` are permissive strings and unknown keys are retained.
- <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseResponse</a></code> -- `markdown`, `structure`, `grounding`, `metadata` (<code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseMetadata</a></code>, which nests <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseBilling</a></code> and carries `output_markdown_chars`, `range_units`, and `openapi_spec`). `structure` is a typed <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseStructure</a></code> tree (`document` → <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParsePage</a></code> → <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseElement</a></code>); each node below the root carries its spatial data inline in a <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseNodeGrounding</a></code> (`page`, <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseRange</a></code>, <code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseBox</a></code>, normalized page coordinates, and an optional `confidence` in `[0, 1]` that is present only on word-granularity `atomic_grounding` segments (`dpt-3-verity`), where it is the lowest per-character OCR confidence in the word, and that is `None` on node-level grounding and on line-granularity models (`dpt-3-pro`)), and leaf elements additionally carry an `atomic_grounding` list. A `V2ParsePage` node is a `page` of a parsed document or a `sheet` of a parsed spreadsheet (`type`); a `sheet` node names itself in `id` (`None` on a `page` node). For spreadsheet content, `V2ParseNodeGrounding.page` and `.box` are `None` -- a workbook has no page number and no visual position -- and the Excel-style `address` (`Sales!C5`, `Sales!C5:F20`) locates the content instead; `address` is `None` for page-based documents. `V2ParseElement.id` is an opaque string: do not parse it or assume a format, and prefer `grounding.address`, which is stable across re-parses of the same spreadsheet. With `options.inline_markdown`, each node also carries its `markdown` slice. The legacy top-level `grounding` tree (<code><a href="./src/landingai_ade/types/v2/parse_response.py">V2ParseGrounding</a></code> → `V2ParseGroundingPage` → `V2ParseGroundingElement` → `V2ParseGroundingEntry`) is retained for older gateway responses. Element `type`/page `status` are permissive strings and unknown keys are retained.
- <code><a href="./src/landingai_ade/types/v2/extract_response.py">V2ExtractResult</a></code> -- `extraction`, `extraction_metadata`, `markdown`, `output_ref`, `schema_violation_error` (set when `strict=False` and the schema had unextractable fields), `warnings`, and `metadata` (<code><a href="./src/landingai_ade/types/v2/extract_response.py">V2ExtractMetadata</a></code>, which carries `model_version`, `input_markdown_chars`, `output_extraction_chars`, `range_units`, `openapi_spec`, and nests <code><a href="./src/landingai_ade/types/v2/extract_response.py">V2ExtractBilling</a></code>).
- <code><a href="./src/landingai_ade/types/v2/build_schema_response.py">V2BuildSchemaResponse</a></code> -- `extraction_schema` (the generated JSON Schema serialized as a string) and `metadata` (<code><a href="./src/landingai_ade/types/v2/build_schema_response.py">V2BuildSchemaMetadata</a></code>: `job_id`, `duration_ms`, `openapi_spec`, `filename`/`org_id`/`version` (retained for compatibility), a `warnings` list of <code><a href="./src/landingai_ade/types/v2/build_schema_response.py">V2BuildSchemaWarning</a></code> (`code`, `msg`), and nested <code><a href="./src/landingai_ade/types/v2/build_schema_response.py">V2BuildSchemaBilling</a></code>).
- <code><a href="./src/landingai_ade/types/v2/ground_response.py">V2GroundResult</a></code> -- `grounding` (a tree mirroring the input `extraction_metadata`, each `{value, ranges}` leaf replaced by the list of `structure` blocks its ranges overlap) and `metadata` (<code><a href="./src/landingai_ade/types/v2/ground_response.py">V2GroundMetadata</a></code>: `job_id`, `duration_ms`, `openapi_spec`, and nested <code><a href="./src/landingai_ade/types/v2/ground_response.py">V2GroundBilling</a></code>).
Expand Down
46 changes: 44 additions & 2 deletions docs/v2-testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,15 +28,18 @@ LANDINGAI_ADE_STAGING_APIKEY=... rye run pytest tests/contract/test_v2_smoke.py
`V2ParseResponse` with:

- `markdown` -- the full document as one Markdown string.
- `structure` (`V2ParseStructure`) -- the `document → page → element` tree.
- `structure` (`V2ParseStructure`) -- the `document → page | sheet → element` tree.
**Every node below the root carries its spatial data inline** in a
`V2ParseNodeGrounding` object (`grounding`):
- `page` -- 1-indexed page number.
- `page` -- 1-indexed page number. `None` for spreadsheet content (see below).
- `range` (`V2ParseRange`) -- `{start, end}` code-point offsets into
`markdown` (`metadata.range_units` names the unit, always
`"unicode_codepoints"`).
- `box` (`V2ParseBox`) -- `{xmin, ymin, xmax, ymax}` as `[0, 1]` fractions of
the page width/height (a page node's box is the full page `{0, 0, 1, 1}`).
`None` for spreadsheet content.
Comment on lines 38 to +40
- `address` -- spreadsheet-only Excel-style reference; `None` for page-based
documents (see below).
- `confidence` -- an optional `[0, 1]` probability. It is present **only** on
word-granularity `atomic_grounding` segments (`dpt-3-verity`), where it is the
lowest per-character OCR confidence in the word -- so a word is only as
Expand All @@ -59,6 +62,45 @@ The legacy top-level `grounding` tree (`V2ParseGrounding` and friends) is retain
on the model for backward compatibility with older gateway responses; current
responses omit it in favor of the inline `grounding` above.

### Spreadsheets: `sheet` nodes and `grounding.address`

The tree under `structure` now covers workbooks as well as page-based documents.
Nothing was renamed — a spreadsheet reuses `V2ParsePage` and
`V2ParseNodeGrounding`, with a different set of fields populated:

- **`V2ParsePage.type` is `"page"` or `"sheet"`.** It was a `const: "page"` in the
previous snapshot and is an enum now. The SDK already typed it as a permissive
`str` (not a `Literal`), so `"sheet"` deserializes without a model change —
but code that *compared* against `"page"` to find the page nodes will silently
skip every sheet.
- **`V2ParsePage.id`** (new) is the sheet name, e.g. `"Sales"`. It is present only
on a `sheet` node and `None` on a `page` node.
- **`V2ParseNodeGrounding.address`** (new) is the Excel-style reference of the
content, sheet name included: `Sales!C5` for a cell, `Sales!C5:F20` for a table,
the anchor cell (`Sales!B2`) for content parsed out of an embedded image. It is
`None` for page-based documents, and unlike `V2ParseElement.id` it is **stable
across re-parses of the same file** — it is the right key to join a re-parse on.
- **`grounding.page` and `grounding.box` are now `Optional`** on the wire, not just
in the model: a workbook has no page number and no visual position, so both come
back `null`/omitted for spreadsheet content. (The model already declared them
`Optional[...] = None`, so this needed no code change; treat missing and `None`
alike.) The exception is content parsed out of an image embedded in a
spreadsheet, where `box` is the fraction *of that image*, not of a page.
- **`V2ParseElement.id` is documented as opaque.** The `<type>-<index>` format
(`text-0`, `table_cell-0`) is gone from the spec — do not parse it or assume a
shape. It is still unique within a response and still unstable across re-parses.

Testing it: the live suite parses a PDF, so only the page-based half is assertable
there, and only as absence — `test_parse_node_locators_match_the_source_kind` in
`tests/contract/test_v2_smoke.py` walks the tree and asserts `address` and the node
`id` are `None` while `page`/`box` are populated. That much is a spec guarantee for
a page-based document. The spreadsheet half needs a workbook fixture staging is not
guaranteed to accept, so it is pinned deterministically against a mocked body
instead: `test_parse_sync_spreadsheet_sheet_nodes_and_addresses` in
`tests/api_resources/v2/test_parse.py` (plus `test_parse_response_spreadsheet_sheet_nodes`
in `tests/test_v2_types.py`) asserts the `sheet` node, its `id`, the per-node
`address`, and the `None` `page`/`box`.

### Encrypted PDFs (`password`)

`options.password` is a **supported** parse option — earlier snapshots documented
Expand Down
35 changes: 27 additions & 8 deletions specs/_generated/v2_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -151,6 +151,15 @@ class Status(Enum):
failed = 'failed'


class Type1(Enum):
"""
The node type. `page` is a page of a parsed document. `sheet` is one sheet of a parsed spreadsheet.
"""

page = 'page'
sheet = 'sheet'


class ServiceTier(Enum):
"""
The service tier the request ran in: `standard` or `priority`. A sync request reports `priority` (same lane, same price).
Expand Down Expand Up @@ -2614,18 +2623,23 @@ class Grounding(BaseModel):
can be lifted out of the tree and still locates its content.
"""

box: Box = Field(
address: Optional[str] = Field(
None,
description='Spreadsheet only. Excel-style reference of the content this grounding covers, with the sheet name: `Sales!C5` for a cell, `Sales!C5:F20` for a table, the anchor cell (`Sales!B2`) for content parsed out of an embedded image. Stable across re-parses of the same file. Omitted for page-based documents.',
title='Address',
)
box: Optional[Box] = Field(
...,
description="Bounding box in normalized page coordinates (`0`–`1` fractions of page width/height, at most 5 decimal places). A page node's box is always the full page `{0, 0, 1, 1}`.",
description="Bounding box in normalized page coordinates (`0`–`1` fractions of page width/height, at most 5 decimal places). A page node's box is always the full page `{0, 0, 1, 1}`. `null` (omitted from the response) when the source is a workbook, which has no visual position. For content parsed out of an image embedded in a spreadsheet, this is the fraction of that image, not of a page.",
)
confidence: Optional[float] = Field(
None,
description='How sure the model is of the text in this segment, in `[0, 1]`. Present only on word-granularity `atomic_grounding` entries (`dpt-3-verity`), where it is the lowest per-character OCR confidence in the word — so a word is only as trustworthy as its weakest character. Omitted on node-level grounding and on models that ground at line granularity.',
title='Confidence',
)
page: int = Field(
page: Optional[int] = Field(
...,
description="1-indexed page number this grounding is on. On a page node, the page's own number.",
description="1-indexed page number this grounding is on. On a page node, the page's own number. `null` (omitted from the response) when the source is a workbook, which has no page.",
title='Page',
)
range: Range = Field(
Expand Down Expand Up @@ -2983,7 +2997,7 @@ class Element(BaseModel):
)
id: str = Field(
...,
description='Semantic element id, unique within the document. Format `<type>-<index>`, where `<index>` is a per-type 0-based counter assigned in reading order — `text-0` is the first text element in the document, `figure-0` the first figure, `table_cell-0` the first cell of the first table. Stable within a response but not across re-parses of the same document.',
description='Element id, unique within the document. An opaque string — do not parse it or assume a format. Stable within a response but not across re-parses of the same document; for spreadsheet content, use `grounding.address` instead, which IS stable across re-parses.',
title='Id',
)
markdown: Optional[str] = Field(
Expand Down Expand Up @@ -3016,7 +3030,12 @@ class Page(BaseModel):
)
grounding: Grounding = Field(
...,
description="The page's spatial data: `page` is the 1-indexed page number in the source document (not contiguous when `options.pages` filters out some pages); `range` covers this page's content in the top-level `markdown` string (zero-length `start == end` for failed pages); `box` is always the full page `{0, 0, 1, 1}`.",
description="The node's spatial data. On a `page` node: `page` is the 1-indexed page number in the source document (not contiguous when `options.pages` filters out some pages); `range` covers this page's content in the top-level `markdown` string (zero-length `start == end` for failed pages); `box` is always the full page `{0, 0, 1, 1}`. On a `sheet` node: `page` and `box` are both `null` (a spreadsheet sheet has no page number or visual position); `range` covers the sheet's content in the top-level `markdown` string.",
)
id: Optional[str] = Field(
None,
description='The sheet name, present only on a `sheet` node. `null` (omitted from the response) on a `page` node.',
title='Id',
)
markdown: Optional[str] = Field(
None,
Expand All @@ -3033,9 +3052,9 @@ class Page(BaseModel):
description='Whether this page was parsed successfully (`ok`) or failed (`failed`).',
title='Status',
)
type: Literal['page'] = Field(
type: Optional[Type1] = Field(
'page',
description='The node type. Identifies this node as a page in the structure tree.',
description='The node type. `page` is a page of a parsed document. `sheet` is one sheet of a parsed spreadsheet.',
title='Type',
)

Expand Down
Loading
Loading