Document Layout & Field Extraction
Annotate PDFs and scanned documents with layout regions, OCR-backed text spans, table structure, and page and document classifications, then map labeled regions to a structured field schema.
- Document
- hOCR
Your documents do not look like Docling's training data: a lab report or an internal invoice has little in common with the public PDFs a converter learned on. This interface gives you a place to see where that breaks.
You draw bounding boxes around document elements, layer reading-order paths across them, and correct merged regions and table structure, with the DocLang XML updating in a sidebar as you work. Predictions from the Docling ML backend load as a starting point, and the result exports as a .dclx archive carrying both the structure and the page images.