The strongest open-source alternatives to Docling for building AI agents — document parsing & extraction ranked by GitHub traction, each with a head-to-head.
Short answer: the closest alternative to Docling is Unstructured (★ 12k), the most-starred document parsing & extraction; LlamaParse and Reducto are also strong.
Docling (★ 27k) is IBM's MIT-licensed document parser — runs DocLayNet layout and TableFormer table models locally on commodity hardware, no cloud egress. If it is not the right fit, these 3 document parsing & extraction cover the same ground — Unstructured is the most-starred option below. Or browse the best document parsing & extraction and Docling's own page.
Open-source library that turns 25+ file types into semantically labeled elements (title, table, list) with positions — the preprocessing layer for RAG ingestion. Best for mixed file-type ingestion.
LlamaIndex's managed document parser — per-page tiers from fast heuristics to VLM-agentic, with native LlamaIndex ingestion for RAG. Best for agent builders.
Agentic document parsing — layout-aware vision + VLMs + a multi-pass correction loop turn messy PDFs, scans, and spreadsheets into structured, RAG-ready data. Best for agent builders.
Unstructured (★ 12k) is the most-starred document parsing & extraction alternative to Docling. Open-source library that turns 25+ file types into semantically labeled elements (title, table, list) with positions — the preprocessing layer for RAG ingestion.
The strongest document parsing & extraction alternatives to Docling are Unstructured, LlamaParse, Reducto — each with a head-to-head comparison.
Yes — the alternatives listed are open source and free to self-host; you bring your own model keys.
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.