Turn messy PDFs, scans, and spreadsheets into structured, LLM-ready data for RAG and extraction pipelines. Ranked by community traction, with live GitHub stars and what each is best at.
Short answer: the best document parsing & extraction for AI agents by community traction is Docling (★ 27k), followed by Unstructured and LlamaParse.
IBM's MIT-licensed document parser — runs DocLayNet layout and TableFormer table models locally on commodity hardware, no cloud egress. Best for on-prem/air-gapped parsing.
Open-source library that turns 25+ file types into semantically labeled elements (title, table, list) with positions — the preprocessing layer for RAG ingestion. Best for mixed file-type ingestion.
LlamaIndex's managed document parser — per-page tiers from fast heuristics to VLM-agentic, with native LlamaIndex ingestion for RAG. Best for agent builders.
Agentic document parsing — layout-aware vision + VLMs + a multi-pass correction loop turn messy PDFs, scans, and spreadsheets into structured, RAG-ready data. Best for agent builders.
By community traction, Docling (★ 27k) leads the document parsing & extraction in our directory. IBM's MIT-licensed document parser — runs DocLayNet layout and TableFormer table models locally on commodity hardware, no cloud egress.
Docling is the most-starred open-source option; Unstructured and LlamaParse are strong runners-up.
Docling, at ★ 27k (live count).
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.