A side-by-side of two document parsing & extraction for building AI agents — live GitHub data, languages, and what each is best at.
Short answer: Docling leads Unstructured vs Docling by community traction (★ 27k vs ★ 12k). Pick Unstructured for mixed file-type ingestion; pick Docling for on-prem/air-gapped parsing.
| Unstructured | Docling | |
|---|---|---|
| GitHub stars | ★ 12k | ★ 27k |
| Language | Python | Python |
| Category | Document parsing & extraction | Document parsing & extraction |
| Best for | mixed file-type ingestion | on-prem/air-gapped parsing |
| Repository | Unstructured-IO/unstructured | docling-project/docling |
Unstructured and Docling are both credible choices. By community traction, Docling leads (★ 27k). Pick Unstructured for mixed file-type ingestion; pick Docling for on-prem/air-gapped parsing.
Both are credible document parsing & extraction. By community traction Docling leads (★ 27k). Pick Unstructured for mixed file-type ingestion; pick Docling for on-prem/air-gapped parsing.
Unstructured is Open-source library that turns 25+ file types into semantically labeled elements (title, table, list) with positions — the preprocessing layer for RAG ingestion.. Docling is IBM's MIT-licensed document parser — runs DocLayNet layout and TableFormer table models locally on commodity hardware, no cloud egress..
Docling has more — ★ 27k vs ★ 12k (live counts).
Often yes — many teams combine document parsing & extraction. Check each tool's docs for interop; they solve overlapping but not identical problems.
Unstructured is primarily Python; Docling is primarily Python.
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.