A side-by-side of two document parsing & extraction for building AI agents — live GitHub data, languages, and what each is best at.
Short answer: Docling leads Docling vs Unstructured by community traction (★ 27k vs ★ 12k). Pick Docling for on-prem/air-gapped parsing; pick Unstructured for mixed file-type ingestion.
| Docling | Unstructured | |
|---|---|---|
| GitHub stars | ★ 27k | ★ 12k |
| Language | Python | Python |
| Category | Document parsing & extraction | Document parsing & extraction |
| Best for | on-prem/air-gapped parsing | mixed file-type ingestion |
| Repository | docling-project/docling | Unstructured-IO/unstructured |
Docling and Unstructured are both credible choices. By community traction, Docling leads (★ 27k). Pick Docling for on-prem/air-gapped parsing; pick Unstructured for mixed file-type ingestion.
Both are credible document parsing & extraction. By community traction Docling leads (★ 27k). Pick Docling for on-prem/air-gapped parsing; pick Unstructured for mixed file-type ingestion.
Docling is IBM's MIT-licensed document parser — runs DocLayNet layout and TableFormer table models locally on commodity hardware, no cloud egress.. Unstructured is Open-source library that turns 25+ file types into semantically labeled elements (title, table, list) with positions — the preprocessing layer for RAG ingestion..
Docling has more — ★ 27k vs ★ 12k (live counts).
Often yes — many teams combine document parsing & extraction. Check each tool's docs for interop; they solve overlapping but not identical problems.
Docling is primarily Python; Unstructured is primarily Python.
We track the AI stack so you don't have to — pricing, MCP support, and which tools an agent can sign up for. Free.