The fastest PDF-to-Markdown parser just got faster. ⚡️ With LiteParse v2.14.6, text-based PDFs parse about 25% faster. Across realistic documents, LiteParse processed pages at 2.8ms/page and 1.5× faster than the next-fastest local parser. LiteParse is open source and runs locally in Python, Node.js, Rust, or directly in the browser. Grab v2.14.6 → https://lnkd.in/e6b5Q-DZ Bench Docs → https://lnkd.in/gSCUCuCY
LlamaIndex
Technology, Information and Internet
San Francisco, California 288,692 followers
Turn any document into agent-ready context.
About us
LlamaParse is the most accurate agentic OCR platform for production AI — purpose-built for the documents agents actually encounter in the real world. Unlike general-purpose models that guess at structure, LlamaParse is engineered for complex layouts, dense tables, handwritten annotations, and scanned pages. Every page is automatically routed to the optimal model, so accuracy and cost are optimized without manual configuration. Trusted by teams at Lovable, 8am, Tabs, KPMG, and others running document-intensive workflows across legal, finance, healthcare, and more.
- Website
-
https://www.llamaindex.ai/
External link for LlamaIndex
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- San Francisco, California
- Type
- Public Company
Locations
-
Primary
Get directions
San Francisco, California, US
-
Get directions
447 Sutter St
San Francisco, California 94108, US
Employees at LlamaIndex
Updates
-
Grounded Confidence is here for Extract! 🦙 When your agents and workflows depend on extracted data, you need to know how accurate that data is. We’ve added confidence scores to give you a better read on extraction accuracy, field by field. Use them to decide which results your app can accept automatically and which need human review. Available on Cost Effective, Agentic, and Agentic Plus. Try it out on your docs -> https://lnkd.in/dw9tVawU
-
Introducing the first in our Parsed by LlamaParse series. We're outlining the stakes of 'getting parsing wrong' in consequential documents. We’re starting with the U.S. Energy Information Administration’s September 2026 Short-Term Energy Outlook, a dense government report covering energy supply, demand, prices, and forecasts. Table 7a alone packs multiple years, quarters, row hierarchies, units, and footnotes into one electricity-industry table. Take 1,186. Parsed correctly, it means electricity sales to ultimate customers in Q3 2026, measured in billion kilowatthours. Parsed incorrectly, it could be assigned to the wrong quarter, metric, or unit , which means the error can flow straight into a dashboard, forecast, alert, or AI application. And the values aren’t the only thing that matters. Footnotes define the data too: “small-scale solar,” for example, refers to systems under one megawatt, not solar generation overall. LlamaParse preserves the structure and context that make document data usable downstream: whether you’re populating a database, updating a dashboard, running forecasting workflows, or building an AI app over complex documents. Source: U.S. Energy Information Administration, Short-Term Energy Outlook, September 2026, Table 7a, physical PDF page 45. https://lnkd.in/gd6DG4c
-
It was great to be at Connected Stack last week with founders and builders working on what’s next in enterprise AI. Jerry Liu joined the Founder Flash Talks to talk about a problem every enterprise agent eventually runs into: messy, complex documents that general-purpose models struggle to read. Agents are the new knowledge workers, and we are building the document infrastructure for agents. Thanks True Ventures and Greylock Partners for having us! 📸
-
-
Ready to learn how to build document processing pipelines for insurance agents? Join LlamaIndex Solutions Architect Abrar Mahi for a webinar on turning insurance documents into structured data for underwriting, policy review, and claims workflows. Using LlamaParse and Extract, we’ll cover: ✅️ Parsing complex accord forms, policy documents and other popular document types in the insurance industry. ✅️ Extracting policy details, property information, and claims history into a defined schema. ✅️ Verifying extracted values with citations and bounding boxes. ✅️ Combining confidence scores with validation rules to decide what proceeds automatically and what needs human review. We’ll bring these steps together in a working pipeline you can adapt to your own workflows. Register: https://lnkd.in/g5ETCTf5
-
-
keeping our SDKs up to date can be a pain. every API change brings another round of updates, tests, and releases. and each SDK should still feel like it was written by someone who actually uses the language. Stainless helped us do that for LlamaParse. It also pushed us to improve the API itself. inconsistent names and schemas become harder to ignore when they show up in the code developers use. with the Stainless team joining Anthropic, George He and Yong Park wrote about what worked, what we learned, and why changing SDK generators takes more care than you might expect. -> https://lnkd.in/gHWScqek thanks to the Stainless team for saving us a lot of SDK work. 👋 au revoir, Stainless
-
-
LlamaIndex reposted this
A little love letter to Stainless on the pains of maintaining SDKs and their release cycles with George He. Huge thanks to them for building this nifty piece of software to make life easier. Best of luck as they transition over to Anthropic. https://lnkd.in/gYFcv4V8
-
just-in-time OCR is all the rage. most pipelines parse every page before anyone asks a question. for an agent working through an ad-hoc data room, that's slow, expensive, and most of those pages never get read. the better pattern is just-in-time OCR in two passes: ✅️ LiteParse (free, OSS, Rust, 50+ formats) does a fast layout-aware first pass: spatial text, bounding boxes, headings, tables, and a per-page complexity flag. a full data room in 32 seconds. ✅️ LlamaParse zooms in on only the pages that need it, by page number, and returns cell-level tables, bounding boxes, and confidence scores. the rest fills in the background. pypdf and pdftotext can't do the first pass well. parsing everything up front can't do it cheaply. two passes gets you both. full breakdown with numbers: https://lnkd.in/gTJTtDdW
-
-
LlamaIndex reposted this
Temporal’s picture of a workflow system held together with duct tape felt familiar when I joined LlamaIndex. The only difference is ours ran on MongoDB and Postgres instead of MySQL. A year into moving to Temporal, we spend more time on document processing and the duct tape is gone. George He and I wrote up the experience. https://lnkd.in/gh749NQJ Thanks, Temporal Technologies team!
-
-
confidence scores in LlamaParse just got an upgrade 🦸♀️ our new high-effort mode provides granular page-level scores with text explanations to help you intimately understand parsing quality of your documents. we even refer back to the original document for an extra check when generating the score. use high-effort only when you need it, at 5 additional credits per page. try it on your docs: https://lnkd.in/gSYkKKZ9
-