LlamaIndex’s cover photo
LlamaIndex

LlamaIndex

Technology, Information and Internet

San Francisco, California 288,692 followers

Turn any document into agent-ready context.

About us

LlamaParse is the most accurate agentic OCR platform for production AI — purpose-built for the documents agents actually encounter in the real world. Unlike general-purpose models that guess at structure, LlamaParse is engineered for complex layouts, dense tables, handwritten annotations, and scanned pages. Every page is automatically routed to the optimal model, so accuracy and cost are optimized without manual configuration. Trusted by teams at Lovable, 8am, Tabs, KPMG, and others running document-intensive workflows across legal, finance, healthcare, and more.

Website
https://www.llamaindex.ai/
Industry
Technology, Information and Internet
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Public Company

Locations

Employees at LlamaIndex

Updates

  • The fastest PDF-to-Markdown parser just got faster. ⚡️ With LiteParse v2.14.6, text-based PDFs parse about 25% faster. Across realistic documents, LiteParse processed pages at 2.8ms/page and 1.5× faster than the next-fastest local parser. LiteParse is open source and runs locally in Python, Node.js, Rust, or directly in the browser. Grab v2.14.6 → https://lnkd.in/e6b5Q-DZ Bench Docs → https://lnkd.in/gSCUCuCY

    • No alternative text description for this image
  • Grounded Confidence is here for Extract! 🦙 When your agents and workflows depend on extracted data, you need to know how accurate that data is. We’ve added confidence scores to give you a better read on extraction accuracy, field by field. Use them to decide which results your app can accept automatically and which need human review. Available on Cost Effective, Agentic, and Agentic Plus. Try it out on your docs -> https://lnkd.in/dw9tVawU

  • Introducing the first in our Parsed by LlamaParse series. We're outlining the stakes of 'getting parsing wrong' in consequential documents. We’re starting with the U.S. Energy Information Administration’s September 2026 Short-Term Energy Outlook, a dense government report covering energy supply, demand, prices, and forecasts. Table 7a alone packs multiple years, quarters, row hierarchies, units, and footnotes into one electricity-industry table. Take 1,186. Parsed correctly, it means electricity sales to ultimate customers in Q3 2026, measured in billion kilowatthours. Parsed incorrectly, it could be assigned to the wrong quarter, metric, or unit , which means the error can flow straight into a dashboard, forecast, alert, or AI application. And the values aren’t the only thing that matters. Footnotes define the data too: “small-scale solar,” for example, refers to systems under one megawatt, not solar generation overall. LlamaParse preserves the structure and context that make document data usable downstream: whether you’re populating a database, updating a dashboard, running forecasting workflows, or building an AI app over complex documents. Source: U.S. Energy Information Administration, Short-Term Energy Outlook, September 2026, Table 7a, physical PDF page 45. https://lnkd.in/gd6DG4c

  • It was great to be at Connected Stack last week with founders and builders working on what’s next in enterprise AI. Jerry Liu joined the Founder Flash Talks to talk about a problem every enterprise agent eventually runs into: messy, complex documents that general-purpose models struggle to read. Agents are the new knowledge workers, and we are building the document infrastructure for agents. Thanks True Ventures and Greylock Partners for having us! 📸

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Ready to learn how to build document processing pipelines for insurance agents? Join LlamaIndex Solutions Architect Abrar Mahi for a webinar on turning insurance documents into structured data for underwriting, policy review, and claims workflows. Using LlamaParse and Extract, we’ll cover: ✅️ Parsing complex accord forms, policy documents and other popular document types in the insurance industry. ✅️ Extracting policy details, property information, and claims history into a defined schema. ✅️ Verifying extracted values with citations and bounding boxes. ✅️ Combining confidence scores with validation rules to decide what proceeds automatically and what needs human review. We’ll bring these steps together in a working pipeline you can adapt to your own workflows. Register: https://lnkd.in/g5ETCTf5

    • No alternative text description for this image
  • keeping our SDKs up to date can be a pain. every API change brings another round of updates, tests, and releases. and each SDK should still feel like it was written by someone who actually uses the language. Stainless helped us do that for LlamaParse. It also pushed us to improve the API itself. inconsistent names and schemas become harder to ignore when they show up in the code developers use. with the Stainless team joining Anthropic, George He and Yong Park wrote about what worked, what we learned, and why changing SDK generators takes more care than you might expect. -> https://lnkd.in/gHWScqek thanks to the Stainless team for saving us a lot of SDK work. 👋 au revoir, Stainless

    • No alternative text description for this image
  • just-in-time OCR is all the rage. most pipelines parse every page before anyone asks a question. for an agent working through an ad-hoc data room, that's slow, expensive, and most of those pages never get read. the better pattern is just-in-time OCR in two passes: ✅️ LiteParse (free, OSS, Rust, 50+ formats) does a fast layout-aware first pass: spatial text, bounding boxes, headings, tables, and a per-page complexity flag. a full data room in 32 seconds. ✅️ LlamaParse zooms in on only the pages that need it, by page number, and returns cell-level tables, bounding boxes, and confidence scores. the rest fills in the background. pypdf and pdftotext can't do the first pass well. parsing everything up front can't do it cheaply. two passes gets you both. full breakdown with numbers: https://lnkd.in/gTJTtDdW

    • No alternative text description for this image
  • confidence scores in LlamaParse just got an upgrade 🦸♀️ our new high-effort mode provides granular page-level scores with text explanations to help you intimately understand parsing quality of your documents. we even refer back to the original document for an extra check when generating the score. use high-effort only when you need it, at 5 additional credits per page. try it on your docs: https://lnkd.in/gSYkKKZ9

    • No alternative text description for this image

Similar pages

Browse jobs