Extract v2.5 flips LlamaIndex tier accuracy
Extract v2.5 launched Oct 1 with higher ExtractBench scores and Advanced Citations. LlamaIndex says Cost Effective now beats old Agentic.

- LlamaIndex's October 1 post and Unite.AI's write-up both date Extract v2.5 to October 1 and name Advanced Citations for Agentic and Agentic Plus.
- On ExtractBench, LlamaIndex prints Cost Effective rising from 87.1 to 93.9 overall value F1, Agentic from 89.8 to 95.8, and Agentic Plus from 95.1 to 96.4.
- LlamaIndex says Cost Effective now beats the previous Agentic tier, and Agentic now beats the previous Agentic Plus. That ladder flip is the company's own chart, kept attributed.
in this block
Extract v2.5 is LlamaIndex's October 1, 2026 bump to its schema-based document agents, and the company blog plus Unite.AI's coverage agree that accuracy rose across tiers without a per-page price hike. The interesting part is not another "AI eats PDFs" slogan. It is which tier leapfrogged which.
What actually happened
Extract v2.5 keeps the same product shape: you bring a schema, the agent returns structured fields, and higher tiers spend more effort. LlamaIndex says architectural work includes a new agent harness tuned for document failure modes across vision, reasoning, and verification. Unite.AI's coverage repeats the accuracy-and-grounding pitch and the no-price-increase line. Shared launch facts stay shared. Deep examples stay with the primary blog.
Long lists are the demo that sticks. LlamaIndex says Cost Effective previously returned 87 of 238 holdings on a 17-page fund filing and that Extract v2.5 returns all 238, with that document's extraction score moving from 52.8 to 98.7. Across long-list tasks, Cost Effective rises from 82.0 to 92.6 and Agentic from 86.1 to 95.5 on the company's bench. Those are ExtractBench numbers from LlamaIndex, not an external auditor.
Records that span pages and scanned forms get their own before-and-after stories on the blog: a grant purpose that used to stop at a page break, a permit field that used to swallow reviewer ink. The point of those examples is failure-mode coverage, not a claim that every enterprise schema is solved.
Grounding and citations
Advanced Citations are the other headline. LlamaIndex says grounding scores rise from 46.8 to 80.6 for Agentic and from 46.4 to 82.2 for Agentic Plus, with bounding boxes for evidence that may not match the extracted string word for word. Unite.AI echoes the grounding theme. Confidence scores remain a separate toggle for human-in-the-loop review. Native spreadsheet extraction is called out on the LlamaIndex post as agents working on workbook cells rather than a flattened dump. If your pipeline is Excel-heavy, that line matters more than another PDF screenshot.
How to pin the version
LlamaIndex's sample configuration uses an explicit version string for this topic. Pin something reproducible for the October 1 cut. Do not float on a moving latest alias if you need the same extracted fields next month.
Structural Reasoning is LlamaIndex's name for spending more effort on dense, messy documents and less on simple ones. That adaptive spend is how the company explains lower latency on easy files without giving up the hard ones. Unite.AI does not need to re-derive the harness for the launch to count as covered. Readers comparing other October model drops can open the Clef decision models note or the Gemini 4 Argon note. Readers building agent glue can open the OpenClaw Enterprise note.
Primary sources: LlamaIndex's this topic blog and Unite.AI's this topic coverage.
What still fails in the old ways
LlamaIndex is careful to show scanned-form mistakes the prior tier made: mixing a reviewer date into a permit number, or appending a note into a depth field. The new cut separates those layers in the examples. That does not mean your particular annotated PDF is solved. It means the vendor is finally demoing the failure mode that wrecks real back offices.
Schemas with thousands of fields get a passing mention too. The blog says agents were tested with tasks up to 3,200 fields and that you should call out record IDs in the prompt when matching to a database matters. That is practical. It is also a reminder that prompt craft still sits outside the model card.
What to do as a reader
If you already pay for Extract, rerun your hardest schemas on this version and compare citation boxes, not just F1 vanity. If you are evaluating from scratch, test Cost Effective first because LlamaIndex claims it now clears the old Agentic bar on ExtractBench. Nothing here is a procurement order. Not financial advice. DYOR, ser.
Not financial advice. DYOR, ser.