Published collection
ExtractBench finds VLMs miss rows when extracting long lists
Enterprise document extraction evaluations can expose missed-line issues more directly, rather than only checking whether returned values are correct.
- Published entries
- 1
- Sources
- 1
- Date range
- 2026-08-12
- Primary labels
- Research · 论文 · Evaluation Benchmarks · ExtractBench · Meta
Published evidence
Every entry keeps its summary and a path back to the source context.
ExtractBench finds VLMs miss rows when extracting long lists
LlamaIndex released ExtractBench, covering 370 enterprise documents and 14 systems, with a focus on testing long-list completeness; it says Frontier VLMs achieve F1 scores of 8.9–35.8% on the longest documents.
Original source