Skip to content

Evidence brief

ExtractBench finds VLMs miss rows when extracting long lists

LlamaIndex released ExtractBench, covering 370 enterprise documents and 14 systems, with a focus on testing long-list completeness; it says Frontier VLMs achieve F1 scores of 8.9–35.8% on the longest documents.

Published
Updated
Editorial
Frontline Lab
Source
X
Source author
@llama_index
Related topics
1
Collected
2026-08-13

Frontline Lab summary and source

Editorial summary

LlamaIndex released ExtractBench, covering 370 enterprise documents and 14 systems, with a focus on testing long-list completeness; it says Frontier VLMs achieve F1 scores of 8.9–35.8% on the longest documents.

This brief preserves the original source so the summary and editorial context can be checked independently.

Source attributionX · @llama_index

Open the original source

Related published evidence

Relationships are derived from shared topics, entities, categories, tags, and community context; every result remains independently source-linked.