Frontline Lab summary and source
LlamaIndex released ExtractBench, covering 370 enterprise documents and 14 systems, with a focus on testing long-list completeness; it says Frontier VLMs achieve F1 scores of 8.9–35.8% on the longest documents.
This brief preserves the original source so the summary and editorial context can be checked independently.
Source attributionX · @llama_index
Open the original source