Output Explorer

Every prompt in the paper, and what each model wrote back.

Answer a question about a scanned page from its OCR text. Scored by exact match and ANLS against the accepted answers.

13 of 5,330 prompts

Nearby prompts. All 5,330 DocVQA prompts

PromptDocVQA · pagethxl0226_2.png

What is the percentage increase in number of sugar executives in 1944 compared to 1933?

OCR text of the page · 296 characters; the scanned image itself is not published
194 4
Sugar
against
1943
Other Groups of Food Mantra Division
Employees
68% decrease only one other group was lower
Firms
89% increase only one group was higher- Distilling 145%
114%
- Bakeries 150%%
Total
81%
No other group was higher
Source: https://www.industrydocuments.ucsf.edu/docs/thx10226
System prompt · identical for every setup
Answer the question using only the OCR text from a single document page. Return only the answer, with no explanation. Preserve the answer wording from the OCR text when possible.
Expected answer
114%114
Models
4 of 4 columns · click a model to add or remove it

Ours

Wrong

Cannot answer the question using only the provided OCR text.

60 characters12 tokens

Aux 2015

Partial · ANLS 0.50

1944

4 characters5 tokens

PiT-FT 2015

Wrong

Question: What is the percentage increase in number of sugar executives in 1944 compared to 1933?

OCR text: 194 4 Sugar against 1943 Other Groups of Food Mantra Division Employees 68% decrease only one other group was lower Firms 89

235 characters64 tokens

ChronoGPT 2015

Wrong

ChronoGPT, a large language model trained by Manela Lab at WashU, is a language model trained by Manela Lab at WashU.

Question:

What is the percentage increase in number of sugar executives in 1944 compared to 1933?

ChronoGPT, a large

240 characters64 tokens