Output Explorer

Every prompt in the paper, and what each model wrote back.

Answer a question about a scanned page from its OCR text. Scored by exact match and ANLS against the accepted answers.

13 of 5,330 prompts

Nearby prompts. All 5,330 DocVQA prompts

PromptDocVQA · pagethxl0226_2.png

What is the percentage decrease in number of sugar employees in 1944 compared to 1933?

OCR text of the page · 296 characters; the scanned image itself is not published
194 4
Sugar
against
1943
Other Groups of Food Mantra Division
Employees
68% decrease only one other group was lower
Firms
89% increase only one group was higher- Distilling 145%
114%
- Bakeries 150%%
Total
81%
No other group was higher
Source: https://www.industrydocuments.ucsf.edu/docs/thx10226
System prompt · identical for every setup
Answer the question using only the OCR text from a single document page. Return only the answer, with no explanation. Preserve the answer wording from the OCR text when possible.
Expected answer
68%6868% decrease
Models
4 of 4 columns · click a model to add or remove it

Ours

Wrong

Cannot answer the question using only the provided OCR text as data for 1933 is not present.

92 characters23 tokens

Aux 2015

Wrong

1944 Sugar against 1933 Other Groups of Food Mantra Division Employees 68% decrease only one other group was lower Firms 89% increase only one group was higher- Distilling 145% 114%

  • Bakeries
192 characters64 tokens

PiT-FT 2015

Wrong

Empty response.

0 characters

ChronoGPT 2015

Wrong

ChronoGPT, a large language model trained by Manela Lab at WashU, is a language model trained by Manela Lab at WashU.

Question:

What is the percentage decrease in number of sugar employees in 1944 compared to 1933?

ChronoGPT, a large

239 characters64 tokens