Output Explorer

Every prompt in the paper, and what each model wrote back.

    Look-ahead bias

    Each model forecasts from a fixed date. A leak is any mention of what actually happened after it.

    Extraction

    Named entities from financial news (FIRE) and questions over scanned documents (DocVQA).

    Summarization

    Ten investor-relevant facts from an earnings call. Transcripts are redacted; model outputs are shown in full.

    Long-context

    RULER needle-in-a-haystack from 4K to 128K tokens. Prompts are shortened for display; the needle and the question are kept.

    How to read this

    • Add or remove models with the pills; the URL keeps your selection.
    • Responses are the first attempt each model made, even where a benchmark sampled several.
    • “Ours” is feature-steered for look-ahead bias, summarization and long-context; unsteered for extraction.