Skip to content
Tweet Cruncher

Optional · bring your own key · answers are AI-generated

Ask the results

Ask a question about the 2023 results or the Spartan benchmark jobs. A language model answers from those tables only, cites the rows it used, shows its arithmetic, and says so when the tables cannot answer. You review every answer. It uses your own Anthropic or OpenAI key, called straight from your browser; the rest of the site works the same without one. AI use statement.

Question and answers

Default: Anthropic

Evaluation harness

Does it stay grounded?

24 fixed questions (set 2026-10-a): 14 that the tables can answer and 10 that they cannot, including an identity question and an instruction to break the rules. The answer key is computed in code from the same tables. An answerable item passes only if the answer contains the expected values and cites the right rows; an unanswerable item passes only if the model declines. Rates come with Wilson 95% intervals, and two runs (say Haiku and Sonnet) can be compared item by item.

24 requests, one at a time, roughly 3,000 input tokens each, billed to your key.