2025
Instructions for Form 1040-AI
TaxCalcBench Leaderboard
What's New
Tax Year 2025 (v2 edition). The 2025 edition is harder than the 2024 edition in three ways: inputs are realistic PDF documents (W-2s, 1099s, and so on) instead of structured data, cases include state returns as well as federal returns, and the cases cover much more complex tax and financial situations.
Tax Year 2024 (v1 edition). Federal-only returns for relatively simple situations, with inputs provided as structured JSON. Each case was run four times and the scores were averaged (pass@1). Use the tax year boxes at the top of the form to switch editions.
General Instructions
What is TaxCalcBench?
TaxCalcBench measures whether frontier AI models can do the calculation step of tax preparation: given everything about a taxpayer, produce the completed return. Each test case pairs a taxpayer's inputs with the expected, correctly computed return, and the model's answer is compared to it line by line.
Tax calculation has traditionally been done by hand-built, deterministic tax engines that encode tens of thousands of pages of rules. This benchmark asks whether a model can do the same job on its own.
Who must file
Model providers who want their model on this form can email team@columntax.com. Anyone can also run the open-source harness from the repository.
Line Instructions
Column (a)—Correct returns (strict). The share of test cases where every evaluated line exactly matches the expected return. This is the number that matters: a tax return has to be completely correct to be filed.
Column (b)—Correct returns (lenient). The share of cases where every evaluated line is within ±$5 of the expected value. Many small misses come from computing tax with bracket math instead of the official tax tables.
Column (c)—Correct (by line). The average percentage of evaluated lines that exactly match. A single early mistake can cascade through the rest of a return, so this is usually much higher than column (a).
Column (d)—Correct (by line, lenient). The same as column (c), counting lines within ±$5 as correct.
Column (e)—Cost per return. Average API cost to produce one return. A blank means cost wasn't available; blanks are never treated as $0. 2025 edition only.
Column (f)—Time per return. Average generation time for one return. 2025 edition only.
Thinking level. Each model is tested across its supported reasoning levels, and each column reports the best setting for that column. Two numbers on the same line can come from different settings.
Web search. Lines marked Web search had a web search tool available, so the model could look up current forms and instructions.
Partial coverage. Lines marked with a case count, such as 40/50, are scored only on the cases that completed.
Specific Instructions
Model-specific notes are attached to each model's Schedule M. Click a line on page 1 to open it.
Paperwork Reduction Act Notice
We ask for the information on this form to find out whether a model can do your taxes. Models are not required to provide the information requested unless they want to be on the leaderboard. The average time burden for completing this form varies by model; see column (f). If you have suggestions for making this form simpler, we would be happy to hear from you at team@columntax.com.
Results synced from the TaxCalcBench repository on 10/01/2026.