Key Takeaways
- The top is a close race: Claude Fable 5.1 leads at 77.64%, with Claude Opus 5 (75.06%) and GLM 5.3 (73.09%) within five points. No model clears 80% overall, so even the best agents still miss required elements on a meaningful share of corporate tax questions.
- Forms & Filings and Controversy & Precedence Analysis are the hardest categories, with no model above 74%: producing filing-ready work product and determining which rules and rulings apply to a given task remain harder than computing the right figure. Numbers & Calculations is the easiest category for nearly every model (top score 87.97%).
- Top performing agents shift by category. GLM 5.3 leads Rule & Source Lookup (81.29%) and Controversy & Precedence Analysis (71.08%) despite ranking third overall, and GPT-5.6 Sol is second on Forms & Filings despite ranking seventh.
Background
Tax Agent Bench evaluates AI agents on realistic, research-grade US tax questions, with a focus on corporate tax. Rather than static Q&A, tasks reflect the multi-step research process a tax practitioner actually follows: identifying the governing authority, synthesizing statutes, regulations, IRS guidance, and forms, applying the rules to a fact pattern, and producing an accurate, well-supported answer.
Each task is a single written question. The agent researches it with a set of tools over as many steps as it needs, then submits one final written answer with its citations. Only that final answer is graded, against the question’s expert rubric and a check that its cited sources are real; the intermediate trajectory (the sequence of tool calls) is recorded and shown on this page for analysis but is not scored.
The questions center on the work of an enterprise tax department: federal corporate income tax (roughly 60% of the test set), with the remainder spanning partnership tax, state and local tax (SALT), transfer pricing, and corporate employment tax. Fact patterns involve C corporations, consolidated groups, reorganizations, reportable-transaction disclosure, penalties and procedure, and multi-year or multi-jurisdiction comparisons.
The benchmark tests whether models can conduct the kind of research that tax associates perform daily, where a single matter frequently requires chaining several lookups and reasoning steps, and where an early error in a threshold, date, or citation propagates to the final answer.
All questions, gold-standard answers, and rubrics are authored and reviewed by tax professionals.
Scoring
Each question carries a rubric of required checks written by tax experts (3 to 89 checks per question, mean 23.6, mode 10). Every check carries a weight from 1 to 3 reflecting how much it matters to the answer, and roughly 40% of checks are marked must-pass: elements a correct answer cannot omit, such as identifying the controlling Code section or the required form.
Our primary metric, Accuracy, is a weighted partial-credit score. For each question, the rubric score is the share of weighted rubric points earned across its checks, set to 0 if any must-pass check fails. Citation quality is the fraction of sources cited in the answer that resolve to a real, authoritative document. The task score combines the two as
task score = rubric score × (0.7 + 0.3 × citation quality)
so an answer whose citations all resolve keeps its full rubric score, one whose citations all fail keeps 70% of it, and one with half its citations resolving keeps 85%. Citation quality can therefore reduce a task score by at most 30%, and it cannot rescue an answer that failed a must-pass check. Scores are averaged over all 193 questions of the private test set with every task weighted equally; a question the agent did not answer within the time limit scores 0. We also report All-Pass, a strict secondary metric: the proportion of questions on which the model satisfied every rubric check. All grading is performed by an LLM judge; see Methodology for the harness and judge.
Results
Unless otherwise noted, all results on this page are computed on the private 193-question test set. The Pareto chart at the top of the page shows how accuracy trades off against cost and latency across 18 models. Claude Fable 5.1 leads at 77.64% for ~$13.18 per task and ~29 minutes. Claude Opus 5 follows at 75.06% for ~$5.14 per task and ~17 minutes, the fastest of the top three and less than half the cost of the leader, while GLM 5.3 is third at 73.09% for ~$1.80 per task and ~51 minutes. Muse Spark 1.3 is fourth at 71.93% for ~$0.32 per task and ~8 minutes, cheaper and faster than every model above it, making it the strongest budget option ahead of Grok 4.6 (70.79%, ~$0.98, ~11 minutes); Gemini 3.8 Flash (66.77%, ~$0.86, ~3 minutes) sits eighth overall and fourth on Fact-Pattern Analysis (67.95%), while its predecessor Gemini 3.7 Flash answers in about 1.5 minutes at 57.67%.
Performance by Category
Questions fall into one of six categories. The first three are bounded research tasks with a single locatable answer: the right rule, the right figure, or the right form. The latter three require reasoning across procedure, time, and interacting facts, where the answer depends on weighing authority rather than retrieving it. The counts below are for the 193-question test set; the validation set follows the same distribution (31/30/30/30/30/42):
- Rule & Source Lookup (30 test questions): identifying the governing tax rule, threshold, or requirement and locating the most authoritative source among statutes, regulations, forms, and IRS guidance.
- Numbers & Calculations (30): determining applicable percentages, dollar thresholds, and due dates, and performing multi-step computations where an early error propagates to the final answer.
- Forms & Filings (30): identifying filing requirements, disclosure obligations, and information returns, and producing work product such as memos or completed forms.
- Controversy & Precedence Analysis (30): audits, penalties, appeals, statutes of limitation, elections, and procedural requirements, including determining which rules and rulings apply to a given task.
- Current & Temporal Analysis (30): recent law changes, new forms, and recent guidance, or comparing rules across tax years, jurisdictions, or alternative fact patterns.
- Fact-Pattern Analysis (43): applying one or more tax rules to a specific fact pattern, from a single rule applied to a clearly stated set of facts to scenarios where multiple concepts interact and require synthesis across sources.
Use the task selector on the results table to view each category. Forms & Filings and Controversy & Precedence Analysis are the hardest categories, with best scores of 73.18% (Claude Fable 5.1) and 71.08% (GLM 5.3) respectively. Numbers & Calculations is the easiest: the top five all exceed 83%, and even the weakest model reaches 54.9%.
No model leads everywhere. Claude Fable 5.1 tops Fact-Pattern Analysis (79.29%), Current & Temporal Analysis (81.21%), Numbers & Calculations (87.97%), and Forms & Filings (73.18%); GLM 5.3 leads Rule & Source Lookup (81.29%) and Controversy & Precedence Analysis (71.08%). The spread within a category is wide: on Rule & Source Lookup, GPT 5.5 scores 44.23% while GLM 5.3 scores 81.29%, so locating the right authority is far from solved even for frontier models.
All-Pass: Fully Correct Answers
Accuracy awards partial credit for each rubric check an answer passes, weighted by importance. All-Pass asks the stricter question a client would: was the answer fully right? A question counts only if every rubric check passes.
Under All-Pass, no model reaches 50%. Claude Fable 5.1 leads at 49.22%, followed by Claude Opus 5 (45.60%), GLM 5.3 (39.90%), and Muse Spark 1.3 (36.79%); every other model is below 34%. The gap between the two metrics is between 26 and 43 points for every model, so the typical answer earns most of its rubric points but still omits at least one required element. The ranking also shifts: Kimi K3 (33.68%) edges past Grok 4.6 (33.16%) on All-Pass despite trailing it by two points on Accuracy, and GPT 5.5 drops to 21.76%, below Gemini 3.7 Flash, because its answers more often miss one check outright. Switch the metric selector above to All-Pass to view it overall or per category.
Tool Use
The harness exposes five tools: web_search, lookup_authority, fetch_document, retrieve_information, and calculator (each is defined under Methodology). Every model uses lookup_authority, which returns statutory and regulatory text by citation, on every question, but the overall research style differs sharply: the three leaders make it their most-used tool, while search-heavy agents such as Grok 4.6, Kimi K3, and the OpenAI 5.6 checkpoints issue more web_search calls than authority lookups.
The three OpenAI 5.6 checkpoints make far more calls than anyone else, 113 tool calls per question for GPT-5.6 Sol, about 100 for GPT-5.6 Terra, and 91 for GPT-5.6 Luna, roughly double the 41 to 49 of the three leaders. The extra calls are mostly web searches (43, 37, and 33 per question versus 6 to 8 for the two Claude leaders) and do not translate into accuracy: they score 61% to 68%. Grok 4.6 and Kimi K3 are also search-heavy, with 21 and 25 web searches per question, while the top three models instead spend their calls on lookup_authority (10 to 16 per question) and on fetching and querying primary documents. Errors during tool calls are rare across the board, between 0.2 and 3.5 per question, and are highest for the OpenAI 5.6 checkpoints, so the extra volume also carries more failed calls. Across all 18 models, 3.8% of tool calls fail (6,192 of 164,123), and the failures follow a consistent pattern: about three-quarters are fetch_document calls that hit a page the harness could not read, mostly URLs that return 404 or sites that block automated access, with a smaller share of PDFs that yield no text or time out; 13% are lookup_authority requests for a citation that does not resolve, most often a regulation section number that does not exist in that form (for example a call for Treas. Reg. §1.338 without the subsection); and 9% are retrieve_information queries that name no saved document or one the agent never stored, nearly half of them from Inkling.
The radar chart shows the per-tool mix. The Claude models and GLM 5.3 have the most balanced profiles, pairing authority lookups with fetch_document and retrieve_information to read full regulations and rulings, which is what the Controversy and Fact-Pattern categories require. Gemini 3.7 Flash is the most economical, averaging 20 tool calls and only 1.7 document fetches per question, which explains its 1.5-minute latency but also its 57.7% accuracy. Muse Spark 1.3 shows that a light footprint need not cost accuracy: it averages 24 calls per question, about the same as Inkling and MiniMax-M3 and fewer than every other model except Gemini 3.7 Flash, yet ranks fourth overall. calculator is used between roughly two and eight times per question by every model, most heavily by the OpenAI 5.6 checkpoints.
Answer Length
Tax practitioners expect a research memo, not a one-line answer, and the rubrics reward supporting analysis and citations. The chart below plots each model’s mean final-answer length in words against its Accuracy; the reference line marks the mean length of the expert-written reference answers in the dataset, 601 words.
The three most accurate models also write the longest answers: Claude Opus 5 averages 2,757 words, GLM 5.3 2,434, and Claude Fable 5.1 2,366, versus 950 to 2,000 for the rest of the field. Length is not sufficient on its own, however: Inkling writes 1,809 words for a 43.5% score, about the same length as Grok 4.6 (1,769 words, 70.8%). The shortest answers come from GPT-6 Astra (956 words) and GPT 5.5 (1,101 words), whose All-Pass rates (20.7% and 21.8%) are the two lowest, consistent with answers that omit required elements rather than getting them wrong. Every model writes far more than the experts did: the shortest model average (956 words) is about one and a half times the 601-word reference mean, and the longest (2,757) is more than four times it. The reference answers satisfy every rubric check in a fraction of the space, so the extra length is not what the rubrics require, and no model is close to the expert length: the verbosity is a property of the agents, not of the task.
Trajectories on a Public Question
To show how the agents actually work through a problem, the chart below traces the tool-call sequence of all 18 models on public question P-004, a Current & Temporal Analysis task in State & Local Tax: for tax year 2025, how do New York’s and California’s pass-through entity taxes (PTET) differ in how the credit is allocated to an individual partner, and is that credit refunded if it exceeds the partner’s state income tax for the year? Each bar is one tool call, in order, colored by tool. For the 14 models run on the full dataset the trajectories come from the same runs that produce the scores above, where this question appears under its full-dataset ID; Gemini 3.8 Flash, GPT-6 Astra, GPT-5.6 Sol, and Muse Spark 1.3 were run on the public split separately from their test-set runs.
Only Grok 4.6 passes all 18 checks, and it does so with one of the longest trajectories (88 calls, 59 of them web searches). At the other extreme, Muse Spark 1.3 scores 79.4% in 19 calls, the shortest trajectory of any model here, clearing every must-pass check and dropping four lower-weight details. The overall leader Claude Fable 5.1 scores 91%, clearing every must-pass check and missing two lower-weight points (the New York statutory citation and one category of California owner that cannot claim the credit). Kimi K3, GPT-5.6 Luna, and GPT 5.5 each land at 76.5%, getting the refundability answer right for both states but dropping the same cluster of supporting details: the governing statutes, the California claim form, what the electing entity must report, and which owners California excludes. The remaining twelve, including Claude Opus 5 and GLM 5.3, score 0. Eleven of them fail the same must-pass check, how New York treats a partner who receives PTET credits from more than one electing entity, and Inkling fails a different one, which California tax the credit offsets. Effort does not decide it: GPT-5.6 Terra makes 95 calls and scores 0, while GPT 5.5 reaches 76.5% in 40 and GPT-6 Astra scores 0 in the same 40, and Muse Spark 1.3 beats both in half as many calls. Research style differs as much as length: Grok 4.6, Gemini 3.8 Flash, Gemini 3.7 Flash, and GPT-5.6 Terra lean on web_search for more than half their calls, while Claude Fable 5.1, Claude Opus 5, Kimi K3, and DeepSeek V4 Pro 0813 work almost entirely through fetch_document and retrieve_information. On a single question with a hard threshold, the overall leaderboard is not predictive; the aggregate ranking emerges only across the full 193 questions.
For calendar tax year 2025, how do the tax credit allocation and liquidity mechanics differ between New York State's pass-through entity tax (PTET) and California's pass-through entity tax (PTET) regarding whether the resulting credit passed through to an individual partner is fully refundable if it exceeds that partner's personal state income tax liability for the year?
91%
Missed checks: the governing New York statute; which California owners are excluded from the credit (disregarded entities).
The question above is a Current & Temporal Analysis task in State & Local Tax: for 2025, the agent must compare how New York’s and California’s pass-through entity tax credits reach an individual partner and whether each state refunds a credit that exceeds the partner’s personal income tax, citing the governing authority in each state. The feedback under each answer summarizes the topics of the rubric checks it missed; most turn on how New York handles credits from more than one electing entity and on supporting details such as the statutory citations, the California claim form, and which owners California excludes from the credit. Missed checks are paraphrased from the rubric. A check fails when the answer does not state the required point, even if it touches the topic.
Timeouts
Each task has a three-hour limit. If the agent has not submitted a final answer when the limit is reached, the run is killed and the question scores 0. Five of the 18 models hit the limit on at least one test question:
| Model | Timeouts (of 193) | Share of test set |
|---|---|---|
| GPT-5.6 Terra | 18 | 9.3% |
| GPT-5.6 Luna | 8 | 4.1% |
| GPT-5.6 Sol | 5 | 2.6% |
| Kimi K3 | 4 | 2.1% |
| Qwen 3.8 Max | 1 | 0.5% |
The other thirteen models, including all three leaders, answered every test question within the limit. The timeouts are concentrated in the three OpenAI 5.6 checkpoints, which are also the most tool-heavy agents (see Tool Use): their logs end mid-research, typically inside a retrieve_information call dozens of turns in, rather than at a stalled or crashed step. The timed-out questions cluster in Fact-Pattern Analysis (16 of the 36 timeouts across the five models), Current & Temporal Analysis (10), and Forms & Filings (8), the categories that require reading several full documents and reconciling them. For GPT-5.6 Terra the 18 zeros hypothetically cost about seven points: its Accuracy on the 175 questions it did answer is 71.9%, against 65.2% overall, which would move it from tenth to fifth had it finished every task at that rate. GPT-5.6 Luna gains 2.6 points on the same basis (63.4% vs. 60.8%), GPT-5.6 Sol 1.8 (69.8% vs. 68.0%), and Kimi K3 1.4 (70.1% vs. 68.7%).
Methodology
Agents are evaluated on a shared harness with access to five tools. Each task has a three-hour time limit with no fixed step count; the agent may iterate with tools until it submits a final answer or the limit is reached.
web_search: searches the web for relevant sources and returns results with titles, URLs, and excerpts from each pagelookup_authority: retrieves the text of statutes, regulations, and public laws directly by citation, returning the cited provision along with its canonical source URLfetch_document: downloads a web page or PDF, converts it to plain text, and saves it to the agent’s data storage, allowing the agent to work with documents larger than its context windowretrieve_information: queries documents saved in data storage, extracting or summarizing the relevant content without loading the full document into contextcalculator: evaluates mathematical expressions for exact arithmetic

Grading
All responses are graded by GPT 5.4 as a judge against each task’s rubric (see Scoring for the weighted metric and must-pass structure). Cited sources are independently verified, and the fraction that resolve to real authoritative documents scales the rubric score.
Dataset
The dataset comprises 391 expert-authored questions, each paired with a gold-standard answer, authoritative sources, and a detailed grading rubric. It is divided into three parts: Public (5 open samples), Private Validation (193 samples available for license), and Test (193 samples).
- The Public set and agent harness are open and can be accessed here.
- The Private Validation set is available for license. Interested parties are encouraged to contact us directly for access.
- The Test set will remain private. Unless otherwise noted, all results reported on this page are based solely on the Test set to prevent overfitting; the only exceptions are the trajectory and sample-answer panels for public question P-004, which are illustrative and do not enter any score.
The Test and Validation splits were sampled to preserve the distribution of question categories.
Citation
If you use this benchmark in your research, please cite:
Citation (BibTeX)
@misc{valsai2026taxagentbench,
title = {Tax Agent Bench: Evaluating Agents on US Corporate Tax Research},
author = {Vals AI},
year = {2026},
month = sep,
howpublished = {Vals AI},
url = {https://www.vals.ai/benchmarks/tax_agent_bench},
}