Web Search Index

Proprietary

Updated: 7/16/2026

Comparing native provider search against independent web-search tools on legal-research and finance-analysis tasks

Key Takeaways

  • The better search tool depends on the domain: controlling for task difficulty with a mixed-effects model, Exa beats native by about 6 points on finance (p < 0.001), while on legal the two are statistically even (p = 0.86).
  • Exa’s finance edge concentrates in information-gathering tasks like Market and Earnings Analysis (+14 to +18 points over native) and fades on computation-heavy categories.
  • Native is cheaper on average, but total cost is dominated by model inference, so the search fee is negligible, and Exa is faster in both domains.

Background

Web search is a critical information-retrieval tool for a wide range of tasks, from simple fact-checking to in-depth academic research. A wave of search tools built for AI agents has emerged recently, from frontier model providers’ native search tools to independent tools like Exa, each with its own pricing and capabilities.

The Web Search Index evaluates these tools end-to-end across two established benchmarks: FAB v2 (Finance Agent Benchmark) and Legal Research Benchmark.

To isolate each tool’s contribution, we swap the search tool the model has access to while holding the model and the rest of the harness constant. The difference in performance then reflects the marginal boost the search tool provides.

We also remove the domain-specific search tools the original benchmarks provide (such as SEC EDGAR for financial filings and CourtListener for legal cases), so the general web-search tool is the model’s only source of external information.

We’re eager to evaluate other search API providers. If you’re interested in having your product evaluated, get in touch with us here.


Results

Overall

Score / Cost
$0.50$1.0$2.0$5.0$10$20303540455055Cost per task (log scale)Score (%)Gemini 3.5 FlashGrok 4.5GPT-5.6 SolClaude Fable 5Grok 4.5Gemini 3.5 FlashClaude Fable 5GPT-5.6 Sol
Native
Exa
Pareto efficiency frontier

We analyze the cost-accuracy tradeoff through the lens of the Pareto efficiency frontier: for a given budget, which model-tool combination achieves the highest accuracy. Cost here is the total per task: model inference cost plus the Exa search fee for Exa runs. Overall, Claude Fable 5 with Exa has the highest accuracy of any combination (48.5%), and two of the four Pareto-efficient combinations use Exa search. The split is clearer in the Finance Analysis domain: five combinations are Pareto-efficient, three of them on Exa, again led by Fable 5 with Exa (53.6%). In the Legal Research domain, all three combinations on the frontier use native search, topped by Fable 5 with native search (49.0%).

Taken together, Exa makes up more of the finance frontier (three of five combinations), while native search holds the entire legal frontier; overall the frontier is split roughly evenly between the two.

It is also worth pointing out that GPT-5.6 Sol with Exa is the cost outlier: it is consistently the most expensive option in every domain (about $13.32 overall, up to $19.28 on legal) yet never falls on the efficiency frontier. The higher cost comes from a small set of tasks where the model falls into a runaway search loop, firing hundreds of Exa calls before answering (one legal task made over 1,900).

Cost per task ($): model + external search (mean)
$0.00$4.0$8.0$12$16$20Gemini 3.5 FlashGrok 4.5Claude Fable 5GPT-5.6 Sol$0.40Native$1.18Exa$1.01Native$0.75Exa$4.03Native$3.48Exa$1.17Native$13.32ExaCost per task ($)
Native model costExa model costExa search fee

Next, we break down the total cost into model inference cost and Exa search fee. Model cost dominates in every combination; the Exa search fee accounts for only a minor part.

Cost per Task Distribution: Native vs. Exa (total cost, log scale)
Finance Analysis
$0.01$0.10$1$10$100$1000Gemini 3.5 FlashGrok 4.5Claude Fable 5GPT-5.6 SolNativeExaNativeExaNativeExaNativeExaCost / task ($, log)
Legal Research
$0.01$0.10$1$10$100$1000Gemini 3.5 FlashGrok 4.5Claude Fable 5GPT-5.6 SolNativeExaNativeExaNativeExaNativeExaCost / task ($, log)

Here we show the distribution of total cost across models and domains. Notably, when Exa is used as the search tool, GPT-5.6 Sol shows high cost variance in both domains.

Finance Analysis, by Category: Native vs. Exa

Exa scores higher than native in most finance task categories and roughly matches it on the rest. The margin is widest where a task depends heavily on retrieving data: Market Analysis (about 54% for Exa vs. 36% for native), Earnings Analysis (73% vs. 59%), and Adjustments (44% vs. 36%), the categories where a stronger search tool has the most to work with. The two tools converge where retrieval matters least. On Comparables they land within about a point of each other, and on Financial Modeling both sit at roughly 18%. These tasks hinge on exact, multi-step computation over a small, fixed set of documents rather than on finding more sources, so extra search barely moves them.

Legal Research, by Question Type: Native vs. Exa

The legal domain exhibits a different pattern. Across task categories the two tools have similar performance: native performs slightly better in a majority of the categories, but the pooled effect is not statistically significant. Native search performs best on tasks requiring the application of precedent to a fact pattern (about 35% vs. 32% for Exa) and a few other tasks involving case-law reasoning. Exa’s clearest edge is Regulatory Framework Interpretation (54% vs. 48%), where finding the right regulatory text matters a lot for successful reasoning. Both tools bottom out together on the same hard reasoning types, such as Consensus vs. Disagreement Across Courts, where neither clears 26%.

Search Calls per Task, Native vs. Exa (box plot, log scale)
Finance Analysis
1101001000Search calls / task (log)18Native14Exa
Legal Research
110100100010000Search calls / task (log)16Native12Exa

At the median, models make more native search tool calls than Exa calls in both domains: the median legal run makes 16 native calls vs. 12 with Exa, and the median finance run makes 18 vs. 14 calls. The two distributions differ mainly in the upper tail: a handful of Exa runs reach 1,000+ searches, the same GPT-5.6 Sol runaway loops noted above, while native counts stay bounded, peaking near 90 in legal and under 400 in finance. That long Exa tail is what lifts its mean above native’s even though its median sits lower.

Answer Length, Native vs. Exa
Legal Research
Finance Analysis
Sources Cited per Answer, Native vs. Exa
Legal Research
Finance Analysis

Models generate longer responses with Exa than with the native search tool across both domains. Average response length increases from 1,380 to 1,590 words in legal and from 735 to 790 words in finance. They also cite about 15 to 18% more distinct sources with Exa, rising from 8.5 to 10.0 in legal and from 5.3 to 6.1 in finance.

Where do those sources come from? The tables below list the five most-cited domains for each tool, as a share of that tool’s citations.

Finance Analysis

RankNativeExa
1sec.gov53.7%sec.gov57.6%
2yahoo.com4.1%yahoo.com3.9%
3q4cdn.com3.9%q4cdn.com2.5%
4google.com2.1%macrotrends.net1.2%
5macrotrends.net1.3%fool.com1.0%
Top 565.1%66.2%

Legal Research

RankNativeExa
1justia.com25.9%justia.com19.2%
2cornell.edu10.3%cornell.edu11.5%
3virginia.gov6.2%virginia.gov5.9%
4texas.gov5.7%vlex.com5.8%
5findlaw.com5.1%texas.gov5.3%
Top 553.1%47.8%

As it turns out, both search tools cite similar, authoritative sources. In legal, both tools lead with Justia and Cornell’s Legal Information Institute; in finance, both are dominated by SEC. The five most-cited domains alone account for roughly half of all legal citations and two-thirds of finance citations.


Methodology

Agents run on the same harness across both domains, with the web-search tool (native/Exa) as the variable under test. Two things distinguish these runs from the original FAB v2 and Legal Research Benchmark evaluations: the tool is swapped (native vs. Exa), and the domain-specific search tools those benchmarks provide (SEC EDGAR for finance, CourtListener for legal, and the like) are removed, leaving the web-search tool as the agent’s only means of gathering external information. Runs are orchestrated through Valkyrie with model calls routed through model-library.

Grading

  • Legal Research: each task in the Legal Research Benchmark carries a rubric of required elements written by legal experts; rubrics range from 1 to 22 items per task (mean 9.47). All-pass is our primary metric: a task counts as correct only if every rubric item is satisfied. Grading is performed by a single LLM judge (GPT-5.4).
  • Finance Analysis: tasks use FAB v2’s dealbreaker-gated Partial Credit as the primary metric, category-balanced across the dataset’s nine question categories. A subset of checks in each question is flagged as dealbreakers—load-bearing facts or numbers required for a satisfactory answer—and failing any dealbreaker means the response scores 0% for that question regardless of everything else. All-pass (100% only if every check passes) is reported as a secondary metric. Grading uses the same three-judge jury as FAB v2: GPT-5.4, Gemini-3.1-Pro, and Claude Sonnet 4.6.

For the full rubrics, jury setup, and scoring rules behind any specific grading decision, refer to the original FAB v2 and Legal Research Benchmark methodologies.

Dataset

  • Legal: 208 expert-authored legal research tasks, each tagged by the reasoning skill it tests (a task can test more than one): Statutory Interpretation, Application of Precedent to a Fact Pattern, Doctrinal Test Identification, Regulatory Framework Interpretation, Interaction of Multiple Legal Regimes, Cross-Jurisdiction Comparison, Timeline/Doctrinal Evolution, Validity of Precedent, Resolving Conflicting Precedent, and Consensus vs. Disagreement Across Courts.
  • Finance: 450 finance-analyst questions across nine categories, scored with FAB v2’s category-balanced methodology and reflecting real equity-research workflows: General Qualitative Analysis, General Quantitative Analysis, Market Analysis, Comparables, Precedents, Adjustments, Earnings Analysis, Disclosure Analysis, and Financial Modeling.

Statistical Analysis

To establish whether the two search tools differ statistically, we fit a mixed-effects model per domain, with search tool as a fixed effect (native as the reference level) and a random intercept per task absorbing task-level difficulty: a logistic model for legal’s All-pass metric (all_pass ~ search_tool + (1 | task_id)) and a linear model for finance’s primary metric (score ~ search_tool + (1 | task_id)). Each model pools the per-task, per-model observations underlying the Results charts above. The domain-level significance cited in the Key Takeaways comes from these models.

The fixed effect of search tool (Exa relative to native) in each model:

DomainOutcome (model)Exa effectStd. errorStatisticp
FinanceWeighted score, 0-100 (linear)+6.51 points0.98t = 6.62< 0.001
LegalAll-pass, 0/1 (logistic)−0.02 log-odds0.13z = −0.170.86