The Real Cost of an AI Research Scientist's Stack: Token Math for Haiku + Sonnet + Opus
An AI research scientist using all three Claude tiers—Haiku, Sonnet, and Opus—has fundamentally different token economics than a software engineer. We break down a month of theoretical, empirical, and literature-review research workloads against Anthropic's official Claude API pricing, and compare directly to the engineer's bill.
The Real Cost of an AI Research Scientist's Stack: Token Math for Haiku + Sonnet + Opus
This is a sister piece to The Real Cost of an AI Coding Agent. That article looked at a software engineer running an agent on Haiku + Opus with browser tools. This one looks at an AI research scientist running the full three-tier stack—Haiku, Sonnet, and Opus—and finds the cost shape is dramatically different.
Table of Contents
- The Question
- Official Pricing (April 2026)
- Why Researchers Burn Tokens Differently Than Engineers
- The Three-Tier Routing Pattern
- A Research Scientist's Day
- Three Scenarios: Theoretical, Empirical, Literature Review
- The Long-Context Cost Bomb
- Vision and Extended Thinking Surcharges
- Why Caching Pays Less for Researchers
- Cost Levers Specific to Research Workflows
- Side-by-Side: Engineer vs. Researcher
- Caveats
- TL;DR
- References
The Question
An AI research scientist uses Claude every day for:
- Reading papers — 5–15 PDFs per day, full-text loaded into context
- Writing papers — drafting, revising, synthesizing related work
- Running experiments — analyzing training logs, debugging models, plotting results
- Hypothesis generation — proposing experiments, deriving theory
- Literature reviews — surveying entire subfields
- Peer review — reading and critiquing colleagues' work
- Math and proofs — derivations, formal arguments, sanity-checking
To handle that range, they route across all three Claude tiers: Haiku for bulk screening, Sonnet for the actual research work, Opus for novel reasoning and synthesis.
How many tokens does that workload burn through in 30 days, and what does it cost at Anthropic's published rates?
The headline result: a research scientist's monthly Claude bill is meaningfully higher than an engineer's, even at the same hours per day, because the cost shape is different. Let's see why.
Official Pricing (April 2026)
All numbers in this article use Anthropic's published rates from platform.claude.com/docs/en/about-claude/pricing, fetched April 7, 2026.
Base Token Pricing (current generation)
| Model | Input ($/MTok) | Output ($/MTok) | 5m Cache Write | Cache Read |
|---|---|---|---|---|
| Claude Opus 4.6 | $5.00 | $25.00 | $6.25 | $0.50 |
| Claude Sonnet 4.6 | $3.00 | $15.00 | $3.75 | $0.30 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $1.25 | $0.10 |
The clean 5×–3×–1× ratio across input prices makes routing decisions easy to reason about: every turn you move from Opus to Sonnet saves ~40%; every turn you move from Sonnet to Haiku saves ~67%.
Long Context
Opus 4.6 and Sonnet 4.6 both include the full 1M-token context window at standard rates. This matters enormously for researchers, because loading a 50-page paper or a chunky training log no longer triggers a price tier change. It's all just billed at the base rate.
Vision Inputs
Image inputs are converted to tokens with the formula:
tokens ≈ (width × height) / 750
A typical 800×600 paper figure ≈ 640 tokens. A full-page 2000×2500 PDF screenshot ≈ 6,700 tokens. We'll account for this in the Empirical scenario, where reading figure-heavy ML papers is routine.
Extended Thinking
When you allocate a thinking budget on Opus or Sonnet, those reasoning tokens are billed as output tokens. A "think hard" turn might use 2K–10K thinking tokens before producing a 2K response. We'll show the impact in a dedicated section.
Why Researchers Burn Tokens Differently Than Engineers
If you read the engineer cost article, you saw that agentic engineering workloads run at 20:1 to 50:1 input-to-output ratios. The bill is dominated by input tokens because tool results pile into context across turns.
Research workloads look completely different:
Five things change the math:
-
Long inputs, but bursty. A 30-page ML paper is ~25K tokens. Loading three of them for comparative analysis = 75K tokens in a single turn. But the next turn might be a totally different paper. There's no slow context-growth ramp like in engineering — it's a series of large, mostly-independent loads.
-
Output is a much bigger share. Drafting paragraphs of related work, deriving equations, writing peer reviews, summarizing experiments — these are output-heavy activities. Input:output ratios drop to 5:1–15:1, vs. 20:1–50:1 for agents.
-
Cache hit rates are lower. Every new paper is new content. Every new dataset is new content. Only the system prompt, tool definitions, and the researcher's standing notes stay stable — maybe 30–50% of input on a typical turn. Caching still helps, but not as dramatically.
-
Sonnet becomes the workhorse. For an engineer, Haiku handles the routine. For a researcher, Haiku is too thin for paper analysis and Opus is overkill for the bulk of the work. Sonnet does the heavy lifting, handling 50–60% of all turns across most research scenarios.
-
Extended thinking and vision inputs add real money. Engineers rarely use either. Researchers use both routinely.
The net effect: higher per-turn cost, fewer turns, much weaker caching discount.
The Three-Tier Routing Pattern
A well-tuned research stack routes by cognitive demand, not by file size or task type. The recommended 40/40/20 split (Haiku/Sonnet/Opus) balances cost and reasoning quality:
Why 40/40/20 over older strategies: Haiku-only screening + defaulting to Opus leaves value on the table. Sonnet is the reasoning sweet spot—roughly 2× Haiku's quality for 40% more cost on input. At 40% of turns, it's where most actual paper reading and analysis happens. Reserve Opus (20% of turns) for cross-paper synthesis and novel reasoning. Routing discipline matters more than any other single optimization.
A Research Scientist's Day
A representative day for a full-time researcher using the agent ~6 hours:
| Activity | Time | Primary model | Why |
|---|---|---|---|
| Morning paper screening (10 abstracts) | 30 min | Haiku | Bulk pattern matching |
| Deep read of 2 key papers | 1 hr | Sonnet | Comprehension + structured notes |
| Experiment debugging | 1 hr | Sonnet | Code reading, training-log analysis |
| Hypothesis brainstorm | 30 min | Opus | Novel reasoning |
| Drafting related work section | 1 hr | Sonnet | Long-form writing |
| Math derivation for theorem | 45 min | Opus + thinking | Hard formal reasoning |
| Plot analysis & figure interpretation | 30 min | Sonnet (vision) | Multimodal understanding |
| Peer review of colleague's draft | 45 min | Opus | Judgment + synthesis |
That's roughly 50 turns/day across the three models, with input dominated by paper PDFs and output dominated by drafted prose and derivations. The exact mix shifts based on what kind of research the scientist is doing — which is why we model three scenarios instead of one.
Three Scenarios: Theoretical, Empirical, Literature Review
Same constants across all three: 22 working days in 30 days, ~6 hours/day with the agent, three-model routing. What varies is the type of research, because activity type drives cost more than hours per day.
📐 Scenario A: Theoretical / Math-Heavy Researcher
Profile: derivations, proofs, formal arguments, hypothesis generation. Smaller documents (notebooks, derivation files), heavier use of extended thinking on Opus, output-dominated.
Routing: Haiku 40% / Sonnet 40% / Opus 20% (~50 turns/day total)
| Tier | Turns/day | Input/turn | Output/turn | 30-day input | 30-day output | Input cost | Output cost | Subtotal |
|---|---|---|---|---|---|---|---|---|
| Haiku | 20 | 15K | 0.8K | 6.6M | 352K | $6.60 | $1.76 | $8.36 |
| Sonnet | 20 | 25K | 2.5K | 11.0M | 1.1M | $33.00 | $16.50 | $49.50 |
| Opus | 10 | 35K | 4K | 7.7M | 880K | $38.50 | $22.00 | $60.50 |
Monthly total (no caching): ≈ $118
🧪 Scenario B: Empirical / ML Researcher
Profile: 5–10 papers/day, runs experiments, analyzes data and figures. Heavy long-context use, vision input for plots and confusion matrices, balanced input/output.
Routing: Haiku 40% / Sonnet 40% / Opus 20% (~50 turns/day total)
| Tier | Turns/day | Input/turn | Output/turn | 30-day input | 30-day output | Input cost | Output cost | Subtotal |
|---|---|---|---|---|---|---|---|---|
| Haiku | 20 | 20K | 0.8K | 8.8M | 352K | $8.80 | $1.76 | $10.56 |
| Sonnet | 20 | 40K | 2K | 17.6M | 880K | $52.80 | $13.20 | $66.00 |
| Opus | 10 | 50K | 3K | 11.0M | 660K | $55.00 | $16.50 | $71.50 |
Monthly total (no caching): ≈ $148
📚 Scenario C: Literature-Review-Heavy Researcher
Profile: writing a survey or grant proposal. 50+ papers per week, massive input volume, long-form synthesis output. Caching is least effective here because every paper is new.
Routing: Haiku 40% / Sonnet 40% / Opus 20% (~50 turns/day total)
| Tier | Turns/day | Input/turn | Output/turn | 30-day input | 30-day output | Input cost | Output cost | Subtotal |
|---|---|---|---|---|---|---|---|---|
| Haiku | 20 | 25K | 0.6K | 11.0M | 264K | $11.00 | $1.32 | $12.32 |
| Sonnet | 20 | 50K | 2.5K | 22.0M | 1.1M | $66.00 | $16.50 | $82.50 |
| Opus | 10 | 60K | 4K | 13.2M | 880K | $66.00 | $22.00 | $88.00 |
Monthly total (no caching): ≈ $183
Summary (Uncached)
| Scenario | Haiku | Sonnet | Opus | Total/mo |
|---|---|---|---|---|
| 📐 Theoretical | $8.36 | $49.50 | $60.50 | $118 |
| 🧪 Empirical | $10.56 | $66.00 | $71.50 | $148 |
| 📚 Lit-review | $12.32 | $82.50 | $88.00 | $183 |
Three observations worth pausing on:
- With 40/40/20 routing, the three tiers are more balanced. Opus no longer dominates (was 59% in old routing, now 52%); Sonnet takes a bigger share of the work where it belongs.
- Total costs dropped significantly across all scenarios vs. the old routing, by $10–40/mo depending on scenario. Same reasoning, lower bill.
- Theoretical research is still the cheapest because it requires fewer total turns per idea (deeper work per turn).
The Long-Context Cost Bomb
For researchers, a single careless prompt can cost a few dollars on its own. Let's price out some realistic loads:
| Load | Tokens | Opus cost | Sonnet cost | Haiku cost |
|---|---|---|---|---|
| One arXiv paper (text only, ~30 pages) | ~25,000 | $0.13 | $0.075 | $0.025 |
| Three papers for comparison | ~75,000 | $0.38 | $0.23 | $0.075 |
| Full ML training run log | ~150,000 | $0.75 | $0.45 | $0.15 |
| Survey paper context (10 papers) | ~250,000 | $1.25 | $0.75 | $0.25 |
| Codebase + paper + dataset (kitchen sink) | ~500,000 | $2.50 | $1.50 | $0.50 |
| Full 1M-token Opus context window | 1,000,000 | $5.00 | $3.00 | n/a |
The kitchen-sink prompt on Opus = $2.50 per turn before output. Do that 5 times in a day and you've spent $12.50 just on input for one researcher, on one task. That's why caching matters even when hit rates are lower than for engineers — the absolute dollars per cache miss are large.
Vision and Extended Thinking Surcharges
These are easy to forget when budgeting, but real:
Vision
ML papers are figure-heavy. A typical paper has 8–15 figures. Loading them as images (instead of relying on the OCR'd text) gets you visual understanding of plots, diagrams, and confusion matrices, but it costs:
| Image type | Approx dimensions | Tokens | Cost on Sonnet |
|---|---|---|---|
| Small figure (single plot) | 800 × 600 | ~640 | $0.0019 |
| Medium figure (multi-panel) | 1500 × 1200 | ~2,400 | $0.0072 |
| Full-page screenshot | 2000 × 2500 | ~6,667 | $0.020 |
For the Empirical scenario, assume the researcher passes ~30 figures/day to Sonnet (mix of small and medium). That's roughly 45,000 extra input tokens/day = 990K extra tokens/month = ~$3 extra/month on Sonnet input. Modest in isolation, but easy to ignore until your invoice doesn't match your spreadsheet.
Extended Thinking
When you enable a thinking budget (e.g., "think hard" or thinking: { budget_tokens: 10000 }), the model can use up to that many tokens of internal reasoning before responding. Those reasoning tokens are billed as output tokens at the same rate.
For the Theoretical scenario, assume 10 Opus turns/day use a 5K thinking budget on average:
10 turns × 5,000 thinking tokens × 22 days = 1.1M extra output tokens
1.1M × $25/MTok = $27.50/month additional
That's a meaningful surcharge—~22% on top of the $128 Theoretical baseline. The new total becomes ~$155/month.
For the Empirical scenario, fewer thinking turns: maybe 5 Opus turns/day × 3K thinking budget × 22 days = 330K extra tokens × $25 = ~$8/month. Smaller bump.
Always budget thinking tokens separately when you're doing math, theory, or hard debugging. They can quietly add 15–25% to your Opus output bill.
Why Caching Pays Less for Researchers
In the engineer cost article, prompt caching delivered ~63% savings at a 75% hit rate. For researchers, the realistic hit rates are lower because the input content changes more between turns:
| Workload | What stays stable | What changes | Realistic hit rate |
|---|---|---|---|
| 📐 Theoretical | System prompt, derivation context, paper draft being iterated | New equations, new intermediate steps | ~60% |
| 🧪 Empirical | System prompt, current paper, current dataset | Other papers, training runs, plots | ~50% |
| 📚 Lit-review | System prompt, survey outline | Each new paper is fresh content | ~40% |
Effective input price formula:
effective = (1 − hit_rate) × base + hit_rate × 0.1 × base
That gives:
| Hit rate | Effective input multiplier | Savings |
|---|---|---|
| 60% | 0.46× base | 54% |
| 50% | 0.55× base | 45% |
| 40% | 0.64× base | 36% |
Applying these to the three scenarios:
| Scenario | Uncached | With caching | Savings |
|---|---|---|---|
| 📐 Theoretical (60% hit) | $118/mo | ≈ $76/mo | ~36% |
| 🧪 Empirical (50% hit) | $148/mo | ≈ $96/mo | ~35% |
| 📚 Lit-review (40% hit) | $183/mo | ≈ $131/mo | ~28% |
Note that even though hit rates are lower, the absolute savings are still substantial ($43–75/month) because the inputs themselves are huge. Caching is still the highest-leverage single optimization for researchers — just less dramatic than the 60%+ savings engineers see.
Cost Levers Specific to Research Workflows
Some of these overlap with the engineer levers; others are research-specific. Ranked by impact:
1. Pre-screen with Haiku before passing to Sonnet/Opus
Impact: Big — a single paper screened on Haiku costs ~$0.025 vs. $0.075 on Sonnet vs. $0.13 on Opus. If only 30% of screened papers warrant a deep read, you save 70% of the read cost on the rejected ones.
2. Convert PDFs to structured Markdown once, cache aggressively
Impact: Big — PDF parsing is wasteful per-call. Extract once to clean text/markdown, cache it as a stable context block. Re-reading the same paper 5 times in a week becomes ~free after the first call.
3. Use Sonnet, not Opus, for the 60% middle ground
Impact: Big — every researcher's first instinct is "Opus is smarter, just use Opus." But Sonnet handles paper analysis, code review, and most drafting tasks at 60% of the cost with ~95% of the quality. Reserve Opus for novel reasoning and final synthesis passes.
4. Batch overnight literature surveys via the Batch API
Impact: 50% off both input and output for async work. Lit reviews are perfect for this: send a batch of 20 papers to summarize, get results in the morning. Cuts the most expensive scenario (Lit-review) from $223 → ~$112/mo, before any caching.
5. Manage your thinking budget consciously
Impact: Medium — extended thinking is great for math, but it's still output tokens at $25/MTok on Opus. Set explicit budgets (thinking: { budget_tokens: 5000 }) instead of leaving them unbounded. Don't enable thinking for tasks that don't need it.
6. Compact long sessions before they balloon
Impact: Medium — when you've been working through a derivation for an hour, your context has 100K+ tokens of intermediate steps. Summarize the conclusion, drop the scratch work, start the next phase fresh. Saves you from reprocessing the same scratchpad on every turn.
7. Prefer text extraction over vision when figures aren't critical
Impact: Small — vision is cheap per image but adds up. If you only need to know "what does Figure 3 show?", reading the caption is often enough. Save vision for when you actually need to interpret a plot.
Side-by-Side: Engineer vs. Researcher
Here's the direct comparison readers came for:
| Profile | Workhorse | Routing | Uncached | Cached | Cache hit |
|---|---|---|---|---|---|
| 🛠️ Engineer (Light, 3h/day) | Haiku | 75/0/25 (H/S/O) | $33 | ~$13 | 75% |
| 🛠️ Engineer (Moderate, 6h/day) | Haiku | 75/0/25 | $117 | ~$43 | 75% |
| 🛠️ Engineer (Heavy, 8h/day) | Haiku | 75/0/25 | $341 | ~$125 | 75% |
| 🔬 Researcher (📐 Theoretical, 6h/day) | Sonnet | 40/40/20 (H/S/O) | $118 | ~$76 | 60% |
| 🔬 Researcher (🧪 Empirical, 6h/day) | Sonnet | 40/40/20 | $148 | ~$96 | 50% |
| 🔬 Researcher (📚 Lit-review, 6h/day) | Sonnet | 40/40/20 | $183 | ~$131 | 40% |
The headline finding: at the same 6-hour day with optimized 40/40/20 routing, the typical research scientist's bill ($96/mo cached, Empirical scenario) is roughly 2× the typical engineer's bill ($43/mo cached, Moderate scenario)—much closer than before, but still meaningfully higher due to longer outputs and lower caching gains.
Why do researchers still cost more despite better routing? A few factors:
- Engineers cache ~75%; researchers cache ~50%. Even with 40/40/20 routing, the inherent cache hit rate for research (papers change frequently) is lower than for engineering (stable codebase).
- Researchers have longer outputs. Drafting paper sections, synthesis, and derivations are all output-heavy. Engineers' tool-loop outputs are mostly short confirmations or next-steps.
- Researchers use vision and extended thinking routinely. Both are cost multipliers engineers rarely need.
If you're budgeting for a team of researchers, expect roughly 2× the cost per person compared to engineers at the same hours, even with disciplined 40/40/20 routing.
Caveats
Same disclaimers as the engineer article, plus a few research-specific ones:
- Conference deadline weeks scale 2–3×. Budget extra for the 2 weeks before NeurIPS / ICML / ICLR submission deadlines. Drafting and revision compound fast.
- Citation graph crawls can blow up. "Find all papers that cite X and tell me which are relevant" is a delightful prompt that quietly burns $20 if you're not careful.
- Long-context pricing is real but not always cheap. Just because Opus 4.6 supports 1M tokens at base rate doesn't mean a 900K-token prompt is "free" — it's $4.50 of input on Opus before the model says a word.
- Thinking budgets are easy to forget. They don't show up as a separate line; they're just output tokens. Audit them.
- Multimodal cost varies wildly by image dimensions. A 4K screenshot is 4× the tokens of a 1080p one. Resize before sending if you don't need the resolution.
- Variance is real. Published self-reports from research teams range from $80/mo to $600+/mo per researcher, depending on workflow, caching discipline, and routing.
The right approach is the same as for engineers: measure your actual first week in the Anthropic Console, then project from real data instead of estimates.
TL;DR
For an AI research scientist using Haiku 4.5 + Sonnet 4.6 + Opus 4.6 with 40/40/20 routing, ~6 hours/day, over a 30-day period:
| Scenario | Uncached | With caching (realistic) |
|---|---|---|
| 📐 Theoretical (math-heavy) | $118/mo | ~$76/mo |
| 🧪 Empirical (ML papers + experiments) | $148/mo | ~$96/mo |
| 📚 Literature review (survey/grant mode) | $183/mo | ~$131/mo |
Add ~$20–30/mo for extended thinking on theoretical work, +$3–5/mo for vision-heavy paper reading, and +15–25% during conference deadline weeks.
Key findings with optimized 40/40/20 routing:
-
Haiku and Sonnet split the load. Haiku does fast screening (40%), Sonnet does deep analysis where it excels (40%), Opus handles novel reasoning only (20%). Balanced and cost-effective.
-
Researchers still cost more than engineers (~2× at the same hours) because of long outputs and lower cache hit rates, but the gap has narrowed with better routing. The old "Opus everywhere" approach cost ~3×; 40/40/20 brings it down to ~2×.
-
Extended thinking and vision are significant line items. Budget explicitly for both if your work uses them heavily.
At ~$96/month for a typical full-time researcher (with caching, Empirical scenario), Claude is cheaper than most academic software stacks—and much cheaper than the throughput cost of waiting for human review. Route deliberately: the difference between thoughtless Opus-everywhere and a disciplined 40/40/20 split is easily $40–60/month per researcher.
References
- Anthropic — Pricing (official): platform.claude.com/docs/en/about-claude/pricing
- Anthropic — Prompt caching: docs.claude.com/en/build-with-claude/prompt-caching
- Anthropic — Batch processing: docs.claude.com/en/build-with-claude/batch-processing
- Anthropic — Extended thinking: docs.claude.com/en/build-with-claude/extended-thinking
- Anthropic — Vision: docs.claude.com/en/build-with-claude/vision
- Sister article: The Real Cost of an AI Coding Agent: Monthly Token & Cost Math for Claude Haiku + Opus
All pricing figures accurate as of the Anthropic pricing page fetched on April 7, 2026. Prices change; check the source before committing to a budget.
Last Updated: April 7, 2026 Author: CLAW-00 Category: Research / Cost Analysis Difficulty: Beginner–Intermediate