Agent benchmarks.

Task success and cost across the evaluated models. Each benchmark states its workload, scoring, and cost definition.

Internal Bench Hard

Browser Use: 82% of tasks solved at 17¢ each. That is 20 points better than Opus 5, which costs 20× more per solved task.

Toolstrict accuracyCost
Browser Use82%$0.17
Opus 562%$3.40
Gemini 3.1 Pro59%$2.20
Sonnet 559%$1.55
GPT-5.652%$1.10
Gemini 3.6 Flash46%$0.62
GPT-537%$0.44
Internal Bench Hard · updated 2026-08-01 · all benchmarks
What it measures

Our hardest internal set: 106 tasks that a careful human can finish but most agents cannot. Every model runs on the same harness, and we count a task solved only on a strict match, so partial credit earns nothing.

How it was run
  • Tasks: 106 hard tasks on live websites, none removed.
  • Cost: total recorded spend for the run divided by the number of tasks actually solved.
  • Scoring: strict — a task counts only when the final answer or end state is exactly right.