MODELS 13 configurations evaluated

Model Directory

Every agent and model configuration FlutterBench has evaluated, with its full accuracy, cost, and latency profile across each Critical User Journey.

ANTHROPIC

Claude 3.7 Sonnet

EVAL KEY: claude-code__claude-3-7-sonnet__adhoc

DEVELOPER
Anthropic
AGENT
claude-code
MODEL ID
claude-3-7-sonnet
VARIANT
adhoc
DART TOOLING
Enabled
TRIALS
5
TOKENS (IN/OUT)
470k / 23k
PASS@1
0.80
Data coming soon Release date, context window, max output tokens, published token pricing, weight availability, and input modalities.
OVERALL SCORE 0.89 Range 0.71–0.98
COST / TRIAL $0.351 Total $1.753
LATENCY / TRIAL 2m 28s Mean wall clock
View hyperparameter settings
Data coming soon Temperature, top-p, reasoning effort, and tool configuration for each run.

Reward on a 0–1 scale. Higher is better. Rankings compare this configuration against every other one scored on the same task.

Benchmark Score Accuracy Ranking
Custom RenderObject & Canvas 0.71 4 / 11
Build Command-Line CLI App 0.97 4 / 13
Offline SQLite Sync Repository 0.93 1 / 13
Manage State with BLoC 0.98 2 / 12
Adaptive Material & Cupertino UI 0.85 2 / 13