MODELS
13 configurations evaluated
Model Directory
Every agent and model configuration FlutterBench has evaluated, with its full accuracy, cost, and latency profile across each Critical User Journey.
ANTHROPIC
Claude 3.7 Sonnet
EVAL KEY: claude-code__claude-3-7-sonnet__adhoc
- DEVELOPER
- Anthropic
- AGENT
- claude-code
- MODEL ID
- claude-3-7-sonnet
- VARIANT
- adhoc
- DART TOOLING
- Enabled
- TRIALS
- 5
- TOKENS (IN/OUT)
- 470k / 23k
- PASS@1
- 0.80
Data coming soon
Release date, context window, max output tokens, published token pricing, weight availability, and input modalities.
OVERALL SCORE
0.89
Range 0.71–0.98
COST / TRIAL
$0.351
Total $1.753
LATENCY / TRIAL
2m 28s
Mean wall clock
View hyperparameter settings
Data coming soon
Temperature, top-p, reasoning effort, and tool configuration for each run.
Reward on a 0–1 scale. Higher is better. Rankings compare this configuration against every other one scored on the same task.
| Benchmark | Score | Accuracy | Ranking |
|---|---|---|---|
| Custom RenderObject & Canvas | 0.71 | 4 / 11 | |
| Build Command-Line CLI App | 0.97 | 4 / 13 | |
| Offline SQLite Sync Repository | 0.93 | 1 / 13 | |
| Manage State with BLoC | 0.98 | 2 / 12 | |
| Adaptive Material & Cupertino UI | 0.85 | 2 / 13 |