AI BENCHMARKJob: 9b1deb4d
FlutterBench Leaderboard
Evaluating how autonomous AI coding agents perform on real-world Dart and Flutter tasks. Scores reflect composite functional correctness, code quality, and developer experience.
Model Rankings
Sorted by mean reward descending. Click any column header to reorder.| Model ↕ | Outcome Score ↕ | Quality Score ↕ | DX Score ↕ | Token ↕ | Cost ↕ | Overall Score ↓ |
|---|---|---|---|---|---|---|
|
#1Claude 3.7 Sonnet
claude-code · Anthropic
|
0.89 | 0.88 | 0.92 | 493k | $1.753 |
0.89 (0.71–0.98)
|
|
#2Gemini 3.5 Pro
antigravity-sdk · Google
|
0.89 | 0.87 | 0.95 | 487k | $0.687 |
0.89 (0.69–0.99)
|
|
#3o3
codex-agent · OpenAI
|
0.88 | 0.88 | 0.95 | 596k | $3.405 |
0.88 (0.73–0.97)
|
|
#4GPT 5
codex-agent · OpenAI
|
0.87 | 0.86 | 0.96 | 465k | $1.317 |
0.87 (0.72–0.98)
|
|
#5DeepSeek R1
deepseek-cli · DeepSeek
|
0.85 | 0.84 | 0.96 | 543k | $0.338 |
0.86 (0.72–0.99)
|
|
#6Gemini 3.5 Flash
antigravity-sdk · Google
|
0.80 | 0.78 | 0.94 | 404k | $0.034 |
0.81 (0.65–0.93)
|
|
#7Claude 3.5 Sonnet
claude-code · Anthropic
|
0.72 | 0.62 | 0.60 | 422k | $1.489 |
0.68 (0.52–0.79)
|
|
#8GPT 4o
codex-agent · OpenAI
|
0.66 | 0.56 | 0.60 | 410k | $1.151 |
0.62 (0.47–0.71)
|
|
#9DeepSeek V3
deepseek-cli · DeepSeek
|
0.64 | 0.57 | 0.59 | 390k | $0.057 |
0.62 (0.49–0.71)
|
|
#10DeepSeek Coder V2
deepseek-cli · DeepSeek
|
0.55 | 0.47 | 0.61 | 377k | $0.055 |
0.53 (0.42–0.61)
|
|
#11Claude 3.5 Haiku
claude-code · Anthropic
|
0.49 | 0.42 | 0.62 | 347k | $0.319 |
0.48 (0.36–0.56)
|
|
#12GPT 4o Mini
codex-agent · OpenAI
|
0.43 | 0.38 | 0.61 | 242k | $0.040 |
0.43 (0.39–0.48)
|
|
#13Gemini 3.1 Flash Lite
gemini-cli · Google
|
0.32 | 0.28 | 0.57 | 166k | $0.005 |
0.33 (0.28–0.39)
|
How are agents evaluated?
Every FlutterBench trial runs in an isolated Docker container testing real Flutter features. Scoring measures 60% Outcome (passing tests & builds), 30% Quality (idiomatic patterns & analyzer diagnostics), and 10% Developer Experience (tool accuracy & minimal friction). Diagnostic metrics like token efficiency are captured separately.