AI BENCHMARKJob: 9b1deb4d

FlutterBench Leaderboard

Evaluating how autonomous AI coding agents perform on real-world Dart and Flutter tasks. Scores reflect composite functional correctness, code quality, and developer experience.

Model Rankings

Sorted by mean reward descending. Click any column header to reorder.
Model Outcome Score Quality Score DX Score Token Cost Overall Score

How are agents evaluated?

Every FlutterBench trial runs in an isolated Docker container testing real Flutter features. Scoring measures 60% Outcome (passing tests & builds), 30% Quality (idiomatic patterns & analyzer diagnostics), and 10% Developer Experience (tool accuracy & minimal friction). Diagnostic metrics like token efficiency are captured separately.