AI / FlutterBench / Tasks / dart-build-cli-app / dart-build-cli-app__t21-gpt-5
Trial Evaluation

dart-build-cli-app__t21-gpt-5

PASS
98 % Composite Reward
Model GPT 5 OpenAI
Agent Harness codex-agent Dart MCP + Skills
Tokens 99.3k 94.9k in / 4.4k out
Cost $0.2813 USD estimated
Execution Phase Durations
Environment Setup10.0s
Agent Setup10.0s
Agent Execution2.2m
Verifier32.0s

Scoring Rubric Breakdown

Composite Reward = 0.60×Outcome + 0.30×Quality + 0.10×DX

Detailed criteria evaluation produced by automated test graders and LLM rubrics:

OUTCOMEWeight 60%
99%
dart_compile:exe wt: 30% 1.00

Executable compiled with dart compile exe.

cli_flag_tests wt: 70% 0.99

Flag parsing tests for --help, --version, --verbose, and --format=json.

QUALITYWeight 30%
98%
StaticAnalysisGrader wt: 50% 0.98

Zero static analysis issues found.

structural_validation wt: 30% 1.00

Clean separation of concerns and architectural layers.

idiomatic_review wt: 20% 0.98

Dart 3 idioms (pattern matching, records, sealed classes) adherence.

DXWeight 10%
96%
tool_usage wt: 40% 0.96

Dart MCP tools (analyze_files, dtd, hot_reload) utilized efficiently.

trajectory wt: 30% 0.96

Linear problem-solving without excessive tool backtracking.

error_recovery wt: 30% 1.00

Fast turnaround from initial compile warnings to clean build.

Agent Execution Trajectory

Sequential step timeline captured during autonomous agent execution:

  1. 1
    step0ms
  2. 2
    step0ms
  3. 3
    step0ms
  4. 4
    step0ms
  5. 5
    step0ms
  6. 6
    step0ms

Generated Code Artifacts

Files modified or generated by the agent during this trial:

📄 artifacts/workspace/bin/main.dart Generated
// Generated implementation for dart-build-cli-app
// Model: gpt-5
// Verification score: 0.98 (pass)
📄 artifacts/workspace/lib/cli_runner.dart Generated
// Generated implementation for dart-build-cli-app
// Model: gpt-5
// Verification score: 0.98 (pass)
📄 artifacts/workspace/lib/command_options.dart Generated
// Generated implementation for dart-build-cli-app
// Model: gpt-5
// Verification score: 0.98 (pass)
📄 artifacts/workspace/pubspec.yaml Generated
// Generated implementation for dart-build-cli-app
// Model: gpt-5
// Verification score: 0.98 (pass)

Execution & Verifier Logs

Verifier Standard Output (test-stdout.txt)
00:00 +0: loading tests/graders.dart
00:01 +1: Environment setup verification passed.
00:02 +2: Task codebase compilation check.
00:03 +3: All Build Command-Line CLI App unit tests passed.
00:04 +4: Static analysis: 0 warnings, 0 errors.
00:05 +5: Rubric grader verification completed.
Overall result: PASS (Score: 0.98)