LLM Agent Evaluation Dashboard

Compare agent/model runs with deterministic scoring, token metrics, provider context, and links back to the full generated reports.

Reference Implementations

Known-good examples for visual comparison.

Total Runs2
Average Score69.0
Local Runs1
Cloud Runs1

2D Elevator Prompt Scores Scores prioritize working browser simulations. Runtime verification can contribute up to 40 points for zero startup errors, a nonblank canvas, animation frames, scene complexity, and visible changes over time. Prompt-specific implementation signals, required files, completion metadata, and efficiency make up the remaining points. Hard caps prevent browser-dead, blank, missing-file, or unverified runs from ranking as excellent.

Self-contained canvas elevator dispatch simulations, scored by browser viability, one-based floor geometry, bounded smoke behavior, dispatch signals, and efficiency.

#1
88
Codex CLI
gpt-5_6-sol
Elevator Prompt 2D / OpenAI
#2
50
Pi Coding Agent
qwen3_8-27b_q4_k_s
Elevator Prompt 2D / Local (LM Studio)
Runtime not verified

Comparison Matrix

C=completion, F=files, I=implementation signals, R=runtime verification, E=efficiency.

Score Agent / Model Provider Tokens Cost Time Breakdown Links
88
Excellent
Codex CLI gpt-5_6-sol
OpenAI 527,981 - -
C 10 F 8 I 35 R 33 E 2
summary report present, artifact files present, machine-readable result metrics present
50
Weak
Pi Coding Agent qwen3_8-27b_q4_k_s
Local (LM Studio) 623,060 - -
C 10 F 8 I 24 R 8 E 0
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
Total Runs22
Average Score88.5
Local Runs14
Cloud Runs6

3D Elevator Prompt Scores Scores prioritize working browser simulations. Runtime verification can contribute up to 40 points for zero startup errors, a nonblank canvas, animation frames, scene complexity, and visible changes over time. Prompt-specific implementation signals, required files, completion metadata, and efficiency make up the remaining points. Hard caps prevent browser-dead, blank, missing-file, or unverified runs from ranking as excellent.

Local-model focused 3D elevator simulations, scored primarily by runtime viability, then file completeness, Three.js behavior cues, and efficiency.

#1
98
Antigravity CLI
gemini-3_1-pro-preview
Elevator Prompt V2 / Google
#2
98
OpenCode CLI
muse-spark-1_2
Elevator Prompt V3 / unknown
#3
98
OpenCode CLI
nemotron-3_5-lightning
Elevator Prompt V3 / Local-B70
#4
98
Pi Wiggum
agents-a1
Elevator Prompt V3 / Local-B70
#5
96
OpenCode CLI
agents-a1
Elevator Prompt V3 / unknown
#6
96
Pi Coding Agent
coder-next
Elevator Prompt V3 / Local (LM Studio)

Comparison Matrix

C=completion, F=files, I=implementation signals, R=runtime verification, E=efficiency.

Score Agent / Model Provider Tokens Cost Time Breakdown Links
98
Excellent
Antigravity CLI gemini-3_1-pro-preview
Google 184,098 - 3.2m
C 10 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
98
Excellent
OpenCode CLI muse-spark-1_2
unknown 46,219 $0.10 -
C 8 F 15 I 30 R 40 E 5
summary report present, artifact files present, machine-readable result metrics present
98
Excellent
OpenCode CLI nemotron-3_5-lightning
Local-B70 54,848 - -
C 10 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
98
Excellent
Pi Wiggum agents-a1
Local-B70 39,312 - 3.7m
C 10 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
OpenCode CLI agents-a1
unknown 50,921 - -
C 8 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Pi Coding Agent coder-next
Local (LM Studio) 23,464 - -
C 8 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Qoder CLI Qwen3_8-Max-Preview
Qoder ≈891,371 N/A 6.4m
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Claude Code Opus_4_8
Anthropic 2,409,128 $3.99 18.6m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
94
Excellent
Pi Wiggum qwen3_6-35b-a3b
Local-B70 31,092 - 5.1m
C 10 F 15 I 24 R 40 E 5
summary report present, artifact files present, machine-readable result metrics present
93
Excellent
Mistral Vibe mistral-medium-3_5
Mistral AI 1,478,410 $2.45 -
C 8 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
93
Excellent
OpenCode CLI google_gemma-4-26b-a4b-qat
Local (LM Studio) 575,839 - -
C 8 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
93
Excellent
Pi Coding Agent deepreinforce-ai_ornith-1_0-35b
Local (LM Studio) 729,661 - -
C 8 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
93
Excellent
Pi Wiggum muse-glimmer-30b
Local-B70 44,559 - 30.2m
C 10 F 15 I 24 R 40 E 4
summary report present, artifact files present, machine-readable result metrics present
92
Excellent
Pi Wiggum agents-a1
Local-B70 61,286 - 7.0m
C 10 F 15 I 24 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
90
Excellent
Pi Wiggum gemma-4-e4b
Local-B70 276,208 - 42.5m
C 10 F 15 I 24 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
89
Excellent
Pi Wiggum qwen3_8-27b
Local (LM Studio) 53,792 - 39.6m
C 10 F 15 I 20 R 40 E 4
summary report present, artifact files present, machine-readable result metrics present
89
Excellent
Qoder CLI Qwen3_8-Max
Qoder ≈671,641 N/A 3.9m
C 10 F 15 I 23 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
70
Strong
Charmbracelet Crush unsloth_gemma-4-26b-a4b-it
Local (LM Studio) 20,210 - -
C 10 F 15 I 30 R 12 E 3
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
70
Strong
OpenCode CLI upstage_solar-pro4
Openrouter 118,553 $0.00 -
C 10 F 15 I 30 R 35 E 4
summary report present, artifact files present, machine-readable result metrics present
No runtime motion detectedCapped at 70: no runtime motion detected
70
Strong
OpenCode CLI zai-org_glm-4_7-flash
Local (LM Studio) 118,994 - -
C 8 F 15 I 30 R 35 E 4
summary report present, artifact files present, machine-readable result metrics present
No runtime motion detectedCapped at 70: no runtime motion detected
68
Partial
Pi Coding Agent qwen3_6-35b-a3b
Local-B70 72,002 - -
C 8 F 15 I 30 R 12 E 3
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
67
Partial
Pi Coding Agent nemotron-3-nano-omni
Local-B70 23,229 - -
C 8 F 15 I 27 R 12 E 5
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
Total Runs37
Average Score88.9
Local Runs7
Cloud Runs26

Office Prompt Scores Scores prioritize working browser simulations. Runtime verification can contribute up to 40 points for zero startup errors, a nonblank canvas, animation frames, scene complexity, and visible changes over time. Prompt-specific implementation signals, required files, completion metadata, and efficiency make up the remaining points. Hard caps prevent browser-dead, blank, missing-file, or unverified runs from ranking as excellent.

Frontier/cloud office simulations, scored primarily by runtime viability, then office-world behavior, elevator scheduling cues, and efficiency.

#1
98
Antigravity CLI
gemini-3_6-flash-high
Office Prompt V3 / Google
#2
98
Antigravity CLI
gemini-3_7-flash
Office Prompt V3 / Google
#3
97
Codex CLI
gpt-6-astra
Office Prompt V3 / OpenAI
#4
97
DeepSeek Harness
deepseek_deepseek-v4-pro-0813
Office Prompt V3 / openrouter
#5
97
Pi Wiggum
qwen3_8-27b
Office Prompt Wiggum / Local (LM Studio)
#6
96
OpenCode CLI
glm-5_3
Office Prompt V3 / Openrouter

Comparison Matrix

C=completion, F=files, I=implementation signals, R=runtime verification, E=efficiency.

Score Agent / Model Provider Tokens Cost Time Breakdown Links
98
Excellent
Antigravity CLI gemini-3_6-flash-high
Google 48,809 - 1.9m
C 10 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
98
Excellent
Antigravity CLI gemini-3_7-flash
Google 61,735 - 3.4m
C 10 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
97
Excellent
Codex CLI gpt-6-astra
OpenAI 1,822,028 - -
C 10 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
97
Excellent
DeepSeek Harness deepseek_deepseek-v4-pro-0813
openrouter ≈0 N/A -
C 10 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
97
Excellent
Pi Wiggum qwen3_8-27b
Local (LM Studio) 135,796 - 2.2h
C 10 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
OpenCode CLI glm-5_3
Openrouter 371,245 $3.66 -
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
OpenCode CLI Muse_Spark_1_3
Openrouter 279,690 $1.24 -
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
OpenCode CLI ox-alpha
Openrouter 409,187 $4.19 -
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Pi Coding Agent qwen3_8-27b
Local (LM Studio) 370,638 - -
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Pi Coding Agent qwen3_8-27b-think
Local (LM Studio) 242,711 - -
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Pi Wiggum Agents-A1-MTPLX-Q4
Local (LM Studio) 401,670 - 1.0h
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
96
Excellent
Pi Wiggum qwen3_8-27b-think
Local (LM Studio) 271,228 - 1.8h
C 10 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Antigravity CLI Gemini_3_8_Flash
Google 67,555 - 5.1m
C 7 F 15 I 30 R 40 E 3
summary report present, artifact files present, machine-readable result metrics present
Result flagged error
95
Excellent
Claude Code Fable_5
Anthropic 4,268,781 $18.20 53.1m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Claude Code Opus_5
Anthropic 11,452,403 $13.93 47.3m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Claude Code Sonnet
Anthropic 15,683,159 $11.32 48.4m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Codex CLI gpt-5_6-luna
OpenAI 1,487,407 - -
C 8 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Codex CLI gpt-5_6-sol
OpenAI 2,184,142 - -
C 8 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Codex CLI gpt-5_6-terra
OpenAI 2,965,162 - -
C 8 F 15 I 30 R 40 E 2
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
OpenCode CLI DeepSeek_V4_1_Flash
Openrouter 640,551 $0.19 -
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
OpenCode CLI GLM_5_3_Flash
Openrouter 677,168 $0.26 -
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
OpenCode CLI qwen3_8-27b-think
Local (LM Studio) 969,064 - -
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI Cantus
Qoder ≈4,644,975 N/A 19.4m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI DeepSeek-V4-Flash
Qoder ≈4,576,933 N/A 9.5m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI GLM-5_2
Qoder ≈14,532,447 N/A 22.2m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI MiniMax-M3
Qoder ≈6,168,387 N/A 11.3m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI Qwen3_8-Max
Qoder ≈20,936,897 N/A 1.0h
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
95
Excellent
Qoder CLI Qwen3_8-Max-Preview
Qoder ≈13,626,252 N/A 40.1m
C 10 F 15 I 30 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
94
Excellent
OpenCode CLI moonshotai_kimi-k3
unknown 225,317 $4.43 -
C 8 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
90
Excellent
Pi Wiggum agents-a1
Local (LM Studio) 4,339,565 - 1.8h
C 10 F 15 I 25 R 40 E 0
summary report present, artifact files present, machine-readable result metrics present
70
Strong
Qoder CLI DeepSeek-V4-Pro
Qoder ≈4,478,998 N/A 22.0m
C 10 F 15 I 30 R 35 E 0
summary report present, artifact files present, machine-readable result metrics present
No runtime motion detectedCapped at 70: no runtime motion detected
67
Partial
OpenCode CLI x-ai_grok-4_6
Openrouter 501,237 $1.95 -
C 10 F 15 I 30 R 12 E 0
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
66
Partial
OpenCode CLI gemma-4-31b-qat
unknown 336,888 - -
C 8 F 15 I 30 R 12 E 1
summary report present, artifact files present, machine-readable result metrics present
Runtime not verified
64
Partial
OpenCode CLI muse-spark-1_2
unknown 526,497 $1.00 -
C 7 F 15 I 30 R 12 E 0
summary report present, artifact files present, machine-readable result metrics present
Runtime not verifiedResult flagged error
55
Partial
Codex CLI gpt-5_5
OpenAI 448,065 - -
C 8 F 15 I 30 R 25 E 3
summary report present, artifact files present, machine-readable result metrics present
Static JS reference check failedCapped at 55: static JS reference check failed
Runtime errorsevals/codex_gpt-5_5_office_prompt_v2/world.js:159:27 'sitTargets' is not definedevals/codex_gpt-5_5_office_prompt_v2/world.js:159:44 'desks' is not definedevals/codex_gpt-5_5_office_prompt_v2/world.js:212:9 'sitTargets' is not definedevals/codex_gpt-5_5_office_prompt_v2/world.js:213:9 'sitTargets' is not definedevals/codex_gpt-5_5_office_prompt_v2/world.js:214:9 'sitTargets' is not defined
55
Partial
OpenCode CLI Kimi_K2_7_Code
unknown 180,651 $1.51 -
C 8 F 15 I 30 R 25 E 1
summary report present, artifact files present, machine-readable result metrics present
Static JS reference check failedCapped at 55: static JS reference check failed
Runtime errorsevals/opencode_Kimi_K2_7_Code_office_prompt_v2/elevator_logic.js:213:73 'e' is not definedevals/opencode_Kimi_K2_7_Code_office_prompt_v2/sim.js:925:38 'e' is not definedevals/opencode_Kimi_K2_7_Code_office_prompt_v2/sim.js:932:34 'e' is not defined
50
Weak
OpenCode CLI Hy4_preview
Openrouter 280,436 $0.91 -
C 7 F 15 I 30 R 40 E 1
summary report present, artifact files present, machine-readable result metrics present
Safety stop: time limitCapped at 50: time limit reachedResult flagged error
Antigravity CLI 4 evaluations
Google 98
gemini-3_1-pro-preview
Elevator Prompt V2
Google 98
gemini-3_6-flash-high
Office Prompt V3
Google 98
gemini-3_7-flash
Office Prompt V3
Google 95
Gemini_3_8_Flash
Office Prompt V3
Charmbracelet Crush 1 evaluation
Local (LM Studio) 70
unsloth_gemma-4-26b-a4b-it
Elevator Prompt V3
Claude Code 4 evaluations
Anthropic 95
Fable_5
Office Prompt V3
Anthropic 95
Opus_4_8
Elevator Prompt V2
Anthropic 95
Opus_5
Office Prompt V3
Anthropic 95
Sonnet
Office Prompt V3
Codex CLI 6 evaluations
OpenAI 55
gpt-5_5
Office Prompt V2
OpenAI 95
gpt-5_6-luna
Office Prompt V3
OpenAI 88
gpt-5_6-sol
Elevator Prompt 2D
OpenAI 95
gpt-5_6-sol
Office Prompt V3
OpenAI 95
gpt-5_6-terra
Office Prompt V3
OpenAI 97
gpt-6-astra
Office Prompt V3
DeepSeek Harness 1 evaluation
openrouter 97
deepseek_deepseek-v4-pro-0813
Office Prompt V3
Mistral Vibe 1 evaluation
Mistral AI 93
mistral-medium-3_5
Elevator Prompt V2
OpenCode CLI 18 evaluations
Unknown 96
agents-a1
Elevator Prompt V3
Openrouter 95
DeepSeek_V4_1_Flash
Office Prompt V3
Unknown 66
gemma-4-31b-qat
Office Prompt V3
Openrouter 96
glm-5_3
Office Prompt V3
Openrouter 95
GLM_5_3_Flash
Office Prompt V3
Local (LM Studio) 93
google_gemma-4-26b-a4b-qat
Elevator Prompt V3
Openrouter 50
Hy4_preview
Office Prompt V3
Unknown 55
Kimi_K2_7_Code
Office Prompt V2
Unknown 94
moonshotai_kimi-k3
Office Prompt V3
Unknown 98
muse-spark-1_2
Elevator Prompt V3
Unknown 64
muse-spark-1_2
Office Prompt V3
Openrouter 96
Muse_Spark_1_3
Office Prompt V3
Local-B70 98
nemotron-3_5-lightning
Elevator Prompt V3
Openrouter 96
ox-alpha
Office Prompt V3
Local (LM Studio) 95
qwen3_8-27b-think
Office Prompt V3
Openrouter 70
upstage_solar-pro4
Elevator Prompt V3
Openrouter 67
x-ai_grok-4_6
Office Prompt V3
Local (LM Studio) 70
zai-org_glm-4_7-flash
Elevator Prompt V3
Pi Coding Agent 7 evaluations
Local (LM Studio) 96
coder-next
Elevator Prompt V3
Local (LM Studio) 93
deepreinforce-ai_ornith-1_0-35b
Elevator Prompt V3
Local-B70 67
nemotron-3-nano-omni
Elevator Prompt V3
Local-B70 68
qwen3_6-35b-a3b
Elevator Prompt V3
Local (LM Studio) 96
qwen3_8-27b
Office Prompt V3
Local (LM Studio) 96
qwen3_8-27b-think
Office Prompt V3
Local (LM Studio) 50
qwen3_8-27b_q4_k_s
Elevator Prompt 2D
Pi Wiggum 10 evaluations
Local-B70 98
agents-a1
Elevator Prompt V3
Local-B70 92
agents-a1
Elevator Prompt Wiggum
Local (LM Studio) 90
agents-a1
Office Prompt Wiggum
Local (LM Studio) 96
Agents-A1-MTPLX-Q4
Office Prompt Wiggum
Local-B70 90
gemma-4-e4b
Elevator Prompt Wiggum
Local-B70 93
muse-glimmer-30b
Elevator Prompt Wiggum
Local-B70 94
qwen3_6-35b-a3b
Elevator Prompt Wiggum
Local (LM Studio) 89
qwen3_8-27b
Elevator Prompt Wiggum
Local (LM Studio) 97
qwen3_8-27b
Office Prompt Wiggum
Local (LM Studio) 96
qwen3_8-27b-think
Office Prompt Wiggum
Qoder CLI 9 evaluations
Qoder 95
Cantus
Office Prompt V3
Qoder 95
DeepSeek-V4-Flash
Office Prompt V3
Qoder 70
DeepSeek-V4-Pro
Office Prompt V3
Qoder 95
GLM-5_2
Office Prompt V3
Qoder 95
MiniMax-M3
Office Prompt V3
Qoder 89
Qwen3_8-Max
Elevator Prompt V3
Qoder 95
Qwen3_8-Max
Office Prompt V3
Qoder 96
Qwen3_8-Max-Preview
Elevator Prompt V3
Qoder 95
Qwen3_8-Max-Preview
Office Prompt V3