Brood War Bench

Which model wins at Brood War?

Key takeaways

  • None of the models played beyond a beginner level.
  • Codex Astra is the clear leader beating all other models consistently.
  • Grok models are not smart enough to play Brood War yet.
  • Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking.

Powered By Freestyle

Leaderboard

RankSystemWinsLossesAPMCost / gameWin rate
🥇
Codex Astra / xhigh
18012.6$10.54100.0%
🥈
Codex Astra / medium
16217.2$15.1188.9%
🥉
Claude Fable
15312.6$12.2483.3%
4
Codex Astra / low
14425.7$21.0777.8%
5
Codex 5.6 Sol / medium
13510.1$5.1272.2%
6
Codex 5.6 Sol / low
12618.1$9.2366.7%
7
Claude Opus 5
12610.5$20.7866.7%
8
Codex 5.6 Sol / xhigh
1178.0$3.2361.1%
9
Codex 5.6 Luna / low
9923.8$0.4250.0%
10
Codex 5.6 Terra / xhigh
9915.8$2.1050.0%

Brood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn't done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.

What I observed

01

Codex found cheese before it found macro

Codex's strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.

The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.

I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn't communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build.

This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing. In games where I helped direct Codex, it was much better at planning those moments and getting its subagents to work together.

The persistence was real. In G009, after losing its army and main base, Codex 5.6 Terra / medium lifted its last Command Center and moved it toward the opposite corner. It survived for another six minutes.

Six Probes cross the map
A Probe first, then Zealots in drips
The last Command Center runs

02

Grok spent the game between actions

Grok 4.6 frequently produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit.

The actions it did take rarely developed into a working control loop. In G003, Grok / xhigh made three Marines and never reached the enemy base. In G002, Grok / medium made two Zealots and also never crossed the map. These looked less like bad strategies than failures to keep observing and acting.

Forty-three minutes, no army

03

Fable earnestly tried to play the game

I found myself rooting for Claude Fable in more than a few games. Fable usually tried to build an economy and climb the tech tree instead of stopping at the first unit available. It seemed more interested in actually playing the game than any of the other models.

In G007 it reached a Lair, Spire, and Mutalisks and won. In G027 it added a Robotics Facility, Citadel of Adun, Observatory, and Templar Archives before winning. Ambition did not guarantee execution: in G036 Fable reached a Factory and Academy but Opus 5 overran it.

Fable gets Mutalisks
Fable keeps climbing
The build does not become an army

No agent here played beyond beginner level

Even Astra and Fable were unable to build complex army's, defend simple attacks or play concrete strategies. A beginner playing photon rush would win every single one of these games.

That said, watching the agents play made me more excited than I have been in a while. This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do. I look forward to watching them get there.

Choose up to five
5:00 / 15:00

Technology investment

Completed research + upgrade levels

00.511.520:005:0010:0015:00
Codex Astra0.3Claude Fable0.2Grok 4.60

Workers

Completed workers alive

081624320:005:0010:0015:00
Codex Astra14.7Claude Fable16.7Grok 4.67.4

Army size

Completed army and support units

0204060800:005:0010:0015:00
Codex Astra8.3Claude Fable7.6Grok 4.62

Structures

Completed buildings, including add-ons

071421280:005:0010:0015:00
Codex Astra6.2Claude Fable7.4Grok 4.64.5

Minerals in the bank

Unspent minerals, not income

05001,0001,5002,0000:005:0010:0015:00
Codex Astra240.6Claude Fable248.6Grok 4.6536.6

Gas in the bank

Unspent gas, not income

05001,0001,5002,0000:005:0010:0015:00
Codex Astra151Claude Fable326.9Grok 4.6244.5

Supply used

Includes production in progress

03060901200:005:0010:0015:00
Codex Astra26.5Claude Fable26.6Grok 4.611.3

When games ended

Share of games ending per 5-minute window

Codex Astra: median 8:10. 0:00 to before 5:00: 3.9%; 5:00 to before 10:00: 54.9%; 10:00 to before 15:00: 31.4%; 15:00 to before 20:00: 5.9%; 25:00 to before 30:00: 2%; 30:00 to before 35:00: 2%. Claude Fable: median 10:37. 5:00 to before 10:00: 38.9%; 10:00 to before 15:00: 38.9%; 15:00 to before 20:00: 22.2%. Grok 4.6: median 9:15. 0:00 to before 5:00: 2%; 5:00 to before 10:00: 51%; 10:00 to before 15:00: 23.5%; 15:00 to before 20:00: 13.7%; 20:00 to before 25:00: 2%; 40:00 to before 45:00: 7.8%.Codex AstraMedian 8:1060%Codex Astra: 3.9% (2 games) ended from 0:00 to before 5:00Codex Astra: 54.9% (28 games) ended from 5:00 to before 10:00Codex Astra: 31.4% (16 games) ended from 10:00 to before 15:00Codex Astra: 5.9% (3 games) ended from 15:00 to before 20:00Codex Astra: 0% (0 games) ended from 20:00 to before 25:00Codex Astra: 2% (1 games) ended from 25:00 to before 30:00Codex Astra: 2% (1 games) ended from 30:00 to before 35:00Codex Astra: 0% (0 games) ended from 35:00 to before 40:00Codex Astra: 0% (0 games) ended from 40:00 to before 45:00Claude FableMedian 10:3760%Claude Fable: 0% (0 games) ended from 0:00 to before 5:00Claude Fable: 38.9% (7 games) ended from 5:00 to before 10:00Claude Fable: 38.9% (7 games) ended from 10:00 to before 15:00Claude Fable: 22.2% (4 games) ended from 15:00 to before 20:00Claude Fable: 0% (0 games) ended from 20:00 to before 25:00Claude Fable: 0% (0 games) ended from 25:00 to before 30:00Claude Fable: 0% (0 games) ended from 30:00 to before 35:00Claude Fable: 0% (0 games) ended from 35:00 to before 40:00Claude Fable: 0% (0 games) ended from 40:00 to before 45:00Grok 4.6Median 9:1560%Grok 4.6: 2% (1 games) ended from 0:00 to before 5:00Grok 4.6: 51% (26 games) ended from 5:00 to before 10:00Grok 4.6: 23.5% (12 games) ended from 10:00 to before 15:00Grok 4.6: 13.7% (7 games) ended from 15:00 to before 20:00Grok 4.6: 2% (1 games) ended from 20:00 to before 25:00Grok 4.6: 0% (0 games) ended from 25:00 to before 30:00Grok 4.6: 0% (0 games) ended from 30:00 to before 35:00Grok 4.6: 0% (0 games) ended from 35:00 to before 40:00Grok 4.6: 7.8% (4 games) ended from 40:00 to before 45:000:0015:0030:0045:00

Time-series charts show means of recorded player-runs at each game time. Finished games drop out; missing samples are not filled. Models pool their effort settings. Units and buildings count only once completed; army excludes workers, Overlords, eggs, larvae, and ammunition.

Win rate vs. cost

Average cost per game, using the same prices as the leaderboard. Codex and Sonnet costs are token-based estimates.

CodexClaudeGrok
Win rate0%25%50%75%100%$0.1$0.5$1$5$10$20Cost per game (USD, log scale)

How the benchmark ran

We built a round-robin matrix of model and effort configurations and had every configuration play every other. The harness ran those matchups in parallel across Freestyle VMs, saving game-engine data and both agents' harness logs for each match.

Head-to-head matrix

Read across a row. W is a win, L is a loss, and T is a match that reached the benchmark time limit.

Open the full 19 × 19 matrix
W win L loss T time limit
System12345678910111213141516171819
1Codex Astra / xhigh-WWWWWWWWWWWWWWWWWW
2Codex Astra / mediumL-WWWWWWWWWWLWWWWWW
3Codex Astra / lowLL-WWWWWWWWWLLWWWWW
4Codex 5.6 Sol / xhighLLL-LWWWLWWWLLWWWWW
5Codex 5.6 Sol / mediumLLLW-WWWLWWWLWWWWWW
6Codex 5.6 Sol / lowLLLLL-WWWWWWWWLWWWW
7Codex 5.6 Luna / xhighLLLLLL-WLWLWLLLWWWW
8Codex 5.6 Luna / mediumLLLLLLL-LLLLLWWWWWW
9Codex 5.6 Luna / lowLLLWWLWW-WLLLLLWWWW
10Codex 5.6 Terra / xhighLLLLLLLWL-WWLWWWWWW
11Codex 5.6 Terra / mediumLLLLLLWWWL-LLLWWWWW
12Codex 5.6 Terra / lowLLLLLLLWWLW-LLWWWWW
13Claude FableLWWWWLWWWWWW-LWWWWW
14Claude Opus 5LLWWLLWLWLWWW-WWWWW
15Claude SonnetLLLLLWWLWLLLLL-WWWW
16Claude HaikuLLLLLLLLLLLLLLL-LTT
17Grok 4.6 / xhighLLLLLLLLLLLLLLLW-WT
18Grok 4.6 / mediumLLLLLLLLLLLLLLLTL-W
19Grok 4.6 / lowLLLLLLLLLLLLLLLTTL-

Play your own match

Bring your agent and play Brood War with friends.

Play Brood War