Back to Blog
Tools & Resources6 min readSeptember 29, 2026

Pac-Bench Asked 30 AI Coding Setups to Build Pac-Man From One Line

Pac-Bench gave 30 model and harness setups one sentence and scored the Pac-Man games they built. What the results tell you about choosing a model, and how to run your own bake-off.

Sarah Chen

Sarah Chen

Content at NeedBase

A Show HN post on today's Hacker News front page gave 30 model and harness setups one sentence, "Create a Pac-Man game in a single html page", and scored what came back. The project is called Pac-Bench. It puts Claude, GPT, Gemini, Grok, Qwen, GLM, DeepSeek, Kimi and MiniMax runs side by side, with cost, wall time and tokens logged for most of them, and every result is a playable game you can open yourself. The spread in results is huge, and so are the caveats.

What was tested, and how

Each entry is one model inside one harness: Claude Code, Claude Code routed through OpenRouter, OpenAI's Codex, Google's Antigravity, Grok Build or Cursor Cloud. The model got the prompt and no follow-ups. The submitter says they first tried it about a year ago, found that no model handled it well, and have kept re-running it with each new release.

According to the project README, scores out of 100 come from an automated re-test run on 28 September, combining a 90-second play test with an audit of the source code and maze. The rubric weights Controls 20, Ghosts 25, Pac-Man getting stuck 20, Maze 20 and Sound 15. The re-test was run by Claude Opus 5.5, which is also the model that scored highest.

The results

Claude Opus 5.5 in Claude Code topped the table with 99, at a logged cost of about $1.99 and roughly nine minutes of wall time. Claude Fable 5.1 scored 96 and Claude Sonnet 5.5 scored 95. Grok 4.7 in Grok Build scored 94. GPT-5.6 Sol in Codex scored 90 at about $0.72 and just under five minutes, one of the best score-per-dollar results on the board.

Lower down, cost stopped predicting quality. Claude Fable 5 through OpenRouter scored 95 but cost about $17.28 and took over half an hour. GPT-6 Luna cost about $0.01 and scored 45. Qwen 3.8 Max cost about $4.15 and scored 74. MiniMax M3 came last on 2, with the grader's note reading "syntax error, blank canvas". Some newer models also scored below older ones from the same family: Gemini 3.8 Flash got 17 against 38 for Gemini 3.7 Flash.

Not every cost is a billed amount. The README says Codex costs are rate-card maths on captured usage, Cursor costs come from its dashboard, and Gemini costs are estimates at public rates.

What commenters pushed back on

Commenters on Hacker News raised several fair objections. One argued the prompt is so vague that the test mostly measures how a model reads an under-specified request. Another replied that this is exactly what makes it useful, because it shows how well a model fills in missing context. Several pointed out that Pac-Man clones are all over the web, so a good result may partly reflect memorised code rather than reasoning. Others noted that the details that make Pac-Man feel right, especially ghost behaviour, are easy to get subtly wrong. One predicted the benchmark would saturate soon, and another asked for the same models to be tested across different harnesses.

Two more caveats are worth adding. First, the grader is also the winner. That does not make the scores wrong, but a model marking its own homework deserves some suspicion. Second, the harness matters as much as the model. The same model in Claude Code, Codex or Cursor gets different tools, system prompts and retry loops, so these are really scores for a model and harness pair.

What one-shot benchmarks do and don't tell you

Pac-Bench is good at one thing: showing whether a model and harness can plan, write and self-check a complete small program without help. If you run an AI app builder where users type one sentence and expect something that works, that is close to your real task, and the gap between 99 and 2 is meaningful.

It does not tell you how a model handles your codebase, your framework, a long multi-turn session, a vague bug report or a migration across 40 files. It is a single well-known task graded by a single judge, and as far as the published data shows, each card is one run. One good run tells you what a model can do, not what it usually does.

How to run your own bake-off

You do not need a leaderboard. An afternoon and a spreadsheet are enough.

1. Pick five to ten real tasks. Use tickets you have actually shipped: a Stripe webhook handler, a settings page in your component library, a flaky test, a SQL migration. For an app builder, use real user prompts with personal details removed.

2. Fix the prompt and the harness. Write each prompt once and don't touch it during the test. Run every candidate in the tool you will actually use, whether that is Claude Code, Codex, Cursor or your own API wrapper through OpenRouter.

3. Run each task at least three times. Model output varies between runs. Record the worst run as well as the best, because your users will sometimes get the worst one.

4. Score blind. Remove model names from the outputs before you or a teammate mark them against a checklist written in advance: does it build, do the tests pass, did it touch files it shouldn't have, would you merge it. If you use an LLM as the judge, don't let it judge its own output.

5. Track cost and latency per task. Log tokens, cost and wall time the way Pac-Bench does. A setup that scores 90 in five minutes for under a dollar may suit you better than one that scores 95 in over half an hour for $17.

6. Re-run when models change. Pac-Bench's own table shows newer versions sometimes scoring below older ones. Keep your task set, and re-run it whenever you are considering a switch.

The bottom line

Pac-Bench is a useful, honest snapshot. On one small, well-known build, the top Claude, Grok and GPT setups now produce playable games, and several others still fail badly. Treat it as a quick filter, not a verdict. Before you commit your product or your budget to a model, run your own tasks, several times, graded blind, with cost and time logged next to each score.

Found this useful?

Share it with a founder who needs it.

Ready to launch your product?

Join thousands of makers who launched on NeedBase.

Submit Your Product โ†’

Compare the Tools in This Post