INDEX 47 ▼3 todaySPLIT OF THE DAY OpenAI annualized revenue reported at $50bn not $70bn71 STORIES · 212 REACTIONSANTI-AI 75% · MIDDLE GROUND 17% · PRO-AI 8%LATEST Anthropic launches free AI security scans for open-source software
1 source0 reactions

Developers benchmark AI decision models playing Pac-Man

60 BoomStory toneCommunity benchmark, framed as capability comparison
1 source · Hacker News front page (AI)
  • Boom: Six AI decision models each played 100 Pac-Man games; results published as a leaderboard
  • Boom: OpenAI launched a decisions endpoint, cited as motivation for the benchmark
  • Boom: Cloudflare launched Clef, a decision model included in the Pac-Man test
  • Neutral: Benchmark repo was open-sourced, allowing public replication of results
The story in full

A team posted a project to Hacker News on October 8, 2026, testing several AI decision models by having them play Pac-Man against bot ghosts. The models tested include Jev 1.13, Kev, Clef, Clef Flash, GPT-6 Luna, and Laya. Each model played 100 games, and the results were published as a leaderboard with the repository open-sourced.

The project was prompted by recent launches in the AI decision model space, including OpenAI's decisions endpoint and Cloudflare's Clef model. The creators describe Pac-Man as a benchmark for simple and fast decision making, citing the low latency of these models as enabling real-time play.

Analysis

314 words

On October 8, 2026, a team posted a project called Jevman to Hacker News, pitting six AI decision models against bot ghosts in Pac-Man. The models tested were Jev 1.13, Kev, Clef, Clef Flash, GPT-6 Luna, and Laya, each playing 100 games. Results were published as a public leaderboard, and the repository was open-sourced so that anyone can replicate the runs. The project was prompted by two recent launches: OpenAI's decisions endpoint and Cloudflare's Clef model, both arriving within weeks of the benchmark going live.

The context here is a rapidly filling field of AI decision models, meaning systems designed to make discrete, low-latency choices rather than generate long-form text. The creators framed Pac-Man as a practical stress test for that specific capability, pointing to the real-time demands of the game as a natural fit for models built around speed. What is genuinely in dispute is whether a classic arcade game constitutes a meaningful benchmark for decision-making quality, or whether it captures only a narrow slice of what these models are actually being built to do in production environments. Open-sourcing the repo at least makes the methodology auditable, which shifts the argument from trust to interpretation.

No reactions from the Pro-AI, Anti-AI, or Middle Ground camps have been published yet. Pro-AI voices would typically welcome a concrete, reproducible leaderboard as evidence that the decision model space is maturing and that competition between providers is producing measurable progress. Anti-AI voices would likely question whether gaming benchmarks flatter models in controlled conditions while obscuring risks in higher-stakes real-world deployments. A middle-ground position would probably praise the open-source transparency while urging caution about reading too much into a single, narrow task.

The natural next development to watch is whether other researchers adopt or challenge the Pac-Man benchmark, and whether the model providers themselves, particularly OpenAI and Cloudflare, respond with their own evaluations or contest the leaderboard results.

Where do you stand?

Add your take

0 reader votes

Sign in with Google to pick a side and post. Your vote moves the story's Doom / Boom score.

Sources

1 article from 1 outlet
  1. Hacker News front page (AI)Show HN: Jevman – AI decision models play Pac-Man