jonclegg.github.io

Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?

thefourthchime · 78 points · 50 comments · ieri · Open original

Benchmarks how well Harness+models can create a Pac-Man game from a single prompt: “Create a Pac-Man game in a single HTML page” Each model gets one shot — no follow-up prompts or fixes.

Comments

5 preview comments · loading full thread
yambamieri

This reminds me of a couple decades ago when some friends and I set ourselves a challenge to, individually, each create as much of a Pac-Man clone as possible in 10 hours. None of us had any experience in games programming or graphics coding. It was great fun and we all learned a lot. No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing. I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.

drcxdieri

Interesting, recently I am working on my own clone of Pac-Man. LLM implementations lose lots of details. They are not 1:1 replication of the original game. For example, the behavior of the ghost is not the same as the original. If I have not implemented the game myself, I can not tell the differences. What LLM produced look like the original game, but they are not.

strataspaceieri

I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering. The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.

jmathaiieri

This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving. Remember when people considered you a genius for prompting with "You are a skilled writer....".

_matthew_ieri

I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.