LLM benchmark
Plan every move.
Before the first one.
Snakemark gives a model the full map and every future food, then asks for the whole game as one string. No feedback. No second chances. Just planning.
The loop
- 01
Generate
Pick a seed. Every board is 32×32 with 256 obstacles and 512 food. Same seed, same board, every time.
- 02
Reveal
The model sees everything up front: walls, its body, and every food that will ever spawn.
- 03
Plan
It replies with one compressed string like 6L1R2L1L6. No retries, no feedback.
- 04
Simulate
The plan runs deterministically. Score is food eaten before the first mistake.
Snake is easy to play.
It is hard to plan blind.
- Long-horizon planningDozens of foods, hundreds of moves, one answer.
- Spatial reasoningCoordinates, walls and turns relative to facing.
- State trackingThe body grows and trails behind. Never bite it.
- OptimizationFewer moves breaks ties. Detours cost points.