Docs
Snakemark measures how well a language model plans a long sequence of actions with no feedback. The model sees a whole Snake game up front, the board, the walls, its body and the full ordered list of food that will ever spawn, and must answer with one movement string. The simulator then plays that string exactly as written.
It tests long-horizon planning, spatial reasoning, state tracking and optimization at once.
Benchmark flow
- Seed. Pick a seed. Everything else is fixed. A live preview shows the exact board.
- Prompt. Copy the generated prompt into any LLM, then paste its answer back.
- Simulate. Watch the plan play out at 1×, 2× or 5×, or skip straight to the end.
- Score. See food collected, survival, efficiency and how it stacks up against the reference solver.
Board and coordinates
- Cells are
(x,y).(0,0)is the top-left.xgrows right,ygrows down. - The snake starts with length 3 at the centre column, head at
(⌊n/2⌋, ⌊n/2⌋), body trailing downward, facing up. - Obstacles never sit on the snake or the two cells straight ahead of it, and never cut off part of the board.
Command syntax
A response is a sequence of tokens:
N, a positive integer: move forwardNcells.L/R: turn 90° left or right, relative to the current facing. Turning does not move the snake.
Example: 6L1R2L1L6 means forward 6, turn left, forward 1, turn right, forward 2,
and so on. Whitespace and backticks are stripped and letters are case-insensitive. Anything
else, or a 0, makes the response invalid.
Rules and edge cases
- One food at a time. Only the current food is on the board. The next appears the moment it is eaten.
- Growth. Eating grows the snake by 1: the tail stays put on that move.
- Tail chasing. The tail vacates its cell during a move, so moving into the tail's cell is safe, except on a move that eats.
- Food placement. Food never spawns on an obstacle or inside a pocket the board could only reach through a single cell. It can spawn under the snake's body, and is only eaten when the head enters that cell. Consecutive foods never share a cell.
- Death. Leaving the grid, entering an obstacle, or entering your own body ends the game immediately. The fatal move does not count.
- Running out. The game also ends when commands are exhausted, all food is eaten, or the move limit (
food × grid × 4) is reached. - Malformed responses score 0 food and an F. Nothing is guessed or repaired.
- No tools. The model must plan by reasoning alone. The prompt forbids writing or running code, solvers, simulators, search or any other tool. Run it with tools and code execution turned off. A tool-assisted answer is not a valid result.
- Trailing turns after the last move are allowed and do nothing.
Scoring
- Food collected is the primary score.
- Crash penalty. Dying by wall, obstacle or own body keeps only 75% of the food eaten, rounded down. Running out of commands is free, so a model should stop where it can no longer track the board instead of guessing. Grades, vs reference and the leaderboard all use this score.
- Moves survived counts every successful move.
- Efficiency is the sum of Manhattan distances between consecutive food pickups divided by the moves actually used to reach the last one. 100% means no detours at all, which walls and the body rarely allow.
- vs reference compares the score to a built-in solver: BFS to each food, only taking paths that keep its tail reachable, otherwise following its tail until a safe path opens. It is a solid baseline, not an optimal player.
| Grade | Meaning |
|---|---|
| S | Ate every food |
| A | ≥ 90% of the reference solver's food |
| B | ≥ 70% |
| C | ≥ 50% |
| D | ≥ 25% |
| F | Below 25%, or an invalid response |
Fixed settings
Every game uses the same settings. Only the seed changes, so any two results on the same seed can be compared directly. The settings are deliberately brutal so that even the strongest models have room to improve. Report results averaged across several seeds.
| Grid | Obstacles | Food | Move limit |
|---|---|---|---|
| 16×16 | 64 | 192 | 12288 |
Determinism
Boards and food come from a seeded mulberry32 generator, so the same seed
always builds the same game in every browser. The simulator has no randomness. To compare
models fairly, give each one the same seed, in a fresh conversation, with the prompt unchanged.
Leaderboard runs
The leaderboard only lists results we collect ourselves. Once a day, every free text model from NVIDIA, OpenCode Zen and OpenRouter that isn't on the board yet gets one attempt.
- The same fixed settings as every game, on one seed we keep private so nobody can tune against the board.
- The prompt is exactly what the benchmark page copies. The model's final message is scored as-is.
- The harness is stock OpenCode with every tool denied, in a sandbox with no internet. Only the model's own API is reachable.
- Runs over 15 minutes score as invalid. Provider errors are not recorded and are retried the next day.