Comparison · decision agents · multi-agent
Same game, different decisions.
The recorded replays on this site put three systems side by side on the same Maze and Snake runs: Jev, NanoJev, and Untuned Qwen3-0.6B. They share a controller; they diverge on how the next move is chosen. This page collects the comparison data.
On this page
01The three systems
| System | Backbone | Decision head | Output |
|---|---|---|---|
| Jev | Qwen3-0.6B + decision heads | Choice / Boolean / Score | Typed probability distribution |
| NanoJev | Qwen3-0.6B + decision heads | Choice / Boolean / Score | Typed probability distribution |
| Untuned Qwen3-0.6B | Qwen3-0.6B (raw LM) | Language-model head | Free-form text, constrained to A–D |
Jev and NanoJev are siblings: same architecture, different training. The "Untuned Qwen" row is a control — a pretrained LM with no decision-head training, conditioned on the offered answer tokens.
02Decision interface
- Jev / NanoJev
- Given a state and a question, the model emits a vector over the candidate set.
P(action | state, question). The argmax is one statistic of that vector; the full distribution is the output. - Untuned Qwen
- Given a state and a question phrased as a multiple-choice prompt, the LM picks one of the offered tokens (A–D). The argmax is the output; the distribution over A–D is not directly returned by the head.
- What this means
- The decision-agent output is a distribution. The chat-agent output is a token. The composition patterns on the homepage rely on the first; the second has to be coerced.
03Maze (50×50)
Recorded on the same map, the same seed, the same controller code. The controller filters immediate collisions and remembers which edges have been tried; the model picks among the remaining directions.
| System | Attempts | Collisions | Outcome |
|---|---|---|---|
| NanoJev | 244 | 36 | Goal reached |
| Jev | 2,738 | 1,044 | Goal reached |
| Starting NanoJev | 171 | 43 | Goal reached |
All three reach the goal. NanoJev does it in fewer attempts and fewer collisions than Jev on this map. See the side-by-side replay for the actual trajectories.
04Snake (12×12)
Same controller, same food sequence, same seed (61005), same horizon (256 steps).
| System | Food collected | Steps | Outcome |
|---|---|---|---|
| NanoJev | 27 | 256 | Alive at horizon |
| Jev | 30 | 256 | Alive at horizon |
| Untuned Qwen3-0.6B | 25 | 211 | Trapped |
The chat-conditioned Untuned Qwen reaches the same food count in fewer steps — it plays greedily and dies earlier. The structured systems survive the full horizon. See the decision arcade for the per-step probabilities.
0540-map navigation benchmark
Earlier benchmark: 20 test maps + 20 OOD maps, controller T=1 probability sampling.
| System | 4×4 test | 6×6 OOD |
|---|---|---|
| NanoJev | 19/20 — 95% | 18/20 — 90% |
| Jev | 20/20 — 100% | 19/20 — 95% |
| Untuned Qwen3-0.6B | 7/20 — 35% | 3/20 — 15% |
The OOD drop is informative: the chat-conditioned system collapses; the decision-agent systems hold. See the benchmark viewer.
06What scales, what doesn't
Scales well
Batching many states and questions in a single forward. NanoJev reports 6 states × 18 questions × 44 candidates in one backbone pass — the kind of throughput a multi-agent loop needs at every step.
Scales poorly
Autoregressive token decoding. Even a 0.6B model, called many times across N agents × M questions, dominates latency. NanoJev's "no output-token decoding" is the engineering argument for switching shapes.
Scales with calibration, not argmax
Composition (vote / mixture / marginalise / Stackelberg) only works correctly if each agent's distribution is calibrated. Without calibration, voting reduces to argmax noise. The probability-learning pilot reported in the upstream README shows the calibration path.
See the comparison in motion.
Open the three-system replay