Three-system side-by-side
Jev, NanoJev and Untuned Qwen play the same Maze and Snake runs, synchronized by environment step. Probability bars on the last decision.
Open the replay →Decision agents · multi-agent systems · research notes
A decision agent is a model that returns structured probability distributions instead of free-form text. A multi-agent system is many such agents coordinating. NanoJev proves the first is tractable at 0.6B parameters; this site collects the evidence for the second.
Most AI demos today start with a chatbot and end with a tool call. Decision agents are a different shape: they take a state, ask a question, and return a typed distribution over candidates. Stack enough of them on the same world and you get a multi-agent system without ever leaving structured outputs.
Free-form text in, free-form text out. The agent argues with itself, then commits to a string.
Same chat, plus structured calls to functions. Decisions are still emitted as JSON; reasoning is text.
State + question in, probability distribution over 2–255 candidates out. No output-token decoding. Choice, Boolean, Score.
One forward, many independent decisions. NanoJev handles 6 states × 18 questions × 44 candidates in a single backbone pass.
Multiple decision agents share a world, ask different questions, return distributions. Coordination is a function of these distributions.
Agents' distributions are calibrated on observed events, so the system-of-systems can compute joint beliefs, expected utility, and counterfactual regret.
github.com/TianyuCodings/NanoJev.Words used elsewhere with fuzzy meanings, sharpened for this site. The goal is that any two sentences below map to the same model in your head.
Jev, NanoJev and Untuned Qwen play the same Maze and Snake runs, synchronized by environment step. Probability bars on the last decision.
Open the replay →Watch NanoJev solve a 50×50 maze and grow through a complete Snake run. Action probability distribution shown for every step.
Open the arcade →Earlier 40-map benchmark — 20 test maps and 20 OOD maps. Three systems compared head-to-head with frozen trajectories.
Open the benchmark →Four short-form content pages add framing around the replays. Each one is a different lens on the same evidence.
The mission, principles, and what is and isn't on this site.
NanoJev vs Jev vs Untuned Qwen — the comparison data behind the replays.
Five paths into decision agents — 10 minutes to a full weekend.
Case studies and candidate multi-agent compositions.
Given N decision agents returning distributions, the system needs a rule to commit to an action. The table lists the patterns that fall out naturally when the outputs are already distributions — no extra parsing.
| Pattern | What it does |
|---|---|
| Voting | Each agent votes its argmax; the system takes the majority. Robust when agents are independent but loses the gradient information in each distribution. |
| Mixture | Average the distributions weighted by agent reliability or recency. The most common pattern when agents are roughly exchangeable. |
| Marginalisation | One agent conditions on another's distribution. Useful for hierarchical agents (planner + executor). |
| Stackelberg | A leader agent commits to a distribution, the followers best-respond. Required when one agent moves first. |
| Joint argmax | Compute the joint distribution across agents (assumes independence), then take the argmax. Only correct when agents' world-views are conditionally independent given the shared state. |
| Learned aggregator | A second model takes the agents' distributions as input and outputs a system distribution. Most expressive, hardest to train. |
| Quorum | The system acts only when at least K agents agree above threshold τ. Used in safety-critical multi-agent loops. |
When agents coordinate, the natural handoff is a distribution, not a string. NanoJev's Choice head already returns that — agents can vote, marginalise, or condition without parsing text.
NanoJev batches 6 states × 18 questions in one backbone forward. Multi-agent tasks look exactly like this: each agent sees the same world, asks its own questions, returns its own distribution.
If each agent's distribution is calibrated, the system-of-systems can compute joint beliefs, expected utility, or counterfactual regret. Without calibration, voting reduces to argmax noise.
Once you have N agents × M questions per step, autoregressive token decoding dominates latency. NanoJev's "no output-token decoding" is the engineering argument that structured decisions are the only path to multi-agent at production speed.
Every agent output is a distribution over a known candidate set. The system can log, replay, and inspect each decision at every step — essential when many agents cooperate and a single bad distribution can cascade.
A model that, given a state and a question, returns a probability distribution over a fixed candidate set. The output is typed (Choice, Boolean, or Score) — not free-form text. NanoJev is the smallest open example at 0.6B parameters.
Multiple decision agents that share (or partially share) a world and coordinate through their returned distributions, using rules like voting, mixture, marginalisation, Stackelberg, or a learned aggregator.
A tool-calling agent still produces free-form text as its main output; tools are interrupts. A multi-agent system is composed of decision agents whose outputs are themselves distributions — the system function consumes these directly.
Not yet. This site is the research evidence layer for the single-decision end of the spectrum. The recorded replays show three decision systems (Jev, NanoJev, Untuned Qwen) playing the same games — a controlled comparison, not yet a composed multi-agent loop.
Because the coordination function must consume the agent's output. If the output is text, the function has to parse and trust; if the output is a typed distribution, the function can compute. At scale, parse-and-trust is the bottleneck.
Yes — for the question distributions it was trained on. NanoJev reports 77.84% accuracy on local safety test questions and 76.56% on 50×50 OOD questions, with the local model handling 6 states × 18 questions × 44 candidates in one forward.
A distribution over a candidate set. If two agents can both emit that shape, they can be wired together with any of the patterns in §06. If one of them emits text, you have built a parser, not a multi-agent system.
NanoJev is published by TianyuCodings — code, checkpoints, dataset, and the pipeline runbook are at github.com/TianyuCodings/NanoJev.
The original "System One" model description is at typesafe.ai.
All replay data and trained models come from the upstream open-source project. This site mirrors the recorded evidence only; no upstream code or weights are modified here.
Start with the three-system replay — the cleanest demonstration of structured decisions in motion.
Open the recorded replay