multi-agents.site

Decision agents · multi-agent systems · research notes

From decisions to multi-agent systems.

A decision agent is a model that returns structured probability distributions instead of free-form text. A multi-agent system is many such agents coordinating. NanoJev proves the first is tractable at 0.6B parameters; this site collects the evidence for the second.

Open the recorded replay Arcade Benchmark → /side-by-side.html

01The decision-agent spectrum

Most AI demos today start with a chatbot and end with a tool call. Decision agents are a different shape: they take a state, ask a question, and return a typed distribution over candidates. Stack enough of them on the same world and you get a multi-agent system without ever leaving structured outputs.

L0

Chat-style agent

Free-form text in, free-form text out. The agent argues with itself, then commits to a string.

L1

Tool-calling agent

Same chat, plus structured calls to functions. Decisions are still emitted as JSON; reasoning is text.

L2

Decision agent (NanoJev)

State + question in, probability distribution over 2–255 candidates out. No output-token decoding. Choice, Boolean, Score.

L3

Multi-decision batch

One forward, many independent decisions. NanoJev handles 6 states × 18 questions × 44 candidates in a single backbone pass.

L4

Multi-agent system

Multiple decision agents share a world, ask different questions, return distributions. Coordination is a function of these distributions.

L5

Self-calibrating multi-agent

Agents' distributions are calibrated on observed events, so the system-of-systems can compute joint beliefs, expected utility, and counterfactual regret.

02Research arc

  1. Structured single-model decisions. NanoJev (0.6B) returns typed probability distributions — Choice, Boolean, Score — without output-token decoding. Source: github.com/TianyuCodings/NanoJev.
  2. Recorded environments. Maze and Snake replays show the same controller code, the same game, three different decision systems — Jev, NanoJev, and Untuned Qwen.
  3. From one to many. Structured decisions are the substrate for multi-agent coordination: each agent emits distributions, the system merges or votes. No free-form text in the loop.
  4. Calibration under composition. When several agents share a question, the answer is no longer a single greedy argmax. Probability learning and proper-reward training become first-class.

03Decision-agent vocabulary

Words used elsewhere with fuzzy meanings, sharpened for this site. The goal is that any two sentences below map to the same model in your head.

decision agent
A model that, given a state and a question, returns a probability distribution over a candidate set. Output is typed (choice / boolean / score), not free-form text.
state
The current observation the agent sees. For a game, it is a local view; for a multi-agent task, it is a shared or partially-shared world.
question
The query the agent is asked. Independent of the state; multiple questions can be asked about the same state.
candidate
One of the possible answers the agent can return. Typed as Choice (2–255), Boolean (2), or Score (2–10 ordered levels).
distribution
The full probability vector over the candidate set. This is the agent's output — not the argmax.
multi-agent
A system of two or more decision agents that share (or partially share) a world and coordinate through their distributions.
coordination
The function that maps a set of agent distributions to a system decision: vote, marginalise, condition, weighted sum, or learned aggregator.
calibration
The property that an agent's reported probability matches observed event frequency. Trained via CE, Brier, or paired proper-reward losses.
zero decoding
Producing the decision from a single forward pass without autoregressive token decoding. Mandatory at multi-agent scale.

04Live demonstrations

DEFAULT

Three-system side-by-side

Jev, NanoJev and Untuned Qwen play the same Maze and Snake runs, synchronized by environment step. Probability bars on the last decision.

Open the replay →
SINGLE MODEL

Decision arcade

Watch NanoJev solve a 50×50 maze and grow through a complete Snake run. Action probability distribution shown for every step.

Open the arcade →
BENCHMARK

40-map navigation

Earlier 40-map benchmark — 20 test maps and 20 OOD maps. Three systems compared head-to-head with frozen trajectories.

Open the benchmark →

05Pages on this site

Four short-form content pages add framing around the replays. Each one is a different lens on the same evidence.

06Coordination patterns

Given N decision agents returning distributions, the system needs a rule to commit to an action. The table lists the patterns that fall out naturally when the outputs are already distributions — no extra parsing.

PatternWhat it does
VotingEach agent votes its argmax; the system takes the majority. Robust when agents are independent but loses the gradient information in each distribution.
MixtureAverage the distributions weighted by agent reliability or recency. The most common pattern when agents are roughly exchangeable.
MarginalisationOne agent conditions on another's distribution. Useful for hierarchical agents (planner + executor).
StackelbergA leader agent commits to a distribution, the followers best-respond. Required when one agent moves first.
Joint argmaxCompute the joint distribution across agents (assumes independence), then take the argmax. Only correct when agents' world-views are conditionally independent given the shared state.
Learned aggregatorA second model takes the agents' distributions as input and outputs a system distribution. Most expressive, hardest to train.
QuorumThe system acts only when at least K agents agree above threshold τ. Used in safety-critical multi-agent loops.

07Why structured decisions matter for multi-agent

1. Probability is the interface

When agents coordinate, the natural handoff is a distribution, not a string. NanoJev's Choice head already returns that — agents can vote, marginalise, or condition without parsing text.

2. Shared state, independent questions

NanoJev batches 6 states × 18 questions in one backbone forward. Multi-agent tasks look exactly like this: each agent sees the same world, asks its own questions, returns its own distribution.

3. Calibration composes

If each agent's distribution is calibrated, the system-of-systems can compute joint beliefs, expected utility, or counterfactual regret. Without calibration, voting reduces to argmax noise.

4. Zero-decoding is mandatory at scale

Once you have N agents × M questions per step, autoregressive token decoding dominates latency. NanoJev's "no output-token decoding" is the engineering argument that structured decisions are the only path to multi-agent at production speed.

5. Decisions are auditable

Every agent output is a distribution over a known candidate set. The system can log, replay, and inspect each decision at every step — essential when many agents cooperate and a single bad distribution can cascade.

08Frequently asked questions

What is a decision agent?

A model that, given a state and a question, returns a probability distribution over a fixed candidate set. The output is typed (Choice, Boolean, or Score) — not free-form text. NanoJev is the smallest open example at 0.6B parameters.

What is a multi-agent system, in one sentence?

Multiple decision agents that share (or partially share) a world and coordinate through their returned distributions, using rules like voting, mixture, marginalisation, Stackelberg, or a learned aggregator.

How is a multi-agent system different from an agent with many tool calls?

A tool-calling agent still produces free-form text as its main output; tools are interrupts. A multi-agent system is composed of decision agents whose outputs are themselves distributions — the system function consumes these directly.

Does this site actually run a multi-agent system?

Not yet. This site is the research evidence layer for the single-decision end of the spectrum. The recorded replays show three decision systems (Jev, NanoJev, Untuned Qwen) playing the same games — a controlled comparison, not yet a composed multi-agent loop.

Why does structured output matter for multi-agent?

Because the coordination function must consume the agent's output. If the output is text, the function has to parse and trust; if the output is a typed distribution, the function can compute. At scale, parse-and-trust is the bottleneck.

Can a single 0.6B model really serve as a decision agent?

Yes — for the question distributions it was trained on. NanoJev reports 77.84% accuracy on local safety test questions and 76.56% on 50×50 OOD questions, with the local model handling 6 states × 18 questions × 44 candidates in one forward.

What's the smallest unit of composition?

A distribution over a candidate set. If two agents can both emit that shape, they can be wired together with any of the patterns in §06. If one of them emits text, you have built a parser, not a multi-agent system.

09Where to read more

Upstream source

NanoJev is published by TianyuCodings — code, checkpoints, dataset, and the pipeline runbook are at github.com/TianyuCodings/NanoJev.

Jev paradigm

The original "System One" model description is at typesafe.ai.

Attribution

All replay data and trained models come from the upstream open-source project. This site mirrors the recorded evidence only; no upstream code or weights are modified here.

Start with the three-system replay — the cleanest demonstration of structured decisions in motion.

Open the recorded replay