multi-agents.site

Projects · case studies · compositions

Decisions, composed.

This page collects concrete projects built on NanoJev-style decision agents. Some are shipped (the replays on this site). Some are compositions that haven't been built yet, listed here so the path from "one model" to "many agents" is visible.

01Shipped: replays on this site

CASE 1

Three-system side-by-side

Jev, NanoJev, and Untuned Qwen play the same Maze and Snake runs. Synchronised by environment step, with action probability bars on the last decision that produced the displayed state.

Open the replay →
CASE 2

Decision arcade

50×50 maze + 12×12 Snake, single-system view. Recorded from the trained NanoJev checkpoint. Per-step probability bars on every action.

Open the arcade →
CASE 3

40-map benchmark

Earlier evaluation: 20 test maps + 20 OOD maps. Three systems compared head-to-head with frozen trajectories and shared controller code.

Open the benchmark →

02Candidate: 2-agent compositions

Compositions that fit naturally inside NanoJev's decision interface. Each example picks one coordination pattern from the homepage §05 and sketches the wiring.

PatternSketch
Voting Two NanoJev instances with different seeds ask the same question on the same state. The system takes the majority argmax. Distribution information is discarded — robust but lossy.
Mixture Two NanoJev instances weighted by recent calibration error. The system returns the weighted average distribution. Useful when agents are exchangeable.
Marginalisation A planner agent emits a goal distribution; an executor agent conditions its action distribution on that goal. Hierarchical composition.
Stackelberg A leader agent commits to a distribution (e.g., target region); a follower agent best-responds. Useful for tasks with sequential commitment.
Joint argmax Two agents with conditionally independent world-views emit distributions; the system computes the joint and takes the argmax. Correct only when independence holds.
Learned aggregator A second model takes both agents' distributions as input and outputs the system distribution. Most expressive; hardest to train.
Quorum Two safety agents vote on the same question; the system acts only when both probabilities exceed τ. Used in safety-critical loops.

03Research frontier

The L5 rung on the homepage spectrum — self-calibrating multi-agent — is where the open questions sit. Below are the threads this site is watching as the upstream project evolves.

Calibration under composition

When agents' distributions are individually calibrated, can the composed system-of-systems be guaranteed calibrated? The upstream RLCD experiment is the first step; full propagation bounds are open.

Agent topology

Single-model → mixture → hierarchical → graph → competitive. Each topology is a different multi-agent regime. Most demos today are at the second rung.

Decision interfaces across vendors

Structured outputs are appearing across model providers in different shapes. The site's vocabulary is shape-agnostic; the engineering reality isn't yet.

Evaluation

A multi-agent system is harder to evaluate than a single model: the action space grows, the credit assignment spreads, and the recorded trajectory becomes the only honest ground truth.

04Add a project

Building something on NanoJev or on a related decision-agent framework? The site mirrors upstream evidence only; it does not host external projects. For now, the canonical place to list a new composition is the upstream issue tracker. As more decision-agent projects land, this page will grow.

See the compositions in motion.

Open the replay