How a simulation works, start to finish
Language-model agents, each playing a faction inside a real actor, act over a multi-year horizon in a world your experts designed. The same design runs many times against different draws of disruptive events, and findings are read as frequencies across those runs.
One design pack, three surfaces
All domain content lives in one versioned design pack: actors, factions, state descriptors, options, events, sources, and the disagreement register. The authoring surface, the engine, and the analysis surface read from that pack, and none of them changes between engagements, so a design can be reviewed, frozen, hashed, and cited independently of the software.
One quarter inside a run
Time advances in ticks, typically quarters, each with two decision rounds in which every agent acts simultaneously. Agents read the state, their memory, and their messages; they decide, messages are delivered, and they decide again. Events are sampled, actions and event effects are adjudicated into the next state, and everything is written to the record before memory is updated.
Agents and red lines
Actors are grouping units and do not act. Factions act, one agent per faction, which is how internal politics are played rather than assumed. Each persona carries documented positions, red lines written as acts the faction takes, behavioral notes, relationships, and grounding sources; its memory is explicit, curated, and inspectable at any quarter of any run. A decision is a structured object: situation read, reasoning, action, posture, intensity, expected reaction, and messages.
Options as starting boards
An option is a starting configuration of the state plus constraints on what each faction may do, marked permitted, prohibited, or conditional. Constraints are enforced softly, so an attempt at a forbidden act is observable and counted rather than filtered away. A status-quo option runs as a full arm, and every option runs against matched random seeds, so a difference between two options is attributable to the option and not to the event draw.
Shocks, organic and scripted
Each event carries a base probability per quarter, conditional modifiers, effects on descriptors, an occurrence cap, and a calibration note for your experts. On the organic path an event is sampled each quarter, conditioned on state and history. On the scripted path it is forced at a fixed quarter in a fixed share of runs, which measures sensitivity to an event without waiting for chance to supply it. Events with known timing and unknown resolution run as timed branching events.
Adjudication: rule, ensemble, mixed
Adjudication turns a round's actions and any fired events into the next state, and the mode is declared per descriptor at the workshop. Rule mode handles mechanical descriptors; in ensemble mode, models from two or three vendor families judge contextual ones in parallel, with disagreement logged and routed to your experts. The referee carries inertia, decay toward baseline, and expert-set plausibility bounds, because a referee without them ratchets the state one way and produces a storyline that repeats under every option and looks like a finding. The ensemble protects against single-vendor idiosyncrasy; it does not correct bias shared across models.
Design elicitation with your experts
The protocol is adapted from the IDEA method in structured expert judgment. SYG pre-loads the full design in your own terminology; every expert reviews every entity in silence and answers confirm, not my area, revise, or flag; a workshop resolves what the review surfaced; every change enters a register with its evidence; and the pack passes a validator, is hashed, and is signed.
A recorded expert split is a legitimate outcome of the workshop. Positions are anonymized, and where a split defines a range it becomes a sweep parameter, so your experts' own unresolved disagreements define the uncertainty space the sensitivity analysis explores.
The four gates
Build does not start until the design is frozen, and production does not start until calibration clears. If your experts cannot accept that agents behave recognizably across the pilot trajectories, production does not proceed.
| Gate | Condition | Who decides |
|---|---|---|
| Design frozen | Validator passes; pack hashed; confirmation signed | Expert group |
| Pilot passed | Mechanics reliable; plausibility reviewed on sampled trajectories; bias battery answered | Design owner and expert group |
| Calibration cleared | Corrections applied and verified on fresh runs; historical-anchor check reviewed | Expert group |
| Validation review | Sample of production trajectories read; residual bias measured and disclosed | Expert group |
Human-in-the-loop modes
No established typology exists for the human role in a multi-agent language-model simulation; the four modes are offered as an adaptation.
| Mode | Description | Status |
|---|---|---|
| Fully autonomous | Agents and adjudication run without intervention between gates | Production batches |
| Expert gates | Design frozen, pilot passed, calibration cleared, validation review | Default |
| Adjudicator override | Experts rule on flagged disagreements; each ruling is recorded as a design change | Calibration; available in production |
| Expert plays a faction | A human seated against the full agent board | Least evidenced; offered as untested |
The pre-registered bias battery
Design-time mitigation reduces bias without showing that it did. The battery shows it: ten families are fixed before the pilot data is opened, each traced to a design mechanism, and the same scripts re-run after every correction. The residual that survives calibration is stated in the method note.
Nine sources of bias, and what reaches each
Model priors on the domain; escalation propensity; persona drafting; prompt framing; model-strength asymmetry between agents; state-transition rules; message quotas; free-text corruption that survives schema validation; and the Western, English-language normalization of the evidence base. Each has a named design response except the last, which only disclosure and a heavier expert-calibration burden address, and we say so.
Models, tested before they are selected
Candidate models are compared on pre-registered metrics, on fixed seeds, before the fleet is set; a uniform-model configuration is tested the same way, because tier and role are confounded and only a paired test settles it. Versions are pinned per batch and recorded in every call, and the provider allow-list runs with fallbacks disabled, so a call reaches the pinned model or fails visibly.
Traceability
Every claim traces along one chain: claim, run, quarter, decision, persona, source. Every model call is logged with its full input and output, and every state change carries a causal field naming the actions and events that produced it. Reproducibility of process rests on the frozen pack with its hash, the audit log, the trajectory database, and a replication kit with pinned model versions.
Six lenses on one database
Outcome frequencies per option, conditional probabilities given an event sequence, the scenario tree, tipping points, a sensitivity ranking of assumptions, and traceability. All six read the same record, so a new question needs no re-run.
- Reproducibility of process, never of outcome. No run is a forecast.
- The method is a measurement instrument that follows expert elicitation; it does not replace expert judgment.
- Resettability is its one categorical advantage over the human wargame; the rest are differences of degree.
- Scale improves reliability, not validity.
- Agent bias and expert bias are symmetrical in kind; the simulation's bias can be measured, which makes it auditable.
- The evidence base is Western-normed and English-language, which argues for a heavier expert-calibration burden.
- Known gaps are stated: no published simulation of this kind has made a calibrated out-of-sample forecast, and no controlled study seats humans and language-model agents together.
See what a simulation can hold.
Actors, options, shocks, scenarios, and the probability layer on top, in capability ranges.