What a simulation can hold
The ranges a design can carry, and the design rules our testing made standard.
Actors
Hundreds of actors, each decomposed into the factions that make its decisions, so internal politics are played rather than assumed. Each faction is defined by what we decide defines it with your experts: positions, red lines, relationships, and sources.
Options
As many policy options as the question needs, each a starting board of state settings and constraints on what each faction may do. A status-quo arm always runs in full, and every option runs on matched event draws, so a difference between options is attributable to the option.
Shocks
Hundreds of shocks and disruptive events, each with a base probability, conditional modifiers, and effects on the state. Organic events are sampled each quarter; scripted events are forced at a fixed quarter in a share of runs to test sensitivity; timed branching events carry known timing and a distribution over outcomes.
Scenarios and narratives
Tens of thousands of runs across options and event draws. The big narratives are read out of that population as frequencies and conditional probabilities, and the scenario tree shows where futures diverge.
The Monte Carlo layer
On top of the runs, a probability layer: outcome classes with their frequencies per option, conditional probabilities given an event sequence, tipping points, and a sensitivity ranking of the assumptions the findings depend on.
Model experiments
Candidate models are compared on pre-registered metrics before the fleet is set, versions are pinned per batch, and contextual judgments are adjudicated by two or three vendor families.
What our testing taught us
- A referee with inertia, decay, and bounds. A rule without them ratchets one way and produces a storyline that repeats under every option and looks like a finding.
- Restraint written into every brief as a costed option. Telling a model not to escalate is itself framing that mentions escalation.
- Content validation of free text. A decision can parse as valid output while carrying scaffold fragments in the text a reviewer reads.
- A uniform-model test before adoption. Tier and role are confounded, and only a paired test on pre-registered metrics settles it.
- Opinion proxies with fatigue. Population proxies carry a fatigue state, a hold base rate, and a cap on how far one proxy moves the state.
Read the method behind the capabilities.
The full walk from design pack to traceability, with the claim boundary stated.