SYG Background Paper 2

Methodological foundations of agentic simulation for policy comparison

Design, elicitation, bias measurement, and validation in a multi-agent language-model method for comparing policy options in contested multi-actor systems

Shay Hershkovitz, PhD · SYG ConsultingSYG Background Paper 239 min read
Abstract

This paper describes the method behind SYG Consulting's agentic simulation service and states the claims it can and cannot support. A population of language-model agents, each grounded in the documented positions of one faction inside a real-world actor, acts over a multi-year horizon in a world defined by declared state variables. Policy options are operationalized as constrained starting boards, run against matched draws of disruptive events, and compared on frequencies read across hundreds of trajectories. The method is a measurement instrument that follows structured expert elicitation; its one categorical advantage over the human wargame is resettability. We set out the ontology, time structure, agent architecture, hybrid adjudication, and outcome measurement; model tiering and prompt architecture; the elicitation protocol and the points at which human judgment is decisive; a nine-source taxonomy of bias and the pre-registered ten-family battery used to measure it; the validation vocabulary and four gates; and the artifacts that deliver reproducibility of process. Pilot figures from the first application document the instrument. The pilot showed that the largest distortions arose from the state-transition rule and from the design of population-opinion proxies, that a cross-vendor adjudication ensemble works as a diagnostic and does not correct shared bias, and that corrective prompt language can raise the behavior it targets. Each finding is now a design rule.

All figures come from the pilot of the first application, a comparison of four policy options for a policy research institute. Fourteen of twenty-six pilot runs were executed on the frozen design and form the basis of the figures. They document the measurement apparatus, they are not findings about the domain, and every per-option figure is directional at pilot scale.

1. The question class

Institutions that advise on policy in contested environments face a recurring problem: several discrete courses of action, many actors whose internal politics matter, a horizon long enough for exogenous shocks to dominate any single storyline, and a comparative question. The two conventional instruments have limits their own practitioners describe (Perla, 1990; Perla and McGrady, 2011; NATO ACT, 2023; UK MoD DCDC, 2017; GAO, 2023; Lin-Greenberg, Pauly and Schneider, 2022). The expert narrative produces one trajectory and cannot express the distribution its assumptions imply. The wargame produces one trajectory per play, cannot be reset with the world restored and the players' memories cleared, is documented unevenly, and draws on a small, self-selected player pool.

Agentic simulation addresses one of these limits directly. Persistent, memory-equipped language-model agents produce emergent social behavior (Park et al., 2023), and conditioned language models can reproduce distributions of human responses, unevenly across populations (Argyle et al., 2023). A population of such agents, each grounded in documented positions, can be run through the same configuration hundreds of times with the world reset between runs and the event draw varied. For each option this yields a population of trajectories from which frequencies can be read, and a complete record from which any claim can be traced to the decisions and sources that generated it.

The method fits questions with five properties: several discrete options; a multi-actor system in which internal politics matter; a horizon of one to five years; a high rate of exogenous shocks; and a client whose product is a comparison of options rather than a point estimate. It is a poor fit for single-actor optimization, for questions that turn on actor knowledge unavailable to the design team, and for horizons short enough that one expert narrative is adequate. It does not resolve the validity problem both human instruments share, which is whether the generative model of the world, in the heads of experts or in the weights of a language model, is correct. It improves reliability, and the design reflects that distinction throughout.

2. Claim boundary

Seven positions govern every claim made on the method's behalf.

  1. The method claims reproducibility of process, never of outcome. Any run can be re-executed from the frozen design and the recorded inputs. No run is a forecast.
  2. The method is a measurement instrument that follows expert elicitation. The experts define the world, the actors, the options, the events, and the outcome questions; the simulation executes that design at scale and returns a record. It does not replace expert judgment.
  3. Resettability is the single categorical advantage over the human wargame. Breadth of interaction, completeness of record, and cost per trajectory are advantages of degree, and overstating them invites legitimate critique.
  4. Scale improves reliability. More runs tighten the frequency estimate; they do not make the generative model more correct.
  5. Agent bias and expert bias are symmetrical in kind. They differ in that the simulation's bias can be measured in principle, which makes it auditable rather than absent.
  6. The published evidence base for language-model agents in political simulation is almost entirely Western-normed and English-language. This argues for a heavier expert-calibration burden.
  7. Known gaps in the evidence base are stated. No published multi-agent language-model political simulation has made a calibrated, scored, out-of-sample forecast. The one serious historical backtest in the literature produced war in every run. No controlled study seats humans and language-model agents together. No established typology exists for placing humans in the loop of such a simulation.

3. Design principles

Six principles were fixed before the first application.

Separation of design content from execution. All domain content (actors, factions, state variables, dimensions, options, events, sources, the disagreement register, the glossary) lives in one versioned design artifact, the scenario pack. The engine, authoring surface, and analysis tooling carry no engagement-specific content, so a design can be reviewed, frozen, hashed, and cited independently of the software that executes it.

Frozen design as a gate. A design is frozen when it passes a validator that checks referential integrity, completeness, structural expectations, and closure of the disagreement register. A content hash is computed and the artifact is committed and tagged. Any later change requires a new version and a change-log entry linked to the disagreement item that motivated it. Build does not start before freeze; production does not start before the calibration gate.

Hybrid adjudication declared per variable. Every state variable declares whether its transitions are deterministic (rule mode), contextual (a language-model ensemble), or mixed. Pure rules are brittle in a contested domain; pure model adjudication is opaque and hard to defend. The split is a workshop decision, recorded in the design.

Counterfactual symmetry and the status-quo arm. Options run against matched random seeds, so a difference between two options is attributable to the option and not to the event draw. A status-quo option runs as a full parallel arm; without it the comparison cannot say whether any option is worth pursuing relative to the present course.

Traceability as a deliverable. Every claim in a report traces along the chain claim → run → tick → decision → persona → source. Stable identifiers propagate from the design artifact through the trajectory database and the audit log. Every model call is logged with its model identifier, full input, full output, and reasoning.

Measurement of bias over assertion of neutrality. The design does not claim its agents are unbiased. It builds the instruments to measure bias before results are read, pre-registers the measures, and discloses what survives calibration.

4. The simulation model

4.1 Ontology

Actors are grouping units: states, organizations, institutions, publics. Actors do not act. Factions are the decision units; one faction is one agent is one persona. Decomposing actors into factions is what allows internal politics to be played rather than assumed: a governing coalition against its security establishment, a leadership against its armed wing. Below roughly fifty factions the decomposition collapses; above roughly seventy-five the interaction load grows without proportional gain. The first application ran 47 factions across 15 actors, with a ceiling of 60, under one workshop-recorded constraint: adding agents adds noise, so a request for a new actor defaulted to a note in an existing faction's profile or to a sweep parameter, and became an agent only where the workshop reached consensus or the entity takes decisions the model must route.

State variables (descriptors, in client-facing material) are the tracked properties of the world, each with a data type, range, initial baseline, unit of analysis, and adjudication mode. The first application tracked 33 in three units of analysis (territorial, institutional-governance, security-regional). The variables matrix is the longest-lead dependency in a design, because dimensions, event effects, option deltas, and thresholds all reference it. Dimensions are outcome measures computed from named state variables, each with anchors defining what a 1, a 3, and a 5 mean; without anchors a score is unanchored opinion. Weight schemes are applied to stored scores after the run.

Alternatives are the policy options: starting deltas on state variables, action constraints scoped to a faction, an actor, or all (permitted, prohibited, or conditional), thresholds, and a pass rule. Shocks are exogenous events with a per-tick base probability, conditional modifiers triggered by a prior event or a state condition, effects on state variables in rule or adjudicated mode, a maximum occurrence count per run, and flags for forced injection and sweep candidacy. Sources are the registered grounding documents, typed as client-internal, public, or historical record; every faction references at least one, and the first application registered 57. The disagreement register records expert splits from the workshop, with anonymized positions, a resolution or recorded-split status, and, where the split defines a range, a sweep parameter.

4.2 Time structure

Time advances in ticks; the first application used quarters, twelve over a three-year horizon. Each tick contains two decision rounds. Within a round all agents decide simultaneously, and moves become visible from the next round. Simultaneity was chosen over sequential play because sequential ordering introduces a first-mover artifact with no counterpart in the world being modeled; two rounds per tick allow a response cycle within the quarter without multiplying call volume fourfold.

Tick loop: one quarter inside a runTHE ENGINEOne quarter inside a run1ReadState, own memory,messages received2Round oneAll agents decidesimultaneously3MessagesPrivate and broadcastdelivery4Round twoResponse cycle within thequarter5Events sampledBase probability ×conditional modifiers6AdjudicationReferee rules andensemble judgement7PersistState, decisions,messages, events,adjudications8Memory updateCurated summary per agentnext quarter: quarterly steps over the horizon set per engagementTHREE FEEDBACK PATHSAgent to worldActions change state that every agent reads next round.Agent to agentPrivate messages to named counterparts and broadcast signals.World to publicOpinion proxies move with visible actions and events; decision-making factions read them.Moves are simultaneous within a round and visible from the next; every state change traces to a rule or a logged adjudication.
Figure 1.One quarter inside a run: two simultaneous decision rounds, sampled events, hybrid adjudication, a persisted record.

The sequence within a tick is fixed: agents read the current state, their memory, and messages received since their last action; round one; message delivery; round two; sampling of exogenous events; adjudication of actions and event effects into the next state; persistence of state, decisions, messages, events, and adjudications; memory update. Three feedback paths operate. Agent-to-world: actions change state that every agent reads next round. Agent-to-agent: private messages to named counterparts and broadcast signals. World-to-public: population-opinion proxies move with visible actions and events and are read in turn by the decision-making factions.

4.3 Agent architecture

Each agent is a persona record in the design artifact, executed as one language-model call per round with a constructed prompt and a structured output. The persona carries a role summary, a narrative objective, documented positions, red lines, behavioral notes, an information-access note, relationships to named counterparts (ally, rival, principal–agent, channel), grounding sources, and, for top-tier factions, historical anchors: dated episodes with the behavior the record shows, used in calibration.

Three choices deserve explanation. Narrative objectives in place of utility functions. A utility function would require a preference ordering that the experts do not hold and the record does not support. A narrative objective grounded in documented positions lets the model reason as the faction reasons in its own statements, and keeps the design reviewable by non-technical experts. Documented non-rationality. Behavioral notes record patterns the historical record shows: escalation reflexes, identity commitments, revenge and honor dynamics, a preference for inaction over decision. Agents follow the behavioral logic they are calibrated against; the method does not model irrationality from first principles, and says so. Explicit, curated memory. Each agent has a persistent memory store, updated after every round and held in the database, with a structured summary of its prior actions, key observed events, and recent messages. The context window is never relied on across ticks; full state in the prompt raises cost, degrades reasoning, and is unreadable to a reviewer. Any agent's memory at any tick of any run is inspectable.

The action space is role-constrained. A decision is a structured object with a situation read, reasoning, an action (type, target, description), a posture (escalate, hold, de-escalate, initiate, respond, signal), an intensity, an expected reaction, and messages to named counterparts. Population-opinion proxies emit only signal or hold. A mandatory message quota drives inter-agent traffic, with a side-effect discussed in Section 7.

4.4 Options as constrained starting boards

An option is the frame within which the system is allowed to behave; the simulation does not test it for success. It is a starting configuration of state variables with constraints on what each faction may do. Constraints are enforced softly: the option's summary, the agent's permitted and prohibited actions, and the current stage are placed in the prompt as context, replicated in four positions, and never applied as hard filters on output. Hard filters would remove the possibility of observing whether an agent attempts a forbidden action, which is itself a calibration signal. The audit counts forbidden actions after the fact. In the pilot the count was zero across 15,792 decisions, and 1.9 percent of decisions carried a red-line flag routed for expert reading.

4.5 Exogenous events

Events enter by two paths. On the organic path each event is sampled per tick from its base probability, conditioned by modifiers that reflect state and recent history. On the scripted path an event is forced at a fixed tick in a fixed share of runs, which is how sensitivity to a specific event is measured without waiting for chance to supply it. Provisional probabilities are authored by the design team and flagged for expert calibration; the experts accept, adjust, or re-elicit them. Structural events with known timing and unknown resolution, such as a leadership succession or an election inside the horizon, are implemented as timed baseline events with a distribution over branches.

4.6 Adjudication

Adjudication turns a round's actions and any fired events into the next world state. It is hybrid, declared per variable and per event effect.

Rule mode covers variables whose logic is mechanical: a territorial transfer following a policy decision, a fiscal balance following a funding decision. The first pilot's rule set the step equal to the sign of the summed signed pressure on the variable. It had no inertia, no decay toward baseline, no scaling to magnitude, and no plausibility bound, and it produced the largest artifact the pilot found (Section 7.4). The corrected rule scales the step to summed intensity, decays toward baseline when nothing pushes, applies per-option plausibility bounds authored by the experts (what each option makes impossible), and gives weight to the initiate and respond postures, which the original rule counted as zero.

Ensemble mode covers contextual judgments (legitimacy, alignment, sentiment), issued in parallel by models from two or three vendor families. Agreement is recorded as consensus; disagreement is logged with each judge's reasoning and routed to expert review during calibration. Per-variable movement guidance and anchors are what make ensemble adjudication move at all: adding anchors to one variable reduced inter-judge disagreement from 31.8 percent to 1.0 percent, and the same treatment was applied to every immobile variable. What an ensemble buys should be stated precisely. Error correlation across models rises with capability (Kim et al., 2025), so a frontier cross-vendor ensemble is the most correlated configuration available. It does not correct shared bias, and the design does not claim that it does. It protects against single-vendor idiosyncrasy, the more common and more tractable failure, and it produces a fragility signal, disagreement, that routes contested calls to humans. Its value is diagnostic.

Mixed mode is a rule-based transition with ensemble-adjudicated modifiers.

Every state change is attributable to a rule or to a specific adjudication with full reasoning, and the adjudication record carries a causal field naming the actions and events that produced the change. Causal edges cannot be reconstructed from an empty field after a production run without methodological disclosure, so populating it natively is a precondition for production.

4.7 Outcome measurement

The design as originally specified scored each trajectory at the final tick on eight dimensions and evaluated per-option thresholds. The pilot showed that the 33 state variables saturated within a few ticks under every option, so that dimension scores computed from them were nearly identical across options; eleven of the 33 had an end-state spread of zero on a 1–5 scale. The scoring layer was set aside as an outcome measure, the state variables were kept as the world the agents read each tick, and results are reported as the record read against outcome questions, common and per-option, as frequencies across runs. Event-based questions are computed directly from the record; judged questions are read by a reader with traceability to the decisions it read. Defining the outcome questions is the single most consequential decision the commissioning experts make, and fewer, sharper questions serve end users better than many.

5. Model and prompting choices

5.1 Model tiering

Factions are assigned to model tiers: a top tier for principal decision nodes and for adjudication, a mid tier for sub-factions and secondary state actors, and a base tier for background actors and opinion proxies. Fifty to sixty agents over twelve ticks with two rounds generate on the order of 1,200 to 1,500 agent decisions per run before adjudication, and running every agent on a frontier model at hundreds of runs is neither affordable nor obviously better.

Tiering has a documented cost. In the pilot, escalation was graded by tier: top-tier agents escalated in 56.8 percent of decisions, held in 9.1, and responded in 15.7; mid-tier agents 81.3, 3.3, and 3.9; base-tier agents 92.8, 1.0, and 1.2. Tier and role are confounded in that comparison (four of six same-role pairs followed the gradient, two reversed), so the only test that settles it is paired: the same design run with a uniform model for all agents against the tiered configuration, on pre-registered secondary metrics (posture entropy, combined hold-plus-de-escalate share per tier, a text-integrity rate). That test runs in the first production runs, and the uniform configuration is adopted only if it performs equivalently. If dominance follows model strength, the tier assignment is a hidden parameter with a first-order effect on results, and it must be either eliminated or measured.

Model versions are pinned for a batch and recorded in every call; no model changes mid-batch. Extended reasoning modes are disabled for agent calls, so the recorded reasoning field is the only reasoning trace and the record is comparable across tiers. Temperature is fixed, recorded, and treated as a parameter under test, given replication evidence that it moves escalation propensity (Meerveld et al., 2026; Solopova et al., 2026). Agent calls route through a gateway with an explicit provider allow-list and fallbacks disabled, so a call either reaches the pinned model or fails visibly, and a fallback assignment is drafted before any batch so that loss of an endpoint cannot silently change the model population. The platform supports a multi-vendor fleet, including non-Western model families and regionally hosted endpoints where data-handling constraints require them, with three stated limitations: no model originating in the region under study is available; the non-Western portion of the fleet comes in practice from a single country; and heterogeneous model sizes across tiers confound any comparison by vendor origin.

5.2 Prompt architecture

Prompts are generated from the design artifact by a generator, never authored by hand, so that every change to a faction's profile propagates and a parity check confirms that assembled prompts match the reviewed design. Each agent prompt composes a shared instruction template with a faction brief.

Prompt section Content Rationale
Shared instruction The rules of play: time structure, simultaneity, the decision schema, the message rule (messages only to named counterparts), the one-signal rule for population proxies, the information cut-off One template for all agents keeps the rules identical across tiers and makes any rule change auditable in one place
Who you are Role summary, written in role terms where the actor is modeled structurally instead of as a named incumbent Structural framing lets composition (coalition type, veto players) be a per-run parameter
What you want Narrative objective and documented positions Grounds reasoning in the faction's own stated ends
How you act Behavioral notes, including documented non-rational patterns; for a faction whose activity rises under pressure, the inverted pressure logic is stated explicitly Prevents the model's default assumption that pressure produces restraint
Red lines Written as acts the faction takes Event-phrased red lines (things that happen to the faction) produced no behavioral consequence in the pilot; act-phrased ones are executable
Counterparts Named allies, rivals, principals, channels; an empty list where the faction acts and does not negotiate The message rule depends on it; forcing a counterpart on a non-negotiating faction produced artifactual messaging
Option brief The option in force: summary, this faction's permitted and prohibited actions, the current stage Without it agents learn the option only if the round briefing supplies it; replicated in four positions
Round briefing Current state-variable values, events this tick, messages received, the agent's memory summary The only source of world knowledge; agents receive briefs grounded in sources, never the source texts
Output schema The structured decision object Structured output is requested from the model interface; the schema is validated with a tolerance pass

Three further choices are recorded. Information cut-off. Agents are told they know nothing after the scenario's start, so the model's training knowledge of later events is unusable inside the world. This is a soft constraint, checked in calibration by reading sampled decisions for anachronism. Restraint as a named, costed option. A word count of the first pilot's brief corpus found 433 escalation-related words against 24 restraint-related words. Corrective rewrites that added cautionary language raised escalation across design versions (69.4 to 77.6 to 78.8 percent), consistent with the prompt-sensitivity literature: telling a model not to escalate is itself framing that mentions escalation. The standard response is to write hold and de-escalate into every brief as named options with stated costs and conditions, so that restraint is an action with content. Content validation of free text. A quarter of top-tier decisions in the pilot (820 of 3,360) carried scaffold fragments, markup, or the literal word "placeholder" inside the free-text fields of an otherwise valid decision object; the mid and base tiers produced one such decision in 12,432. Every one passed the schema gate. A content validator on the four free-text fields, with one retry and then an engine-recorded flag, now runs before any other correction, because every downstream metric is read through those fields.

5.3 Prompt sensitivity and the stability probe

The evidence that format and framing move outputs (Sclar et al., 2024) and that outputs vary at fixed inputs (Atil et al., 2024) is treated as a design constraint. The shared template holds all rule text, so a format change is one change measured across all agents. Every design version is run on the same seeds and the posture distribution compared across versions, so a rewrite's effect is measured. And a stability probe, repeat runs at fixed seed and design, is part of the calibration battery; its run-to-run variance is reported as a floor below which no between-option difference is interpreted.

6. Design elicitation and the human role

6.1 The workshop as structured elicitation

The design is elicited from domain experts in a protocol adapted from IDEA (Hemming et al., 2018) and from the center-body-range framing of SSHAC (US NRC, NUREG-2213). The protocol rests on the expert-elicitation literature, which has cross-validated evidence that structure improves group judgment and calibration (Cooke, 1991; Colson and Cooke, 2017; Rowe and Wright, 1999; Tetlock, 2005; Mellers et al., 2014), and does not depend on structured analytic techniques, which the peer-reviewed literature criticizes as largely untested. It has five stages.

Elicitation protocol: pre-load to freezeDESIGN ELICITATIONThe workshop protocol, adapted from IDEA1Pre-loadDesign team drafts actors, factions with sources and anchors, variables, options, events withprovisional probabilities, glossary.2Silent reviewINVESTIGATEEach expert answers per entity: confirm · not my area · revise · flag, plus free-text suggestions.Every response is logged per expert and per field3WorkshopDISCUSSFour hours: table sessions by domain, then a clinic on the questions the tables surfaced. Splits goto the register, anonymised.4Change registerESTIMATEEvery revise, flag and suggestion becomes a change item with evidence, current content, proposedchange, decision required.Rule: adding agents adds noise5Decisions fileOpen questions answered by the design owner, recorded against the pack version they govern.6FreezeAGGREGATEValidator pass, SHA-256 hash, tagged version; the design confirmation document is generated forsignature.Gate: design frozenThe register of expert splits is the distinctive output: recorded splits are legitimate outcomes and become sweep parameters.Rests on the expert-elicitation literature (IDEA, Cooke, SSHAC centre-body-range, ICD 203), not on structured analytic techniques.
Figure 2.The workshop protocol, adapted from IDEA; the disagreement register is its distinctive output.

Pre-load. The design team drafts the full content model into an authoring application: actors, factions with sources and anchors, the variables matrix, dimensions with anchors, options, events with provisional probabilities, sources, and a glossary seeded from the commissioning institution's own terminology. Where experts work in another language, machine translation with a locked glossary produces the canonical text, and the original input is stored immutably alongside it with a provenance flag.

Silent review (Investigate). Before the room, each expert reviews entities in the application and responds per entity with one of four verdicts (confirm, not my area, revise, flag) plus free-text suggestions. The "not my area" verdict lets support be counted only among experts who claim the domain. In the first application this stage produced 393 responses (148 confirm, 164 not my area, 63 revise, 18 flag) and 84 suggestions before the workshop convened.

Workshop (Discuss). Table sessions by domain, then a clinic on the questions the tables surfaced. Decisions, risks, and open items are recorded live; positions entered in the disagreement register are anonymized in live views.

Change register and decisions (Estimate). Every revise or flag response and every suggestion is assigned to a change item with evidence, current content, proposed change, and whether a decision is required. Confirm counts per entity gauge support. Open questions are answered by the design owner in a decisions file that names the design version it governs.

Freeze (Aggregate). The design passes the validator, is hashed, committed, and tagged, and a design confirmation document is generated from it for the experts to sign.

6.2 The disagreement register as the uncertainty model

The workshop's distinctive output is the register of expert splits. A recorded split is a legitimate outcome. Each is reviewed as a sweep candidate and, where it defines a range, becomes a sweep parameter, so the experts' own unresolved disagreements define the model's uncertainty space. This anchors sensitivity analysis in elicited judgment instead of in the design team's guesses about which parameters matter, following the SSHAC principle that the object of elicitation is the range of defensible interpretations and the ICD 203 standard that alternatives be recorded (ODNI, ICD 203).

6.3 Human-in-the-loop configurations

No established typology exists for placing humans in a multi-agent language-model simulation. The four-mode spectrum below is offered as an adaptation.

Four human-in-the-loop configurationsHUMAN IN THE LOOPFour configurations, presented as an adaptationless human intervention between gatesmoreMODE 0Fully autonomousAgents and adjudicationrun without interventionbetween gates.Production batchesMODE 1Expert gatesDesign freeze, pilotreview, calibrationsign-off, validation review.The defaultMODE 2Adjudicator overrideExperts rule on flaggedensemble disagreements;rulings become designchanges.Used in calibrationMODE 3Human on the boardOne faction played bya person against a fullagent board.Least evidencedBeyond the gates, humans are decisive at three points: plausibility bounds, decay rates and fatigue modelsare signed by experts; red-line flags are routed for expert reading; the outcomequestions are the experts' to define.No established typology exists for this application; no controlled study seats humans and language-model agents together.
Figure 3.Four configurations for the human role, presented as an adaptation; no established typology exists.
Mode Description Evidence status
0 Fully autonomous: agents and adjudication run without intervention between gates Used for production batches; humans act at the gates only
1 Expert gates: design frozen, pilot passed, calibration cleared, validation review The default configuration; Section 8.2
2 Adjudicator override: experts review flagged ensemble disagreements and rule on them, and the ruling is recorded as a design change Used during calibration; available as a production configuration at higher cost
3 A human seated against a full agent board, playing one faction The least-evidenced configuration: no controlled study seats humans and language-model agents together, and none has experts blind-rate model-generated moves for plausibility

Beyond the gates, humans are decisive at three points the pilot made explicit. First, domain judgments the substrate cannot supply, meaning what each option makes impossible (plausibility bounds), how fast a situation returns toward normal when nothing pushes it (decay rates), and how a public behaves when it tires or turns (the fatigue model for opinion proxies), are drafted by the design team and signed by the experts as design content. Second, red-line flags (1.9 percent of decisions in the pilot) are routed for expert reading. Third, the outcome questions are the experts' to define.

Two practices complete the picture. The expert group is asked to include researchers across the relevant political spectrum, because systematic skew in calibration outputs is most likely to be caught by reviewers who do not share the same priors; this is a control on the review, and it is stated as such. And no simulation replaces the expert's understanding of the actors, the judgment of plausibility, or the responsibility for the conclusion. Where both instruments are available, the recommended practice is to run them in sequence: the simulation to map the distribution and its conditional structure, then the human wargame or expert panel to interrogate the trajectories that matter.

7. Bias: taxonomy, mitigation, and measurement

7.1 Nine sources of bias

Bias in a multi-agent language-model simulation is several distinct things, of which only the first is the one usually meant. The pilot led us to distinguish nine, each with a different mitigation.

Source Description Design response
Substrate priors The model's training-data priors on the domain, including documented asymmetries on contested topics (Santurkar et al., 2023; Rozado, 2024) Persona grounding; ensemble adjudication addresses part of it
Escalation propensity A general tendency of language-model agents toward escalation in crisis settings (Rivera et al., 2024; Lamparth et al., 2024), moved by temperature and framing, alongside documented social sycophancy (Cheng et al., 2025) Restraint written as a costed option; residual propensity measured and disclosed
Persona bias Bias introduced by the design team's drafting of personas, including which patterns are documented and which omitted Grounding in registered sources; expert review; historical anchors
Prompt-framing bias Effects of wording and format on behavior independent of content (Sclar et al., 2024), including the paradox that cautionary language raises escalation Shared template; every rewrite measured across design versions on fixed seeds
Tier asymmetry Behavior that follows model strength instead of the faction's documented behavior The paired uniform-model test
State-transition bias Artifacts of the referee rule that move the world in one direction regardless of what agents do The corrected referee: inertia, decay, bounds, magnitude
Quota artifacts Structural effects of design rules such as mandatory message quotas, which make some network measures uninformative Reporting rule: in-degree is never reported as salience
Record-integrity bias Corruption inside free-text fields that survives schema validation and biases any reader of the record Content validator, run before any other correction
Cultural and linguistic normalization An evidence base and model population that are largely Western and English-language, affecting every actor modeled and non-Western actors more Disclosure and a heavier calibration burden; nothing in the design removes it

7.2 Mitigation at design time

Personas are drafted from three source classes (the commissioning institution's internal assessments, public documented positions, and the historical record), and every faction references at least one registered source. This is the differentiator from generic role-play: the agent reasons from the faction's documented positions, and the reviewer can see which documents produced which behavior. Top-tier factions carry historical anchors; during calibration the agent is placed in the anchor's context and its decision compared with the record, as a check on behavioral grounding and with no claim to reproduce specific paths. Hybrid adjudication, the ensemble as diagnostic, matched seeds with a status-quo arm, and expert diversity in review complete the design-time mitigations (Sections 4.6, 3, and 6.3).

7.3 The pre-registered bias battery

Mitigation at design time is necessary and insufficient, so a bias battery runs on the pilot before results are read. The families are fixed before the data is opened, each is traced to a design mechanism, runs of earlier design versions serve as controls, and the same scripts re-run after every correction, which is how corrections are shown to have worked. The ten families in the first application:

Ten pre-registered bias families, mechanism and statusMETHOD · BIASThe pre-registered bias batteryFAMILYMECHANISM EXAMINEDSTATUSState-direction ratchetReferee ruleFound · correctedVariable mobilityReferee rule · ensemble anchorsFound · correctedPopulation-proxy behaviourProxy design · one-signal ruleFound · correctedTier gradientModel tieringFound · correctedText integrityStructured-output handlingFound · correctedAttention concentrationMessage quotaReporting ruleRecord compositionRoster and quotaReporting ruleOption differentiationOption constraintsClearedEvent behaviourEvent library and samplingClearedRepertoire parityAction-space constraintsClearedFamilies are fixed before the data is opened, each traced to a design mechanism; runs of earlier designversions serve as controls. The same scripts re-run after each correction, which is how corrections are shown to work.Fixed before the data is opened; re-run after every correction; residual bias disclosed in the method note.
Figure 4.Ten families fixed before the data was opened; each traced to a design mechanism.
Family Measure Mechanism examined Pilot result
State-direction ratchet Up versus down step counts per variable and per adjudication mode; floor and ceiling dwell Referee rule Found; corrected
Variable mobility Actions targeting each variable against its movement; end-state spread across options Referee rule; ensemble anchors Found; corrected
Population-proxy behavior Posture shares, target concentration, single-action-combination degeneracy, share of record Proxy design; one-signal rule Found; corrected
Tier gradient Posture shares by tier and by role; same-role paired comparison Model tiering Found; paired test scheduled
Text integrity Rate of scaffold, markup, and placeholder fragments in free-text fields Structured-output handling Found; validator added
Attention concentration Message in-degree and out-degree per faction Message quota Artifact; reporting rule adopted
Record composition Share of decisions by actor class; repertoire breadth by class Roster and quota Reading rule adopted
Option differentiation Narrative similarity: same option and different seed, against same seed and different option, against unrelated pairs Option constraints Cleared
Event behavior Firings per event against base probability; share of events that fired Event library and sampling Cleared
Repertoire parity Range of action types used by each side Action-space constraints Cleared

The pilot base was fourteen complete runs on the frozen design: 15,792 decisions, 26,836 messages, 38,674 events, 5,544 state observations, and 3,067 adjudications, with seventeen runs on earlier design versions as controls.

7.4 What the battery found, and the design rules that resulted

The faction the design team and the experts had suspected of distorting the results, an armed movement that attacked in 320 of 336 decisions, was not the largest distortion. In order of effect:

The state layer was a one-way ratchet. 83.3 percent of state steps went in one direction (315 up, 63 down). Twelve variables only rose, four only fell, six never moved in 168 run-quarters. A coalition-stability variable reached its floor in every run, in ten of them by the first quarter, and never recovered. An institutional-reform variable reached "certified complete" in every run, including under the option that abolishes the institution; the adjudication record read, verbatim, "deterministic transition: net pressure +3 → step +1". Rule mode produced 222 up-steps against 30 down; ensemble mode was directionally balanced (26 up, 23 down) but almost immobile, 49 steps against 252. The storyline that repeated in every option (attacks by the armed movement, a security response, coalition strain, elections) was substantially a product of this rule. The resulting rule is the corrected referee of Section 4.6 and a verification gate: any variable with more than 100 targeted actions and zero movement fails.

The option-distinguishing variables were the inert ones. The variable most targeted by agent action (1,354 actions) never moved. The mechanism variable that defines one option ended at 1.90 of 5; the mechanism variable that defines another ended near its floor. The instrument was most responsive where options behave alike and least where they should diverge. The resulting rule is per-variable movement guidance and anchors in ensemble mode and expert-authored plausibility bounds in rule mode.

Population-opinion proxies were a larger artifact than the armed movement. By design the proxies emit only signal or hold. In practice they escalated in 89 to 95 percent of decisions and de-escalated in under 3 percent; four of them aimed most of their output at coalition stability; they were more degenerate on the single-action-combination measure than the armed movement (0.86–0.94 against 0.89); and they generated 21 percent of the record. The resulting rule is a fatigue or salience state for each proxy, a real hold base rate, and a cap on how far one proxy can move one variable per tick.

Escalation was graded by model tier and rose across corrective rewrites (Sections 5.1 and 5.2). The resulting rules are restraint as a named, costed option in every brief and the paired uniform-model test on pre-registered metrics.

Record integrity (Section 5.2). The resulting rule is the content validator, run first.

Attention concentration was a quota artifact. Five of 47 factions received 45.4 percent of all messages and four received none in fourteen runs; out-degree was near-identical for every agent because the quota is mandatory. In-degree is therefore not evidence of salience and is not reported as such.

Record composition. External actors were 28 of 47 factions and 59.6 percent of decisions; the two principal parties together were a quarter. Repertoire breadth was adequate on all sides, so this is a volume imbalance, and outcome questions are read against the principal parties' record.

What the battery cleared matters as much. Differentiation between options in behavior is real: narrative similarity was 0.955 for same option and different seed, 0.902 for same seed and different option, and 0.899 for unrelated pairs, so the option explains roughly fifteen times as much of the difference between two runs as the seed does. The event library behaved as specified: 128 firings in 168 run-quarters, 26 of 30 events fired, at rates near base probabilities. Repertoire parity was adequate. Red-line flags (1.9 percent) and first-attempt schema validity (97.5 to 100 percent) held. No agent took a forbidden action.

Two lessons generalize. A one-way referee produces a storyline that looks like a finding, so the state layer needs inertia, decay, and bounds before anyone reads outcomes. And corrective prompt rewrites can make a bias worse, so the effect of every rewrite must be measured across design versions on fixed seeds.

7.5 Residual bias and disclosure

A residual escalation bias in language-model agents is expected to survive calibration. It is measured on the first tranche of production runs and stated in the method note, with the posture distribution by tier and by role. The pilot's own limits are disclosed alongside it: fourteen runs with three or four per option make every per-option figure directional, and where a branch was scripted by seed in the pilot, its shares were not sampled.

8. Validation and verification

8.1 Vocabulary

The vocabulary is borrowed from agent-based modeling, including its notions of model alignment and docking (Axtell et al., 1996). Following Windrum, Fagiolo and Moneta (2007), we distinguish input validation (are the personas, variables, events, and options a faithful operationalization of the experts' design), process validation (does the engine execute the design as specified), and output validation (do the trajectories exhibit the patterns the record supports). Following Grimm et al. (2005), output validation is pattern-oriented: several patterns at once. Following Epstein (2008), the model's purposes are explanatory and exploratory, and prediction is not among them. To this vocabulary we add three items specific to language-model agents: persona fidelity against historical anchors; escalation propensity treated as a parameter under test, with its sensitivity to temperature and framing measured (work on language models in strategic wargaming treats both as design variables that move agent behavior; Meerveld et al., 2026; Solopova et al., 2026); and record integrity of free-text fields.

8.2 The four gates

Engagement lifecycle with four expert gatesENGAGEMENT LIFECYCLEEight phases, four expert gates0FramingQuestion-class fit, alternatives, horizon, configuration, deliverables1DesignPre-load, silent review, workshop, change register, decisions fileGATE · DESIGN FROZENValidator passes; hash computed; design confirmation signed — expert group2BuildEngine adapted to the pack; prompts generated and checked for parity3PilotStaged runs with automatic gates and a spend ceiling; bias batteryGATE · PILOT PASSEDMechanics reliable; plausibility reviewed; battery responded to — design owner and experts4Client confirmationDesign-and-pilot report, decision form, player and event book, record5Corrections and verificationContent validator first, then referee and ensemble; four-run verificationGATE · CALIBRATION CLEAREDVerification gates pass, including mobility; anchors reviewed — production does not start before this6ProductionHundreds to thousands of runs per option; extended where not convergedGATE · VALIDATION REVIEWSampled production trajectories read by experts; residual bias measured and disclosed7HandoverDatabase, record, frozen pack, method note, replication kit; 90-day supportBuild does not start until the design is frozen; production does not start until calibration clears. Experts re-confirmed at every gate.
Figure 5.Build does not start until the design is frozen; production does not start until calibration clears.
Gate Condition Who decides
Design frozen Validator passes; hash computed; design confirmation signed Expert group
Pilot passed Mechanics reliable (schema validity, forbidden-action count, event rates, runtime); plausibility reviewed on sampled trajectories; bias battery run and responded to Design owner and expert group
Calibration cleared Corrections applied; four-run verification passes its gates, including the mobility gate; historical-anchor check and, where used, a historical backtest exercise reviewed Expert group; production does not start before this
Validation review A sample of production trajectories read by the experts; residual bias measured and disclosed Expert group

The gates are the method's falsification structure. The clearest falsification is at calibration: if the experts cannot accept that agents behave recognizably across the pilot trajectories, behavioral fidelity has failed and production does not proceed. A weaker falsification is a sensitivity analysis that flips the conclusion under small parameter changes. The method is falsifiable in its parts even where it cannot be falsified as a whole, and the design is structured to fail visibly.

8.3 Verification metrics

At every gate the same metrics are computed: first-attempt schema validity; content-validator failure rate; forbidden-action count; red-line flag rate; event firing rates against base probabilities; wall-clock per run; state-step direction ratio per variable and per mode; variable mobility against targeting; posture distribution by tier and role; text-integrity rate; and the narrative-similarity triad. Reusable scripts extract the runs and run the battery; re-running them after each correction is how corrections are shown to have worked.

8.4 Historical backtest exercise

Where a comparable past decision episode exists, the design is run with agents constrained to period-appropriate information and the resulting decisions are read for recognizability against the record. This checks behavioral grounding and makes no claim of replay. The first application ran two such exercises on one past episode alongside its pilots on the live question. The published literature contains, to our knowledge, one serious historical backtest of a multi-agent language-model crisis simulation, and it produced war in every run; the bibliographic entry for that study is under verification in our source base and is not cited by name here. The exercise here is therefore also a test of whether the design's restraint mechanisms have any effect at all.

8.5 Sensitivity and convergence

The primary sweep crosses a small number of high-weight axes drawn from the disagreement register with the options; the first application used two archetypes of one external power, switched deterministically at a fixed tick, at a fixed 50/50 split within each option. Forced-injection sets test specific events at fixed shares of runs. The run plan is 150 per option in the first tranche, extended to 250 only where the outcome-question frequencies have not converged. Below 100 runs per option, ordinal claims are supportable and rare pathways are not. At 250, a pathway occurring in two percent of runs is resolved well enough to characterize. Beyond that, returns diminish until Monte Carlo scale, which is a different epistemic commitment. An extended sweep across parameter conditions is the upgrade path; it delivers a ranked list of the assumptions the findings depend on, with effect sizes, which answers the most common critique of an exercise of this kind.

9. Traceability and reproducibility

Reproducibility of process is delivered by four artifacts. The frozen design artifact, with its content hash, change log, and decisions file. The audit log, in which every model call carries its model identifier, full input, full output, and reasoning. The trajectory database, in which every run, tick, round, decision, message, event, state observation, and adjudication is a row keyed by the design's stable identifiers, handed over as a self-contained file browsable without a server. And a replication kit of deployment scripts, dependency manifests, and pinned model versions, sufficient for the commissioning institution or an external auditor to re-execute any run from the recorded inputs.

Traceability chain: claim to run to tick to decision to persona to sourceTRACEABILITYFrom any claim to its sourceClaimA sentence in the findings brief, with the runs it rests onRunRun identifier, option, random seed, pack version and hashTickQuarter and round; the state the agents readDecisionStructured object: situation read, reasoning, action, posture, messagesPersonaThe faction brief and the agent's memory at that tickSourceA registered grounding document: internal assessment, public position or historical recordStable identifiers propagate from the design pack through the trajectory database and the audit log.Every model call is logged with model identifier, full input, full output and reasoning.Client documents strip the internal codes; the database keeps them.
Figure 6.Every claim traces to the run, quarter, decision, persona and source that produced it.

Two disclosures qualify the claim. Vendor model versions are pinned but are not under our control; a deprecation mid-project is handled by moving the affected agents to the equivalent tier and documenting the swap, and the audit log makes the change visible. And run-to-run variance at fixed inputs is non-zero (Atil et al., 2024); the stability probe reports it, and no between-option difference smaller than it is interpreted. Client-facing documents strip internal identifiers and engineering terms; the database keeps them. Sample databases distributed to reviewers are built from the full record, because a truncated extract cuts decision descriptions and biases any reader of it.

10. Limitations and threats to validity

Validity of the generative model. The simulation's world is the experts' design executed by language models. It is not a validated forecasting instrument, and findings hold under the assumptions of the simulation. Nothing in the method establishes that the assumptions are correct; the elicitation protocol establishes only that they are the experts' and that their disagreements are recorded.

Substrate bias and cultural normalization. Persona grounding, ensemble adjudication, and diverse expert review make bias traceable; they do not remove it. A coordinated bias shared across vendors would affect every agent and every adjudicator alike, and the evidence base is Western-normed.

Escalation propensity. A residual tendency toward escalation is expected, measured, and disclosed.

Tier asymmetry. Until the uniform-model test is run, the tier assignment is a candidate explanation for any behavioral pattern that follows it.

Adjudication. This is the component the design team is least sure of. Hybrid design, logged reasoning, per-variable anchors, and expert review of flagged disagreements are the mitigations. The pilot exists in part to expose adjudication failure before production, and in the first application it did.

Soft constraint enforcement. Option constraints are prompt context; the audit checks compliance after the fact.

Sample size at pilot. Every per-option figure from a fourteen-run pilot is directional.

Human-in-the-loop evidence. Modes 2 and 3 have almost no controlled evidence behind them; mode 3 is offered as an option the client should understand to be untested.

Analytical mediocrity. The method can produce defensible findings that are not decision-relevant. The sensitivity sweep, the outcome-question discipline, and the long-term value of the database to the commissioning institution are the responses.

11. Conclusion

The method turns an expert-elicited design of a contested strategic environment into a population of traceable trajectories under each of several policy options. Its claims are deliberately narrow: process reproducibility, resettability, completeness of record, and measurable bias. Its defense against methodological critique rests on making every choice visible, from the ontology to the residual bias disclosed, and on measuring the effect of every correction. The first application's pilot showed that the largest distortions in a simulation of this kind can arise from the state-transition rule and from the design of background agents, that a cross-vendor ensemble is a diagnostic, and that corrective prompt language can worsen the behavior it targets. Each of those findings is now a design rule. Whether the method's findings in any application are decision-relevant is a question the commissioning experts answer at the gates, which is where the method places them.

References

Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.

Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., and Baldwin, B. (2024). Non-determinism of "deterministic" LLM settings (earlier versions titled "LLM stability: A detailed analysis with some surprises"). arXiv:2408.04667.

Axtell, R., Axelrod, R., Epstein, J. M., and Cohen, M. D. (1996). Aligning simulation models: A case study and results. Computational and Mathematical Organization Theory, 1(2), 123–141.

Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., and Jurafsky, D. (2025). ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995.

Colson, A. R., and Cooke, R. M. (2017). Cross validation for the classical model of structured expert judgment. Reliability Engineering & System Safety, 163, 109–120.

Cooke, R. M. (1991). Experts in Uncertainty: Opinion and Subjective Probability in Science. Oxford University Press.

Epstein, J. M. (2008). Why model? Journal of Artificial Societies and Social Simulation, 11(4), 12.

Grimm, V., Revilla, E., Berger, U., Jeltsch, F., Mooij, W. M., Railsback, S. F., Thulke, H.-H., Weiner, J., Wiegand, T., and DeAngelis, D. L. (2005). Pattern-oriented modeling of agent-based complex systems: Lessons from ecology. Science, 310(5750), 987–991.

Hemming, V., Burgman, M. A., Hanea, A. M., McBride, M. F., and Wintle, B. C. (2018). A practical guide to structured expert elicitation using the IDEA protocol. Methods in Ecology and Evolution, 9(1), 169–180.

Kim, E. M., Garg, A., Peng, K., and Garg, N. (2025). Correlated errors in large language models. Proceedings of the 42nd International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, 267, 30038–30066. arXiv:2506.07962.

Lamparth, M., Corso, A., Ganz, J., Mastro, O. S., Schneider, J., and Trinkunas, H. (2024). Human vs. machine: Behavioral differences between expert humans and language models in wargame simulations. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES).

Lin-Greenberg, E., Pauly, R. B. C., and Schneider, J. G. (2022). Wargaming for international relations research. European Journal of International Relations, 28(1), 83–109.

Meerveld, H., Brighton, H., Lindelauf, R., and Šafář Postma, M. (2026). Effective and responsible use of large language models in strategic wargaming. Journal of Defense Modeling and Simulation, published online 19 April 2026. https://doi.org/10.1177/15485129261438519

Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., Scott, S. E., Moore, D., Atanasov, P., Swift, S. A., Murray, T., Stone, E., and Tetlock, P. E. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106–1115.

NATO Allied Command Transformation (2023). NATO Wargaming Handbook. Headquarters, Supreme Allied Commander Transformation, Norfolk, VA. September 2023.

Office of the Director of National Intelligence. Intelligence Community Directive 203: Analytic Standards.

Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST).

Perla, P. P. (1990). The Art of Wargaming: A Guide for Professionals and Hobbyists. Naval Institute Press. (Revised edition: Peter Perla's The Art of Wargaming, edited by J. Curry, History of Wargaming Project, 2012.)

Perla, P. P., and McGrady, E. (2011). Why wargaming works. Naval War College Review, 64(3), 111–130.

Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., and Schneider, J. (2024). Escalation risks from language models in military and diplomatic decision-making. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT).

Rowe, G., and Wright, G. (1999). The Delphi technique as a forecasting tool: Issues and analysis. International Journal of Forecasting, 15(4), 353–375.

Rozado, D. (2024). The political preferences of LLMs. PLOS ONE, 19(7), e0306621.

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023). Whose opinions do language models reflect? Proceedings of the 40th International Conference on Machine Learning (ICML).

Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2024). Quantifying language models' sensitivity to spurious features in prompt design. International Conference on Learning Representations (ICLR).

Solopova, V., Skorik, V., Tereshchenko, M., Haidun, A., and Vykhopen, O. (2026). LLMs as strategic actors: Behavioral alignment, risk calibration, and argumentation framing in geopolitical simulations. arXiv:2603.02128.

Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press.

United Kingdom Ministry of Defence, Development, Concepts and Doctrine Centre (2017). Wargaming Handbook.

United States Government Accountability Office (2023). Defense Analysis: Additional Actions Could Enhance DOD's Wargaming Efforts. GAO-23-105351. April 2023.

United States Nuclear Regulatory Commission. NUREG-2213: Updated Implementation Guidelines for SSHAC Hazard Studies.

Windrum, P., Fagiolo, G., and Moneta, A. (2007). Empirical validation of agent-based models: Alternatives and prospects. Journal of Artificial Societies and Social Simulation, 10(2), 8.

Discuss a simulation for your question

A short conversation is enough to establish whether your question fits the method. If it does not, we will say so.