SYG Background Paper 1

What the literature says

Language-model agents, bias, and the limits of strategic prediction

Shay Hershkovitz, PhD · SYG ConsultingSYG Background Paper 127 min read
Abstract

This paper reviews the published evidence that bears on one question: what can a simulation built from language-model agents be trusted to say about a contested, multi-actor strategic environment? It draws on five literatures. The wargaming and intelligence-analysis literature explains why single-trajectory instruments fail under strategic surprise and why their own practitioners say so. The agent-based modeling literature supplies a validation vocabulary that separates input, process, and output validity and that names purposes for a model other than prediction. The generative-agent literature establishes that language-model agents can produce recognizable social behavior and warns that fidelity is uneven, that over-prescriptive prompting predetermines outcomes, and that believability is not empirical realism. The bias literature documents political priors, prompt sensitivity, run-to-run instability, sycophancy, escalation propensity, and correlated errors across models. The expert-elicitation literature shows which procedures for aggregating human judgment have cross-validated evidence behind them. The paper then states what the evidence base cannot yet support and closes with the design consequences for a service that reads frequencies across hundreds of traceable runs and declines to sell forecasts.

1. Why prediction fails in contested multi-actor systems

Institutions that advise on policy in contested environments face a recurring analytical problem. Several discrete courses of action are available; the environment contains many actors whose internal politics matter; the horizon is long enough for exogenous shocks to dominate; and the question put to the analyst is comparative. Under these conditions the record of expert prediction is poor, and the reasons are well documented.

Prediction versus simulation: one narrative line against a fan of trajectoriesPREDICTION VERSUS SIMULATIONOne narrative, or a distribution of futuresNarrative assessmentone trajectory, its assumptions implicittimestatethe expected coursea shock that was not written inAgentic simulationhundreds of trajectories per option, shocks sampledtimestateshockshockThe simulation returns frequencies across runs, conditional probabilities given an event sequence, and tipping points.It does not resolve whether the generative model of the world is correct; it improves reliability, not validity.Conceptual. No run is a forecast; findings hold under the assumptions of the simulation and are stated as such.
Figure 1.One narrative, or a distribution of futures: the simulation returns frequencies, not a forecast.

The cognitive part of the problem is the better studied. Tetlock's (2005) twenty-year study of expert political judgment found that specialists' probability estimates about political and economic outcomes were only modestly better than simple extrapolation, that confidence was weakly related to accuracy, and that the style of reasoning mattered more than the depth of subject knowledge. The subsequent forecasting-tournament work (Mellers et al., 2014) showed that calibration can be improved and named the mechanisms: training in probabilistic reasoning, working in teams that argue, and tracking performance against outcomes. More expertise on its own did not help. Expert judgment is therefore a measurable quantity with known failure modes and known correctives, and unstructured expert opinion is the weakest form of it.

The organizational part of the problem is the more consequential for warning. Strategic surprise more often follows from a failure to weight signals against the noise of routine reporting, to record the alternatives that were set aside, and to express uncertainty in a usable form than from a failure to collect the signals at all. The analytic standards codified in Intelligence Community Directive 203 (ODNI) exist because of institutional failures of this kind. They require that products describe the quality of sources, express uncertainty in defined terms, distinguish underlying information from analytic judgment, and record alternative hypotheses. Each requirement answers a documented way in which an assessment can be confidently wrong without any single analyst having erred.

Two instruments are conventionally used for the comparative question, and both are valuable. The expert narrative assessment produces a reasoned account of how a situation is likely to develop under a given course of action. Its limit is structural: a narrative produces one trajectory and cannot express the distribution of outcomes its own assumptions imply. If the assumptions admit an escalation path that occurs one time in five, the narrative either reports the modal path and loses the tail or reports the tail and misstates the mode. The reader cannot recover the frequency from the text, because the text never contained it.

The wargame answers part of this limit by putting adversarial human players into a structured environment and letting the trajectory emerge from their play. Its practitioners are candid about what the instrument can bear. Perla (1990) and Perla and McGrady (2011) locate the value of the wargame in the narrative engagement of the players, which is also its weakness as evidence: the trajectory depends on who played. Lin-Greenberg, Pauly and Schneider (2022) set out the conditions under which wargames can serve as research instruments and identify repeatability and player selection as the two central threats to validity. A game cannot be reset with the world restored and the players' memories cleared, so a replication is never a replication of the same game, and the player pool is small and self-selected. Doctrine handbooks (NATO ACT, 2023; UK MoD DCDC, 2017) and an audit of Department of Defense practice (GAO, 2023) add a third limit: documentation and adjudication transparency are uneven, so the reasons a game went one way rather than another are often not recoverable afterward. We draw these limits from wargamers rather than from advocates of computational methods so that the comparison is on the instrument's own terms.

Table 1 sets the three instruments side by side on the properties the literature identifies as decisive. On validity, the property that matters most, the three are on the same footing: each depends on a generative model of the world that lives in the heads of experts, in the play of participants, or in the weights of a language model, and none can establish that the model is correct. What differs is repeatability, the completeness of the record, and the marginal effort of producing one more trajectory. Of these, only repeatability is a categorical difference.

Table 1. Three instruments for the comparative strategic question

Property Expert narrative assessment Human wargame Agentic simulation
Trajectories per exercise One One per play; plays are not replications of each other Hundreds per option; each run resets the world and the agents' memories
Repeatability Not applicable Not achievable; players remember and the pool is small (Lin-Greenberg et al., 2022) Achievable for process; the same design and inputs re-execute any run
Record of reasoning The text; alternatives often not recorded (ICD 203) Uneven; adjudication often not transparent (GAO, 2023; NATO ACT, 2023) Complete by construction; every decision, message, and adjudication is stored with its reasoning
Marginal effort per additional trajectory A new assessment A new game with new players Low; a difference of degree, since design and calibration effort is fixed and large
Source of behavior The analyst's model of the actors The players' engagement (Perla, 1990) Documented positions rendered by language models; subject to the biases of Section 4
Validity of the generative model Unverifiable in advance Unverifiable in advance Unverifiable in advance
What scale buys Nothing; there is one trajectory Little; each game adds one non-replicable trajectory Reliability of the frequency estimate; it does not improve validity

The last row states the boundary that governs everything that follows. More runs tighten the estimate of how often a path occurs under the simulation's assumptions and do nothing to establish that the assumptions are right.

2. Agent-based modeling and its validation vocabulary

Agent-based modeling (ABM) is the tradition in which heterogeneous agents with bounded rationality interact locally and produce system-level patterns that are not derivable from aggregate equations. Epstein (1999) described its epistemology as generative: a macroscopic regularity is explained when a population of agents following stated rules can grow it. The tradition is the closest established methodology to what a language-model simulation does, and over three decades it has developed a vocabulary for saying what a model has and has not shown.

The first item in that vocabulary is purpose. Epstein (2008) listed sixteen reasons to build a model other than prediction, among them explaining structure, bounding the range of outcomes, revealing sensitivities and the parameters that matter, illuminating core uncertainties, disciplining the policy dialogue, and exposing which assumptions a conclusion depends on. Prediction is one purpose among many and often the wrong one for the question at hand. A comparative question about policy options in a contested system is better served by a model that bounds outcomes and reveals sensitivities than by one that issues a point estimate it cannot defend.

The second item is validation. Windrum, Fagiolo and Moneta (2007) distinguished input validation, the correspondence between the model's assumptions and what is known about the target system; output validation, the correspondence between the model's behavior and observed patterns; and the indirect calibration of parameters against stylized facts. Grimm et al. (2005) added pattern-oriented modeling, in which a model earns credibility by reproducing several observed patterns at once, because a single pattern can be matched by many wrong models and several constrain the space far more tightly. Axtell et al. (1996) defined docking, the alignment of two independently built models on the same problem, as a test of whether a result belongs to the phenomenon or to the implementation. These distinctions carry over directly. Input validation becomes the question of whether the personas, state variables, events, and options faithfully operationalize the experts' design; process validation, whether the engine executes that design as specified; output validation, whether the trajectories exhibit the patterns the historical record supports, where Grimm's discipline says that one recognizable pattern is not enough.

The third item is the record of what ABM has achieved in a policy setting, which is instructive in both directions. Farmer and Foley (2009) argued in a widely cited Nature commentary that representative-agent equilibrium models exclude, by construction, the financial instability, distributional effects, and emergent dynamics that policy most needs to see. Poledna et al. (2023) then built a fully calibrated agent-based model of a national economy with forecasting performance comparable to conventional econometric models and with distributional resolution retained. Dignum et al. (2020) deployed a combined health, social, and economic agent-based simulation as a government decision-support tool during the pandemic. Each success came with the same caveat. Gill et al. (2024), reporting a one-to-one-scale agent-based model of a national economy, state that calibration requires micro-level data unavailable in most studies and that accuracy degrades materially when aggregates are used as proxies. The field's thirty-year history (Tesfatsion and Judd, 2006; Dawid et al., 2012) records the same constraint at every stage: methodological progress has repeatedly outrun the data needed to calibrate it. Before language models, a second bottleneck was that behavioral rules had to be hand-coded by domain experts, which limited how rich the agents could be.

For a strategic-political application, the calibration bottleneck takes a different form. There is no micro-data set of faction decisions; there is a documentary record of positions, statements, and past episodes, and there are experts who have read it. Calibration therefore becomes a question of behavioral recognizability judged by people, which is where the elicitation literature of Section 5 enters.

3. Generative agents and generative agent-based modeling

The generative-agent literature is short and recent. Park et al. (2023) demonstrated that language-model agents equipped with persistent memory, reflection, and planning could populate a small simulated town and produce emergent social behavior that observers found believable: information spread, relationships formed, and coordinated activity arose without being scripted. Argyle et al. (2023) showed something more directly relevant to policy analysis. Conditioned on demographic profiles, a language model could reproduce the distribution of survey responses across human sub-populations, which the authors called silicon samples. The same study showed that the reproduction was uneven: fidelity was higher for some groups and questions than for others, so the model's approximation of a population is a property to be measured.

Ghaffarzadegan et al. (2024) formalized the synthesis of these capabilities with the ABM tradition as generative agent-based modeling (GABM), in which a language model serves as the decision engine inside each agent through a profile, memory, planning, and action loop. This removes the hand-coding bottleneck of Section 2: an agent's behavioral rules can be expressed as documented positions in natural language, and the model does the interpretive work of applying them to a situation. Gao et al. (2024) surveyed the growing body of work that uses language models in agent-based modeling and reached the same conclusion about its promise and the same list of open problems, with evaluation and reproducibility at the head of it.

Three critical contributions temper the promise, and each has become a design constraint for any serious application.

Li and Wu (2025) analyzed twenty-two GABM studies and identified what they call the over-control paradox. Highly prescriptive prompts predetermine agent behavior and suppress the emergent dynamics that make agent-based modeling worth doing; under-specified prompts let the model's defaults, rather than the documented behavior of the actor, drive the outcome. There is no neutral setting. Prompt design is a scientific variable whose effects must be measured across versions.

Larooij and Törnberg (2025a; 2025b) reviewed the use of language models in agent-based modeling and concluded that validation is the central unresolved challenge of generative social simulation. The field has mostly demonstrated believability, which is a property of the observer's reaction, and has rarely demonstrated empirical realism, causal validity, or operational usefulness, which are properties of the model's relation to the world. They also note that performance shifts materially across languages, populations, and settings, which converts Argyle's uneven fidelity from an observation into a warning.

Münker et al. (2026) made the same point empirically for a bounded case. Generative agents asked to mimic communication on social networks did not reproduce the empirical patterns of real communication unless their realism had been benchmarked and tuned, and the authors' title advises against trusting such agents until that benchmark has been run. A population of language-model agents produces output that looks like the phenomenon; whether it behaves like the phenomenon is a separate question that requires a measurement.

Read together, these works establish feasibility and set the terms of use. Language-model agents can carry documented positions into a multi-actor interaction and produce trajectories that experts recognize. Whether the trajectories are informative depends on prompt design that neither over-controls nor abdicates, on a validation program that measures realism rather than believability, and on a statement of the populations and settings for which fidelity has and has not been checked.

4. Bias, instability, and escalation in language models

The bias literature on language models is larger than the generative-agent literature and bears on the design in more specific ways. Four strands can be separated.

Nine bias sources mapped to mitigationsMETHOD · BIASNine sources of bias, and what addresses each1 Substrate priors2 Escalation propensity3 Persona drafting4 Prompt framing5 Tier asymmetry6 State-transition rule7 Quota artefacts8 Record integrity9 Cultural normalisationPersona grounding in registered sourcesEnsemble adjudication (diagnostic, notcorrective)Restraint as a named, costed option; everyrewrite measured on fixed seedsPaired uniform-model testCorrected referee: inertia, decay, boundsReporting rule: in-degree is not salienceContent validator on free-text fieldsDisclosure and a heavier expert-calibrationburdenOnly the first source is the one usually meant by 'model bias'. Sources 6 and 3 produced the largest distortions in our testing, not source 1.Source 9 (dashed) is addressed only by disclosure; the evidence base is Western-normed and English-language.
Figure 2.Nine sources of bias, and what addresses each; the largest pilot distortions arose from the state-transition rule and the design of background agents.

The first is political and opinion bias. Santurkar et al. (2023) compared model outputs on public-opinion questions with the responses of demographic groups in a large survey and found that the models' distribution of opinions aligned with some groups far more than others; the models had identifiable priors that were not representative of the population. Rozado (2024) administered political-orientation instruments to a range of models and found a consistent lean across most of them. Any agent that is not thoroughly grounded in its faction's documented positions will drift toward the model's defaults, and those defaults are not neutral.

The second is instability. Sclar et al. (2024) showed that semantically irrelevant features of prompt format, such as spacing, punctuation, and the order of options, move outputs by amounts large enough to change the ranking of models on a benchmark. Atil et al. (2024) documented run-to-run variation at fixed inputs even where the sampling temperature is set to zero. Cheng et al. (2025) measured social sycophancy, the tendency of models to preserve the user's self-image and agree with the framing they are given. Format sensitivity means that a rewrite of an agent's brief can change behavior for reasons unrelated to its content. Run-to-run instability sets a floor: no difference between two options smaller than the variance at fixed inputs can be interpreted. Sycophancy means that an agent told, in cautionary language, what not to do may absorb the framing rather than the instruction.

The third is escalation, the strand most directly relevant to a defense or security audience. Rivera et al. (2024) placed language-model agents in simulated crisis and diplomatic decision-making and found a systematic tendency toward escalation, in some configurations toward extreme actions, with agents justifying escalation in arms-race and deterrence terms. Lamparth et al. (2024) compared expert human players with language models in a wargame setting and found systematic behavioral differences, so the model cannot be treated as a stand-in for the expert without adjustment. Subsequent work on language models in strategic wargaming (Meerveld et al., 2026; Solopova et al., 2026) treats temperature and prompt framing as design variables that move agent behavior, which is why the method records temperature and measures each rewrite's effect rather than assuming neutrality. Escalation propensity is a parameter of the configuration, and a simulation that does not measure it under its own settings has no basis for reading escalation frequencies as a property of the actors.

The fourth concerns what an ensemble of models can do about the first three. Kim et al. (2025) showed that errors across language models are correlated and that the correlation rises with capability, presumably because the strongest models share training data, architectures, and alignment procedures. A cross-vendor ensemble of frontier models is therefore the most correlated configuration available. It cannot correct a bias the vendors share; it protects against the idiosyncrasy of a single vendor and produces a signal, disagreement, that identifies contested judgments. An ensemble is a diagnostic rather than a corrective, and a design that presents it as a corrective overstates what it buys.

Table 2 maps these findings onto the sources of bias a multi-agent simulation must account for. The taxonomy is wider than the model's priors because the largest distortions can arise elsewhere: in the rule that moves the world state, in the design of background agents, and in the structural rules of the exercise.

Table 2. Sources of bias in a multi-agent language-model simulation and the evidence behind each

Source What the literature shows Consequence for a simulation
Substrate priors Models carry identifiable political and demographic priors (Santurkar et al., 2023; Rozado, 2024) Ungrounded agents drift toward the model's defaults; grounding in documented positions is a necessity
Escalation propensity Agents tend to escalate in crisis settings; temperature and prompt framing move the propensity (Rivera et al., 2024; Meerveld et al., 2026; Solopova et al., 2026) Escalation must be treated as a parameter under test and measured under the configuration in use
Prompt framing and format Irrelevant format features move outputs; cautionary framing can raise the behavior it targets (Sclar et al., 2024; Cheng et al., 2025) Every rewrite of an agent brief must be measured across versions on fixed seeds
Run-to-run instability Outputs vary at fixed inputs (Atil et al., 2024) A stability probe sets the floor below which no between-option difference is interpretable
Over-control Prescriptive prompts predetermine outcomes (Li and Wu, 2025) Constraints belong in context and are audited after the fact; hard filters remove the signal
Uneven population fidelity Silicon-sample fidelity varies by group and setting (Argyle et al., 2023; Larooij and Törnberg, 2025a) Fidelity must be checked per actor class, with heavier checking for non-Western actors
Persona authorship Which patterns are documented and which omitted is a design choice Persona drafting is itself a bias source and is reviewed by experts who do not share one set of priors
Correlated errors across models Error correlation rises with capability (Kim et al., 2025) An ensemble flags contested calls; it does not remove shared bias
Cultural and linguistic normalization The evidence base is largely Western and English-language (Larooij and Törnberg, 2025a; Münker et al., 2026) A heavier expert-calibration burden, disclosed rather than assumed away

Two points about this table deserve emphasis. First, bias in the agents and bias in the experts are symmetrical in kind. Both come from a generative model of the actors formed by a particular history and a particular selection of sources. They differ in one respect: the simulation's bias can be measured in principle, across thousands of recorded decisions, whereas an expert's bias can only be inferred. Measurability makes the simulation's bias auditable, and auditable is a weaker word than absent. Second, the literature's numbers are population averages over benchmark tasks. What they imply for a particular design, with particular personas and rules, is unknown until it is measured on that design. That is why a pre-registered measurement of bias belongs inside the method rather than in a caveat at the end of the report.

5. Structured expert elicitation as the defensible foundation

If the generative model of the actors cannot be validated by the simulation, it has to be supplied by people, and the question becomes which procedures for eliciting and aggregating human judgment have evidence behind them. This is a mature literature, and its findings are clearer than those of Sections 3 and 4.

Cooke (1991) established the classical model of structured expert judgment, in which experts are scored on calibration questions with known answers and their judgments on the questions of interest are weighted by performance. Colson and Cooke (2017) cross-validated the approach across a large body of applications and showed that performance-weighted aggregation outperforms equal weighting out of sample, which is the test that matters. Rowe and Wright (1999) reviewed the evidence on the Delphi technique and found that structured, iterated, anonymized elicitation improves group judgment over unstructured discussion, while noting the variability of implementations. Hemming et al. (2018) set out the IDEA protocol, in which experts investigate the question individually, discuss their estimates as a group, re-estimate individually, and have their estimates aggregated, with evidence that this sequence outperforms both unstructured groups and single experts. The structure does the work: private judgment before discussion protects against anchoring on the first speaker, discussion surfaces information and interpretation, and a second private estimate prevents the group from collapsing to a dominant voice.

Two bodies of practice guidance extend these findings into institutional settings. The SSHAC guidance for hazard studies (US NRC, NUREG-2213) formalizes the object of elicitation as the center, body, and range of technically defensible interpretations: the purpose is to capture the distribution of expert views, including the disagreements, rather than to force a consensus. ICD 203 (ODNI) requires sourcing, defined expressions of uncertainty, the recording of alternatives, and a clear separation of information from judgment. Tetlock (2005) and Mellers et al. (2014) supply the evidence on calibration described in Section 1.

One family of methods is conspicuously absent from this list. Structured analytic techniques (SATs), the toolkit of devil's advocacy, analysis of competing hypotheses, and related exercises widely taught in intelligence organizations, are criticized in the peer-reviewed literature as largely untested: they are plausible and widely used, and they have not been shown in controlled study to improve judgment. A workshop whose defense rests on SATs rests on assertion, whereas one that rests on Cooke, IDEA, Delphi, and SSHAC rests on cross-validated evidence.

The implications for a language-model simulation are concrete. The experts define the world, the actors, the options, the events, and the outcome questions; the simulation executes that design at scale and returns a record. Elicitation should proceed through private review before group discussion, with a verdict that lets an expert decline a domain outside their competence so that support is counted only among those who claim it. Disagreements among experts are the uncertainty model: a recorded split is a legitimate output, and where it defines a range it becomes a parameter the simulation sweeps. The expert group should be diverse in its priors, because systematic skew in a calibration output is most likely to be caught by a reviewer who does not share it. And the whole design should be frozen, hashed, and signed before execution begins, so that the object being simulated is a documented artifact of expert judgment.

6. What the evidence base cannot yet support

A review that reported only what the literature supports would be incomplete. The following gaps bound the claims that can be made, and a reader evaluating any service in this class should know them.

No published multi-agent language-model political simulation has, to our knowledge, made a calibrated out-of-sample forecast. The generative-agent studies report believability and, in bounded cases, distributional match to survey data; the escalation studies report propensities under controlled settings; none has issued probabilities about future political events and been scored against outcomes. This is why a service in this class cannot describe its output as a forecast and describes it instead as a distribution of futures under stated assumptions.

The published literature contains, to our knowledge, one serious historical backtest of a multi-agent language-model crisis simulation, and it produced war in every run. The bibliographic entry for that study is under verification in our source base and is not cited by name here. The finding matters for what it implies: a historical exercise is as much a test of whether a design's restraint mechanisms have any effect as it is a check on behavioral grounding, and a design that has not run such an exercise does not know whether its agents can hold.

No controlled study seats humans and language-model agents together in the same exercise, and no study has experts blind-rate model-generated in-game moves for plausibility against human-generated ones. Lamparth et al. (2024) compared humans and models playing separately. A configuration in which a human plays one faction against a board of agents is an option a client should understand to be untested.

No established typology exists for where humans should be placed in a multi-agent language-model simulation. A systematic review of agentic digital twins in industrial settings finds that human-agent collaboration frameworks and adaptive-autonomy mechanisms are open research gaps even there, where the systems are far more constrained (Miadowicz et al., 2026; Heininger, Jost and Stary, 2024). Any human-in-the-loop scheme for a political simulation is an adaptation offered as such.

The evidence base is Western-normed and English-language. The models were trained predominantly on English text; the bias studies used predominantly Western survey instruments and political scales; the generative-agent studies simulated predominantly Western populations; and Larooij and Törnberg (2025a) and Münker et al. (2026) both report that fidelity shifts across languages and populations. Thin evidence on non-Western actors argues for a heavier calibration burden on them, since the checks the literature has run elsewhere have not been run there.

7. Implications for practice: what a defensible service can claim

The literature reviewed here converges on a small number of positions that any service offering agentic simulation for strategic decision support should adopt and be judged against. They are stated below as they inform SYG's design rules; the method itself is described in the companion paper.

The service claims reproducibility of process and never of outcome. Any run can be re-executed from the frozen design and the recorded inputs. No run is a forecast. The output is a distribution of trajectories under stated assumptions, read as frequencies, conditional frequencies, tipping points, and sensitivities, in the manner Epstein (2008) describes as the proper purposes of a model.

The service is a measurement instrument that follows structured expert elicitation. The experts define the world; the workshop follows IDEA and SSHAC rather than SATs; the disagreement register is the uncertainty model; and human judgment is decisive at four gates: design frozen, pilot passed, calibration cleared, and validation review. The simulation adds breadth of interaction, completeness of record, and repeatability. It adds no validity, and the expert keeps responsibility for the conclusion.

The single categorical advantage over the human wargame is resettability. Everything else is a difference of degree, and the service says so, because the wargaming literature's own account of its limits (Lin-Greenberg, Pauly and Schneider, 2022) is precise enough that overstating the comparison invites a critique that is easy to make and hard to answer.

Bias is measured rather than asserted away. The service carries the taxonomy of Table 2 into a pre-registered battery run on the pilot before any result is read; treats escalation propensity as a parameter under test that temperature and framing are known to move; measures the effect of every prompt rewrite across design versions on fixed seeds, because the prompt-sensitivity literature says a rewrite can worsen what it targets; and describes its cross-vendor adjudication ensemble as a diagnostic that flags contested calls, because Kim et al. (2025) show that shared error is exactly what an ensemble of strong models cannot remove.

Every claim traces to its evidence. The validation vocabulary of Windrum, Fagiolo and Moneta (2007) and the pattern-oriented discipline of Grimm et al. (2005) are adopted as stated, with three additions specific to language-model agents: persona fidelity against dated historical anchors, escalation propensity as a measured parameter, and the integrity of free-text fields in the record. The chain from claim to run to tick to decision to persona to source is a deliverable, because ICD 203's requirements for sourcing and the recording of alternatives apply to a simulation's output as much as to an analyst's. The gaps of Section 6 are disclosed in every method note.

The literature, in short, supports a narrow and useful claim. A population of grounded language-model agents can be run through an expert-designed world hundreds of times, with the world reset between runs and the event draw varied, to produce a traceable record from which the frequencies of outcomes under each option can be read and compared, subject to biases that can be measured and disclosed. The literature does not support the claim that such a simulation predicts what will happen. A service that keeps to the first claim and refuses the second is consistent with the evidence. That is the position SYG takes.

References

Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., and Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.

Atil, B., Aykent, S., Chittams, A., Fu, L., Passonneau, R. J., Radcliffe, E., Rajagopal, G. R., Sloan, A., Tudrej, T., Ture, F., Wu, Z., Xu, L., and Baldwin, B. (2024). Non-determinism of "deterministic" LLM settings (earlier versions titled "LLM stability: A detailed analysis with some surprises"). arXiv:2408.04667.

Axtell, R., Axelrod, R., Epstein, J. M., and Cohen, M. D. (1996). Aligning simulation models: A case study and results. Computational and Mathematical Organization Theory, 1(2), 123–141.

Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., and Jurafsky, D. (2025). ELEPHANT: Measuring and understanding social sycophancy in LLMs. arXiv:2505.13995.

Colson, A. R., and Cooke, R. M. (2017). Cross validation for the classical model of structured expert judgment. Reliability Engineering & System Safety, 163, 109–120.

Cooke, R. M. (1991). Experts in Uncertainty: Opinion and Subjective Probability in Science. Oxford University Press.

Dawid, H., Gemkow, S., Harting, P., van der Hoog, S., and Neugart, M. (2012). The Eurace@Unibi model: An agent-based macroeconomic model for economic policy analysis. Bielefeld Working Papers in Economics and Management No. 05-2012, Bielefeld University.

Dignum, F., et al. (2020). Analysing the combined health, social and economic impacts of the coronavirus pandemic using agent-based social simulation. Minds and Machines, 30, 177–194. https://doi.org/10.1007/s11023-020-09527-6

Epstein, J. M. (1999). Agent-based computational models and generative social science. Complexity, 4(5), 41–60.

Epstein, J. M. (2008). Why model? Journal of Artificial Societies and Social Simulation, 11(4), 12.

Farmer, J. D., and Foley, D. (2009). The economy needs agent-based modelling. Nature, 460, 685–686. https://doi.org/10.1038/460685a

Gao, C., et al. (2024). Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanities and Social Sciences Communications, 11(1). https://doi.org/10.1057/s41599-024-03611-3

Ghaffarzadegan, N., et al. (2024). Generative agent-based modeling: An introduction and tutorial. System Dynamics Review, 40(1). https://doi.org/10.1002/sdr.1761

Gill, A., Lalith, M., Hori, M., and Ogawa, Y. (2024). Analysis of postdisaster economy using high-resolution disaster and economy simulations. Risk Analysis, 45(6), 1254–1270 (published online October 2024). https://doi.org/10.1111/risa.17662

Grimm, V., Revilla, E., Berger, U., Jeltsch, F., Mooij, W. M., Railsback, S. F., Thulke, H.-H., Weiner, J., Wiegand, T., and DeAngelis, D. L. (2005). Pattern-oriented modeling of agent-based complex systems: Lessons from ecology. Science, 310(5750), 987–991.

Heininger, R., Jost, T. E., and Stary, C. (2024). Developing and operating digital process twins while preserving autonomy in CPS. Proceedings of the IEEE Conference on Business Informatics, 50–59. https://doi.org/10.1109/cbi62504.2024.00016

Hemming, V., Burgman, M. A., Hanea, A. M., McBride, M. F., and Wintle, B. C. (2018). A practical guide to structured expert elicitation using the IDEA protocol. Methods in Ecology and Evolution, 9(1), 169–180.

Kim, E. M., Garg, A., Peng, K., and Garg, N. (2025). Correlated errors in large language models. Proceedings of the 42nd International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, 267, 30038–30066. arXiv:2506.07962.

Lamparth, M., Corso, A., Ganz, J., Mastro, O. S., Schneider, J., and Trinkunas, H. (2024). Human vs. machine: Behavioral differences between expert humans and language models in wargame simulations. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES).

Larooij, M., and Törnberg, P. (2025a). Validation is the central challenge for generative social simulation: A critical review of LLMs in agent-based modeling. Artificial Intelligence Review, 59, 15. https://doi.org/10.1007/s10462-025-11412-6

Larooij, M., and Törnberg, P. (2025b). Do large language models solve the problems of agent-based modeling? A critical review of generative social simulations. arXiv:2504.03274.

Li, Z., and Wu, Q. (2025). Let it go or control it all? The dilemma of prompt engineering in generative agent-based models. System Dynamics Review, 41(3). https://doi.org/10.1002/sdr.70008

Lin-Greenberg, E., Pauly, R. B. C., and Schneider, J. G. (2022). Wargaming for international relations research. European Journal of International Relations, 28(1), 83–109.

Meerveld, H., Brighton, H., Lindelauf, R., and Šafář Postma, M. (2026). Effective and responsible use of large language models in strategic wargaming. Journal of Defense Modeling and Simulation, published online 19 April 2026. https://doi.org/10.1177/15485129261438519

Mellers, B., Ungar, L., Baron, J., Ramos, J., Gurcay, B., Fincher, K., Scott, S. E., Moore, D., Atanasov, P., Swift, S. A., Murray, T., Stone, E., and Tetlock, P. E. (2014). Psychological strategies for winning a geopolitical forecasting tournament. Psychological Science, 25(5), 1106–1115.

Miadowicz, I., Kuhl, M., Quinto, D. M., Pitz-Paal, R., and Felderer, M. (2026). From automation to autonomy: A digital twin framework for transparent agent and human collaboration in industrial multi-agent systems. Systems, 14(1), 76. https://doi.org/10.3390/systems14010076

Münker, S., Schwager, N., and Rettinger, A. (2026). Don't trust generative agents to mimic communication on social networks unless you benchmarked their empirical realism. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Volume 1: Long Papers, 1141–1151. https://doi.org/10.18653/v1/2026.eacl-long.51

NATO Allied Command Transformation (2023). NATO Wargaming Handbook. Headquarters, Supreme Allied Commander Transformation, Norfolk, VA. September 2023.

Office of the Director of National Intelligence. Intelligence Community Directive 203: Analytic Standards.

Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST).

Perla, P. P. (1990). The Art of Wargaming: A Guide for Professionals and Hobbyists. Naval Institute Press. (Revised edition: Peter Perla's The Art of Wargaming, edited by J. Curry, History of Wargaming Project, 2012.)

Perla, P. P., and McGrady, E. (2011). Why wargaming works. Naval War College Review, 64(3), 111–130.

Poledna, S., Miess, M. G., Hommes, C., and Rabitsch, K. (2023). Economic forecasting with an agent-based model. European Economic Review, 151, 104306. https://doi.org/10.1016/j.euroecorev.2022.104306

Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., and Schneider, J. (2024). Escalation risks from language models in military and diplomatic decision-making. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT).

Rowe, G., and Wright, G. (1999). The Delphi technique as a forecasting tool: Issues and analysis. International Journal of Forecasting, 15(4), 353–375.

Rozado, D. (2024). The political preferences of LLMs. PLOS ONE, 19(7), e0306621.

Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T. (2023). Whose opinions do language models reflect? Proceedings of the 40th International Conference on Machine Learning (ICML).

Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. (2024). Quantifying language models' sensitivity to spurious features in prompt design. International Conference on Learning Representations (ICLR).

Solopova, V., Skorik, V., Tereshchenko, M., Haidun, A., and Vykhopen, O. (2026). LLMs as strategic actors: Behavioral alignment, risk calibration, and argumentation framing in geopolitical simulations. arXiv:2603.02128.

Tesfatsion, L., and Judd, K. L. (Eds.) (2006). Handbook of Computational Economics, Volume 2: Agent-Based Computational Economics. North-Holland.

Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton University Press.

United Kingdom Ministry of Defence, Development, Concepts and Doctrine Centre (2017). Wargaming Handbook.

United States Government Accountability Office (2023). Defense Analysis: Additional Actions Could Enhance DOD's Wargaming Efforts. GAO-23-105351. April 2023.

United States Nuclear Regulatory Commission. NUREG-2213: Updated Implementation Guidelines for SSHAC Hazard Studies.

Windrum, P., Fagiolo, G., and Moneta, A. (2007). Empirical validation of agent-based models: Alternatives and prospects. Journal of Artificial Societies and Social Simulation, 10(2), 8.

Discuss a simulation for your question

A short conversation is enough to establish whether your question fits the method. If it does not, we will say so.