SYG Background Paper 3

Twins without sensors

What transfers from industrial digital twins to geopolitical and macroeconomic decision-making, and what does not

Shay Hershkovitz, PhD · SYG ConsultingSYG Background Paper 328 min read
Abstract

The term digital twin began as the name for a synchronized virtual replica of a physical asset and now covers autonomous, language-model-driven systems that plan and act. A recent PRISMA review of 69 studies finds that three quarters were published in 2024–2026 and that the evidence is concentrated in manufacturing, energy, and smart cities, where state is measured by sensors, dynamics follow physics, and the plant supplies ground truth within hours. This paper asks what that evidence licenses for the strategic environments in which governments, defense institutions, and macro analysts make decisions, where none of those conditions hold. It argues that a strategic-environment twin is a twin of the experts' documented design of the environment, executed at scale, and never a twin of the environment itself; that the agent-based economics lineage from Farmer and Foley to ASSOCC shows the same calibration bottleneck recurring for thirty years; and that the challenges the industrial literature names, above all validation of emergent behavior and the black-box problem, intensify rather than dissolve when the twin has no sensor feed. It then sets out the design practices the review recommends, shows how a strategic twin applies each of them, and states what such an instrument cannot do.

1. What a digital twin is, and how the term has moved

The systematic review that anchors this paper defines digital twins as "virtual replicas of physical systems that enable real-time monitoring, simulation, and optimization" and traces their evolution "from passive visualization tools to intelligent, autonomous systems capable of independent decision-making" (SYG Consulting, 2026, §1.1). The word replica carries the original meaning. A twin was a model kept in step with a specific pump, turbine, production line, or building by a flow of measurements, so that what happened to the asset happened, with a lag, to the model (Rathore et al., 2021).

The review's own inclusion rules mark how far the term has moved. It excludes "static 3D models, data visualization dashboards, or traditional digital twins that lack autonomous agency and dynamic simulation logic" (SYG Consulting, 2026, §3.1). What qualifies today is a twin that contains agents with goals, reasons about its own state, and proposes or takes actions. Yu, Zhou, and Fu (2025) name the end point of this drift the "digital twin agent": a twin with perception, reasoning, and action capabilities that "operate independently, learn from experience, and collaborate with humans and other agents." Sprock and Sprock (2026) distinguish cognitive from agentic twins along the same axis, Zhou et al. (2026) extend the vocabulary from large language models to world models, and Lehmann et al. (2023) call the combination of agents and twins an "Internet of Digital Twins."

The growth is recent and steep. Of the 69 studies the review retained from 2,309 records, 52 (75 percent) were published in 2024–2026, with 21 in 2024 and 26 in 2025 (SYG Consulting, 2026, §7.4). The domain distribution is the fact a decision-maker should hold onto. Manufacturing and production account for 28 studies (41 percent), energy systems and smart grids for 12 (17 percent), and smart cities and urban systems for 9 (13 percent); robotics, healthcare, supply chain, and transportation make up the remainder (§7.2). By study design, 46 percent are prototype developments, 26 percent simulation studies, and 13 percent case studies of deployments; long-term deployment studies are described as scarce (§7.3, §11.2).

At the far edge of the distribution the word covers systems with no physical referent at all. Nechesov, Dorokhov, and Ruponen (2025) describe "virtual cities" populated by autonomous AI agents, and Palma-Borda, Guzmán, and Belmonte (2025) build an agent-based model of urban crime that they call, with some precision, a digital shadow rather than a twin, since data flows one way, from the city into the model. In economics, Wang et al. (2022) label a regional agent-based model of China a "socioeconomic digital twin." When a vendor says "digital twin" to a policy institution, the word may mean any point on this spectrum, from a synchronized replica of a gas turbine to an agent population that has never touched a sensor. The rest of this paper sorts out which parts of the evidence travel across it.

2. The industrial evidence, and why its settings are bounded

The strongest results in the review come from three families of task. In production scheduling, Siatras et al. (2024) report a bicycle-manufacturing case in which multi-agent scheduling over a twin reduced cycle times and improved on-time delivery against static schedules, and Bakopoulos et al. (2024) train reinforcement-learning scheduling agents inside a twin before deployment. In predictive maintenance, Sakha et al. (2026) propose a safety-constrained framework whose three-valued safety test (safe, unsafe, uncertain) intercepts actions before they reach equipment. In energy, Naeini et al. (2025) report 25–30 percent energy-efficiency improvements and 96.5 percent fault-detection accuracy in a smart building; Natarajan et al. (2025) report 18–30 percent cost reductions and a 15 percent decrease in decision latency for industrial control; Madaminov et al. (2025) report 25–30 percent energy savings without loss of output; and Zhang et al. (2025) report 61 percent higher accumulated returns for agentic control in vehicular edge computing.

The most instructive single result is narrower than any of these. Gill et al. (2025) tested language-model agents on fault handling in a simulated process plant, using an architecture with an action agent, a twin simulation, a validator, and a reprompting loop. In one text-grounded condition a frontier model produced 15 of 15 correct actions. The synthesis that reports this result is explicit that it "remains a narrow laboratory result, not evidence for unrestricted plant autonomy" (Review synthesis, 2026). The architecture matters as much as the score: the language model interprets and plans, the twin supplies state and simulates consequences, and a validator checks every proposed action against the twin before anything is executed. Xia et al. (2025) generalize this division of labor: language models handle interpretation, planning, decomposition, explanation, and coordination, while twins provide synchronized state, engineering semantics, simulation, constraints, and controlled interfaces.

Four properties of these settings account for the results, and none of them holds in a strategic environment. The state is measurable: industrial twins are fed by OPC UA, MQTT, IO-Link, and ROS streams from instrumented equipment (SYG Consulting, 2026, §8.1.4), so the temperature of reactor three has a numerical answer with a known error bar, refreshed in seconds. The dynamics follow physics: physics-informed learning constrains models to respect conservation laws, which improves sample efficiency and generalization (Naeini et al., 2025), and the twin does not have to guess how a heat exchanger responds to a valve change. Ground truth arrives quickly: a schedule meets its deliveries or does not, a fault is cleared or not, and the plant scores the twin within hours, which is what allows Dittler et al. (2022, 2023) to build continuous model-adaptation mechanisms that correct drift between twin and asset as the asset wears. And the loop is closed: the agent acts, the sensors report, the twin updates, and simulation-first development works because physical deployment then supplies a second, decisive round of validation (Bakopoulos et al., 2024; Liang, Wang, Zhang, and Yong, 2025; Pantelidakis and Mykoniatis, 2024).

The review is candid that even here the evidence has limits. Most studies report prototypes or simulations rather than production deployments; there is no standardized set of metrics, so meta-analysis was infeasible; and publication bias likely overrepresents successes (SYG Consulting, 2026, §10). The percentages quoted above are best read as demonstrations that agentic control can improve on static control in bounded systems, and as nothing more general.

3. The strategic environment is a different object

A government deciding among policy options in a contested region, a defense institution stress-testing a posture, or a macro analyst weighing scenarios for a currency crisis works in an environment that fails every one of the four conditions. There is no sensor feed: nobody measures the cohesion of a governing coalition or the intent of an armed movement with an instrument, and the quantities that matter are estimated by analysts who estimate them differently. The state is contested: in a plant, disagreement about the state of a bearing is resolved by inspection, while disagreement about whether a rival's mobilization is defensive is the substance of the strategic problem and may never be resolved. The actors have internal politics: a turbine has no faction that wants it to fail, whereas a state contains a leadership and its security establishment, a movement and its armed wing, a bureaucracy and its political principals, each with documented positions that pull in different directions. Shocks dominate the horizon: over several years, exogenous events from leadership successions to external interventions can matter more than any option under comparison, and their timing is unknown. And ground truth arrives once, and late: the environment cannot be reset and run again with a different option, so by the time a policy has been tested against reality the decision has been taken, and the evidence is a single realization from which no distribution can be read.

The word twin can still be used, with one restriction. What can be twinned is the experts' design of the environment: the actors and factions they identify, the positions they document, the state variables they choose to track, the events they consider possible and the probabilities they assign, and the options they want compared. That design is a written artifact. It can be versioned, frozen, executed many times, and audited. The twin synchronizes with this artifact and with nothing else. It is a twin of a design of the world, run at a scale no human panel can reach, and any claim that it is a twin of the world itself imports the industrial meaning into a setting where the mechanism that earned the word, the sensor feed, is absent.

3.1 The agent-based economics lineage and its recurring bottleneck

Macroeconomics offers the longest-running attempt to build something like a twin of a social system, and its history is a record of one constraint recurring. Farmer and Foley (2009) argued in Nature that the economy needs agent-based modeling because representative-agent models exclude financial instability, distributional effects, and emergent dynamics by construction. The EURACE project had already found, on the compendium's account, that calibration failed for want of harmonized micro-level data across jurisdictions (Dawid et al., 2012). Poledna et al. (2023) built a fully calibrated agent-based model of the Austrian economy validated against national statistics, and Wang et al. (2022) did the same at regional scale for China, both preserving the distributional resolution that aggregate models lose. Gill et al. (2024) report a one-to-one-scale model of Japan with 130 million agents whose output quality degraded when macro aggregates had to stand in for micro-level data, and state that "calibration requires micro-level data unavailable in most studies." The first use of such a model as a live government decision-support instrument came during the pandemic: Dignum et al. (2020) describe ASSOCC, a combined health, social, and economic agent-based simulation used to support government decisions on COVID-19 measures, and the compendium's reading of the episode is that the limiting factor was data rather than modeling technique. From the Santa Fe models (Tesfatsion and Judd, 2006; Epstein, 1999) to the present, method has repeatedly outrun the data available to calibrate it (Agentic Economic Simulation Research Compendium, 2026, §3).

Language-model agents change the shape of this constraint without removing it. Park et al. (2023) showed that memory-equipped generative agents produce emergent social behavior, and Ghaffarzadegan et al. (2024) formalized generative agent-based modeling, in which the language model replaces the hand-coded behavioral rules that had been the modeling bottleneck since the 1990s. Li and Wu (2025) then documented, across 22 studies, the over-control paradox: prompts prescriptive enough to make agents behave as intended can predetermine the outcome and suppress the emergence that made agent-based modeling worth doing. Prompt design is now a scientific variable, and the calibration problem has moved from the data needed to fit rules to the evidence needed to ground personas and the discipline needed to keep that grounding from becoming a script.

3.2 The structural limits of the static and aggregate tools

The instruments most policy institutions use today were built for a different question. The regional input-output packages common in budget offices (REMI, IMPLAN) rest on multiplier tables: a user specifies a shock and the model returns multiplied aggregate effects across industries (Miller and Blair, 2009). They assume one household type and one firm per industry, so the distributional question of who gains and who pays is invisible by construction; they model a snapshot equilibrium with no behavioral feedback; and practitioners report that different analysts obtain materially different results from the same model and data (Agentic Economic Simulation Research Compendium, 2026, §7.1). The dynamic stochastic and computable general equilibrium models used at national level are rigorous, and their representative-agent assumption is a deliberate choice suited to macro questions, which is why Baccini et al. (2024) find heterogeneous-agent models systematically better on distributional impact assessment. None of these tools contains actors. A question of the form "how does a security establishment respond if the coalition adopts option B and an external power intervenes in year two" has no variable in an input-output table and no representative agent in a general-equilibrium model. This is the gap the agent-based lineage and its language-model extension are meant to fill, and the gap into which the word twin has been pulled.

Table 1 sets out the two objects side by side.

Table 1. Industrial twin and strategic-environment twin compared

Attribute Industrial agentic twin Strategic-environment twin
State observability Sensor streams refreshed in seconds; known error bars No instrument; state estimated by experts from documents; contested
Ground truth Supplied by the asset within hours Arrives once, after the decision, as a single realization
Update cycle Continuous synchronization; adaptation corrects drift Versioned redesign: the design pack is re-elicited and re-frozen
Agent grounding Physics, engineering semantics, equipment constraints Documented positions, historical anchors, red lines, behavioral notes
Validation route Simulation-first, then hardware-in-the-loop and deployment Input, process, and pattern-oriented output validation; historical-anchor checks; no out-of-sample test before the decision
Primary output An action or control setting, executed A distribution of trajectories per option, with frequencies and a traceable record
Failure mode Reality gap: the agent trained in the twin fails on the plant Believability mistaken for validity: a storyline produced by a referee rule or substrate prior is read as a finding
Human role Oversight of autonomy; override on unsafe actions Authorship of the design; decision at every gate; ownership of the conclusion

4. The challenges the industrial literature names, and how they intensify

The review's third research question catalogs the obstacles its 69 studies report (SYG Consulting, 2026, §8.3), and its list of research gaps repeats most of them (§11.2). Reading that list from the strategic side, every item is harder, and two of them change character.

Validation of emergent behavior. The review states that validating agentic twins "presents unique challenges due to emergent behavior, non-deterministic decision-making, and complex agent interactions," that verification methods designed for deterministic systems are insufficient, and that reproducibility suffers from stochastic algorithms and incomplete documentation (SYG Consulting, 2026, §8.3.3; Marah and Challenger, 2024; Cosenza et al., 2021; Gao et al., 2024). In industry the difficulty is partly relieved by the plant, which scores the outcome. In a strategic environment the plant is missing, and the difficulty becomes the central problem. Larooij and Törnberg (2025a, 2025b) conclude that validation is the central unsolved issue for generative social simulation and that believability does not establish empirical realism, causal validity, or operational safety. Münker et al. (2026) show that generative agents should not be trusted to mimic communication unless their empirical realism has been benchmarked, and that performance shifts across languages, populations, and settings. A simulation whose trajectories read plausibly to the experts has passed a necessary test and no sufficient one. The agent-based modeling tradition supplies the vocabulary for what can be done instead: input, process, and output validation kept distinct (Windrum, Fagiolo, and Moneta, 2007), output validation against several patterns at once (Grimm et al., 2005), and an explicit statement that prediction is not among the model's purposes (Epstein, 2008).

Black-box models and explainability. The review reports that black-box models "lack explainability, hindering trust and adoption," and that natural-language interfaces introduce risks of misinterpretation, hallucination, and unintended action (SYG Consulting, 2026, §8.3.6; Miadowicz et al., 2026; Shahzad, Ferreira, and Deschamps, 2025). In a plant, an opaque agent that reliably reduces energy use can be tolerated because the outcome is measured. In a strategic setting the outcome is not measured before the decision, so the reasoning is the only thing a reviewer can inspect, and a decision that cannot be traced to the persona and sources that produced it is evidence of nothing.

Human-agent collaboration and adaptive autonomy. The review finds that "excessive autonomy may lead to unsafe actions, while excessive oversight negates efficiency benefits," and points to adaptive autonomy frameworks that adjust agent independence to context and confidence (Miadowicz et al., 2026; Heininger, Jost, and Stary, 2024). Li and Wu's (2025) over-control paradox is the same tension seen from the prompt: control the agents tightly and the simulation returns the designer's assumptions; release them and the substrate's priors take over. No established typology places humans inside a multi-agent language-model simulation, and no controlled study has seated human players alongside language-model agents, so any human-in-the-loop configuration offered for strategic use is an adaptation and should be labeled as one.

Reality gap. In robotics the sim-to-real gap is measured on deployment and narrowed by domain randomization and online adaptation (Liang, Wang, Zhang, and Yong, 2025; Miao, Ge, Li, and Guo, 2024; Zhang et al., 2025). A strategic twin's reality gap can never be measured this way, because the twin is never deployed against the world. What can be measured is the gap between the twin and the design: whether the engine executes the frozen artifact as written, and whether the agents behave recognizably against dated historical episodes with period-appropriate information.

Substrate bias. The review lists algorithmic bias among under-explored issues (SYG Consulting, 2026, §11.2). For language-model agents in political settings the evidence is specific: Rivera et al. (2024) found that such agents tend toward escalation in simulated crisis decision-making, and Kim et al. (2025) show that errors across models are correlated and that the correlation rises with capability, which limits what a multi-vendor ensemble can correct. In a plant a biased agent is caught by the sensors. In a strategic twin it is caught only if someone built the instrument to look for it.

Security. Industrial twins face data manipulation, denial of service, and adversarial inputs (Wen et al., 2024; Kang, 2026). A strategic twin adds a second exposure: the design pack is itself a record of an institution's assessments of named actors, and the audit log holds every reasoning trace. Provider allow-lists with fallbacks disabled, pinned model versions, regionally hosted endpoints where required, and a self-contained record the client holds are the responses; none removes the exposure.

Cost and expertise. The review notes that personnel spanning AI, simulation, domain knowledge, and systems integration are scarce (Latsou et al., 2024; Rathore et al., 2021). A strategic twin depends on a scarcer combination: domain experts willing to submit their assessments to structured elicitation, and a design team that can turn those assessments into an executable artifact without scripting the result.

5. Design implications: the review's practices, applied to a strategic twin

The review's fourth research question distills best practices from the 69 studies (SYG Consulting, 2026, §8.4, §11.3). Each of them has a counterpart in a strategic-environment twin, and the counterpart is usually stricter, because the sensor feed that would otherwise catch a design error is absent. SYG's method is described in full in White Paper 2; this section names only the correspondences.

Architecture: three surfaces on one design packARCHITECTUREThree surfaces, one design packAuthoring surfaceWeb application for the workshopPre-loaded design contentSilent review per entityDisagreement registerFreeze validator, hash, tagSimulation enginePython serviceGraph-structured tick loopDecision schema and validatorReferee and ensemble adjudicatorOrchestrator, gates, spend capAnalysis surfaceTrajectory database and toolingBrowsable simulation recordPilot checks and bias batteryOutcome-question readerSix analytical lensesEngine configurationmodel pins, tiers, concurrency, spend ceilingDESIGN PACKOne versioned JSON document carries all domain content; the platform code is engagement-agnostic.Actors · factions · state variables · dimensions · options · events · sources · disagreement register · glossaryFrozen at the workshop: validator pass, SHA-256 content hash, tagged version. Stable identifiers carry into the record.Three layers (inputs, engine, outputs); the same three surfaces run every engagement. Frozen pack flows in; trajectories flow out.
Figure 1.The same three surfaces run every engagement; only the design pack changes.

Modularity and separation of concerns. The review recommends encapsulating functionality in independent components and decoupling perception, reasoning, and action (Dittler et al., 2022; Lehmann et al., 2023; Yu, Zhou, and Fu, 2025). The strategic counterpart is a single versioned design pack holding all domain content (actors, factions, state variables, options, events, sources, and the register of expert disagreements) and an engine, authoring surface, and analysis layer that are engagement-agnostic. The pack is frozen only when a validator passes and a content hash is computed; any later change is a new version with a logged reason. This lets the design be reviewed and cited independently of the software that runs it.

Hybrid reasoning. The review favors combining fast deterministic control with slower deliberative planning (Marah and Challenger, 2024; Miadowicz et al., 2026). The analog is hybrid adjudication: every state variable declares whether its transitions are ruled, judged by a cross-vendor ensemble of language models, or mixed, and the split is a workshop decision recorded in the design. The first application's pilot showed why the declaration matters. A transition rule with no inertia, decay, or bounds produced one-way movement in 83.3 percent of state steps (315 up, 63 down) and a storyline that repeated under every option; the corrected referee carries inertia, decay toward baseline, and expert-authored plausibility bounds. The ensemble functions as a diagnostic: agreement is recorded as consensus, disagreement is routed to experts, and adding scoring anchors to one variable reduced inter-judge disagreement from 31.8 to 1.0 percent.

Simulation-first and incremental development. The review recommends testing in the twin before deployment and adding complexity progressively (Bakopoulos et al., 2024; Pantelidakis and Mykoniatis, 2024; Latsou et al., 2024). Without a physical deployment to follow, the strategic version is a pilot gate: the first application ran 26 pilot runs, 14 on the frozen design, and applied a pre-registered battery of ten bias families to them before any outcome was read. Production does not start until the calibration gate is cleared.

Transparency and explainability. The review's call for interpretable models and natural-language explanations (Miadowicz et al., 2026; Shahzad, Ferreira, and Deschamps, 2025) becomes a requirement of the record rather than of the interface. Every model call is logged with its model identifier, full input, full output, and reasoning; every state change is attributable to a rule or to a specific adjudication with a causal field naming the actions and events that produced it; and every claim in a report traces along the chain claim, run, tick, decision, persona, source. Empty causal fields cannot be reconstructed after the fact, so the adjudication record carries one by design.

Safety-constrained frameworks. The review's three-valued safety tests intercept unsafe actions before execution (Sakha et al., 2026). In a strategic twin nothing is executed, so the constraint serves a different purpose. Options are enforced as soft constraints in the prompt rather than as hard filters, so that attempts at forbidden actions can be observed and counted; in the pilot the count was zero across 15,792 decisions, and 1.9 percent of decisions carried a red-line flag routed for expert reading.

Human-centered design and collaborative autonomy. The review holds that agents should augment human decision-making rather than replace it, and Masson and Villeneuve (2025) argue that twins must be designed around the human decision context. The counterpart is a structured elicitation protocol adapted from IDEA (Hemming et al., 2018): experts review every design entity before the workshop (393 responses in the first application: 148 confirm, 164 not my area, 63 revise, 18 flag), their unresolved disagreements become the model's sweep parameters, and four gates (design frozen, pilot passed, calibration cleared, validation review) place the decision with the expert group.

Continual adaptation. Industrial twins adapt continuously to the asset (Dittler et al., 2023). A strategic twin has no stream to adapt to, and the equivalent is deliberate: when the experts' understanding changes, the design is re-elicited, re-versioned, and re-frozen, with the change log recording what changed and why.

Measurement of bias and a method note. A strategic twin makes bias auditable by pre-registering the families to be measured before the pilot data is opened, running the battery, and disclosing what survives calibration in a method note that states the claim boundary. In the first application the largest distortions came from the state-transition rule and from the design of population-opinion proxies rather than from the substrate's priors on the domain, which argues for measuring rather than assuming where bias will appear.

Table 2. Challenges named in the industrial literature and the strategic-twin response

Challenge (review section) How it intensifies in a strategic environment Strategic-twin response
Validation of emergent behavior (§8.3.3, §11.2) No plant scores the outcome; believability is all that is visible Input, process, and output validation kept distinct; pattern-oriented checks; historical-anchor exercises; pre-registered bias battery; prediction stated as not a purpose
Black-box models (§8.3.6) Reasoning is the only inspectable evidence Full audit log per call; causal field on every adjudication; claim-to-source traceability chain
Human-agent collaboration (§8.3.6, §8.4.5) No typology and no controlled study for humans alongside language-model agents Four-mode spectrum offered as an adaptation; expert gates as default; adjudicator override during calibration
Reality gap (§8.3.5) Never measurable against the world Measured against the design: engine executes the frozen pack; agents checked against dated episodes
Over-control paradox (Li and Wu, 2025) Tight prompts return the designer's assumptions Narrative objectives rather than utility functions; restraint written as a costed option; every prompt rewrite measured on fixed seeds
Substrate bias and escalation (Rivera et al., 2024; Kim et al., 2025) No sensor catches it Persona grounding in registered sources; cross-vendor ensemble as diagnostic; residual escalation measured and disclosed
Security and privacy (§8.3.4) The design pack is itself sensitive assessment Provider allow-list with fallbacks disabled; pinned versions; regionally hosted endpoints where required; client holds the record
Cost and expertise (§8.3.7) Scarce expert time; no asset value to compare against Silent review before the workshop; run plan extended only where frequencies have not converged

6. What a strategic-environment twin cannot do, and what it can

The industrial results in Section 2 should not be read across. A strategic twin does not predict. It has no live synchronization with the world, because there is no feed to synchronize with. It is not a validated forecasting instrument: no published multi-agent language-model political simulation has made a calibrated out-of-sample forecast, and the one serious historical backtest in the literature produced war in every run, which is a warning about substrate escalation propensity and a reminder that restraint mechanisms must be shown to have any effect at all. Validity is open on both sides: nothing in the method establishes that the experts' design of the environment is correct, and nothing establishes that the language models execute that design without importing priors of their own. Scale improves the reliability of the frequencies read from the runs; it does not make the generative model more correct. The evidence base is Western-normed and English-language, which argues for a heavier expert-calibration burden for non-Western actors.

Within those limits, the instrument does several things the human alternatives cannot. It maps the distribution: for each option, hundreds of trajectories against matched draws of events, from which frequencies, conditional probabilities, and tipping points can be read, where a narrative assessment yields one trajectory and a wargame one per play. It surfaces conditional structure: which events, when they fire, change the frequency of an outcome, and which faction decisions precede which state changes. It stress-tests options against shocks by forcing specific events at fixed ticks in a fixed share of runs. It makes bias auditable: the same ten families can be measured on every design version, and what survives calibration is stated. It resets, which is the single categorical advantage over the human wargame; every other advantage is one of degree. And it leaves the client a database, in a self-contained file, in which every run, tick, decision, message, event, and adjudication is a row that can be queried against questions the commissioning experts have not yet asked; the first application's pilot alone produced 15,792 decisions, 26,836 messages, 38,674 events, 5,544 state observations, and 3,067 adjudications, each traceable to the design version that produced it.

The recommended practice, where both instruments are available, is sequence: the simulation to map the distribution and surface the conditional structure, then the expert panel or wargame to interrogate the trajectories that matter.

7. Conclusion

The digital twin literature has moved, in a few years, from synchronized replicas of assets to agent populations that plan and act, and three quarters of its evidence dates from 2024 onward. That evidence is real and bounded. It comes from settings where the state is measured, the dynamics follow physics, the plant supplies ground truth quickly, and the loop is closed, and the review that collects it is candid that even there, long-term deployments are scarce and validation of emergent behavior remains unsolved. None of the four conditions holds in a strategic environment. What can be built there is a twin of the experts' documented design of the environment, executed at a scale and with a completeness of record that no human panel can match, and every design practice the industrial literature recommends has a stricter counterpart in that setting, because the sensor feed that would otherwise catch an error is missing. A decision-maker who hears "digital twin" and "agentic AI" from a vendor should ask which point on the spectrum is meant, what the twin synchronizes with, how validity is tested before the decision, and where the humans decide. The answers for a strategic twin are: a twin of a design; with the design pack; by input, process, and pattern-oriented output validation with bias measured and disclosed; and at every gate.

References

Agentic Economic Simulation Research Compendium (2026). Agentic Economic Simulation and Digital Twins: Research Compendium. Academic literature, market data, patent and competitive analysis, funding, and full references. Compiled March 2026. Internal document.

Baccini, A., et al. (2024). DSGE versus heterogeneous agent models: A systematic comparison. Journal of Economic Surveys. (bibliographic details unconfirmed; the DOI given in the source, 10.1111/joes.70014, belongs to a different article, and the only Baccini et al. article located in this journal on DSGE and agent-based models is "Does cross-fertilization occur in recent macroeconomics?", 39(4), 1758–1794, 2025, doi:10.1111/joes.12674, whose title and findings do not match the entry as sourced)

Bakopoulos, E., Siatras, V., Mavrothalassitis, P., Nikolakis, N., and Alexopoulos, K. (2024). Digital-twin-enabled framework for training and deploying AI agents for production scheduling. In Artificial Intelligence in Manufacturing, pp. 147–179. doi:10.1007/978-3-031-46452-2_9

Cosenza, B., et al. (2021). Easy and efficient agent-based simulations with the OpenABL language and compiler. Future Generation Computer Systems, 116, 61–75. doi:10.1016/J.FUTURE.2020.10.014

Dawid, H., Gemkow, S., Harting, P., van der Hoog, S., and Neugart, M. (2012). The Eurace@Unibi model: An agent-based macroeconomic model for economic policy analysis. Bielefeld Working Papers in Economics and Management No. 05-2012, Bielefeld University.

Dignum, F., et al. (2020). Analysing the combined health, social and economic impacts of the coronavirus pandemic using agent-based social simulation. Minds and Machines, 30, 177–194. doi:10.1007/s11023-020-09527-6

Dittler, D., Lierhammer, P., Braun, D., Müller, T., Jazdi, N., and Weyrich, M. (2022). An agent-based realisation for a continuous model adaption approach in intelligent digital twins. arXiv:2212.03681. doi:10.48550/arXiv.2212.03681

Dittler, D., Lierhammer, P., Braun, D., Müller, T., Jazdi, N., and Weyrich, M. (2023). A novel model adaption approach for intelligent digital twins of modular production systems. IEEE International Conference on Emerging Technologies and Factory Automation (ETFA), pp. 1–8. doi:10.1109/etfa54631.2023.10275384

Epstein, J. M. (1999). Agent-based computational models and generative social science. Complexity, 4(5), 41–60.

Epstein, J. M. (2008). Why model? Journal of Artificial Societies and Social Simulation, 11(4), 12.

Farmer, J. D., and Foley, D. (2009). The economy needs agent-based modelling. Nature, 460, 685–686. doi:10.1038/460685a

Gao, C., et al. (2024). Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanities and Social Sciences Communications, 11(1). doi:10.1057/s41599-024-03611-3

Ghaffarzadegan, N., et al. (2024). Generative agent-based modeling: An introduction and tutorial. System Dynamics Review, 40(1). doi:10.1002/sdr.1761

Gill, A., Lalith, M., Hori, M., and Ogawa, Y. (2024). Analysis of postdisaster economy using high-resolution disaster and economy simulations. Risk Analysis, 45(6), 1254–1270 (published online October 2024). doi:10.1111/risa.17662

Gill, M. S., Vyas, J., Markaj, A., Gehlhoff, F., and Mercangöz, M. (2025). Leveraging LLM agents and digital twins for fault handling in process plants. 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA), pp. 1–8. doi:10.1109/etfa65518.2025.11205597

Grimm, V., Revilla, E., Berger, U., Jeltsch, F., Mooij, W. M., Railsback, S. F., Thulke, H.-H., Weiner, J., Wiegand, T., and DeAngelis, D. L. (2005). Pattern-oriented modeling of agent-based complex systems: Lessons from ecology. Science, 310(5750), 987–991.

Heininger, R., Jost, T. E., and Stary, C. (2024). Developing and operating digital process twins while preserving autonomy in CPS. IEEE Conference on Business Informatics (CBI), pp. 50–59. doi:10.1109/cbi62504.2024.00016

Hemming, V., Burgman, M. A., Hanea, A. M., McBride, M. F., and Wintle, B. C. (2018). A practical guide to structured expert elicitation using the IDEA protocol. Methods in Ecology and Evolution, 9(1), 169–180.

Kang, S. (2026). Blockchain-secured multi-agent digital twin framework for distributed manufacturing with fuzzy PID control. IEEE Communications Standards Magazine, pp. 1–9. doi:10.1109/mcomstd.2025.3648845

Kim, E. M., Garg, A., Peng, K., and Garg, N. (2025). Correlated errors in large language models. Proceedings of the 42nd International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, 267, 30038–30066. arXiv:2506.07962.

Larooij, M., and Törnberg, P. (2025a). Validation is the central challenge for generative social simulation: A critical review of LLMs in agent-based modeling. Artificial Intelligence Review, 59, 15. doi:10.1007/s10462-025-11412-6

Larooij, M., and Törnberg, P. (2025b). Do large language models solve the problems of agent-based modeling? A critical review of generative social simulations. arXiv:2504.03274. doi:10.48550/arXiv.2504.03274

Latsou, C., Ariansyah, D., Salome, L., Erkoyuncu, J. A., Sibson, J., and Dunville, J. (2024). A unified framework for digital twin development in manufacturing. Advanced Engineering Informatics, 62, 102567. doi:10.1016/j.aei.2024.102567

Lehmann, J. T., et al. (2023). The anatomy of the Internet of Digital Twins: A symbiosis of agent and digital twin paradigms enhancing resilience (not only) in manufacturing environments. Machines, 11(5), 504. doi:10.3390/machines11050504

Li, Z., and Wu, Q. (2025). Let it go or control it all? The dilemma of prompt engineering in generative agent-based models. System Dynamics Review, 41(3), e70008. doi:10.1002/sdr.70008

Liang, Z., Wang, J., Zhang, T., and Yong, X. (2025). DTTF-Sim: A digital twin-based simulation system for continuous autonomous driving testing. Sensors, 25(11), 3447. doi:10.3390/s25113447

Madaminov, B., Saidmurodov, S., Saitov, E., Jumanazarov, D., Alsayah, A. M., and Zhetkenbay, L. (2025). Multi-objective optimization framework for energy efficiency and production scheduling in smart manufacturing using reinforcement learning and digital twin technology integration. International Journal of Industrial Engineering and Management, 16(3), 283–295. doi:10.24867/ijiem-389

Marah, H., and Challenger, M. (2024). Adaptive hybrid reasoning for agent-based digital twins of distributed multi-robot systems. Simulation. doi:10.1177/00375497231226436

Masson, D., and Villeneuve, E. (2025). FIRE: A human-centered framework for digital twin design. Systems Engineering, 29(1), 34–48. doi:10.1002/sys.70014

Miadowicz, I., Kuhl, M., Quinto, D. M., Pitz-Paal, R., and Felderer, M. (2026). From automation to autonomy: A digital twin framework for transparent agent and human collaboration in industrial multi-agent systems. Systems, 14(1), 76. doi:10.3390/systems14010076

Miao, B., Ge, S., Li, Y., and Guo, Y. (2024). Method of motion planning for digital twin navigation and cutting of shearer. Sensors, 24(18), 5878. doi:10.3390/s24185878

Miller, R. E., and Blair, P. D. (2009). Input-Output Analysis: Foundations and Extensions. Cambridge University Press.

Münker, S., Schwager, N., and Rettinger, A. (2026). Don't trust generative agents to mimic communication on social networks unless you benchmarked their empirical realism. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Volume 1: Long Papers, 1141–1151. doi:10.18653/v1/2026.eacl-long.51

Naeini, H. K., et al. (2025). PINN-DT: Optimizing energy consumption in smart building using hybrid physics-informed neural networks and digital twin framework with blockchain security. arXiv:2503.00331. doi:10.48550/arxiv.2503.00331

Natarajan, V. P., Malarvizhi, N., Prabu, S., G, R., and Patel, M. (2025). Implementing digital twins in intelligent control systems for real-time energy optimization. IEEE IACIS 2025, pp. 1–7. doi:10.1109/iacis65746.2025.11211124

Nechesov, A., Dorokhov, I., and Ruponen, J. (2025). Virtual cities: From digital twins to autonomous AI societies. IEEE Access. doi:10.1109/access.2025.3531222

Palma-Borda, J., Guzmán, E., and Belmonte, M. (2025). A digital shadow for modeling, studying and preventing urban crime. arXiv:2501.04435. doi:10.48550/arxiv.2501.04435

Pantelidakis, M., and Mykoniatis, K. (2024). Simulation aspects of a generic digital twin ecosystem for computer numerical control manufacturing processes. Proceedings of the Winter Simulation Conference, pp. 1599–1610. doi:10.1109/wsc63780.2024.10838983

Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST). arXiv:2304.03442

Poledna, S., Miess, M. G., Hommes, C., and Rabitsch, K. (2023). Economic forecasting with an agent-based model. European Economic Review, 151, 104306. doi:10.1016/j.euroecorev.2022.104306

Rathore, M. M., Shah, S. A., Shukla, D., Bentafat, E., and Bakiras, S. (2021). The role of AI, machine learning, and big data in digital twinning: A systematic literature review, challenges, and opportunities. IEEE Access, 9, 32030–32052. doi:10.1109/ACCESS.2021.3060863

SYG Consulting (2026). Agentic simulations and digital twins: A systematic literature review of technical stacks, applications, and best practices. PRISMA 2020-compliant review, 69 included studies, 2020–2026. Internal working paper.

Review synthesis (2026). Narrowed academic synthesis on LLM agents and digital twins: architecture and validation. Internal note, September 2026.

Rivera, J.-P., Mukobi, G., Reuel, A., Lamparth, M., Smith, C., and Schneider, J. (2024). Escalation risks from language models in military and diplomatic decision-making. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT).

Sakha, R., Venkat, S., Sekaran, R., Vaithiyanathan, R., Jeeva, S., and Girisha, G. S. (2026). A safety-constrained agentic digital twin framework for predictive maintenance in smart manufacturing systems. 2026 International Conference on Data Science and Business Systems (ICDSBS), pp. 1–5. doi:10.1109/icdsbs69077.2026.11634862

Shahzad, N., Ferreira, W. D., and Deschamps, F. (2025). Cognitive digital twins: A state-of-the-art review. SSRN. doi:10.2139/ssrn.5085883

Siatras, V., Bakopoulos, E., Mavrothalassitis, P., Nikolakis, N., and Alexopoulos, K. (2024). Production scheduling based on a multi-agent system and digital twin: A bicycle industry case. Information, 15(6), 337. doi:10.3390/info15060337

Sprock, A. S., and Sprock, O. S. (2026). Conceptualizing cognitive and agentic digital twins. International Multidisciplinary Journal of Emerging Technologies and Applications, 1(2), 16–40. doi:10.67294/knpyhb26

Tesfatsion, L., and Judd, K. L. (Eds.) (2006). Handbook of Computational Economics, Volume 2: Agent-Based Computational Economics. North-Holland.

Wang, Z., et al. (2022). A socioeconomic digital twin model for urban economic system. Complexity, 2022, 9805809. doi:10.1155/2022/9805809

Wen, J., Kang, J., Niyato, D., Zhang, Y., and Mao, S. (2024). Sustainable diffusion-based incentive mechanism for generative AI-driven digital twins in industrial cyber-physical systems. arXiv:2408.01173. doi:10.48550/arxiv.2408.01173

Windrum, P., Fagiolo, G., and Moneta, A. (2007). Empirical validation of agent-based models: Alternatives and prospects. Journal of Artificial Societies and Social Simulation, 10(2), 8.

Xia, Y., Jazdi, N., and Weyrich, M. (2025). An architecture for integrating large language models with digital twins and automation systems. 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA), pp. 1–8. doi:10.1109/etfa65518.2025.11205636

Yu, J., Zhou, J., and Fu, J. (2025). From digital twin to digital twin agent. IEEE e-CARGO 2025, pp. 111–116. doi:10.1109/e-cargo65996.2025.11139170

Zhang, R., et al. (2025). Toward edge general intelligence with agentic AI and agentification: Concepts, technologies, and future directions. arXiv:2508.18725. doi:10.48550/arxiv.2508.18725

Zhou, R., et al. (2026). Digital Twin AI: Opportunities and challenges from large language models to world models. arXiv:2601.01321. doi:10.48550/arxiv.2601.01321

Discuss a simulation for your question

A short conversation is enough to establish whether your question fits the method. If it does not, we will say so.