Build the sim, reuse the intelligence
Eval bots and player-facing AI often diverge into separate stacks that cannot transfer learning. A shared legal-action loop over the real engine lets simulation and player modes improve together.
The balance agents improved every sprint. Almost none of that improvement reached the co-pilot.
That split is common inside studios. One group builds bots for balance, regression, and training. Another builds opponents, co-pilots, and new ways for players to issue commands. Both work against the same rules. Both get called AI. Yet they often act through entirely different interfaces.
The organizational split can make this feel natural. Evaluation sits with QA and ML. Player-facing modes sit with design and client teams. But the cost goes beyond duplicated code. The intelligence does not compound.
Headless runs do not improve what players touch. Player sessions do not sharpen the agents used to stress the economy. Every advance has to be rebuilt for the other stack.
The alternative is shared architecture built around one legal-action loop over the real engine. Simulation sits at one end, player experience at the other, and the same legal actions support both.
Why the stacks stop learning from each other
Evaluation and player-facing AI began with different constraints.
Evaluation needed scale and trustworthy results. Teams used vision proxies, scripted bots, or simplified simulations to run thousands of games without driving the full client.
Player-facing AI needed responsiveness and usability. Teams connected models to UI flows, hotkeys, or special-case parsers built for the shipped experience.
Each decision could be reasonable on its own. Together, they produced agents that do not speak the same language as the game or each other.
An evaluation bot working from pixels or a reduced state representation may never have to propose the moves available to a player. A player-facing agent clicking through menus may not exercise the constraints used by the rules engine.
Their measures of success also differ. Evaluation cares about win rates and whether a fallback contaminated the run. Player-facing modes care about latency and whether language maps to something the engine will accept.
Once their state and action interfaces diverge, their learning becomes difficult to transfer.
You cannot promote an evaluation policy into a co-pilot if it never emitted valid engine actions. You cannot use live player transcripts to improve evaluation agents if those agents never saw the same action space. Improvements to planning, validation, and repair stay trapped in the stack where they began.
The studio ends up paying twice for intelligence that should compound.
One legal-action loop, two modes of use
Two technical constraints have loosened.
Models can plan over structured state when the action space is discrete and constrained. Inference can be cheap and fast enough for an agent to inspect the state, plan, check legal moves, and repair failures inside a turn budget players will tolerate.
That combination makes a shared interface practical.
The shared-loop approach starts with a simple constraint: every agent acts through the real engine’s legal-action surface.
Whether the input is a simulation goal or a player command, the loop remains the same:
- Read the command or goal.
- Inspect the current state.
- Plan.
- Break the plan into atomic actions.
- Validate each action.
- Repair failures.
- Commit.
The engine stays the authority.
The agent can propose a move, but it cannot decide that an illegal move succeeded. It does not replace the ruleset with a second interpretation of the game. It works through the same constraints the shipped game enforces.
You also do not substitute a vision proxy for what happened. The engine provides the authoritative account. Statistics show how the run went, while errors and traces show whether those results can be trusted.
That discipline matters in high-scale evaluation. It also matters when a player says, “Push the left flank,” and expects a legal sequence rather than a confident guess.
The input, schedule, and surface can change between simulation and live play. The underlying legal-action loop does not.
Map the game before building the feature
A shared loop depends on exposing the game in terms an agent can reason over without creating a parallel ruleset.
That means mapping the relevant state and the legal actions the engine already enforces. Depending on the game, that can include cards, positioning, movement, terrain, status, abilities, sequencing, and every other constraint that determines what can happen next.
This work is less visible than a conversational interface or adaptive opponent. It is also what allows player-facing experiences to reuse the simulation investment.
Without a shared map, each team creates its own abstraction. Evaluation gets one representation. Player AI gets another. Both start close to the game, then drift as the rules change.
With a shared map, scripted, reinforcement learning, and language-model agents can operate over identical legal moves. Their behavior becomes comparable. Teams can measure win rates with intervals and re-run prior games against a new engine. They can keep a path from every claim back to the game that produced it.
Provider errors, fallbacks, and engine exceptions must be first-class failures. If they contaminate a metric, that cannot be hidden behind an aggregate result. Invalid actions and repairs also become part of the evidence rather than implementation details discarded after the run.
Once that foundation exists, the same planning loop can support several player-facing forms:
- Conversational control
- An onboarding coach
- A sub-unit commander that eats grind
- An adaptive opponent
Language becomes an input layer. Text or voice is translated into constraint-aware actions, while the core game remains intact. Players gain another way to access its depth without requiring a separate interpretation of the rules.
Arena Tactics: one loop across a vast state space
We explored this architecture with Arena Tactics.
Cards, positioning, movement, terrain, status effects, abilities, and sequencing combine to produce possible game states in the septillions. We did not want AI to make that ruleset smaller. We wanted one loop that could stress it at scale and help players express intent through the same legal actions.
The process was explicit: read the command, inspect state, plan, break the plan into atomic actions, validate, repair, and commit.
In simulation, scripted, reinforcement learning, and language-model agents can act through the same legal moves. Results can be compared, prior games can be run against a changed engine, and unexpected outcomes can be traced back to the actions and rules that produced them.
In a player-facing mode, the loop can support conversational play, adaptive opponents, and intelligent co-play. The agent’s role changes, but its relationship to the engine does not.
Language also abstracted some of the game’s most interface-heavy elements. That made it possible to build a mobile companion experience in roughly a month. The game did not become smaller to fit the device. The interface became more expressive.
The architecture still had to meet the demands of live play. Early versions took roughly fourteen seconds to respond. The system worked technically, but it was too slow as a player experience.
After moving to a multi-agent architecture and faster inference, AI-powered modes grew from less than 5 percent of gameplay to more than 60 percent. Speed did not create the shared loop. It made that loop practical for players as well as batch evaluation.
Playthrough and Playable are two returns on the same work
Playthrough and Playable reflect the two uses of this architecture.
Playthrough is high-scale simulation and interrogation. It uses the shared loop to play through the real engine, examine outcomes, and investigate the behavior of the ruleset.
Playable is the player-facing conversational and co-play surface. It uses that same loop to translate player intent into valid actions.
They are two returns on one architecture, not two separate AI stacks connected after the fact.
Improvements to planning can strengthen both simulation and player interaction. Better validation can make evaluation runs more trustworthy and player commands more reliable. Better repair can reduce failed simulations and recover when a requested sequence is not legal.
The feedback can move in both directions. High-scale play can reveal situations that improve player-facing behavior. Live player commands can expose ambiguities and failure cases that harden the agents and validators used in simulation.
This does not require every use case to run the same model. Models can change as cost, speed, and quality change. Transfer is possible because both modes meet at the same boundary: the real engine’s legal actions.
Build once at the boundary that matters
For a studio CTO or AI lead, the important question is not whether evaluation and player-facing AI use the same model.
It is whether they act through the same legal surface.
Models can change. Player experiences can change. Simulation schedules can change. If state, actions, validation, and engine authority remain shared, improvements can continue to transfer across those changes.
If those boundaries differ, the studio keeps maintaining two systems that look related on an architecture diagram but cannot reuse each other’s learning.
If they are shared, the simulation system becomes more than an internal tool. It becomes the foundation for intelligence players can use.
For teams that already have separate evaluation bots and player AI, the cheapest place to start is usually the legal-action boundary, not another model swap.
If you want to assess whether your evaluation stack and player-facing AI can share that boundary, or just want to walk through how Playthrough and Playable sit on the same loop, reach out. We can look at the architecture with you.
