Balance with evidence. QA with receipts.

Build agents that play the real game. Run them headlessly at scale to pressure-test balance changes, reproduce regressions, and trace every finding back to a replay.

Playthrough Console / full film02:54

The loop

From a balance or QA question to an answer you can defend.

Use our agent architecture, bring your own simulation, or work with us to build both. Playthrough keeps the engine at the center of the evidence chain. It never substitutes a vision proxy.

Author

Hand-crafted, reinforcement-learning, and language-driven agents enter through the same legal-action interface.

Play

Agents run the real engine headlessly across thousands of seeded, version-stamped matches.

Interrogate

Ask one run or an entire tournament a question in plain language. The Console writes queries and re-runs the replay.

Verify

Intervals reveal when a difference is signal rather than noise. Replay and error checks reveal whether the result is trustworthy.

Act

Turn a defensible finding into a balance change, prompt revision, regression fix, or next experiment.

What it opens up

Balance the game players actually play.

Agentic simulation turns balance hypotheses and QA questions into repeatable experiments—before scarce human playtest time begins.

01 / BALANCE

Play the change first

Run a card, unit, economy, or rules change through thousands of seeded matches. See who benefits, what breaks, and whether the movement is signal or noise.

02 / EVALUATION

Grade the agent in-game

Compare scripted, RL, and LLM agents on identical legal moves, with win rates and intervals instead of impressions.

03 / QA

Reproduce what changed

Replay prior scenarios against every build and see exactly which outcomes moved, where the regression began, and how to reproduce it.

04 / TRAINING

Train where you grade

Use the simulator as the RL environment so training progress and gameplay results share a single measurement system.

05 / INSIGHT

Ask in plain language

Let production, QA, design, and ML teams explore the same source of truth without learning a specialist query language.

06 / TRUST

Catch contaminated metrics

Find the fallback, provider error, or engine exception that would otherwise turn a bad run into a confident decision.

Playthrough single-run console showing run summary and arena activity
“A balance result is only useful when you can replay the game that produced it.”Playthrough design principle

Bring your hardest balance question

Let the agents play it out.

Plug in your simulation, build an agentic harness with us, or start with a focused proof of concept around one high-value question.

Start a Playthrough conversation