Forensic QA, with receipts.

Build agents that play the real game. Run them headlessly at scale. Ask the results anything. Trace every answer back to the game that produced it.

Playthrough Console / full film02:54

The loop

From a game question to an answer you can defend.

Use our agent architecture, bring your own simulation, or work with us to build both. Playthrough keeps the engine at the center of the evidence chain. It never substitutes a vision proxy.

Author

Hand-crafted, reinforcement-learning, and language-driven agents enter through the same legal-action interface.

Play

Agents run the real engine headlessly across thousands of seeded, version-stamped matches.

Interrogate

Ask one run or an entire tournament a question in plain language. The Console writes queries and re-runs the replay.

Verify

Intervals reveal when a difference is signal rather than noise. Replay and error checks reveal whether the result is trustworthy.

Act

Turn a defensible finding into a balance change, prompt revision, regression fix, or next experiment.

What it opens up

Test how your game really behaves.

Agentic simulation creates experiments where a human playtest alone can only create anecdotes.

01 / BALANCE

Play the change first

Run a card, unit, or economy change through thousands of different games before it reaches a human playtester.

02 / EVALUATION

Grade the agent in-game

Compare scripted, RL, and LLM agents on identical legal moves, with win rates and intervals instead of impressions.

03 / REGRESSION

Replay what changed

Re-run prior games against a new engine version and see exactly which outcomes moved and by how much.

04 / TRAINING

Train where you grade

Use the simulator as the RL environment so training progress and gameplay results share a single measurement system.

05 / INSIGHT

Ask in plain language

Let production, QA, design, and ML teams explore the same source of truth without learning a specialist query language.

06 / TRUST

Catch contaminated metrics

Find the fallback, provider error, or engine exception that would otherwise turn a bad run into a confident decision.

Playthrough single-run console showing run summary and arena activity
“Statistics say how the run went. Errors say whether you can believe them.”Playthrough design principle

Bring your hardest balance question

Let the agents play it out.

Plug in your simulation, build an agentic harness with us, or start with a focused proof of concept around one high-value question.

Start a Playthrough conversation