← All resources

Patch day still runs on stories

After a live patch, clips and charts arrive first. They are useful, and they are the wrong kind of evidence for most balance calls. You need matches played in the real game that a designer can open again.

A patch goes live. Within hours, the team has clips, forum posts, support tickets, and a dashboard that moved. Someone already has a theory.

That is a normal patch day in a live game. It is also a weak basis for deciding whether to revert, hotfix, or wait.

The matches behind those reports happened on player machines. You cannot open them again. You cannot open the full match behind a clip either. A forum post gives you a player’s interpretation. A dashboard confirms that something moved, but not what happened inside the matches that moved it.

The result is a room full of plausible stories. The problem could be one map, one queue, one skill bracket, a client bug, or a genuine change in the rules. Patch-day evidence rarely makes the distinction clear.

Functional QA and balance are different jobs

Part of the problem is that two kinds of work are often treated as one.

Functional QA asks whether the build works:

  • Does it boot?
  • Do the new strings load?
  • Does the event shop take money?
  • Does the client stay synchronized with the server?

Smoke tests and a short regression pass are built for these questions. They are good at answering them.

Design and balance ask something different:

  • Did the new item, unit, or rule change how people should play?
  • Which matchups changed?
  • Did a previously weak strategy become dominant?
  • Is the change acceptable?

Those questions can only be answered by playing under the real rules: maps, fog, cooldowns, loadouts, timing, and the combinations no one thought to turn into a test case.

A live title has to do both jobs on the same day, with less time than a ship build. Content cycles accelerated years ago. The old playtesting model did not.

Human playtests matter, but they are not a population

Human play remains essential. People find problems with feel, sharp bugs, and interactions a designer would never think to try. Five or six people in a room can answer many usability questions.

But a small group cannot cover a systems game after every patch.

The number of legal combinations across cards, items, maps, matchups, and skill levels is larger than any playtest team can reach. Human testers are also likely to play the intended way. They understand the systems, follow the obvious lines, and avoid behavior that looks irrational.

The broken case is often the wrong way to play: an item left in a mid-state, a dropped line, a disconnect, or a strategy that only appears at one rank.

Human playtests give you judgment. They do not give you enough matches to test every claim a live patch will create.

Dashboards show movement, not matches

Telemetry is the other half of the standard live operations toolkit. It is valuable, but easy to ask too much of.

Pick rate, win rate, time-to-kill, and funnel drop-off can tell you that something changed across the population. They cannot show you the game that produced the change.

A new weapon showing a 58% win rate could reflect a real balance problem. It could also come from a client desynchronization that looks like one, a small sample in a particular queue, or the players and matchups selecting the weapon.

The line on the dashboard does not tell you which explanation is right.

Without a way to open the underlying games, a metric starts an investigation. It does not finish one.

Clips and forums are reports, not coverage

Community reports arrive quickly because players are already inside the new rules. Designers should read them. They should not mistake them for coverage.

A clip preserves one striking moment, but you cannot open the whole match behind it. It may omit the setup, earlier decisions, or the state that made the outcome possible. A forum thread reflects the players who chose to post. It can mix “this feels bad,” “this is broken,” and “I lost” into the same signal.

These reports are useful leads. They are not a representative set of matches, and they do not let the team replay the conditions behind a claim.

When the evidence is a clip or a post, the team still has to reconstruct the game around it.

More automation does not automatically produce better evidence

Scripted bots can add volume, but they repeat known paths. They are good at checking scenarios someone already decided to encode. They are less useful for discovering how a changed system behaves when play departs from those scenarios.

Screen-watching agents have a similar problem. A vision layer can observe a game and produce a score overnight, but the first requirement is more basic: can the agent make legal moves in the real build, and can a designer open the match it played?

For patch decisions, a score without an inspectable game is another claim to debate.

The useful unit is a match you can reopen

What a live team needs is enough play under the new rules to ask a design question, find the relevant matches, and open them again.

That is the work Playthrough is built for.

We work with the studio to create agents for its game rather than placing a generic bot on top of it. The agents make legal moves in the real build and do not repeat the same game every time. Openings vary. Lines get abandoned. Unusual interactions have room to appear.

The result is a body of matches played in the build the studio is preparing to ship. Those matches can run before the patch is public and again on the morning it goes live.

That changes the patch-day conversation. Instead of asking whether a clip is representative or whether a chart moved for the reason everyone assumes, the team can find and open the relevant games.

Open the games behind the answer

Matches only help if designers and QA leads can work with them.

In Playthrough Console, teams can write a query, pull the relevant matches, and open the games behind an answer. They can compare two builds, ask why a matchup flipped, inspect what stands out, and identify the next question a designer or QA lead should examine.

The Console is not a score that has to be trusted on faith. It is an assistant for balance and QA that helps a team move from a claim to the games that support or challenge it.

If a number looks wrong, open the matches behind it. If a matchup appears to have flipped, inspect how. If an edge case looks broken, start from a replay rather than trying to reconstruct one from a forum post.

The agents evolve with the title

A live game does not stand still, so its agents cannot be a one-time setup.

New content changes the available decisions. New rules create new lines. A meta that did not exist last season can change how people approach the game.

The agents evolve with the title so they can account for those changes in how they play. That keeps the matches useful as the game grows instead of freezing the simulation around an earlier version of the game.

Human play starts with better questions

None of this removes human play.

It changes where human attention begins. Designers can spend their time judging feel and investigating the cases the match evidence cannot decide. QA can focus on functional failures instead of trying to rediscover the match behind the same clip.

Clips, forums, support tickets, and dashboards will always be part of live operations. The difference is what happens next.

When someone asks whether to revert, hotfix, or wait, the team can pull the relevant matches, open them again, and keep the agents playing the next version of the real build.

Continue the conversation

Bring us the system you are trying to understand.

We help teams turn difficult AI research into products and decisions they can use.

Book a call