Automatic glitch detection in gameplay video
Early ReBlink research on using V-JEPA 2 surprise to find temporal glitches in gameplay footage without labeled bug training, with VideoGlitchBench as the next evaluation.
Gameplay QA often requires people to watch long recordings by hand. A build session can produce hours of video, and someone has to review that footage, find the brief moments when something fails, and mark the relevant seconds so a ticket can be filed.
The world might freeze. A character might jump to the wrong place. Motion might run backward for a moment. Most of the recording may be fine, but finding these short failures still takes time.
We have been studying whether a model can review the same footage automatically and point people toward suspicious moments. The goal is simpler than building a separate glitch classifier for every game. We want to test whether temporal failures can be found without first teaching a model what every possible bug looks like.
Predicting the next part of a video
One approach we have explored uses V-JEPA 2. It does not use a vision-language model or ask a large language model to narrate the clip. It also does not train on labeled glitches or use game-specific data. The model has not been shown a catalog of freezes, teleports, or reverse playback from our titles.
V-JEPA 2 learns from video by predicting missing parts of a sequence. It makes these predictions in an internal representation rather than drawing the next frame pixel by pixel. It is also not trying to guess which button the player will press. Through its training, the model develops a general sense of how visual scenes usually continue over time.
We use that behavior to produce a surprise score. The detection loop has four steps:
- Show the model about one second of gameplay.
- Ask it to predict the internal representation of the next second.
- Compare that prediction with the representation of what actually happened.
- Treat a large difference as surprise.
A spike in surprise becomes a candidate glitch for a person to review.

This method relies on temporal continuity. The model has learned how normal video tends to develop from one moment to the next. When a glitch breaks that continuity, the predicted second and the observed second can disagree. The size of that disagreement provides the detection signal.
Testing a clip with known glitches
We first needed a case where the exact timing of each glitch was known. We took one clean gameplay clip and injected artificial glitches at specific seconds.
In each chart below, the pink band shows where the injected glitch occurs. The blue line shows the model’s surprise score. A useful signal should rise inside the pink band when a glitch is present while remaining near the normal noise floor when the video is unchanged.

The top-left panel is the control condition. It contains no injected glitch, and the blue line remains within a modest noise range. This gives us a baseline for interpreting the other panels.
The clearest results are freeze in the top-right panel and reverse in the bottom-left panel. In both cases, the blue surprise score rises sharply inside the pink band. These are promising early results because freezes and reversed motion are temporal failures that people already notice when reviewing footage by hand.
The green line provides a simple comparison. It measures how much the raw pixels change from one frame to the next. During a freeze, the green line drops because a stuck frame contains almost no pixel change. The blue surprise score still spikes.
This freeze contrast shows why prediction can provide a different signal from raw motion. A pixel-change baseline can interpret a frozen image as a period with little activity. V-JEPA 2 instead compares the frozen result with how the video was expected to continue, so the interruption can produce high surprise.
The other panels contain teleport, jitter, and shake injections. Freeze and reverse are the clearest cases on this clip. Teleport, jitter, and shake need more clips before we rank them against each other.
The orange line is a feature-change baseline. The clearest teaching comparison in this experiment remains the blue surprise score against the green pixel-change baseline during the freeze.
What this result covers
This result is based on one clip with self-injected glitches. On that material, surprise from a general video model lines up with known temporal failures. The next step is measuring those rates across games, studios, and bug taxonomies.
The method focuses on visual problems that emerge over time. These include freezes, teleports, movement that should not happen, reverse playback, and similar breaks in continuity.
Other defects may be visible in a single frame without disrupting the next second of motion. A stretched texture or a floating object that remains still may not surprise a model that is evaluating how one second continues into the next. Those defects require other forms of review.
Intentional discontinuities also matter. Camera cuts, respawns, and menus can produce high surprise because they interrupt temporal continuity. The signal is behaving as designed when it reacts to those changes, even though they are not necessarily bugs. A larger evaluation needs to measure how often these events become false alarms and how well they can be separated from real failures.
Evaluating real annotated gameplay
Self-injected glitches are useful because their exact timing is known. Real gameplay bugs are the next test.
The next evaluation is VideoGlitchBench: 5,238 real gameplay clips from 120 games. Each clip is annotated with a description of the bug and the exact seconds when it occurs. This provides ground truth from real footage rather than from edits we created ourselves.
Applying the same surprise measurement across VideoGlitchBench will help answer several questions. We can test whether freeze and reverse remain comparatively easy to find, identify which intentional cuts produce false alarms, and measure how far an approach without labeled glitch training can go. The results can also show where game-specific training becomes useful and where static defects require a different detector.
Automated video review does not remove the need for a person who understands the game. It can reduce how much clean footage that person has to inspect. If a system can identify suspicious minutes in a long capture, QA can spend more time examining failures and less time scrubbing through unaffected footage.
This research is aimed at making that search less expensive. The broader goal is to make more of a build’s video searchable for moments that deserve human review without assuming that every game has its own labeled glitch dataset.
If your studio wants to explore solutions like this through an applied research engagement, reach out. We care about hard problems, and we want the chance to go deep on them.
