← All resources

Elevating game agent intelligence: our ECCV 2024 research

Our ECCV 2024 CV2 paper turns expert Arena Tactics play into a multimodal imitation-learning benchmark — the research foundation behind ReBlink’s real-engine game agents.

Most game-AI research still trains agents in sandboxes that look nothing like a shipped product. We needed the opposite: agents that learn from real expert play inside a real tactics ruleset, with vision, language, and spatial structure in the same loop.

That work became a peer-reviewed paper at ECCV 2024.

What we built

Elevating Game Agent Intelligence: A Multimodal Benchmark for Strategy Games

Arena Tactics is hard on purpose. Squad composition, deck-building, hidden information, and a large action space make every match different. That is useful for players. It is also a demanding test for imitation learning.

We turned expert play into a public research benchmark:

  • 78.6 hours of expert matches across six players
  • 12,259 state-action pairs from 3,793 turns
  • Game states that combine minimap imagery, unit feature vectors, and natural-language descriptions of operations and protocol cards
  • A multimodal model that predicts what an expert would do next: action type, ability or card choice, and target tiles
  • Spatial-aware attention so unit positioning actually informs targeting, instead of being flattened away

The ablations matter for commercial product decisions as well as for academic ones. Adding ability and card structure beats units-only baselines. Adding natural-language descriptions of those abilities further improves action and target selection. In other words: if you want agents that play a complex game, multimodal state is not a nice-to-have. It is how the agent understands the rules players already speak.

Protocol cards remain the hardest decisions to clone, which is honest. Those actions are rarer and carry real trade-offs. A useful benchmark should surface that rather than hide it.

Download the paper (PDF)

Why this research exists

ReBlink’s product bet depends on a simple loop. Players and experts act inside a real engine. Those trajectories become training signal. Agents learn to act in the same legal action space. Then that intelligence shows up again in co-play, opponents, and large-scale simulation.

The CV2 paper was an early, peer-reviewed proof of that loop. Not a slide about “AI NPCs.” A dataset, a model, and an evaluation protocol other researchers could inspect.

"We've been exploring this space since the fall of 2022, figuring out how to combine AI co-play with strategic depth and high visual fidelity."

— Arunan Sri, Founder & CEO, ReBlink

Language control, agentic QA, and studio-facing systems like Playthrough and Playable all assume the same foundation: an agent that can read a messy, multimodal game state and still choose legal actions. This paper is where we put that foundation under review.

Where it was published

We submitted to CV2 (Computer Vision for Videogames) at ECCV 2024. The paper was peer-reviewed and accepted in August 2024.

CV2 was organized with NVIDIA’s involvement — the workshop hub lives on NVIDIA’s site, NVIDIA researchers helped run the program, and NVIDIA sponsored the academic award. For us, that venue mattered because it put a product-grounded dataset and model in front of the people building vision and learning systems for games, beyond the usual benchmark environments.

Workshop page: CV2 at ECCV 2024

Paper by Cristhian Forigua, Tomas Correa, and Arunan Sri at ReBlink.

View original source

Continue the conversation

Bring us the system you are trying to understand.

We help teams turn difficult AI research into products and decisions they can use.

Book a call