ARTFEED — Contemporary Art Intelligence

VLM-Annotated Datasets Enable Offline RL for Video Game Agents

ai-technology · 2026-08-07

Researchers propose using Vision Language Models (VLMs) to annotate video game datasets with human-defined rewards, enabling offline reinforcement learning (RL) to train conditioned agents. The approach addresses challenges in RL such as requiring game engine access, difficult reward identification and weighting, and sparse rewards. Early experiments show promise but also reveal difficulties and limitations. The work is published on arXiv under the title 'Training a Conditioned Video Game Agent on a VLM Annotated Dataset' and is categorized under Computer Science > Artificial Intelligence.

Key facts

  • Reinforcement Learning (RL) is powerful but not easy to use for policy learning.
  • Access to the game engine is typically required to get rewards for training.
  • Proper identification and weighting of rewards often requires trial-and-error.
  • Rewards are often sparse and understanding their effect on policy is non-trivial.
  • The proposed method annotates a video game dataset with Vision Language Models (VLMs) instructed to extract human-defined rewards.
  • Offline RL can then be used to train a conditioned agent that responds to desired returns.
  • Early experiments revealed difficulties and limitations.
  • The paper is available on arXiv with ID 2608.05954.

Entities

Institutions

  • arXiv

Sources