MirrorCraft: New Benchmark Tests AI Agents Under Hidden Minecraft Rule Changes
A team of researchers has introduced MirrorCraft, a new benchmark aimed at evaluating how well agents powered by large language models (LLMs) can adapt in Minecraft when the game's mechanics change. Unlike existing assessments that focus on fixed rules, MirrorCraft tests agents’ flexibility when familiar elements like recipes and drops are modified. It creates paired 'Mirror' worlds that reflect a 'Vanilla' world but with specific server-side changes made using datapacks. Each pair maintains consistent terrain, resource placement, and objectives. The benchmark features five biomes, six sets of rules, three progression goals, two types of models, and six agent setups, all using a shared Mineflayer interface. Adaptability is measured by task completion. The research is available on arXiv under the ID 2607.29218.
Key facts
- MirrorCraft is a paired benchmark for evaluating LLM-based agents in Minecraft under hidden rule changes.
- Each Mirror world is a copy of its paired Vanilla world with selected server-side rules modified by datapacks.
- Terrain, spawn, resource placement, objective, interface, and action budget remain matched within every Vanilla-Mirror pair.
- MirrorCraft includes five controlled biomes, six rule suites, three progression objectives, two model families, and six agent configurations.
- The benchmark uses a shared Mineflayer interface.
- The paper is available on arXiv with identifier 2607.29218.
- Existing benchmarks typically evaluate agents under fixed game mechanics.
- The goal is to assess whether agents can continue making progress when familiar recipes, drops, and other rules change.
Entities
Institutions
- arXiv