ARTFEED — Contemporary Art Intelligence

Instruction-Conditioned Exploration Enhances RL for LLMs

ai-technology · 2026-08-04

A new paper on arXiv (2608.02087) introduces Instruction-Conditioned Exploration (ICE), a method to improve reinforcement learning (RL) for large language models (LLMs). The authors propose supplementing task prompts with distinct instructions during training to increase behavioral coverage, and combine RL with self-distillation (Asymmetric-RL/SD) to transfer explored behaviors to the test-time policy. Experiments on Qwen3-1.7B show improved held-out pass@1 performance at 4K response length on mathematical reasoning tasks. The paper addresses challenges in LLM action-space structure that differ from classical RL, emphasizing the need for deliberate exploration leveraging pre-trained knowledge.

Key facts

  • arXiv paper 2608.02087 proposes Instruction-Conditioned Exploration (ICE).
  • ICE supplements task prompts with distinct instructions during training.
  • Asymmetric-RL/SD combines reinforcement learning and self-distillation.
  • Method improves Qwen3-1.7B held-out pass@1 at 4K response length.
  • Focus on mathematical reasoning tasks.
  • Addresses LLM action-space challenges in RL.
  • Emphasizes deliberate exploration using pre-trained knowledge.
  • Paper is announced as new on arXiv.

Entities

Institutions

  • arXiv

Sources