ARTFEED — Contemporary Art Intelligence

AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning

ai-technology · 2026-07-29

AdaKP is an online selector that dynamically chooses which atomic knowledge points (KPs) to inject into prompts during reinforcement learning training for large language models. It addresses reward sparsity in competition-level mathematics by using an entropy proxy to score KPs based on next-token entropy reduction, replacing expensive rollout-based estimation. The method involves a single forward pass with a provable bound on truncation bias, supported by three lightweight mechanisms. This approach is introduced in a paper on arXiv (2607.24833).

Key facts

  • AdaKP is an online adaptive knowledge-point selection method.
  • It targets reasoning-oriented reinforcement learning for large language models.
  • The method addresses reward sparsity in competition-level mathematics.
  • Knowledge points are short natural-language hints from gold solutions.
  • An entropy proxy scores KPs by next-token entropy reduction.
  • The selection uses a single forward pass with provable truncation bias bound.
  • Three lightweight mechanisms make the signal usable.
  • The paper is available on arXiv with ID 2607.24833.

Entities

Institutions

  • arXiv

Sources