ARTFEED — Contemporary Art Intelligence

Decaying Reward LinUCB Enhances LLM Iterative Refinement

ai-technology · 2026-08-10

A new research paper on arXiv (2608.06750) proposes a contextual bandit algorithm that explicitly models reward decay to improve iterative refinement in Large Language Models (LLMs). The authors argue that existing methods, such as feedback-based Self-Refine and traditional bandit approaches, often rely on static options or overlook the saturation effect, leading to over-exploitation where repeated use of identical prompts or arms yields diminishing rewards. Their algorithm, which extends the Linear Upper Confidence Bound (LinUCB) framework, uses an Expectation-Maximization (EM) algorithm to simultaneously estimate arm-specific and decay parameters. By embedding prompts as arms, the method facilitates joint learning of arm values, distinguishing it from the disjoint LinUCB approach. Experimental results on Sentiment Reversal and GSM8K benchmarks demonstrate significant performance improvements. The paper is categorized as a cross-type announcement and was published on arXiv.

Key facts

  • Paper ID: arXiv:2608.06750
  • Proposes a contextual bandit algorithm with reward decay modeling
  • Uses Expectation-Maximization (EM) algorithm to estimate arm-specific and decay parameters
  • Embeds prompts as arms for joint learning of arm values
  • Distinguishes from disjoint Linear Upper Confidence Bound (LinUCB)
  • Evaluated on Sentiment Reversal and GSM8K benchmarks
  • Addresses over-exploitation in iterative refinement
  • Published on arXiv with cross-type announcement

Entities

Institutions

  • arXiv

Sources