ARTFEED — Contemporary Art Intelligence

LOFA: Learning Shopping Agents from Online User Feedback

ai-technology · 2026-08-13

The newly introduced framework, LOFA, allows shopping agents powered by large language models to learn directly from actual online interaction logs without the need for human annotation. This innovative method tackles the issues of diverse, sparse, and noisy online feedback by integrating reinforcement learning based on verifiable purchase results with feedback-aware on-policy distillation. It effectively captures users' directives during conversations and transforms them into dependable learning signals, addressing the shortcomings of current methods that depend on offline training signals like user-item interactions or synthetic preference data. A paper outlining this framework was published on arXiv (ID: 2608.11604), showcasing its potential to enhance shopping agents in real-world e-commerce by utilizing the valuable insights from users' natural conversational feedback.

Key facts

  • LOFA is a framework for shopping agents to learn from online interaction logs.
  • It uses reinforcement learning over verifiable purchase outcomes.
  • It employs feedback-aware on-policy distillation.
  • The method identifies users' in-dialogue directives.
  • It converts heterogeneous, sparse, and noisy online feedback into learning signals.
  • Existing approaches rely on offline training signals like user-item interactions.
  • The paper is available on arXiv with ID 2608.11604.
  • The framework requires no human annotation.

Entities

Institutions

  • arXiv

Sources