ARTFEED — Contemporary Art Intelligence

PIHF: Bringing Post-Training RL to In-context Learning

ai-technology · 2026-08-18

A recent paper on arXiv (2608.16831) presents a novel approach called Policy Iteration with Human Feedback (PIHF), which merges reinforcement learning with in-context learning. This method utilizes a pretrained language model as its foundation and features a versioned policy along with a toolkit in natural language. A clinical expert and a language-model critic evaluate the reasoning and tool usage across the complete panel to pinpoint recurring issues and suggest potential improvements. The expert has the power to reinterpret data and manage admissions and rollbacks, while Recall@1 and Recall@5 metrics assess the results post-execution. Testing on cumulative ablations and benchmarks for ultra-rare diseases showed that the PIHF-derived policy enhanced Recall@1 and Recall@5, highlighting the success of combining post-training RL with in-context learning.

Key facts

  • Paper arXiv:2608.16831 introduces Policy Iteration with Human Feedback (PIHF).
  • PIHF builds on generalized policy iteration and language-based task conditioning.
  • Uses a pretrained language model as execution substrate.
  • Maintains a versioned natural-language policy and tool set.
  • A language-model critic and clinical expert review reasoning and tool-use trajectories.
  • Expert has authority over admission and rollback of candidate revisions.
  • Recall@1 and Recall@5 are used for validation.
  • PIHF improved Recall@1 and Recall@5 on ultra-rare-disease benchmarks.

Entities

Sources