ARTFEED — Contemporary Art Intelligence

AI Guilt Calibration from Human Brain Data Enhances Prosocial Multi-Agent Learning

ai-technology · 2026-08-06

A new study posted on arXiv (ID 2608.04663) presents a method for adjusting artificial guilt signals in cooperative multi-agent reinforcement learning using data from human behavior and neural responses. The team analyzed happiness fluctuations from the SoDec responsibility fMRI dataset, which involved 40 participants, applying a subject-fixed-effects regression. They calculated a guilt weight by comparing Partner-negative and Social-negative outcomes (weight = 1.118, Cohen's d = 0.214). This weight was then used in a two-agent Social Lottery framework, training Proximal Policy Optimization actors under various conditions: neurally calibrated, uniform, selfish, and oracle. Over 1,000 evaluation episodes, the calibrated agents achieved social safe-choice rates similar to humans (0.459 vs. baseline), suggesting that human-derived guilt signals could improve the prosocial behavior of AI.

Key facts

  • arXiv preprint 2608.04663 proposes calibrating artificial guilt from human neural data
  • Uses SoDec responsibility fMRI dataset with 40 participants
  • Guilt weight recovered as contrast between Partner-negative and Social-negative outcomes (1.118, Cohen's d=0.214)
  • Embedded in two-agent Social Lottery environment
  • Trained independent Proximal Policy Optimization actor-critics under four shaping regimes
  • Calibrated agents matched human social safe-choice rate most closely (0.459)
  • Evaluated across 1,000 episodes per condition
  • Addresses hand-tuning of social reward terms in multi-agent reinforcement learning

Entities

Institutions

  • arXiv

Sources