Learning Rewards and Worker Reliability from Pairwise Comparisons
A new research paper on arXiv (2608.10045) addresses the challenge of learning from pairwise comparisons when the data is collected from unreliable crowdworkers. The authors propose a method to jointly learn item rewards and worker competencies using a Boltzmann-rational model that extends the Bradley-Terry-Luce model. They derive an EM-based algorithm that introduces Polya-Gamma latent variables to handle the logistic function in the model. The work is motivated by applications in recommendation systems, social choice, and fine-tuning large language models, where pairwise comparisons are often elicited from platforms like Amazon Mechanical Turk and Scale AI. The paper aims to improve the reliability of learning by accounting for worker spamming behavior and limited domain knowledge.
Key facts
- Paper arXiv:2608.10045
- Announce Type: cross
- Focus on learning from pairwise comparisons
- Applications in recommendation systems, social choice, and fine-tuning large language models
- Crowdworkers often unreliable due to limited domain knowledge or spamming
- Adopts Boltzmann-rational model extending Bradley-Terry-Luce
- EM-based algorithm with Polya-Gamma latent variables
- Platforms mentioned: Amazon Mechanical Turk, Scale AI
Entities
Institutions
- arXiv
- Amazon Mechanical Turk
- Scale AI