ARTFEED — Contemporary Art Intelligence

CIDER: New Dataset Aligns LLMs with Human Privacy Preferences

ai-technology · 2026-08-11

A new dataset called CIDER has been developed by researchers to better align large language models (LLMs) with the specific privacy preferences of individuals, moving beyond general standards. It includes 14,850 annotations provided by 169 users, creating 1,650 sets of contextual disclosure boundaries across 60 scenarios of interpersonal communication where privacy norms are breached. Each boundary reflects a real user's decisions regarding disclosure across 9 variations in a scenario, tailored to particular communication roles and AI-mediated contexts. The study shows that in-context personalization enhanced prediction accuracy by as much as 11.41 percentage points with just 6 historical examples across 12 different models. This dataset fills a crucial gap in understanding nuanced privacy preferences for evaluating LLM alignment and is accessible on arXiv with the identifier 2608.09164.

Key facts

  • CIDER dataset introduced for privacy preference alignment in LLMs
  • 14,850 human annotations from 169 users
  • 1,650 contextual disclosure boundary sets
  • 60 interpersonal communication scenarios
  • Each boundary includes 9 sharing variants
  • Task: predict user's disclosure decision from historical boundaries
  • In-context personalization improved accuracy by up to 11.41 percentage points
  • 12 open and proprietary models evaluated

Entities

Sources