LUNAR: New Benchmark for Personalized LLMs Using Real-World Behavioral Logs
LUNAR has been unveiled by researchers as the inaugural benchmark aimed at assessing how large language models (LLMs) tailor their responses based on long-term interaction histories in various everyday domains, including clothing, food, housing, and mobility. This benchmark fills a significant void in current personalized LLM evaluations, which often depend on textual personas or isolated behavioral signals, offering limited insights into cross-domain personalization. LUNAR employs a multi-stage synthesis pipeline that reflects real-world behavioral trends, facilitating scalable benchmark development while addressing data privacy and sparsity issues. Fidelity assessments indicate that LUNAR better mirrors actual behavioral distributions compared to other synthetic benchmarks. Tests conducted on 19 popular LLMs demonstrate that while access to behavioral logs is essential, it alone does not ensure effective personalization. The benchmark's details are documented in a paper on arXiv (arXiv:2608.05246).
Key facts
- LUNAR is the first benchmark for evaluating LLM personalization from longitudinal app interaction histories.
- It covers universal daily-life domains: clothing, food, housing, and mobility.
- The benchmark uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns.
- Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks.
- Experiments were conducted on 19 mainstream LLMs.
- Findings indicate that access to behavioral logs is necessary but not sufficient for deep personalization.
- Neither more context nor larger model size alone guarantees improved personalization.
- The paper is available on arXiv with identifier 2608.05246.
Entities
Institutions
- arXiv