Measuring Ensemble Dispersion in Language Models Without a Unique Answer
A new study released on arXiv (2608.00285) explores how to evaluate the diversity of outputs from groups of large language models when there isn't a clear right answer. The researchers examined sixteen models from ten different families and found that, on average, they produced 1.69 distinct interpretations of a psychotherapy case, compared to 1.43 from a single model. The authors argue that while measuring diversity is well-established using the Vendi Score—calculated from the von Neumann entropy of a similarity matrix—a single score doesn’t show where the diversity originates. To address this, they propose a dissent contribution per model, indicating how different each model is from the others, which helps identify which models boost diversity. The paper is a cross-announcement and can be found on arXiv.
Key facts
- Sixteen language models from ten families were used.
- The average semantic diversity was 1.69 distinct formulations.
- Single-model baseline diversity was 1.43.
- The Vendi Score is used to measure diversity.
- The Vendi Score is the exponential of the von Neumann entropy of a similarity matrix.
- A per-model dissent contribution is defined as the complement of mean similarity to other ensemble members.
- The paper is on arXiv with ID 2608.00285.
- The task involves psychotherapeutic case formulation.
Entities
—