ARTFEED — Contemporary Art Intelligence

LLM Outputs Exhibit Idiolectal Signatures: Study Finds Model-Specific Linguistic Profiles

ai-technology · 2026-08-10

A recent study uploaded to arXiv (ID 2608.06589) disputes the common belief in a unified "AI language," positing that outputs from large language models (LLMs) exhibit unique, model-specific linguistic traits similar to human idiolects. Conducted by researchers, including Improta et al., the analysis focused on two datasets of LLM-generated content concerning societal issues: a 2024 dataset featuring six models and a newly created 2026 dataset using identical prompts with six modern models. Through computational descriptors and stylometric principal component analysis, the team discovered a stylistic evolution between the two cohorts, while each model retained its distinct linguistic identity. Notably, contraction frequencies varied significantly within the 2026 cohort, ranging from over 1,200 to more than 30,000 per million words. These results indicate that viewing LLM outputs as idiolectal may enhance our comprehension of AI-generated texts, impacting areas like digital humanities, stylometry, and AI ethics. This paper, categorized as a cross-announcement, is accessible on arXiv, highlighting the ongoing scholarly exploration of AI systems' linguistic characteristics.

Key facts

  • Paper ID: arXiv:2608.06589
  • Published on arXiv with announcement type 'cross'
  • Analyzes LLM-generated texts on societal topics
  • Two datasets: 2024 corpus (six models) and 2026 corpus (six contemporary models)
  • Uses computational descriptors and stylometric principal component analysis
  • Finds generational shift between 2024 and 2026 cohorts
  • Each model maintains a unique linguistic profile
  • Contraction frequencies range from over 1,200 to over 30,000 per million words in 2026 cohort

Entities

Institutions

  • arXiv

Sources