ChronoLens: A New Framework for Measuring Language Change Across Time and Languages
Researchers have introduced ChronoLens, an innovative computational framework designed to assess language evolution, as outlined in a paper on arXiv (2608.03507). This framework evaluates 44.98 million documents and 17.2 billion tokens from five parliamentary traditions spanning 1803 to 2026, employing frozen multilingual models and feature-aligned crosscoders. ChronoLens fills existing gaps in the study of historical language changes by establishing a cohesive analytical environment for examining morphology, syntax, semantics, and pragmatics. Its sparse representations demonstrate a correlation of 0.72 with linguistic statistics, surpassing the performance of dense embeddings (0.29) and pooled sparse autoencoders (0.28). The findings reveal a synchronized evolution across various linguistic levels, significantly influencing digital humanities, historical linguistics, and computational linguistics.
Key facts
- ChronoLens is a framework for measuring language change across time, languages, and linguistic levels.
- It combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions.
- Applied to 44.98 million documents and approximately 17.2 billion tokens.
- Data from five parliamentary traditions spanning 1803–2026.
- Sparse representations from ChronoLens show a correlation of 0.72 with linguistic statistics, versus 0.29 for dense embeddings and 0.28 for a pooled sparse autoencoder.
- The study addresses the problem of incompatible representations in computational studies of language change.
- Reveals that morphology, syntax, and semantics evolve together across languages.
- Paper available on arXiv with ID 2608.03507.
Entities
Institutions
- arXiv