DialectLLM: Framework for Multi-Dialectal Dialogue Generation
DialectLLM has been unveiled by researchers as the inaugural extensive framework aimed at generating high-quality conversational data across multiple dialects. A significant 80% of the 1.6 billion English speakers do not communicate in Standard American English (SAE), yet large language models frequently misinterpret non-SAE dialects, leading to stereotypical outputs. This framework encompasses three essential aspects of written dialect: lexical, orthographic, and morphosyntactic characteristics. It generates a dialog dataset that is dialect-parallel, covering nine distinct English dialects. To ensure authenticity, native linguists contributed to the creation and validation of transformation rules from SAE to dialect. The framework challenges the norm of using a uniform morphosyntactic feature set for both user inputs and model outputs, revealing that models should not replicate up to 90% of a dialect's grammatical traits. Human assessments validate the quality of the data.
Key facts
- DialectLLM is the first large-scale framework for multi-dialectal conversational data generation.
- More than 80% of 1.6 billion English speakers do not use Standard American English.
- LLMs often fail to identify non-SAE dialects and generate stereotyped responses.
- The framework covers lexical, orthographic, and morphosyntactic features.
- It produces a dataset spanning nine English dialects.
- Native linguists collaborated to design and validate transformation rules.
- Models should not reproduce up to 90% of grammatical features of a dialect.
- Human evaluation confirms data quality.
Entities
—