Kazakh-Russian Code-Switching: Annotation Boundary, Not Model, Key to Identification
Recent research available on arXiv has proposed a shift in how code-switching is analyzed, particularly between Kazakh and Russian languages. The study underscores that the definitions of boundaries between loanwords and code-switching are more critical than the models used to identify them. It introduces a unique language identification dataset based on Kazakh-Russian interactions on social media, categorizing Russian terms as Kazakh and marking certain sentence switches as 'mixed.' Notably, the authors found that the way these boundaries are annotated significantly influences the effectiveness of language identification techniques. The dataset and related code are offered for further exploration in the field.
Key facts
- Paper title: 'Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification'
- Published on arXiv with ID 2608.00581
- Releases a document-level gold LID set for Kazakh-Russian social text
- Guideline treats integrated Russian borrowings as Kazakh, reserves 'mixed' for clause-level switches
- Includes a mixed-only sentiment pool for filter-first cascade
- Benchmarks FastText, Lingua, HeLI (raw and windowed), character-trigram NB, and XLM-R
- Finds annotation boundary is the bottleneck, not model class
- Shared Cyrillic script causes over-labeling of loanwords as code-switching
Entities
Institutions
- arXiv