OSCD: Cross-Distillation for Low-Resource Southeast Asian Language Reasoning
A novel algorithm called Onramp-Sequence Cross-Distillation (OSCD) has been developed to tackle the issue of cross-lingual collapse in Large Language Models (LLMs) when engaging in intricate reasoning within low-resource Southeast Asian languages. This approach translates high-resource reasoning paths into low-resource vocabulary subspaces during generative training rollouts, utilizing an integrated translator agentic loop to ensure stable and effective translation of dynamically produced reference samples for fine-tuning. Additionally, it features joint-embedding semantic alignment between reference and target-language reasoning traces, effectively bridging the cross-lingual divide. The technique seeks to alleviate the cold-start bottleneck for policy optimization and avoid catastrophic forgetting due to shifts in cross-lingual representation. The findings are presented in a paper on arXiv (2608.00533), emphasizing native multilingual chain-of-thought reasoning in low-resource contexts.
Key facts
- OSCD is a post-training algorithm for LLMs.
- It addresses cross-lingual collapse in low-resource Southeast Asian languages.
- Uses an integrated translator agentic loop during generative training rollouts.
- Projects high-resource reasoning trajectories into low-resource vocabulary subspaces.
- Includes joint-embedding semantic alignment of reasoning traces.
- Aims to solve cold-start bottleneck for policy optimization.
- Prevents catastrophic forgetting due to cross-lingual representation drift.
- Paper available on arXiv with ID 2608.00533.
Entities
Institutions
- arXiv