ARTFEED — Contemporary Art Intelligence

MathShikkha: Bangla CoT Supervision Offers No In-Domain Gains for Stronger Models

other · 2026-08-11

A recent preprint on arXiv (2608.08503) explores the impact of teacher-led Chain-of-Thought (CoT) supervision on mathematical reasoning in Bangla, a language with limited resources. The research presents MathShikkha, a dataset containing Bangla math problems accompanied by rationales generated by GPT-5.4, and fine-tunes four small language models (ranging from 4B to 7B parameters) using a consistent protocol. This protocol aligns data splits, response-only loss masking, decoding, and scoring between answer-only and CoT conditions, differing only in training targets. In-domain findings indicate that CoT does not significantly enhance performance for three more robust models, with paired bootstrap 95% confidence intervals including zero and McNemar p-values ≥ 0.17, despite producing 15–52× more tokens. Conversely, the weaker 4B model shows a notable improvement of 18.56 points (p < 0.0001). However, on the larger, contamination-audited BanglaMATH benchmark, CoT significantly surpasses answer-only supervision. The study was published on arXiv under the identifier 2608.08503v1.

Key facts

  • Study compares answer-only and Chain-of-Thought supervision for Bangla mathematical reasoning.
  • Introduces MathShikkha dataset with GPT-5.4-generated rationales.
  • Fine-tunes four student models (4B–7B parameters) under a matched protocol.
  • In-domain, CoT provides no significant improvement for three stronger backbones.
  • Weaker 4B model improves by 18.56 points with CoT (p < 0.0001).
  • On BanglaMATH benchmark, CoT significantly outperforms answer-only.
  • CoT generates 15–52× more tokens than answer-only.
  • Paper announced on arXiv with identifier 2608.08503v1.

Entities

Institutions

  • arXiv

Sources