Context-Injected Fine-Tuning Boosts Small Legal QA Models in Bangladesh
A research paper available on arXiv (2607.23446) investigates whether fine-tuning smaller language models with pertinent legal examples enhances their application of retrieved legal texts. The team compiled 2,165 bilingual question-and-answer records from three schedules and six Bangladeshi acts, subsequently fine-tuning Qwen3.5 with 0.8B, 2B, and 4B parameters. The assessment utilized the Bangladesh Bar Council exams from 2022 and 2023 in both Bangla and machine-translated English, analyzed through strict consistency across three seeded runs without retrieval methods like BM25 or FAISS. Fine-tuning at 0.8B improved the 2022 English FAISS score from 2 to 34 out of 100. While improvements were noted at 0.8B and 2B, the 4B model did not show a net gain, with Bangla performance improving but some English conditions declining. Furthermore, fine-tuning significantly decreased the percentage of answers shifting from Bangla to predominantly English from 44.0–53.2% to 0.2–0.7%, with adjusted p < .001 across all scales. The authors suggest that retrieval quality is not the sole limitation.
Key facts
- Study tests context-injected fine-tuning for legal QA in Bangladesh.
- 2,165 bilingual QA records from six Bangladeshi acts and three schedules.
- Fine-tunes Qwen3.5 at 0.8B, 2B, and 4B parameters.
- Evaluation uses 2022 and 2023 Bangladesh Bar Council exams.
- At 0.8B, fine-tuning raises English FAISS score from 2 to 34 of 100.
- Gains at 0.8B and 2B survive paired testing.
- 4B model shows no detectable net gain; Bangla improves, English regresses.
- Fine-tuning reduces language drift from 44.0–53.2% to 0.2–0.7%.
- Adjusted p < .001 at every scale for language drift reduction.
Entities
Institutions
- Bangladesh Bar Council
- arXiv
Locations
- Bangladesh