ARTFEED — Contemporary Art Intelligence

Context-Injected Fine-Tuning Boosts Small Legal QA Models in Bangladesh

other · 2026-07-29

A research paper available on arXiv (2607.23446) investigates whether fine-tuning smaller language models with pertinent legal examples enhances their application of retrieved legal texts. The team compiled 2,165 bilingual question-and-answer records from three schedules and six Bangladeshi acts, subsequently fine-tuning Qwen3.5 with 0.8B, 2B, and 4B parameters. The assessment utilized the Bangladesh Bar Council exams from 2022 and 2023 in both Bangla and machine-translated English, analyzed through strict consistency across three seeded runs without retrieval methods like BM25 or FAISS. Fine-tuning at 0.8B improved the 2022 English FAISS score from 2 to 34 out of 100. While improvements were noted at 0.8B and 2B, the 4B model did not show a net gain, with Bangla performance improving but some English conditions declining. Furthermore, fine-tuning significantly decreased the percentage of answers shifting from Bangla to predominantly English from 44.0–53.2% to 0.2–0.7%, with adjusted p < .001 across all scales. The authors suggest that retrieval quality is not the sole limitation.

Key facts

  • Study tests context-injected fine-tuning for legal QA in Bangladesh.
  • 2,165 bilingual QA records from six Bangladeshi acts and three schedules.
  • Fine-tunes Qwen3.5 at 0.8B, 2B, and 4B parameters.
  • Evaluation uses 2022 and 2023 Bangladesh Bar Council exams.
  • At 0.8B, fine-tuning raises English FAISS score from 2 to 34 of 100.
  • Gains at 0.8B and 2B survive paired testing.
  • 4B model shows no detectable net gain; Bangla improves, English regresses.
  • Fine-tuning reduces language drift from 44.0–53.2% to 0.2–0.7%.
  • Adjusted p < .001 at every scale for language drift reduction.

Entities

Institutions

  • Bangladesh Bar Council
  • arXiv

Locations

  • Bangladesh

Sources