ARTFEED — Contemporary Art Intelligence

Liquid AI Releases Q4_0 LFM2.5 Checkpoints via Quantization-Aware Distillation

ai-technology · 2026-08-19

Liquid AI has introduced Q4_0 GGUF checkpoints for its LFM2.5 model family using quantization-aware distillation, recovering 97% of the accuracy lost to post-training quantization. The models—LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B—outperform standard Q4_0 on benchmarks including GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4. On edge hardware (MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, Raspberry Pi 5), they achieve 4–33% higher decode throughput while matching the quality of higher-bit quantizations like Q5_K_M and Q4_K_M. The checkpoints are available on Hugging Face today.

Key facts

  • Liquid AI trained Q4_0 checkpoints using Quantization-Aware Distillation (QAD).
  • QAD recovers 97% of the BF16 average accuracy lost to quantization.
  • Benchmark suite includes GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, and BFCLv4.
  • GSM8K used for LFM2.5-230M and LFM2.5-350M; AIME25 for LFM2.5-1.2B-Instruct and LFM2.5-2.6B.
  • QAD checkpoints retain 97.1%, 96.5%, 97.4%, and 96.6% of respective BF16 baseline performance.
  • Decode throughput tested on MacBook Pro, NucBox EVO-X2, Samsung Galaxy S26 Ultra, and Raspberry Pi 5.
  • QAD Q4_0 matches Q5_K_M quality at 4-33% higher throughput for 230M/350M models and Q4_K_M quality at 3-14% higher throughput for 1.2B/2.6B models.
  • Files are available on Hugging Face for all four LFM2.5 model sizes.

Entities

Institutions

  • Liquid AI
  • Hugging Face
  • Unsloth

Sources