APQF: Automated LLM-Guided Compression Framework for Neural Networks
A new framework called APQF (Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning) has been introduced to automate the compression of deep neural networks for resource-constrained edge devices. The method combines structured pruning, mixed-precision quantization-aware training, and accuracy recovery into a single pipeline. A profiling agent measures cost distribution and sensitivity to pruning across the model, and this evidence drives per-layer pruning ratios, bit-widths, and recovery strategies, all proposed by LLM planners and validated before execution. The authors claim APQF is the first framework to combine LLM-guided, profiling-grounded decisions. The paper is available on arXiv with ID 2608.05499, announced as a cross-type submission. The work addresses the challenge of manual, expert-dependent compression choices and the difficulty of applying algorithms across architectures, as well as the accuracy loss from uniform settings that ignore layer-specific responses to compression.
Key facts
- APQF stands for Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning.
- The framework combines structured pruning, mixed-precision quantization-aware training, and accuracy recovery.
- A profiling agent measures cost distribution and sensitivity to pruning.
- LLM planners propose per-layer pruning ratios, bit-widths, and recovery strategies.
- Decisions are validated before execution.
- The paper is available on arXiv with ID 2608.05499.
- The announcement type is cross.
- The framework targets resource-constrained edge devices.
Entities
Institutions
- arXiv