Abstract
Diffusion models (DMs) demonstrate outstanding performance in image generation, but their iterative sampling processes incur substantial computational costs and error accumulation, hindering deployment on resource-constrained IoT devices. In this article, we propose PQCAD-DM, a hybrid compression framework that tightly integrates progressive quantization (PQ) and calibration-assisted distillation (CAD). PQ employs a two-stage strategy with momentum-guided adaptive bit-width transitions to suppress instability, while CAD leverages a dual-calibration dataset to reconstruct robust full-precision guidance for low-bit students. Aligned with an industrial server-to-device deployment paradigm, PQCAD-DM optimizes models on high-resource servers for efficient edge inference. Extensive experiments, including validation on aerial and urban surveillance IoT datasets, demonstrate that PQCAD-DM consistently outperforms fixed-bit quantization baselines and architecture-optimized hybrid models. On a mobile smartphone, our framework achieves a 2.8 × higher throughput and a 48% reduction in model size while maintaining high visual fidelity, making it uniquely suited for latency-sensitive IoT applications.
| Original language | English |
|---|---|
| Pages (from-to) | 33306-33319 |
| Number of pages | 14 |
| Journal | IEEE Internet of Things Journal |
| Volume | 13 |
| Issue number | 15 |
| DOIs | |
| State | Published - 1 Aug 2026 |
Keywords
- Diffusion models (DMs)
- distillation
- model compression
- quantization
Fingerprint
Dive into the research topics of 'PQCAD-DM: Progressive Quantization and Calibration-Assisted Distillation for Efficient Compression of Diffusion Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver