Hugging Face Blog·· 2026-08-25AI 评分40
Quantization-Aware Healing: 压缩 4-bit 模型 GPT-OSS 120B 性能反超全精度原版
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
AI 导读
Hugging Face 提出 Quantization-Aware Healing(QAH)方法,将 GPT-OSS 120B 模型压缩至 60B 参数并量化为 MXFP4 格式。该 4-bit 模型在 9 项基准测试中的 7 项上,性能超越同架构的 bfloat16 全精度版本。
来源:Hugging Face Blog · huggingface.co