跳到正文
电商热点雷达

运营热点解读

Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer

来源:NVIDIA Generative AI · 发布时间:

雷达摘要

NVIDIA Nemotron 3.5 Lightning achieves up to 4x faster throughput with the NVFP4 checkpoint compressed to 22 GB from the 66 GB full-precision version. Quantization-aware distillation (QAD) recovers accuracy lost during aggressive post-training quantization, enabling W4A16 quantization of Mamba linear layers while preserving model quality. The two-stage QAD…

原始信息

本页是基于公开来源生成的摘要与运营提示,不替代原文。请以原始发布方内容为准。

阅读 NVIDIA Generative AI 原文 →