Qwen3.5-9B-MTP-GGUF (unsloth)
Qwen3.5-9B 多模态模型,MTP 投机解码,本地快速运行
~5.4GB 4-bit许可 可商用任务 视觉理解最近核对 2026-06-21
- 格式
- 4-bit
- 文件
- ~4.7GB
- 运行内存
- ~5.4GB–7GB
- 运行时
- CMD
- 下载
- 53,048
快速上手示例
./llama.cpp/build/bin/llama-server -hf unsloth/Qwen3.5-9B-MTP-GGUF:UD-Q4_K_XL -ngl 99 -fa on --spec-type draft-mtp 依赖版本和硬件参数请以源仓库说明为准。
适合与不适合
独立证据与社区反馈
Unsloth 的 Qwen3.5-9B GGUF 量化版本相比官方版本推理速度更快,12GB 显存即可流畅运行且延迟良好;量化鲁棒性强,TQ1_0 等低比特量化几乎无损保留原始精度,但 MMLU Pro 等世界知识类基准上仍有明显退化。
- 12GB 显存即可部署,延迟表现良好
- Unsloth 版本推理速度优于 lm-studio 及官方版本
- TQ1_0 量化几乎无损保留原始模型精度
- 工具调用和编程性能较此前版本有改进
- 兼容 llama.cpp、vLLM、SGLang、Unsloth Studio 等多种推理引擎
- Dynamic 2.0 量化方法进一步降低内存占用
- MMLU Pro 上退化最明显,世界知识有所损失
- 量化后模型与原始 16 位版本不完全等同
- 社区实测Qwen3.5 Unsloth GGUFs Update! - Reddit
- 社区实测(Qwen3.5-9B) Unsloth vs lm-studio vs "official" : r/LocalLLaMA - Reddit
- 社区实测Unsloth Dynamic 2.0 GGUFs - Hacker News
- 社区实测Benjamin Marie on X: "Here's a more complete evaluation of GGUF variants of Qwen3.5 (models by @UnslothAI ), and it's way better than I expected. - Qwen3.5 is very robust to Unsloth quantization - TQ1_0 preserves the original model's accuracy extremely well - Most of the degradation is on MMLU Pro https://t.co/Q5U2P9yq8b" / X
- 独立评测Summary of Qwen3.5 GGUF Evaluations + My Evaluation Method
- 官方/厂商unsloth/Qwen3.5-9B-MTP-GGUF - Hugging Face
可信度53048次下载,46点赞,MTP解码提速1.5-2倍,Unsloth Dynamic 2.0量化
完整规格
下载动量
30天下载 108.1k → 95.7k · likes +42
观测时间线
模型家族
- Qwen3.5-9B
- 量化 MaralGPT-Mythos-9B (MaralGPT)
- 量化 Qwen3.5-9B (Qwen)
- 量化 Qwen3.5-9B-MTP-GGUF (unsloth)
- 量化 Qwen3.5-9B-OptiQ-4bit (mlx-community)
- 微调 Qwythos-9B (Empero)