Qwen3.5-9B-OptiQ-4bit (mlx-community)
Apple Silicon 本地 4-bit 混合精度量化文本生成模型
~9.8GB 8-bit许可 可商用任务 文本生成最近核对 2026-07-22
- 格式
- 8-bit
- 文件
- ~8.6GB
- 运行内存
- ~9.8GB–13GB
- 运行时
- MLX_LM
- 下载
- 11,971
仅 safetensors · 无 pickle 加载风险
快速上手示例
mlx_lm.generate --model mlx-community/Qwen3.5-9B-OptiQ-4bit 依赖版本和硬件参数请以源仓库说明为准。
适合与不适合
独立证据与社区反馈
混合精度量化的 OptiQ 版本在同等磁盘占用下质量明显优于均匀 4-bit 量化,社区实测肯定其速度与质量平衡,是 Mac 本地部署 9B 模型的优选方案之一。
- 混合精度量化在同等磁盘占用下击败均匀 4-bit 量化,Capability Score 达 68.2(+6.4 vs 均匀 4-bit)
- 仅需约 5 GB RAM,适合内存受限的 Mac 本地部署
- 在 Apple Silicon(如 Mac Studio M3)上运行速度极快,同时保持可用输出质量
- 配合 DFlash 投机解码可达 4.13x 加速(30.96→127.07 tok/s),接受率 89.4%
- 通过 mlx-optiq 一行 pip install 即可在 CLI、Web Lab、编程 Agent 三种模式中使用
- 支持 OpenAI 和 Anthropic API 兼容的本地 server
- Tools 得分 83%、Coding 得分 70%,适合工具调用和编程场景
- 磁盘占用约 9 GB 以内,是 OptiQ 系列中磁盘效率最高的模型之一
- 仅支持 MLX 框架,非 Apple Silicon 设备不可用
- 统一内存架构下量化模型仍受带宽限制,投机解码对量化目标的加速有限(27B-4bit 仅 1.90x)
- 纯 attention 模型(如 Qwen3、Gemma)无法享受 GatedDeltaNet 的 tape-replay 投机解码优势
- mlx-community 的量化模型为社区维护,非官方出品
- 社区曾对 mlx-optiq 量化的可靠性存疑,初期缺乏实际使用反馈
- mlx-community 下存在多种量化版本(均匀 4-bit 与 OptiQ),用户需注意区分
- 社区实测Qwen 3.5 for MLX is like its own industrial revolution
- 社区实测Benchmarked 11 MLX models on M3 Ultra — here's which ...
- 社区实测Is mlx-optiq legit? Has anyone tested the new quants for ...
- 社区实测DFlash speculative decoding on Apple Silicon: 4.1x on Qwen3.5-9B, now open source (MLX, M5 Max)
- 独立评测Run LLMs locally on your Mac (Apple Silicon) · mlx-optiq
- 社区实测am.will on X: "I'm not sure anyone is doing more in the MLX niche for local LLMs than Ivan is. If you're into this kind of thing, he's an easy follow. One of the things I've been wanting to see is how well do these quantized versions of models ACTUALLY perform. Obv seeing how fast they are" / X
- 官方/厂商mlx-community/Qwen3.5-9B-OptiQ-4bit · Hugging Face
- 转载/仓库What's the difference between mlx-community model and huggingface original model? · ml-explore/mlx · Discussion #1250 · GitHub
可信度下载量近1.2万,六基准能力分66.77,超过均匀4-bit
完整规格
下载动量
30天下载 11.4k → 9.7k
观测时间线
模型家族
- Qwen3.5-9B
- 量化 MaralGPT-Mythos-9B (MaralGPT)
- 量化 Qwen3.5-9B (Qwen)
- 量化 Qwen3.5-9B-MTP-GGUF (unsloth)
- 量化 Qwen3.5-9B-OptiQ-4bit (mlx-community)
- 微调 Qwythos-9B (Empero)