Gemma-4-31B-it-GGUF (LM Studio)
Google 31B 指令模型 GGUF 量化版,本地消费级硬件推理
~19GB 4-bit许可 可商用最近核对 2026-06-21
- 格式
- 4-bit
- 文件
- ~16GB
- 运行内存
- ~19GB–24GB
- 运行时
- OLLAMA
- 下载
- 340,436
快速上手示例
ollama run hf.co/lmstudio-community/gemma-4-31B-it-GGUF:Q4_K_M 依赖版本和硬件参数请以源仓库说明为准。
适合与不适合
独立证据与社区反馈
Gemma 4 31B 不是最聪明的本地模型,但因其综合可用性和离线能力成为许多人日常首选;一次性编码表现亮眼,工具调用场景则逊色不少;消费级硬件可跑但量化版本选择与 VRAM 策略直接影响体验。
- 本地离线推理,适合通勤、旅行、无网络环境
- 通过 GPU offloading 在消费级笔记本(如 RTX 5080 Laptop 16GB VRAM)上运行大参数量模型
- 多模态输入(文本+图像)
- 工具调用与推理能力
- 256K 长上下文窗口
- 通过推测解码(speculative decoding)加速推理
- 兼容多种推理框架:llama.cpp、Ollama、Unsloth Studio、Docker Model Runner 等
- lmstudio-community 的 GGUF 量化不在 KL 散度 Pareto 前沿(Q8_0 除外),品质不如 unsloth 或 bartowski 的量化
- 长文档和非拉丁脚本在低精度量化下退化最快,即使 Q8_0 也有明显 KL 散度
- 推测解码若草稿模型选择不当,速度反而比不用更慢(实测 7.31 t/s vs 57 t/s 基线)
- LM Studio 中 31B 版本的 thinking mode 开关可能不显示,需通过官网「use this model in LM Studio」按钮或手动修改聊天模板修复
- 编码任务需将 temperature 降至 0.3 或更低,默认 1.0 会产生大量错误
- 工具调用/自定义 harness 场景下表现不如一次性编码
- 密集模型对内存带宽要求高,不适合低带宽设备(如 NVIDIA Spark)
- 低精度量化下滑动窗口注意力与 KV cache 复用可能导致长上下文误差累积
- 社区实测Gemma 4 31B GGUF quants ranked by KL divergence ... - Reddit
- 社区实测Speculative Decoding works great for Gemma 4 31B with E2B draft ...
- 社区实测For anyone having issues with Gemma 4 31b in LM Studio ... - Reddit
- 社区实测Gemma-4-26B-A4B-it-UD-Q4_K_M.gguf : IMHO worst model ever ...
- 社区实测I ran Gemma 4 as a local model in Codex CLI - Hacker News
- 独立评测Gemma 4 31B GGUF Quality Benchmark: unsloth, bartowski ...
- 独立评测Gemma 4 - LM Studio
- 独立评测Practical Gemma 4 Benchmarking with LM Studio - DEV Community
- 独立评测I Spent 3 Nights Testing Gemma 4 (MTP)
- 社区实测Google's Gemma 4 isn't the smartest local LLM I've run ... - Facebook
- 独立评测Slow inference with 31b model Gemma 4? Optimizations?
- 官方/厂商lmstudio-community/gemma-4-31B-it-GGUF - Hugging Face
可信度HuggingFace 下载量 34 万,基于 Google gemma-4-31B-it 量化
完整规格
下载动量
30天下载 115.5k → 100.5k · likes +1
观测时间线
模型家族
- gemma-4-31B-it
- 量化 Gemma-4-31B-it (Google)
- 量化 Gemma-4-31B-it-GGUF (LM Studio)
- 微调 Gemma-4-31B-StyleTune (Gryphe)