进阶 unsloth.ai 2026-10-07 22:27:00 · 6 阅读

第29章 运行与微调 QwQ-32B:Bug 修复指南

unsloth 下载 ☰ 下载 博客 有效运行 QwQ-32B + 修复 Bug 2025年3月7日 • 作者 Daniel & Michael 2025年3月7日•作者 Daniel & Michael Qwen 发布了 QwQ-32B,这是一款性能强劲的推理模型,在多项基准测试中表现可与 DeepSeek-R1 媲美。你可能遇到过无限循环、重复生成、思考 token 报错以及微调困难等问题,但这并不代表模型的真实质量。希望这篇博客能帮你调试并解决大部分问题!查看教程

我们上传的模型已包含 bug 修复,适用于微调、vLLM 和 Transformers。不过,如果你正在使用 llama.cpp 或以其作为后端的引擎,可能会遇到一些问题。要解决这些问题,请遵循下方的教程,或阅读我们文档中的详细指南与分析。

在此查看 Unsloth 修复过的所有 QwQ-32B 模型上传,包括 GGUF 和动态 4-bit 量化版本。
QwQ-32B Bug 修复 我们还发现了一些特别影响微调的问题!EOS token 是正确的,但 PAD token 应该设为 "
dry
typ_p
xtc📖 教程:部署并优化 QwQ-32B
1. 配置 llama.cpp 及引擎
使用 llama.cpp 时,可查阅文档中的完整指南。从 github.com/ggml-org/llama.cpp 获取最新版。以下提供构建步骤;若无 GPU 或仅用 CPU 推理,请将 -DGGML_CUDA=ON 改为 -DGGML_CUDA=OFF。
apt-get update
apt-get install build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggerganov/llama.cpp
cmake llama.cpp -B llama.cpp/build \
-DBUILD_SHARED_LIBS=ON -DGGML_CUDA=ON -DLLAMA_CURL=ON
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-quantize llama-cli llama-gguf-split
cp llama.cpp/build/bin/llama-* llama.cpp
2. 下载模型与测试
先安装 huggingface_hub hf_transfer,然后下载模型。推荐 Q4_K_M 或 BF16 全精度等其他量化版本。其他变体见:huggingface.co/unsloth/QwQ-32B-GGUF
接着运行 Unsloth 的 Flappy Bird 测试,输出将保存至 Q4_K_M_yes_samplers.txt。
# !pip install huggingface_hub hf_transfer
import os
os.environ["HF_HUB_ENABLE_HF_TRANSFER"] = "1"
from huggingface_hub import snapshot_download
snapshot_download(
repo_id = "unsloth/QwQ-32B-GGUF",
local_dir = "unsloth-QwQ-32B-GGUF",
allow_patterns = ["*Q4_K_M*"], # For Q4_K_M
)
3. 测试与评估
通过 --threads 32 设置 CPU 线程数,--ctx-size 16384 设置上下文长度,--n-gpu-layers 99 指定 GPU 卸载层数;若 GPU 内存不足请调整此值,纯 CPU 推理则移除该参数。可尝试调整 --repeat-penalty 1.1 和 --dry-multiplier 0.5。
../llama.cpp/llama-cli \
--model unsloth-QwQ-32B-GGUF/QwQ-32B-Q4_K_M.gguf \
--threads 32 \
--ctx-size 16384 \
--n-gpu-layers 99 \
--seed 3407 \
--prio 2 \
--temp 0.6 \
--repeat-penalty 1.1 \
--dry-multiplier 0.5 \
--min-p 0.01 \
--top-k 40 \
--top-p 0.95 \
-no-cnv \
--samplers "top_k;top_p;min_p;temperature;dry;typ_p;xtc" \
--prompt "<|im_start|>user\nCreate a Flappy Bird game in Python. You must include these things:\n1. You must use pygame.\n2. The background color should be randomly chosen and is a light shade. Start with a light blue color.\n3. Pressing SPACE multiple times will accelerate the bird.\n4. The bird's shape should be randomly chosen as a square, circle or triangle. The color should be randomly chosen as a dark color.\n5. Place on the bottom some land colored as dark brown or yellow chosen randomly.\n6. Make a score shown on the top right side. Increment if you pass pipes and don't hit them.\n7. Make randomly spaced pipes with enough space. Color them randomly as dark green or light brown or a dark gray shade.\n8. When you lose, show the best score. Make the text inside the screen. Pressing q or Esc will quit the game. Restarting is pressing SPACE again.\nThe final game should be inside a markdown section in Python. Check your code for errors and fix them before the final markdown section.<|im_end|>\n<|im_start|>assistant\n 运行命令如下: --model unsloth-QwQ-32B-GGUF/QwQ-32B-Q4_K_M.gguf \
--threads 32 \
--ctx-size 16384 \
--n-gpu-layers 99 \
--seed 3407 \
--prio 2 \
--temp 0.6 \
--repeat-penalty 1.5 \
--repeat-penalty 1.1 \
--dry-multiplier 0.5 \
--min-p 0.1 \
--top-k 40 \
--top-p 0.95 \
-no-cnv \
--prompt " 你或许会想:“难道该用 Q4_K_M?B16 全精度应该没问题吧?”——不对。若开启 Repetition Penalty 却没加上我们的修复参数 `--samplers "top_k;top_p;min_p;temperature;dry;typ_p;xtc"`,输出照样会失败。💡 思考 token 不显示? 已有用户反馈:由于 chat template 默认会插入 标签,部分系统未能正确输出思维轨迹(thinking traces)。你需要手动编辑 Jinja 模板,将模板末尾默认的 `

把 {%- if add_generation_prompt %} {{- '<|im_start|>assistant\n 和往常一样,欢迎加入我们的 Reddit 社区和 Discord 服务器,获取帮助或表达支持!也可以关注我们的 Twitter,并订阅我们的通讯。感谢阅读! Daniel & Michael Han 🦥 2025年3月月7日 立即免费微调 Phi-4! 免费开始使用 加入我们的 Discord

评论 (0)