第46章 Qwen3.8-27B开源模型发布:支持智能体编程与混合推理
Qwen3.8-27B Qwen3.8-27B 是一款用于智能体编程与聊天的混合推理开源模型。您可以通过 Unsloth Dynamic GGUF 和 Unsloth Desktop 在本地运行该模型。
Unsloth Dynamic V3.0 27B 参数 文本 + 视觉 256K 上下文 4-bit 精度下需 17 至 19 GB
内存需求 针对推荐的 4-bit GGUF 格式,建议规划 17 至 19 GB 的总内存(RAM + VRAM)。
核心能力 支持文本与视觉输入、混合推理、智能体编程、嵌套工具调用,以及高达 256K 的上下文长度。
GGUF 量化方案 NVFP4 量化 查看完整指南
本地运行 Unsloth 运行
unsloth run --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
应用场景 使用 unsloth start 命令,针对 Qwen3.8-27B 启动编程智能体。 查看配置指南 →
Claude Code unsloth start claude --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
OpenAI Codex unsloth start codex --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
Hermes Agent unsloth start hermes --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
OpenClaw unsloth start openclaw --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
OpenCode unsloth start opencode --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL
运行选项 大多数计算机建议选用 GGUF 构建版本,NVIDIA Blackwell 架构设备则推荐 NVFP4 版本。
GGUF 4-bit 推荐 本地运行质量与内存占用之间的最佳平衡。
Unsloth 运行 unsloth run --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL 使用 Unsloth Desktop 可自动搜索、下载并调优模型。
NVFP4 速度快约 2.5 倍 针对 RTX 50 系列、DGX Spark、B200 和 B300 进行了优化。
vLLM vllm serve unsloth/Qwen3.8-27B-NVFP4
SGLang python -m sglang.launch_server --model-path unsloth/Qwen3.8-27B-NVFP4 --speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 支持 FP8 KV 缓存校准,可实现更长的上下文。
内存与设置 Unsloth 会自动应用推荐的推理参数。
内存需求 量化精度 2-bit 3-bit 4-bit 6-bit 8-bit BF16 总内存 11–13 GB 13–16 GB 17–19 GB 24 GB 31 GB 56 GB 建议让 RAM 加 VRAM 的总量接近量化版本的大小,以获得最佳速度。若内存不足,也可通过磁盘卸载实现运行,但速度会变慢。
使用 Unsloth Desktop 在本地运行、训练并连接 搜索并下载 GGUF 或 safetensor 模型,利用自我修复工具调用功能,执行网络搜索和代码运行,并让 Unsloth 在 macOS、Windows 或 Linux 上自动调优推理。 下载 Unsloth