在 Apple 芯片上用 MLX 从零手写 Qwen3 推理与 vLLM 级服务系统
MLX 提供数组与扩展运行时作为正确性基准;C++/Metal 用于手写内核;Qwen3 规模适中贴近真实服务场景。
MLX 是 Apple 官方推出、专为 Apple 芯片优化的数组计算与机器学习框架