Microsoft AI 模型现已在 AI Gateway 上线
Vercel 与 Microsoft AI(MAI)达成合作,将 MAI 模型接入 AI Gateway。AI Gateway 是少数支持访问 MAI 模型的平台之一。
MAI 注重安全与零数据保留(ZDR),这与 Vercel 的目标一致:赋予 AI Gateway 用户对其数据的管理权,并明确训练过程中的数据使用透明度。
Copy link to headingAI Gateway 上的 MAI 模型
MAI 最新音频模型现已支持在 AI Gateway 上运行。MAI-Voice-2.1 和 MAI-Voice-2.1-Flash 用于语音生成,而 MAI-Transcribe-2 Streaming 则能在音频流入时实时返回转录更新。
MAI-Voice-2.1(
microsoft/mai-voice-2.1)支持 23 种语言的表达性语音生成,并能在较长段落中保持说话者音色一致。适用于对全程表现力有要求的场景,如旁白、有声读物、播客和教学课程。MAI-Voice-2.1-Flash(
microsoft/mai-voice-2.1-flash)将同样的多语言语音生成能力应用于低延迟交互。适合语音助手、智能体及需要快速响应的语音回复场景。MAI-Transcribe-2 Streaming(
microsoft/mai-transcribe-2-streaming)在音频传输过程中即可返回部分转录结果。应用可实时显示字幕或跟踪正在进行的对话,无需等待完整录音结束。
AI Gateway 按所列费率对这些模型计费,推理过程不收取平台费或加价。
Copy link to heading快速上手
使用 AI SDK 7 中的 generateSpeech 结合 MAI-Voice-2.1-Flash 可生成语音响应。以下代码选择 Microsoft 的 Harper 音色配合 Flash 模型:
1import { experimental_generateSpeech as generateSpeech } from 'ai';2import { writeFile } from 'node:fs/promises';3
4const result = await generateSpeech({5 model: 'microsoft/mai-voice-2.1-flash',6 text: 'Your order is ready for pickup.',7 voice: 'en-US-Harper:MAI-Voice-2.1-Flash',8 outputFormat: 'mp3',9});10
11await writeFile('response.mp3', result.audio.uint8Array);对于旁白等较长音频,请使用 microsoft/mai-voice-2.1。
Copy link to heading实时音频转录
streamTranscribe 接收音频块流,并在音频到达时实时返回转录更新。下例假设 microphoneStream 是一个 16 kHz、16 位 PCM 音频的 ReadableStream:
1import { experimental_streamTranscribe as streamTranscribe } from 'ai';2
3const stream = streamTranscribe({4 model: 'microsoft/mai-transcribe-2-streaming',5 audio: microphoneStream,6 inputAudioFormat: { type: 'audio/pcm', rate: 16000 },7});8
9for await (const part of stream.fullStream) {10 if (part.type === 'transcript-partial') {11 process.stdout.write(`\r${part.text}`);12 }13}14
15console.log(await stream.text);随着更多音频到达,部分转录结果可能会变化,因此收到新的部分结果时应替换已显示的文本。
Copy link to heading更多资源
试用 MAI-Voice-2.1-Flash 进行语音生成,使用 MAI-Transcribe-2 Streaming 进行实时转录。完整模型家族请见 MAI 模型页面,也可以按 语音快速上手指南 开始使用。
AI Gateway 提供统一 API,用于调用模型、追踪用量和成本以及查看请求链路,同时支持在可用供应商之间进行路由、重试和故障转移。