进阶 docs.antigma.ai 2026-10-08 09:54:38 · 6 阅读
第15章 模型与推理力度
模型、提供商与推理力度(Effort)
Ante 支持在会话中途切换模型、提供商和推理力度,无需重启。两个对话框各司其职:/models 用于在当前提供商下调整模型和推理力度,/providers 则用于切换提供商。
选择模型
在 TUI 中输入 /models 即可打开模型选择器,其中列出当前提供商的模型(标题会显示范围,如 Switch Model — Anthropic)。用方向键选择模型,调整推理力度滑块,按 Enter 确认。
/models 可以跳过对话框,直接切换到当前提供商下的指定模型。这个 id 不必出现在选择器列表中——只要提供商支持就能用,因此你可以借此访问目录中未收录的 Open Router 模型。模型会以目录默认的推理力度启动;不带参数重新打开 /models 即可调整滑块。如果尚未选择提供商,id 会像 CLI 上的 --model 一样自动匹配对应提供商。
按 Tab 可以转到提供商选择器,直接切换提供商。
切换提供商
输入 /providers 可列出所有提供商及其连接状态。对已连接的提供商按 Enter,会进入它的模型选择器——同样是模型列表和推理力度滑块,并预先选中你上次在该提供商使用的模型。只有按 Enter 确认后切换才会生效;按 Esc 则退回提供商列表。对未连接的提供商按 Enter,则会启动其登录流程(参见 Connect a provider)。
切换提供商会保留当前对话——消息、标题和权限模式都会延续,新提供商会在你发送下一条消息时接手。
Ante 会记住你在每个提供商上最后使用的模型,因此在提供商之间来回切换时,总能无缝接上之前的进度。
推理力度
推理力度是 Ante 提供的单一调节项,控制模型在回答前投入多少思考,共六个等级:
min · low · medium · high · xhigh · max
min 表示该模型支持的最小努力级别。若模型允许,此级别下将关闭思考功能;对于强制启用推理的模型,则会被限制在其最低级别。各供应商将这一尺度映射到其原生的思考或推理参数上,而支持级别较少的供应商会将请求的努力程度向下取整至最近的支持级别。 在 /models 中,努力程度是一个位于模型列表下方的滑块,每个实际支持的级别对应一个刻度——拥有四个原生级别的模型显示四个刻度,支持全部六个级别的模型则显示六个刻度。没有推理开关的模型不显示滑块。 该阶梯是按供应商和模型家族预先构建的。对于自定义供应商,你可以通过模型目录条目中的 supported_efforts 自行声明(参见目录参考),或者运行 /add-provider,它将探测端点并自动为你写入声明。 Anthropic 在 Claude Sonnet 5、Opus 5、Opus 4.7/4.8 以及始终开启思考的 Fable/Mythos 5 系列中提供 xhigh。Opus 和 Sonnet 4.6 保留了四个活跃自适应级别(从 low 到 high,再到 max)以及 min 关闭位置。Gemini 3.7 和 3.8 仅精确暴露 low、medium 和 high——请求的 xhigh 或 max 会向下取整为 high。 note:Anthropic API 在启用思考时要求 temperature 必须为 1。如果你在目录或通过 MODEL_TEMPERATURE 为 Claude 模型设置了不同的温度,开启思考的轮次将因温度不匹配而失败,并显示相应错误信息——清除覆盖设置即可修复。 你的努力程度选择会按模型保存在 ~/.ante/settings.json 中,你也可以通过 CLI 为单次运行固定设置: ante --effort high Readable reasoning Effort determines the extent of reasoning the model performs; a provider's thinking_display determines how much readable reasoning it returns. Configure it as off, concise, or detailed in ~/.ante/catalog.json. Disabling the display does not disable reasoning, and requesting detailed output does not increase effort. See Thinking display for provider mappings and override behavior. tip You can also set your default model and provider via CLI flags: ante --provider openai --model gpt-6.1-sol, or configure defaults in your preferences. If you pick a provider without naming a model, Ante starts you on that provider's most capable model and may use a lighter, cheaper one for ambient UI hints such as the spinner's thinking phrase and the next-prompt suggestion — see weight classes. WebFetch itself returns page content to the active session model; it does not invoke a helper model. Vision-capable models are detected automatically, so you can drop images straight into the chat.
min 表示该模型支持的最小努力级别。若模型允许,此级别下将关闭思考功能;对于强制启用推理的模型,则会被限制在其最低级别。各供应商将这一尺度映射到其原生的思考或推理参数上,而支持级别较少的供应商会将请求的努力程度向下取整至最近的支持级别。 在 /models 中,努力程度是一个位于模型列表下方的滑块,每个实际支持的级别对应一个刻度——拥有四个原生级别的模型显示四个刻度,支持全部六个级别的模型则显示六个刻度。没有推理开关的模型不显示滑块。 该阶梯是按供应商和模型家族预先构建的。对于自定义供应商,你可以通过模型目录条目中的 supported_efforts 自行声明(参见目录参考),或者运行 /add-provider,它将探测端点并自动为你写入声明。 Anthropic 在 Claude Sonnet 5、Opus 5、Opus 4.7/4.8 以及始终开启思考的 Fable/Mythos 5 系列中提供 xhigh。Opus 和 Sonnet 4.6 保留了四个活跃自适应级别(从 low 到 high,再到 max)以及 min 关闭位置。Gemini 3.7 和 3.8 仅精确暴露 low、medium 和 high——请求的 xhigh 或 max 会向下取整为 high。 note:Anthropic API 在启用思考时要求 temperature 必须为 1。如果你在目录或通过 MODEL_TEMPERATURE 为 Claude 模型设置了不同的温度,开启思考的轮次将因温度不匹配而失败,并显示相应错误信息——清除覆盖设置即可修复。 你的努力程度选择会按模型保存在 ~/.ante/settings.json 中,你也可以通过 CLI 为单次运行固定设置: ante --effort high Readable reasoning Effort determines the extent of reasoning the model performs; a provider's thinking_display determines how much readable reasoning it returns. Configure it as off, concise, or detailed in ~/.ante/catalog.json. Disabling the display does not disable reasoning, and requesting detailed output does not increase effort. See Thinking display for provider mappings and override behavior. tip You can also set your default model and provider via CLI flags: ante --provider openai --model gpt-6.1-sol, or configure defaults in your preferences. If you pick a provider without naming a model, Ante starts you on that provider's most capable model and may use a lighter, cheaper one for ambient UI hints such as the spinner's thinking phrase and the next-prompt suggestion — see weight classes. WebFetch itself returns page content to the active session model; it does not invoke a helper model. Vision-capable models are detected automatically, so you can drop images straight into the chat.