ollama 3小时前 · 2026-09-30 12:20:50 · 1 阅读
Ollama 现在支持 Jev 风格的决策模型
2026年9月29日
Ollama 现已支持决策模型,基于 TypeSafe 的 Jev API,提供快速、强类型的决策能力:
- 无额外费用
- 本地运行时延迟更低
- 即日起可通过 Ollama 使用三款新决策模型
该新 API 从 Ollama 0.35 开始提供,通过新的 /v1/systemone 端点调用。将文本作为 state 连同一组命名问题一起发送,本地运行的模型会在一次请求中回答所有问题。这非常适合需要快速决策的任务,比如工单分诊、模型路由、内容或安全审核。
近乎即时的决策
Ollama 上的决策模型速度很快,因为请求无需经过网络。在下面的吃豆人示例中,Nimble 9B 在 M5 Max 本地运行时平均每次决策仅需 91ms。这个速度足以支撑实时快速决策的场景,比如玩游戏或实时处理内容:
可用模型
目前可以通过 Ollama 运行三款新决策模型:
nimble:由 Bespoke Labs 开发的开源 9B 参数决策模型tev1:来自 Together AI 的实验性 4B 决策模型tev1:0.8b:来自 Together AI 的实验性 0.8B 决策模型
更多决策模型即将推出,包括由 Ollama 云端提供的模型。
快速上手
要开始使用,请先下载或升级至 Ollama 的最新版本。接下来,下载一个决策模型,例如 nimble:
ollama pull nimble
你可以通过 curl 或 TypeSafe 官方 Python SDK 发起请求。
请求
curl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {
"ticket": "I was charged twice. Please refund the extra payment."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}'
响应
{
"model": "nimble",
"answers": {
"team": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
"confidence": 0.922
},
"refund": {"type": "noul", "noul": 0.997},
"urgency": {
"type": "score",
"score": 0.815,
"legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
"probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
"confidence": 0.046
}
},
"usage": {"input_tokens": 841, "output_tokens": 4}
}
配置
uv add typesafe-sdk # 或: pip install typesafe-sdk
export TYPESAFE_BASE_URL=http://localhost:11434
export TYPESAFE_API_KEY=ollama
export TYPESAFE_DEFAULT_MODEL=nimble
请求
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
ticket = "I was charged twice. Please refund the extra payment."
questions = {
"team": Choice(
instructions="Which team should handle this ticket?",
criteria={
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above",
},
),
"refund": Noul(
instructions="Does the customer explicitly ask for a refund?",
),
"urgency": Score(
instructions="How urgent is this ticket?",
criteria=["Routine", "Soon", "Urgent"],
),
}
with TypeSafeClient(timeout=120) as client:
result = client.system_one(
state={"ticket": ticket},
questions=questions,
)
print(result.choices["team"].choice) # billing
print(result.nouls["refund"].noul) # 0.997
print(result.scores["urgency"].score) # 0.815
后续规划
这只是 Ollama 众多决策模型支持更新的第一波。未来更新将包括:
- 由 MLX 赋能,在 Apple Silicon 上实现更快的性能
- 更多专注于不同决策类型的模型
原始来源: ollama