如何用 Policy as Code 和 OPA 治理 AI 生成的基础设施代码(完整手册)
现代大模型生成的代码几乎每次都能保证语法正确。Veracode 的 2026 年报告直言不讳:“语法问题基本上已经解决了。”
这听起来像是一个里程碑,但这恰恰是你面临问题的根源。
该同一份报告测试了一百多个模型,发现平均安全通过率为 56%,“与首份报告的 55% 相比几乎没有变化”,其中约 44% 的生成任务引入了具有风险的安全漏洞。
功能正确性与安全性其实是两个独立的问题,而目前只有其中一项接近解决。
这一结果既不是个例,也不是新现象。早在 2022 年,纽约大学坦登工学院的一个团队在 IEEE Security and Privacy 会议上展示了 GitHub Copilot 在 89 个安全相关场景下的表现,生成了 1689 个程序,结果发现其中约有 40% 存在漏洞,涉及 MITRE 的 CWE Top 25 类别。这篇论文后来还被评为 Communications of the ACM 的研究亮点。
2024 年 11 月,乔治城大学新兴技术与安全中心(CSET)评估了五个 LLM,报告称其中近半数的代码片段包含可能导致被利用的 bug。四年时间,四支独立团队,四种不同的方法论,答案始终一致。
在 2023 年的 ACM CCS 会议上,斯坦福大学的 Neil Perry、Megha Srivastava、Deepak Kumar 和 Dan Boneh 将开发者纳入了实验:47 名参与者,五项与安全性相关的编程任务,三种编程语言,其中 33 人使用了 AI 助手,14 人未使用。使用助手的组写出的代码安全性显著更低,且他们更容易相信自己写的代码是安全的。
虽然这是一项规模较小的研究,但它解释了为什么问题无法自我修正:通常能发现这类问题的机制(开发者仔细审视那些让他们担忧的代码)恰恰是工具正在关闭的机制。
上述研究测量的都是应用代码。但基础设施代码是更难处理的情况,因为错误的安全组配置永远不会“失败”:它会完全按照编写逻辑工作,为任何请求者提供流量服务,唯一能提出异议的是正在阅读 diff 的人。
光靠人眼看那份配置文件,谁也审不过来,我也不例外。能做的,是把规则写成计算机能自动校验的形式,每次变更都检查一遍——这就是 Policy as Code(策略即代码)的含义。
在本手册中,我会带你一步步搭建这套校验机制。我们会拿一个真实存在漏洞的仓库练手,先看它在九条违规中“顺利通过”五条的荒唐表现,然后动手修复。
读完之后,你将学会:
针对
terraform show -json输出的 JSON 编写 Rego 策略。像测试应用代码一样测试策略,配备测试夹具(fixtures)和覆盖率报告。
构建一个命令行门禁,通过退出码约定让 CI 流水线可以放心依赖。
使用 CEL 在 Kubernetes 准入阶段拦截不合规的工作负载。
让模型来写策略,再用
opa check和你自己的测试决定是否采纳。用同一个策略引擎对 AI agent 的工具调用做授权。
Table of Contents
Prerequisites
你需要准备:
一个终端,以及可用的
python3(3.10 或更高版本)。jq,用于在命令行解析 JSON。约 700 MB 磁盘空间,因为 AWS Terraform provider 文件较大。
一个 Anthropic API key,但仅在第七步使用。其他所有步骤均可离线运行。
mkdir policy-lab && cd policy-lab
python3 -m venv .venv
source .venv/bin/activate
pip install anthropic
curl -L -o opa https://openpolicyagent.org/downloads/v1.20.2/opa_darwin_arm64_static
chmod +x opa && sudo mv opa /usr/local/bin/
curl -L -o tf.zip https://releases.hashicorp.com/terraform/1.14.2/terraform_1.14.2_darwin_arm64.zip
unzip tf.zip && sudo mv terraform /usr/local/bin/
在 Windows 上,通过 .venv\Scripts\activate 激活环境,并将上述两个下载链接替换为 opa_windows_amd64.exe 和 terraform_1.14.2_windows_amd64.zip。
以下所有操作均在 macOS 上,基于 OPA 1.20.2、Terraform 1.14.2 和 AWS provider 6.x 运行。OPA 1.x 系列之间的策略语法保持稳定。
如果你使用的是 OPA 0.x,需要在每个规则文件顶部添加 import rego.v1,但我建议直接升级到 1.x。违规项统计仅受 AWS provider 版本影响,具体取决于 plan JSON 的结构,该结构自 provider 5.x 以来一直稳定。
关键术语的通俗解释
Policy as Code(策略即代码):将组织已达成共识的规则编写为程序,输入变更提案,输出判定结果。
Rego:Open Policy Agent 使用的查询语言。它是声明式的,规则主体由必须全部满足的条件列表组成。
Plan JSON:Terraform 即将执行的变更计划之机器可读描述,通过
terraform show -json生成。策略直接读取此内容,而非.tf文件。Admission control(准入控制):Kubernetes API server 内部的一个节点,对象在持久化存储前可在此处被拒绝。
CEL:Common Expression Language,Kubernetes 在 API server 内部原生支持的小型表达式语言,无需部署 webhook。
漏报:本应拦截的资源却通过了策略。这类问题几乎无人察觉,除非专门去排查,否则根本无法量化。
相同的判断逻辑存在于三个环节,其中只有中间那个是 Kubernetes 特有的。
第一步:获取真实的基础设施进行测试
我不想凭空捏造一个有漏洞的 Terraform 文件,因为那样等于先制造 bug,然后策略只能拦住我自己埋下的雷。所以我转而寻找别人编写并发布过的代码。
TerraGoat 是 Bridgecrew 发布的一个故意包含漏洞的 Terraform 仓库。获取其 EC2 模块的指定提交版本:
SHA=729f8da62c6a85ce4af5ad3d123de97776d954c4
curl -s "https://raw.githubusercontent.com/bridgecrewio/terragoat/$SHA/terraform/aws/ec2.tf" \
| sed -n '77,96p'
resource "aws_security_group" "web-node" {
# 安全组对公网开放了 SSH 端口
name = "${local.resource_prefix.value}-sg"
description = "${local.resource_prefix.value} Security Group"
vpc_id = aws_vpc.web_vpc.id
ingress {
from_port = 80
to_port = 80
protocol = "tcp"
cidr_blocks = [
"0.0.0.0/0"]
}
ingress {
from_port = 22
to_port = 22
protocol = "tcp"
cidr_blocks = [
"0.0.0.0/0"]
}
第二行的注释来自 TerraGoat 原文,而 22 端口对公网开放正是该文件指出的风险点。
TerraGoat 的模块在现代 Terraform 上无法初始化,因为它还在用带引号的 type = "string",这种写法在 Terraform 0.12 中已被弃用,1.x 更是直接拒绝。所以我把相关资源搬到了自己的一个精简模块里,只把两处对 TerraGoat 内部 locals 的引用换成了字面值。
创建 main.tf:
...(代码保持不变)
生成 plan JSON:
terraform init
terraform plan -out=tfplan.binary
terraform show -json tfplan.binary > plan.json
这里的 mock 凭证很关键:纯创建型配置在 terraform plan 时不会调用 AWS,因此配合 skip_credentials_validation 及另外三个 skip 参数,provider 不会尝试认证,也不会有任何内容被真正 apply。
第 2 步:先写测试,再写策略
我要实现的规则是:禁止任何安全组向公网暴露管理端口。
这句话听起来简单到只需一行代码,但测试部分恰恰说明了它的复杂性。创建 policy/network_test.rego 文件:
package terraform.network_test
import data.terraform.network
plan(resources) := {"resource_changes": resources}
security_group(ingress) := {
"address": "aws_security_group.web",
"type": "aws_security_group",
"change": {"actions": ["create"], "after": {"ingress": [ingress]}},
}
# 拒绝 SSH 对全网开放
test_denies_ssh_open_to_the_world if {
fixture := plan([security_group({
"from_port": 22,
"to_port": 22,
"protocol": "tcp",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 1 with input as fixture
}
# 简单的 from_port 等值检查会漏掉此情况,但区间检查不会。
test_denies_wide_open_port_range if {
fixture := plan([security_group({
"from_port": 0,
"to_port": 65535,
"protocol": "tcp",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 4 with input as fixture
}
# 拒绝 IPv6 对全网开放
test_denies_ipv6_route_to_the_world if {
fixture := plan([security_group({
"from_port": 22,
"to_port": 22,
"protocol": "tcp",
"ipv6_cidr_blocks": ["::/0"],
})])
count(network.deny) == 1 with input as fixture
}
# 拒绝独立的入站规则
test_denies_standalone_ingress_rule if {
fixture := plan([{
"address": "aws_vpc_security_group_ingress_rule.ssh",
"type": "aws_vpc_security_group_ingress_rule",
"change": {"actions": ["create"], "after": {
"from_port": 22,
"to_port": 22,
"ip_protocol": "tcp",
"cidr_ipv4": "0.0.0.0/0",
"cidr_ipv6": null,
}},
}])
count(network.deny) == 1 with input as fixture
}
# 拒绝已弃用的独立规则
test_denies_deprecated_standalone_rule if {
fixture := plan([{
"address": "aws_security_group_rule.ssh",
"type": "aws_security_group_rule",
"change": {"actions": ["create"], "after": {
"type": "ingress",
"from_port": 22,
"to_port": 22,
"cidr_blocks": ["0.0.0.0/0"],
}},
}])
count(network.deny) == 1 with input as fixture
}
# 允许 SSH 来自私网地址段
test_allows_ssh_from_a_private_range if {
fixture := plan([security_group({
"from_port": 22,
"to_port": 22,
"protocol": "tcp",
"cidr_blocks": ["10.0.0.0/8"],
})])
count(network.deny) == 0 with input as fixture
}
# 允许 HTTPS 对全网开放
test_allows_https_from_the_world if {
fixture := plan([security_group({
"from_port": 443,
"to_port": 443,
"protocol": "tcp",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 0 with input as fixture
}
# “所有协议”规则会开放所有端口,无论其端口字段如何书写。
# 该测试的初版断言方向相反,从而掩盖了 bug。
test_denies_all_protocols_rule_open_to_the_world if {
fixture := plan([security_group({
"from_port": 0,
"to_port": 0,
"protocol": "-1",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 4 with input as fixture
}
# 策略无法读取的端口会被上报,而不是直接放行。
test_reports_a_world_open_rule_with_unreadable_ports if {
fixture := plan([security_group({
"from_port": 22,
"to_port": null,
"protocol": "tcp",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 1 with input as fixture
}
# 字符串类型的端口会被上报
test_reports_string_ports if {
fixture := plan([security_group({
"from_port": "22",
"to_port": "22",
"protocol": "tcp",
"cidr_blocks": ["0.0.0.0/0"],
})])
count(network.deny) == 1 with input as fixture
}
# 非全网开放的规则上,无法读取的端口则静默处理。
test_ignores_unreadable_ports_on_a_private_range if {
fixture := plan([security_group({
"from_port": 22,
"to_port": null,
"protocol": "tcp",
"cidr_blocks": ["10.0.0.0/8"],
})])
count(network.deny) == 0 with input as fixture
}
# `considered` 决定了门禁的通过与否或空判,因此
# 尽管它本身不参与判定逻辑,也需要独立的测试。
test_considers_every_ingress_bearing_type if {
fixture := plan([
security_group({}),
{"address": "aws_vpc_security_group_ingress_rule.a", "type": "aws_vpc_security_group_ingress_rule", "change": {"actions": ["create"], "after": {}}},
{"address": "aws_security_group_rule.b", "type": "aws_security_group_rule", "change": {"actions": ["create"], "after": {}}},
])
count(network.considered) == 3 with input as fixture
}
# 不考虑无关类型
test_does_not_consider_unrelated_types if {
fixture := plan([{
"address": "aws_vpc.main",
"type": "aws_vpc",
"change": {"actions": ["create"], "after": {}},
}])
count(network.considered) == 0 with input as fixture
}
这个文件编码了以下内容:
with input as fixture会在某个表达式里替换为假的 plan,这就是不依赖云账号测试策略的方法。test_denies_wide_open_port_range期望有 四个违规项,每个管理端口一个,因为 ingress 规则描述的是一个范围。一条0-65535的规则等于把 SSH 完全敞开,和显式的 22 端口规则一样宽,还能绕过等值检查。另外三个测试覆盖了 Terraform 表达同一概念的另外三种形式:IPv6 字段、现代的独立资源
aws_vpc_security_group_ingress_rule,以及已弃用的aws_security_group_rule。这些并不是我一开始就写的,Step 9 会解释它们的由来。两个
test_allows_用例和拒绝类测试同样重要。一个拒绝一切的策略能通过所有 deny 测试,但毫无价值。最后三个测试是在策略已经“完成”之后才加入的,Step 9 会解释它们的来历。一条 all-protocols 规则无论端口字段写什么都会开放所有端口,而一条策略无法读取端口信息的规则则必须被上报。
Step 3: 编写策略直到测试通过
Terraform 用四种形式描述 ingress,所以策略先把这四种统一归一化成一组,再对这一组做一次判断。创建 policy/network.rego:
# METADATA
# title: 公网无法访问管理端口
# description: |
# Terraform 中入站规则有四种不同的形式。每个形式首先都会标准化为统一的
# `exposures` 集合,因此下方的判断逻辑只需编写一次,新增一种形式时只需增加一条辅助规则。
#
# 这里的两处设计是刻意为之,而非偶然。全协议规则会覆盖所有端口,无论其端口字段值如何;
# 对于策略无法读取端口的规则,则予以报告而非放行。
package terraform.network
admin_ports := {22, 3389, 3306, 5432}
public_cidrs := {"0.0.0.0/0", "::/0"}
# AWS 提供程序对全协议规则写入 from_port 0 和 to_port 0,这会开放所有端口,
# 因此端口字段不能按字面解读。
all_protocols := {"-1", "all"}
# 形式 1 和 2:内联入站块,IPv4 和 IPv6。
exposures contains exposure if {
some resource in input.resource_changes
resource.type == "aws_security_group"
some ingress in resource.change.after.ingress
some field in ["cidr_blocks", "ipv6_cidr_blocks"]
some cidr in object.get(ingress, field, [])
exposure := {
"address": resource.address,
"protocol": object.get(ingress, "protocol", ""),
"from_port": object.get(ingress, "from_port", null),
"to_port": object.get(ingress, "to_port", null),
"cidr": cidr,
}
}
# 形式 3:AWS 提供程序自 v5 起推荐的独立规则。
exposures contains exposure if {
some resource in input.resource_changes
resource.type == "aws_vpc_security_group_ingress_rule"
some field in ["cidr_ipv4", "cidr_ipv6"]
cidr := object.get(resource.change.after, field, null)
is_string(cidr)
exposure := {
"address": resource.address,
"protocol": object.get(resource.change.after, "ip_protocol", ""),
"from_port": object.get(resource.change.after, "from_port", null),
"to_port": object.get(resource.change.after, "to_port", null),
"cidr": cidr,
}
}
# 形式 4:已弃用的独立规则,仍存在于大多数现有基础设施中。
exposures contains exposure if {
some resource in input.resource_changes
resource.type == "aws_security_group_rule"
resource.change.after.type == "ingress"
some cidr in object.get(resource.change.after, "cidr_blocks", [])
exposure := {
"address": resource.address,
"protocol": object.get(resource.change.after, "protocol", ""),
"from_port": object.get(resource.change.after, "from_port", null),
"to_port": object.get(resource.change.after, "to_port", null),
"cidr": cidr,
}
}
# 规则实际覆盖的端口。当策略无法判断时,此值未定义。
covered_ports(exposure) := [0, 65535] if {
exposure.protocol in all_protocols
}
covered_ports(exposure) := [exposure.from_port, exposure.to_port] if {
not exposure.protocol in all_protocols
is_number(exposure.from_port)
is_number(exposure.to_port)
}
deny contains msg if {
some exposure in exposures
exposure.cidr in public_cidrs
# 若端口落在 [from_port, to_port] 区间内,则认为该规则覆盖此端口。
range := covered_ports(exposure)
some port in admin_ports
port >= range[0]
port <= range[1]
msg := sprintf(
"%s: 入站规则将端口 %d 暴露给 %s",
[exposure.address, port, exposure.cidr],
)
}
# 对公网开放但策略无法读取端口的规则会被报告。
# 放行此类规则相当于策略在宽松方向上进行猜测。
deny contains msg if {
some exposure in exposures
exposure.cidr in public_cidrs
not covered_ports(exposure)
msg := sprintf(
"%s: 针对 %s 的入站规则端口无法被此策略评估(%v 至 %v)",
[exposure.address, exposure.cidr, exposure.from_port, exposure.to_port],
)
}
# 此策略已知如何检查的地址。门控利用此集合来区分
# “无违规”与“未检查”。
considered contains resource.address if {
some resource in input.resource_changes
resource.type in {
"aws_security_group",
"aws_vpc_security_group_ingress_rule",
"aws_security_group_rule",
}
}
从上往下读这段代码:
Rego 中的规则体是合取逻辑。每一行都必须成立,而
some ... in语句负责迭代,因此 OPA 会遍历资源、Ingress 规则、字段和端口的所有组合。三条独立的
exposures规则共同定义了同一个集合,这在 Rego 中称为增量定义。日后若要增加第五种形态,只需多写一个代码块,下方的逻辑完全无需改动。object.get(ingress, field, [])在字段缺失时返回空列表,因此仅包含 IPv4 的规则在策略查找ipv6_cidr_blocks时不会报错。covered_ports在此处充当安全阀门。因为“全协议”规则在计划中报告端口范围从0到0,实际却打开了所有端口,所以不能字面解读端口字段;若某规则的端口为 null 或字符串,该函数将未定义,第二条deny规则会将其判定为违规。considered不参与最终判定。它记录的是该策略实际涉及哪些资源,第 5 步会利用这一点,避免报告策略并未真正覆盖的通过结果。
运行命令:
opa test policy
opa check --strict policy
opa fmt --diff policy
PASS: 20/20
给 opa test 加上 -v 参数,即可显示每条测试的详情。
opa check --strict 能捕获不安全变量和被遮蔽的导入,而 opa fmt --diff 在格式已符合规范时不输出任何内容。OPA 使用制表符格式化 Rego。这两者都应纳入 CI 流程,且优先级高于其他步骤。
第二个策略关于所有权标记,因此创建 policy/tags.rego:
# METADATA
# title: Every managed resource carries ownership tags
# description: |
# Terraform emits `tags: null` for a resource with no tags at all, so a
# policy that reaches into `after.tags` skips exactly the resources with
# the worst tagging. `tags_of` coerces that null to an empty object.
package terraform.tags
required_tags := {"owner", "cost-center", "data-classification"}
# Resource types that genuinely cannot carry tags.
untaggable := {"aws_iam_policy_attachment", "aws_route_table_association"}
in_scope contains resource if {
some resource in input.resource_changes
some action in resource.change.actions
action in {"create", "update"}
not resource.type in untaggable
}
tags_of(resource) := tags if {
tags := resource.change.after.tags
is_object(tags)
} else := {}
deny contains msg if {
some resource in in_scope
some tag in required_tags
value := object.get(tags_of(resource), tag, "")
trim_space(value) == ""
msg := sprintf("%s: missing required tag %q", [resource.address, tag])
}
considered contains resource.address if {
some resource in in_scope
}
这两行代码为什么这样写,原因如下:
真正的检查逻辑是
trim_space(value) == ""。必须这样写,因为在 Rego 中只有false和未定义才是假值,空字符串是真值。如果只做存在性检查,owner = ""也能通过——这正是标签策略本应杜绝的形式主义合规。tags_of的else := {}分支,是为了修复我自己写出的一个 bug——直到第 9 步才被发现。
第 4 步:对着真实计划运行
用 TerraGoat 计划评估一个包:
opa eval --data policy --input plan.json --format pretty 'data.terraform.network.deny'
[
"aws_security_group.web-node: ingress rule exposes port 22 to 0.0.0.0/0"
]
策略只报告了一个违规项,对端口 80 保持沉默。同一个资源里端口 80 也是对外开放的,但公网 Web 服务器本来就该如此。如果检查把两者都标记出来,人们很快就会学会无视它。
每个包跑一条命令没法扩展,好在 Rego 可以用一条查询聚合整个命名空间:
opa eval --data policy --input plan.json --format pretty \
'union({v | v := data.terraform[_].deny})'
[
"aws_security_group.web-node: 入站规则将 22 端口暴露给了 0.0.0.0/0",
"aws_security_group.web-node: 缺少必需标签 \"cost-center\"",
"aws_security_group.web-node: 缺少必需标签 \"data-classification\"",
"aws_security_group.web-node: 缺少必需标签 \"owner\"",
"aws_vpc.web_vpc: 缺少必需标签 \"cost-center\"",
"aws_vpc.web_vpc: 缺少必需标签 \"data-classification\"",
"aws_vpc.web_vpc: 缺少必需标签 \"owner\""
]
{v | v := data.terraform[_].deny} 是一个推导式,用于收集 data.terraform 下所有包中的 deny 集合,union 会将它们扁平化。只需在该命名空间下放入新的策略文件即可自动生效,命令无需更改。
步骤 5:将判定结果转化为退出码
CI 门禁通过退出状态码进行沟通。如果将两种不同性质的失败混用同一个状态码,会导致损坏的流水线在一个月内被误判为通过。该工具使用以下规则:
| 状态码 | 判定结果 | 含义 |
|---|---|---|
| 0 | pass(通过) | 策略已执行,已检查资源,未发现违规 |
| 1 | fail(失败) | 策略已执行,发现违规 |
| 2 | vacuous(空转) | 策略已执行,但未检查到任何相关资源,结果无效 |
| 2 | broken(损坏) | 工具或其输入不可用 |
第四行是大多数门禁最常出错的环节。如果计划文件中没有策略关注的任何资源,报告“通过”便是在提供此次运行无法保证的安全性。它拥有独立的判定名称,且不返回退出码 0。
创建 policy_gate.py:
"""使用 Rego 策略目录评估 Terraform 部署计划。"""
import argparse
import json
import pathlib
import shutil
import subprocess
import sys
PASS, FAIL, BROKEN = 0, 1, 2
DENY_QUERY = "union({v | v := data.terraform[_].deny})"
CONSIDERED_QUERY = "union({v | v := data.terraform[_].considered})"
def die(message: str) -> None:
print(f"policy-gate: {message}", file=sys.stderr)
sys.exit(BROKEN)
def load_plan(path: pathlib.Path) -> dict:
try:
text = path.read_text()
except OSError as exc:
die(f"无法读取 {path}: {exc.strerror}")
try:
return json.loads(text)
except json.JSONDecodeError as exc:
die(f"{path}:{exc.lineno}:{exc.colno}: 无效的 JSON: {exc.msg}")
def query(opa: str, policy_dirs: list[pathlib.Path], plan: pathlib.Path, expr: str) -> list:
command = [opa, "eval", "--input", str(plan), "--format", "raw"]
for directory in policy_dirs:
command += ["--data", str(directory)]
command.append(expr)
result = subprocess.run(command, capture_output=True, text=True)
if result.returncode != 0:
die(f"OPA 执行失败: {result.stderr.strip() or result.stdout.strip()}")
values = json.loads(result.stdout)
# 规则若产生非字符串值属于策略错误;对混合类型列表排序会引发无关联的 TypeError
for value in values:
if not isinstance(value, str):
die(f"{expr} 产生了 {type(value).__name__} 类型值;"
"deny 和 considered 规则必须返回字符串")
return values
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--plan", required=True, type=pathlib.Path,
help="来自 `terraform show -json` 的 JSON 文件")
parser.add_argument("--policy", required=True, nargs="+", type=pathlib.Path,
help="一个或多个包含 .rego 文件的目录")
parser.add_argument("--opa", default="opa", help="opa 可执行文件路径")
args = parser.parse_args()
if shutil.which(args.opa) is None:
die(f"{args.opa} 不在 PATH 中")
for directory in args.policy:
if not directory.is_dir():
die(f"{directory} 不是目录")
plan = load_plan(args.plan)
if "resource_changes" not in plan:
die(f"{args.plan} 缺少 resource_changes 键;是否为 Terraform 计划文件?")
considered = query(args.opa, args.policy, args.plan, CONSIDERED_QUERY)
if not considered:
# 此处报告通过会声称一种本次运行无法保证的安全性
print(f"无效通过: 未有任何策略检查这 {len(plan['resource_changes'])} 个计划资源",
file=sys.stderr)
return BROKEN
violations = sorted(query(args.opa, args.policy, args.plan, DENY_QUERY))
if violations:
print(f"未通过: 在 {len(considered)} 个检查资源中发现 {len(violations)} 个违规项",
file=sys.stderr)
for violation in violations:
print(f" - {violation}", file=sys.stderr)
return FAIL
print(f"通过: 检查了 {len(considered)} 个资源,无违规")
return PASS
if __name__ == "__main__":
sys.exit(main())
这个脚本做了以下几件事:
argparse把--plan和--policy标记为required=True,且--policy使用nargs="+",因此策略列表为空时会在解析阶段直接报错。load_plan会报告 JSON 语法错误的行号和列号,因为json.JSONDecodeError自带lineno和colno——一个只说"JSON 无效"的关卡会让别人排查一下午。所有失败路径都通过
die处理,统一以退出码 2 结束:缺少opa、文件不可读、计划中没有resource_changes键,这些都属于工具本身的故障。considered查询在deny查询之前运行。如果什么都没检查,运行结果就是VACUOUS,根本没有机会输出通过。违规项会排序,因此同一份计划每次运行产生的输出完全一致(逐字节相同),CI 日志之间的 diff 才有意义。
这份计划中的两个资源共有 7 处违规,且退出码可供 CI 直接处理。
同一个工具面对空计划的结果——这种情况下关卡如果回答"通过"就是在撒谎。
GitHub Actions 工作流会先测试这些策略,再用它们去做判定:
name: policy
on: [pull_request]
jobs:
policy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- name: Install OPA
run: |
curl -L -o /usr/local/bin/opa \
https://openpolicyagent.org/downloads/v1.20.2/opa_linux_amd64_static
chmod +x /usr/local/bin/opa
# The policies are code. Lint and test them before trusting them.
- name: Check policy syntax
run: opa check --strict policy
- name: Verify formatting
run: opa fmt --fail --diff policy
- name: Test policies
run: opa test policy --verbose --coverage --format json > coverage.json
# Only now does anything get judged.
- name: Evaluate Terraform plan
run: python3 policy_gate.py --plan plan.json --policy policy
先让这道门禁只上报结果,持续运行两周,再正式启用拦截功能。那些看起来毫无问题的策略,往往会因你现有的基础设施结构而失效。你希望在日志里发现问题,而不是在发布被阻断后才察觉。
Step 6: 在准入阶段强制执行
步骤 5 中的门禁检查的是你打算部署的内容。它无法看到从某台笔记本电脑发起的 kubectl apply、供应商的 Helm 图表,或 Operator 自行创建的 Pod。针对这些场景,你需要准入控制,而 Kubernetes 现已内置该功能。
ValidatingAdmissionPolicy 自 v1.30 起已正式发布,它在 API server 内直接评估 CEL 表达式,无需部署或维护 Webhook:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: require-trusted-registry
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
variables:
# A Pod has three container lists. A policy that reads only
# spec.containers is bypassed by moving the image to an initContainer.
- name: allImages
expression: >-
object.spec.containers.map(c, c.image) +
object.spec.?initContainers.orValue([]).map(c, c.image) +
object.spec.?ephemeralContainers.orValue([]).map(c, c.image)
validations:
- expression: >-
variables.allImages.all(i, i.startsWith('registry.internal.example.com/'))
messageExpression: >-
'images must come from registry.internal.example.com: ' +
variables.allImages.filter(i,
!i.startsWith('registry.internal.example.com/')).join(', ')
reason: Forbidden
策略本身不会生效,直到有绑定将其激活。这种机制允许你先在单个命名空间进行试点:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: require-trusted-registry-binding
spec:
policyName: require-trusted-registry
validationActions: ["Deny"]
matchResources:
namespaceSelector:
matchLabels:
policy.example.com/enforce: "true"
将 validationActions 设为 ["Warn", "Audit"],给一个命名空间打上标签,观察一周,然后切换为 ["Deny"] 并扩大选择器范围。
对 initContainers 的处理很关键,因为绕过镜像来源策略最常见的方式就是把镜像移到 initContainer 中。使用 ? 可选字段语法配合 .orValue([]) 可以安全地读取可能不存在的列表,避免整个表达式报错。
由于手头没有集群,我无法应用这两个清单文件。这些清单已对照 v1 参考架构进行检查,但从未被真实的 API 服务器接受,因此请将其视为起点,并在 Warn 模式下逐步推进(本就应该这样操作)。
另外两个引擎在生产环境中也被广泛使用,第一个是 Kyverno。它已于 2026 年 3 月从 CNCF 毕业,被 Bloomberg、Coinbase、Deutsche Telekom、LinkedIn 和 Spotify 用于生产环境。它的策略用 YAML 编写,平台团队无需学习新语言,而且它还能处理内置策略不管的生成、镜像签名验证和清理工作。
如果你希望用一个 Rego 代码库同时覆盖 Kubernetes、Terraform 和 CI,那么 OPA Gatekeeper 是正确的选择——本教程正是朝着这个方向构建的。Mutation 功能现在也已内置:MutatingAdmissionPolicy 在 v1.36 中转为稳定版。
第 7 步:让模型来写策略
写策略很繁琐,而模型最擅长处理繁琐。所以顺理成章的做法就是让模型来写。
但有一个坑,你自己花一分钟就能验证出来——我在本步骤末尾就是这么做的:公开训练数据中的大量 Rego 代码是 Rego v0 方言,随着 2025 年 1 月 OPA 1.0 发布,这种方言已无法被解析。模型倾向于使用它见过的最常见写法,结果就是写出当前解析器拒绝的方言。
卡拉布里亚大学团队 2025 年的一篇预印本论文 ARPaCCino 在 Terraform 案例中报告了同样的现象:在不使用工具的情况下直接要求生成 Rego,Qwen3-30B 和 GPT-4o 各自生成的 5 条策略中语法正确的都是 0 条。在 OPA 文档上增加检索也没有任何改善。但给模型一个可以运行 opa check 并读取报错的循环后,正确率分别提升到 5 条中的 4 条和 5 条。
这些数字来自一个小型案例研究,所以重点是把握其结论方向,而非具体数字。
解决问题的正是一个反馈循环,而且这个循环不需要任何额外成本——因为你已经在用 opa check --strict、opa fmt 和 opa test 把它搭建好了。
搭建这条闭环的关键,在于一处反转来确保可信度:人来写测试,模型来写策略。测试用例具体且审阅成本低——你扫一眼 JSON 数据就能在三秒内确认“该拒绝”。而包含嵌套推导式的 Rego 规则阅读门槛较高,且极易误读。让人类承担审阅成本低的工作,让机器在输出可被机械化校验的地方发挥作用。
新建 policy_forge.py 文件:
"""从一条英文规则生成 Rego 策略,只有通过工具链验证才保留。"""
import argparse
import pathlib
import re
import subprocess
import sys
import tempfile
WRITTEN, REJECTED, BROKEN = 0, 1, 2
SYSTEM = """你编写 Open Policy Agent 的 Rego v1(OPA 1.0+)策略。
规则:
- 每个规则体都要用 `if`,多值规则用 `contains`。
- 不要输出 `import rego.v1`;在 OPA 1.0+ 上它是多余的。
- 输入是 `terraform show -json` 生成的 JSON。
- 只返回一个 ```rego 代码块,不要输出其他内容。"""
def extract_rego(reply: str) -> str:
blocks = re.findall(r"```rego\n(.*?)```", reply, re.DOTALL)
if not blocks:
raise ValueError("model returned no rego block")
if len(blocks) > 1:
raise ValueError(f"model returned {len(blocks)} rego blocks; expected one")
return blocks[0]
def verify(opa: str, policy: str, tests: pathlib.Path) -> tuple[bool, str]:
with tempfile.TemporaryDirectory() as tmp:
bundle = pathlib.Path(tmp)
# 测试文件保留原文件名;策略文件用一个不会
# 与之冲突的名字,无论调用方把测试文件命名成什么。
(bundle / "candidate_policy.rego").write_text(policy)
(bundle / tests.name).write_text(tests.read_text())
for command in ([opa, "check", "--strict"], [opa, "test"]):
result = subprocess.run(command + [str(bundle)], capture_output=True, text=True)
if result.returncode != 0:
return False, (result.stdout + result.stderr).strip()
return True, "opa check and opa test both passed"
def forge(rule: str, tests: pathlib.Path, ask, opa: str, attempts: int) -> str:
transcript = [{
"role": "user",
"content": (
f"为这条规则编写 Rego 策略:\n\n{rule}\n\n"
f"它必须满足以下测试:\n\n```rego\n{tests.read_text()}```"
),
}]
for attempt in range(1, attempts + 1):
reply = ask(transcript)
policy = extract_rego(reply)
ok, output = verify(opa, policy, tests)
headline = next(iter(output.splitlines()), "no output from the toolchain")
print(f"attempt {attempt}: {'PASS' if ok else 'FAIL'} - {headline}", file=sys.stderr)
if ok:
return policy
transcript += [
{"role": "assistant", "content": reply},
{"role": "user", "content": f"工具链拒绝了上面的策略:\n\n{output}\n\n请修复。"},
]
raise RuntimeError(f"no policy survived {attempts} attempts")
def claude(model: str):
import anthropic
client = anthropic.Anthropic()
def ask(transcript: list[dict]) -> str:
response = client.messages.create(
model=model,
max_tokens=16000,
system=[{
"type": "text",
"text": SYSTEM,
"cache_control": {"type": "ephemeral"},
}],
thinking={"type": "adaptive"},
messages=transcript,
)
return "".join(b.text for b in response.content if b.type == "text")
return ask
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--rule", required=True, help="the policy, in one English sentence")
parser.add_argument("--tests", required=True, type=pathlib.Path,
help="a _test.rego file you wrote by hand")
parser.add_argument("--out", required=True, type=pathlib.Path,
help="where to write the policy, only if it passes")
parser.add_argument("--model", default="claude-opus-5")
parser.add_argument("--attempts", type=int, default=4)
parser.add_argument("--opa", default="opa")
args = parser.parse_args()
if args.attempts < 1:
print("policy-forge: --attempts must be at least 1", file=sys.stderr)
return BROKEN
if not args.tests.is_file():
print(f"policy-forge: {args.tests} does not exist", file=sys.stderr)
return BROKEN
try:
policy = forge(args.rule, args.tests, claude(args.model), args.opa, args.attempts)
except RuntimeError as exc:
print(f"policy-forge: {exc}; nothing written", file=sys.stderr)
return REJECTED
except ValueError as exc:
print(f"policy-forge: {exc}", file=sys.stderr)
return BROKEN
args.out.write_text(policy)
print(f"policy-forge: verified policy written to {args.out}", file=sys.stderr)
return WRITTEN
if __name__ == "__main__":
sys.exit(main())
export ANTHROPIC_API_KEY=...
python3 policy_forge.py \
--rule "No security group may expose an administrative port to the public internet." \
--tests policy/network_test.rego \
--out policy/generated.rego
这个循环提供了以下保证:
验证以子进程方式运行:从不问模型"你的策略是否正确"。由
opa check和opa test来裁决,它们的退出码是循环唯一采信的证据。失败信息以原始工具输出返回,不做摘要:编译错误和测试失败是模型能收到的信息量最大的反馈,改写反而会丢掉最有用的部分。
不通过就不落盘:
forge要么返回经过验证的策略,要么直接抛出异常,因此不会出现"重试次数用完了,未验证的策略就进了仓库"的情况。
我用脚本模拟的模型驱动这个循环,这样即使没有 API key 也能复现结果。三次回复分别是:一个符合 v0 语法的策略、一个看似合理但有错的 v1 策略,以及第 3 步的那个策略:
attempt 1: FAIL - 2 errors occurred during loading:
attempt 2: FAIL - policy/network_test.rego:63:
attempt 3: PASS - opa check and opa test both passed
第一次是 Rego v0,也就是 deny[msg] { ... } 的写法——OPA 1.0 在 2025 年 1 月发布后就不再解析这种语法了。而公开训练数据里绝大多数都是这种写法。opa check --strict 在测试之前就把它拒之门外。
第二次是合法的 Rego v1。大多数工程师审查时都可能放行,但它在测试套件上只拿到了 8 项中的 4 项。原因是它直接把 ingress.from_port 与 admin_ports 比较而忽略了范围,并且只读取了 cidr_blocks。opa check 没有任何报错,因为代码在语法上完全没问题。
一个语法完全正确的策略,仍然可能是错误的策略——而整个循环里唯一知道你真实意图的,就是你亲手写下的测试套件。
修复循环不是单调递进的:每次尝试都是基于错误信息的一次全新生成,之前已经奏效的东西不会延续下来,所以第四次尝试可能丢失第三次已经具备的属性。这套设计里没有任何机制能检测到这一点,因为被检查的只有你写的那个测试套件。 给重试次数设上限,让测试套件持续扩充,并把每一条生成的策略都当作一个 pull request,必须有人批准后才能合并。第 8 步:治理 Agent 本身
AI agent 也是一个行为主体,它会调用工具,因此每次工具调用都是一个必须有人做出的授权决策。 业界很快就对此达成共识:Amazon Bedrock AgentCore Policy 于 2026 年 3 月正式发布,在网关层用 Cedar 评估 agent 的工具调用。常见的开源做法是在 MCP 工具网关前面放一个 OPA sidecar。 研究也在把这条边界往前推:来自华盛顿大学(Defects4J 背后的团队)的一篇 2026 年预印本论文,Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents(Winston、Winston 和 Just),把自然语言策略编译成 SMT 约束,并用 Z3 求解器拦截不合规的调用。 这些工作指向同一个结论:写在 system prompt 里的策略不等于执行。真正的执行,是坐在调用链路上的拦截器,它能够返回"拒绝",让这次调用根本不发生。 创建agent/authz.rego:
# METADATA
# title: Agent 工具调用授权
# description: |
# 在每次工具调用执行前评估一次。决策结果有三个值而不是两个,
# 因为一个值得部署的 agent 有时需要做一些事情,应该由人来批准,
# 而不是由策略直接放行。
package agent.authz
tool_grants := {
"support": {"search_orders", "read_customer", "issue_refund"},
"analytics": {"search_orders", "run_query"},
}
write_tools := {"issue_refund", "run_query"}
refund_ceiling_cents := 10000
# 未映射的角色、未知工具或格式错误的输入都会落到这里。
default decision := {"effect": "deny", "reasons": ["no matching grant"]}
decision := {"effect": effect_for(reasons), "reasons": reasons} if {
count(granted) > 0
reasons := escalations
}
granted contains role if {
some role in input.agent.roles
input.tool in object.get(tool_grants, role, set())
}
effect_for(reasons) := "allow" if count(reasons) == 0
effect_for(reasons) := "require_approval" if count(reasons) > 0
escalations contains reason if {
input.tool in write_tools
not input.session.human_in_loop
reason := sprintf("%q writes state and the session is unattended", [input.tool])
}
# 如果退款请求中读不到金额,就无法对照上限做检查,因此需要升级处理。
# 如果这里静默放行,恰好会放过攻击者精心构造的那类调用。
escalations contains reason if {
input.tool == "issue_refund"
not positive_amount
reason := "refund amount is missing, unreadable, or not positive"
}
# 负数金额就是披着退款外衣的扣款。
positive_amount if {
amount := object.get(input, ["arguments", "amount_cents"], null)
is_number(amount)
amount > 0
}
escalations contains reason if {
input.tool == "issue_refund"
amount := object.get(input, ["arguments", "amount_cents"], null)
is_number(amount)
amount > refund_ceiling_cents
reason := sprintf(
"refund of %d cents exceeds the %d cent ceiling",
[amount, refund_ceiling_cents],
)
}
# 关键词匹配只是粗略的防护,这里主要是展示参数级规则的写法。
# 真正涉及数据的场景需要 SQL 解析器:这种方式能拦住 `DROP TABLE`,
# 但拦不住换一种写法的同类语句。
destructive_sql := `(?i)\b(drop|truncate|delete|alter|grant|revoke)\b`
escalations contains reason if {
input.tool == "run_query"
regex.match(destructive_sql, object.get(input, ["arguments", "statement"], ""))
reason := "statement contains a destructive SQL keyword"
}
决策词汇是封闭的,直接列在这里:
| Effect | 调用方的行为 |
|---|---|
| allow | 执行该工具 |
| require_approval | 暂停,把原因展示给人工,批准后才执行 |
| deny | 拒绝,且不提供人工审批通道 |
策略为什么要这样设计:
default decision默认为 deny:无法识别的工具、忘记映射的角色、格式错误的输入,最终都会落到这里。默认放行的策略恰恰会在没人预料到的输入上失败敞开——而攻击者挑的正是这些输入。三个值:非黑即白的授权只能在“拦住有用的工作”和“放行危险的工作”之间二选一,正是第三个值让高自主 agent 变得可以接受。
原因以集合返回:所有适用的原因都会被收集。凌晨三点收到审批提示时,“退款 250000 美分超过 10000 美分上限”能让人知道该怎么办;“违反策略”则不行。
参数也要检查:同样是
issue_refund,5 英镑是日常操作,2500 英镑就是大事。既然参数是 agent 自己选的,只看工具名的粒度就太粗了。
十五个测试覆盖了整张决策表,包括角色列表为空的 agent、没有任何金额的退款,以及用换行符藏起 DROP 的查询:
opa test agent -v
PASS: 15/15
启动服务,试一次调用:
opa run --server --addr localhost:8181 agent/
curl -s localhost:8181/v1/data/agent/authz/decision \
-d '{"input":{"agent":{"roles":["support"]},"tool":"issue_refund",
"arguments":{"amount_cents":250000},"session":{"human_in_loop":true}}}' | jq .result
{
"effect": "require_approval",
"reasons": [
"refund of 250000 cents exceeds the 10000 cent ceiling"
]
}
在有东西根据策略的答复拒绝继续执行之前,这份策略只是一纸空文。所以还需要一个客户端:
import json
import urllib.request
OPA_URL = "http://localhost:8181/v1/data/agent/authz/decision"
class PolicyDenied(Exception):
pass
class ApprovalRequired(Exception):
pass
def authorize(agent, tool, arguments, session):
payload = json.dumps({"input": {
"agent": agent, "tool": tool,
"arguments": arguments, "session": session,
}}).encode()
req = urllib.request.Request(
OPA_URL, data=payload, headers={"Content-Type": "application/json"}
)
with urllib.request.urlopen(req, timeout=2) as resp:
body = json.load(resp)
# OPA 在查询无匹配结果时会返回 200 和 {}。必须默认拒绝。
decision = body.get("result", {"effect": "deny", "reasons": ["policy unavailable"]})
if decision["effect"] == "deny":
raise PolicyDenied("; ".join(decision["reasons"]))
if decision["effect"] == "require_approval":
raise ApprovalRequired("; ".join(decision["reasons"]))
return decision
search_orders -> 允许
issue_refund -> 需要审批(退款金额 250000 美分,超过了 10000 美分的上限)
delete_account -> 拒绝(没有匹配的授权)
注意 body.get("result", ...) 这里:当查询没有匹配结果时,OPA 会返回状态码 200 和一个空的 {},所以直接写 body["result"] 会抛出 KeyError,而根据你的 agent 框架处理异常的方式,这可能导致系统默认放行(fail open)。每一层都要默认拒绝,包括解析环节。
在你的框架的工具执行钩子中调用 authorize(),位置放在工具函数真正执行之前。总共也就十几行代码,但它能把系统提示词里那些"客气的建议"变成真正的边界。
第 9 步:我踩过的坑
第 2 步和第 3 步展示的是最终的策略版本。我最初尝试的是更简单的写法(也就是大多数教程止步于的版本),而那版和你在上文看到的版本之间的差距,正是这一部分最有价值的内容。
打标签策略漏掉了最棘手的资源
我最初天真的 in_scope 规则以 resource.change.after.tags 结尾,字面意思是"只保留带有标签的资源"。
但它的实际行为比这更糟:Terraform 对完全没有标签的资源会输出 tags: null,而对未定义字段的查询会让规则体判定失败,于是这类资源就彻底脱离了管控范围。
TerraGoat 计划里有两个资源:security group 带 5 个 git_* 标签、没有所有权标签,而 VPC 则一个标签都没有。
jq -r '.resource_changes[] | "\(.address): tags=\(.change.after.tags | type)"' plan.json
aws_security_group.web-node: tags=object
aws_vpc.web_vpc: tags=null
opa eval --data naive --input plan.json --format pretty 'count(data.terraform.tags.deny)'
opa eval --data policy --input plan.json --format pretty 'count(data.terraform.tags.deny)'
3
6
naive 版漏掉的 3 条都来自那个完全没打标签的资源——策略只揪出了带部分标签的那个,却对没标签的那个视而不见。
tags_of 加上 else := {} 分支就是解法,总共三行代码。
网络策略只认四种写法中的一种
我最初的朴素网络策略只读 aws_security_group 和 cidr_blocks,这也是所有教程里的写法。但 Terraform 表达同一条 ingress 规则其实有四种方式。
![图示标题为"Terraform 描述同一条 ingress 规则的四种方式"。图中显示四个方框。左上角是 aws_security_group.ingress[].cidr_blocks,填充淡蓝色并标注为已读取。另外三个用红色轮廓和红色斜线标注为未读取:aws_security_group.ingress[].ipv6_cidr_blocks、aws_vpc_security_group_ingress_rule.cidr_ipv4,以及已弃用的 aws_security_group_rule。图例说明:蓝色实心填充表示策略会检查此处,红色斜线表示不检查。](https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/ab61d36d-78bf-45f7-a073-e7013216e3e1.png align="center")
同一条规则,四种编码,而朴素策略只读了左上角那一种。
我又规划了一个包含三个 security group 的真实配置,它们利用策略读不到的那几种写法把 SSH 对全世界开放。两个版本都在 code/ 目录里,可以复现:
opa eval --data naive --input plan-evasion.json \
--format pretty 'data.terraform.network.deny'
[]
结果是:三个暴露在公网的 SSH 端口,零违规,门禁还会打印 PASS 并以退出码 0 放行。
一个自己留着漏洞的测试
端口范围还有第二个问题,而元凶正是我自己写的测试。我写过一个叫 test_tolerates_null_ports 的用例,断言 protocol: "-1" 的入站规则不会产生任何违规,理由是与空端口做比较时,策略不应该崩溃。
但全协议规则会开放所有端口,AWS provider 会把它记为 from_port: 0 和 to_port: 0。范围检查把这两个值字面地理解成了"端口 0"这一个端口,结果 AWS 里最宽松的规则反而拿到了"零违规"的评价。
对应的 plan 在 code/plan-all-protocols.json,只包含一个资源:
jq -c '.resource_changes[] | select(.type=="aws_security_group")
| .change.after.ingress[0] | {protocol,from_port,to_port,cidr_blocks}' \
plan-all-protocols.json
opa eval --data naive --input plan-all-protocols.json \
--format pretty 'data.terraform.network.deny'
{"protocol":"-1","from_port":0,"to_port":0,"cidr_blocks":["0.0.0.0/0"]}
[]
所有协议、所有端口,对整个互联网敞开,却被我特意写下的一个测试认证为合规。解决办法是引入 covered_ports:把全协议规则映射到完整的端口范围,对无法解析的端口则返回 undefined,并新增一条 deny 规则来报告这种 undefined 情况。原来那个测试删掉了,取而代之的是四个新测试:
opa eval --data policy --input plan-all-protocols.json --format pretty 'data.terraform.network.deny'
[
"aws_security_group.wide_open: ingress rule exposes port 22 to 0.0.0.0/0",
"aws_security_group.wide_open: ingress rule exposes port 3306 to 0.0.0.0/0",
"aws_security_group.wide_open: ingress rule exposes port 3389 to 0.0.0.0/0",
"aws_security_group.wide_open: ingress rule exposes port 5432 to 0.0.0.0/0"
]