← 文章 / 云原生与基础设施
freeCodeCamp 1小时前 · 2026-09-30 06:15:54 · 1 阅读

如何用 Policy as Code 和 OPA 治理 AI 生成的基础设施代码(完整手册)

现代大模型生成的代码几乎每次都能保证语法正确。Veracode 的 2026 年报告直言不讳:“语法问题基本上已经解决了。”

这听起来像是一个里程碑,但这恰恰是你面临问题的根源。

该同一份报告测试了一百多个模型,发现平均安全通过率为 56%,“与首份报告的 55% 相比几乎没有变化”,其中约 44% 的生成任务引入了具有风险的安全漏洞。

功能正确性与安全性其实是两个独立的问题,而目前只有其中一项接近解决。

这一结果既不是个例,也不是新现象。早在 2022 年,纽约大学坦登工学院的一个团队在 IEEE Security and Privacy 会议上展示了 GitHub Copilot 在 89 个安全相关场景下的表现,生成了 1689 个程序,结果发现其中约有 40% 存在漏洞,涉及 MITRE 的 CWE Top 25 类别。这篇论文后来还被评为 Communications of the ACM 的研究亮点。

2024 年 11 月,乔治城大学新兴技术与安全中心(CSET)评估了五个 LLM,报告称其中近半数的代码片段包含可能导致被利用的 bug。四年时间,四支独立团队,四种不同的方法论,答案始终一致。

在 2023 年的 ACM CCS 会议上,斯坦福大学的 Neil Perry、Megha Srivastava、Deepak Kumar 和 Dan Boneh 将开发者纳入了实验:47 名参与者,五项与安全性相关的编程任务,三种编程语言,其中 33 人使用了 AI 助手,14 人未使用。使用助手的组写出的代码安全性显著更低,且他们更容易相信自己写的代码是安全的。

虽然这是一项规模较小的研究,但它解释了为什么问题无法自我修正:通常能发现这类问题的机制(开发者仔细审视那些让他们担忧的代码)恰恰是工具正在关闭的机制。

上述研究测量的都是应用代码。但基础设施代码是更难处理的情况,因为错误的安全组配置永远不会“失败”:它会完全按照编写逻辑工作,为任何请求者提供流量服务,唯一能提出异议的是正在阅读 diff 的人。

光靠人眼看那份配置文件,谁也审不过来,我也不例外。能做的,是把规则写成计算机能自动校验的形式,每次变更都检查一遍——这就是 Policy as Code(策略即代码)的含义。

在本手册中,我会带你一步步搭建这套校验机制。我们会拿一个真实存在漏洞的仓库练手,先看它在九条违规中“顺利通过”五条的荒唐表现,然后动手修复。

读完之后,你将学会:

  • 针对 terraform show -json 输出的 JSON 编写 Rego 策略。

  • 像测试应用代码一样测试策略,配备测试夹具(fixtures)和覆盖率报告。

  • 构建一个命令行门禁,通过退出码约定让 CI 流水线可以放心依赖。

  • 使用 CEL 在 Kubernetes 准入阶段拦截不合规的工作负载。

  • 让模型来写策略,再用 opa check 和你自己的测试决定是否采纳。

  • 用同一个策略引擎对 AI agent 的工具调用做授权。

Table of Contents

Prerequisites

你需要准备:

  • 一个终端,以及可用的 python3(3.10 或更高版本)。

  • jq,用于在命令行解析 JSON。

  • 约 700 MB 磁盘空间,因为 AWS Terraform provider 文件较大。

  • 一个 Anthropic API key,但仅在第七步使用。其他所有步骤均可离线运行。

mkdir policy-lab && cd policy-lab
python3 -m venv .venv
source .venv/bin/activate
pip install anthropic

curl -L -o opa https://openpolicyagent.org/downloads/v1.20.2/opa_darwin_arm64_static
chmod +x opa && sudo mv opa /usr/local/bin/

curl -L -o tf.zip https://releases.hashicorp.com/terraform/1.14.2/terraform_1.14.2_darwin_arm64.zip
unzip tf.zip && sudo mv terraform /usr/local/bin/

在 Windows 上,通过 .venv\Scripts\activate 激活环境,并将上述两个下载链接替换为 opa_windows_amd64.exe 和 terraform_1.14.2_windows_amd64.zip。

以下所有操作均在 macOS 上,基于 OPA 1.20.2、Terraform 1.14.2 和 AWS provider 6.x 运行。OPA 1.x 系列之间的策略语法保持稳定。

如果你使用的是 OPA 0.x,需要在每个规则文件顶部添加 import rego.v1,但我建议直接升级到 1.x。违规项统计仅受 AWS provider 版本影响,具体取决于 plan JSON 的结构,该结构自 provider 5.x 以来一直稳定。

关键术语的通俗解释

  • Policy as Code(策略即代码):将组织已达成共识的规则编写为程序,输入变更提案,输出判定结果。

  • Rego:Open Policy Agent 使用的查询语言。它是声明式的,规则主体由必须全部满足的条件列表组成。

  • Plan JSON:Terraform 即将执行的变更计划之机器可读描述,通过 terraform show -json 生成。策略直接读取此内容,而非 .tf 文件。

  • Admission control(准入控制):Kubernetes API server 内部的一个节点,对象在持久化存储前可在此处被拒绝。

  • CEL:Common Expression Language,Kubernetes 在 API server 内部原生支持的小型表达式语言,无需部署 webhook。

  • 漏报:本应拦截的资源却通过了策略。这类问题几乎无人察觉,除非专门去排查,否则根本无法量化。

  • 标题为“三个决策点,三次否决机会”的图表,包含三行。计划时行:从 Terraform plan JSON 流向 policy_gate.py,最终合并或阻止 PR。准入时行:从 kubectl apply 流向 ValidatingAdmissionPolicy,最终允许或拒绝 Pod。调用时行:从 Agent 选择工具流向 agent.authz 决策,最终允许、拒绝或询问人类。箭头从左至右,从组件指向产物,底部说明文字为:每个边界做一次决策:API 服务器内部使用 CEL,Rego 可在两侧使用。

    相同的判断逻辑存在于三个环节,其中只有中间那个是 Kubernetes 特有的。

    第一步:获取真实的基础设施进行测试

    我不想凭空捏造一个有漏洞的 Terraform 文件,因为那样等于先制造 bug,然后策略只能拦住我自己埋下的雷。所以我转而寻找别人编写并发布过的代码。

    TerraGoat 是 Bridgecrew 发布的一个故意包含漏洞的 Terraform 仓库。获取其 EC2 模块的指定提交版本:

    SHA=729f8da62c6a85ce4af5ad3d123de97776d954c4
    curl -s "https://raw.githubusercontent.com/bridgecrewio/terragoat/$SHA/terraform/aws/ec2.tf" \
      | sed -n '77,96p'
    
    resource "aws_security_group" "web-node" {
      # 安全组对公网开放了 SSH 端口
      name        = "${local.resource_prefix.value}-sg"
      description = "${local.resource_prefix.value} Security Group"
      vpc_id      = aws_vpc.web_vpc.id
    
      ingress {
        from_port = 80
        to_port   = 80
        protocol  = "tcp"
        cidr_blocks = [
        "0.0.0.0/0"]
      }
      ingress {
        from_port = 22
        to_port   = 22
        protocol  = "tcp"
        cidr_blocks = [
        "0.0.0.0/0"]
      }
    

    第二行的注释来自 TerraGoat 原文,而 22 端口对公网开放正是该文件指出的风险点。

    TerraGoat 的模块在现代 Terraform 上无法初始化,因为它还在用带引号的 type = "string",这种写法在 Terraform 0.12 中已被弃用,1.x 更是直接拒绝。所以我把相关资源搬到了自己的一个精简模块里,只把两处对 TerraGoat 内部 locals 的引用换成了字面值。

    创建 main.tf:

    ...(代码保持不变)

    生成 plan JSON:

    terraform init
    terraform plan -out=tfplan.binary
    terraform show -json tfplan.binary > plan.json
    

    这里的 mock 凭证很关键:纯创建型配置在 terraform plan 时不会调用 AWS,因此配合 skip_credentials_validation 及另外三个 skip 参数,provider 不会尝试认证,也不会有任何内容被真正 apply。

    第 2 步:先写测试,再写策略

    我要实现的规则是:禁止任何安全组向公网暴露管理端口。

    这句话听起来简单到只需一行代码,但测试部分恰恰说明了它的复杂性。创建 policy/network_test.rego 文件:

    package terraform.network_test
    
    import data.terraform.network
    
    plan(resources) := {"resource_changes": resources}
    
    security_group(ingress) := {
    	"address": "aws_security_group.web",
    	"type": "aws_security_group",
    	"change": {"actions": ["create"], "after": {"ingress": [ingress]}},
    }
    
    # 拒绝 SSH 对全网开放
    test_denies_ssh_open_to_the_world if {
    	fixture := plan([security_group({
    		"from_port": 22,
    		"to_port": 22,
    		"protocol": "tcp",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 简单的 from_port 等值检查会漏掉此情况,但区间检查不会。
    test_denies_wide_open_port_range if {
    	fixture := plan([security_group({
    		"from_port": 0,
    		"to_port": 65535,
    		"protocol": "tcp",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 4 with input as fixture
    }
    
    # 拒绝 IPv6 对全网开放
    test_denies_ipv6_route_to_the_world if {
    	fixture := plan([security_group({
    		"from_port": 22,
    		"to_port": 22,
    		"protocol": "tcp",
    		"ipv6_cidr_blocks": ["::/0"],
    	})])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 拒绝独立的入站规则
    test_denies_standalone_ingress_rule if {
    	fixture := plan([{
    		"address": "aws_vpc_security_group_ingress_rule.ssh",
    		"type": "aws_vpc_security_group_ingress_rule",
    		"change": {"actions": ["create"], "after": {
    			"from_port": 22,
    			"to_port": 22,
    			"ip_protocol": "tcp",
    			"cidr_ipv4": "0.0.0.0/0",
    			"cidr_ipv6": null,
    		}},
    	}])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 拒绝已弃用的独立规则
    test_denies_deprecated_standalone_rule if {
    	fixture := plan([{
    		"address": "aws_security_group_rule.ssh",
    		"type": "aws_security_group_rule",
    		"change": {"actions": ["create"], "after": {
    			"type": "ingress",
    			"from_port": 22,
    			"to_port": 22,
    			"cidr_blocks": ["0.0.0.0/0"],
    		}},
    	}])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 允许 SSH 来自私网地址段
    test_allows_ssh_from_a_private_range if {
    	fixture := plan([security_group({
    		"from_port": 22,
    		"to_port": 22,
    		"protocol": "tcp",
    		"cidr_blocks": ["10.0.0.0/8"],
    	})])
    
    	count(network.deny) == 0 with input as fixture
    }
    
    # 允许 HTTPS 对全网开放
    test_allows_https_from_the_world if {
    	fixture := plan([security_group({
    		"from_port": 443,
    		"to_port": 443,
    		"protocol": "tcp",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 0 with input as fixture
    }
    
    # “所有协议”规则会开放所有端口,无论其端口字段如何书写。
    # 该测试的初版断言方向相反,从而掩盖了 bug。
    test_denies_all_protocols_rule_open_to_the_world if {
    	fixture := plan([security_group({
    		"from_port": 0,
    		"to_port": 0,
    		"protocol": "-1",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 4 with input as fixture
    }
    
    # 策略无法读取的端口会被上报,而不是直接放行。
    test_reports_a_world_open_rule_with_unreadable_ports if {
    	fixture := plan([security_group({
    		"from_port": 22,
    		"to_port": null,
    		"protocol": "tcp",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 字符串类型的端口会被上报
    test_reports_string_ports if {
    	fixture := plan([security_group({
    		"from_port": "22",
    		"to_port": "22",
    		"protocol": "tcp",
    		"cidr_blocks": ["0.0.0.0/0"],
    	})])
    
    	count(network.deny) == 1 with input as fixture
    }
    
    # 非全网开放的规则上,无法读取的端口则静默处理。
    test_ignores_unreadable_ports_on_a_private_range if {
    	fixture := plan([security_group({
    		"from_port": 22,
    		"to_port": null,
    		"protocol": "tcp",
    		"cidr_blocks": ["10.0.0.0/8"],
    	})])
    
    	count(network.deny) == 0 with input as fixture
    }
    
    # `considered` 决定了门禁的通过与否或空判,因此
    # 尽管它本身不参与判定逻辑,也需要独立的测试。
    test_considers_every_ingress_bearing_type if {
    	fixture := plan([
    		security_group({}),
    		{"address": "aws_vpc_security_group_ingress_rule.a", "type": "aws_vpc_security_group_ingress_rule", "change": {"actions": ["create"], "after": {}}},
    		{"address": "aws_security_group_rule.b", "type": "aws_security_group_rule", "change": {"actions": ["create"], "after": {}}},
    	])
    
    	count(network.considered) == 3 with input as fixture
    }
    
    # 不考虑无关类型
    test_does_not_consider_unrelated_types if {
    	fixture := plan([{
    		"address": "aws_vpc.main",
    		"type": "aws_vpc",
    		"change": {"actions": ["create"], "after": {}},
    	}])
    
    	count(network.considered) == 0 with input as fixture
    }
    

    这个文件编码了以下内容:

    1. with input as fixture 会在某个表达式里替换为假的 plan,这就是不依赖云账号测试策略的方法。

    2. test_denies_wide_open_port_range 期望有 四个违规项,每个管理端口一个,因为 ingress 规则描述的是一个范围。一条 0-65535 的规则等于把 SSH 完全敞开,和显式的 22 端口规则一样宽,还能绕过等值检查。

    3. 另外三个测试覆盖了 Terraform 表达同一概念的另外三种形式:IPv6 字段、现代的独立资源 aws_vpc_security_group_ingress_rule,以及已弃用的 aws_security_group_rule。这些并不是我一开始就写的,Step 9 会解释它们的由来。

    4. 两个 test_allows_ 用例和拒绝类测试同样重要。一个拒绝一切的策略能通过所有 deny 测试,但毫无价值。

    5. 最后三个测试是在策略已经“完成”之后才加入的,Step 9 会解释它们的来历。一条 all-protocols 规则无论端口字段写什么都会开放所有端口,而一条策略无法读取端口信息的规则则必须被上报。

    Step 3: 编写策略直到测试通过

    Terraform 用四种形式描述 ingress,所以策略先把这四种统一归一化成一组,再对这一组做一次判断。创建 policy/network.rego:

    # METADATA
    # title: 公网无法访问管理端口
    # description: |
    #   Terraform 中入站规则有四种不同的形式。每个形式首先都会标准化为统一的
    #   `exposures` 集合,因此下方的判断逻辑只需编写一次,新增一种形式时只需增加一条辅助规则。
    #
    #   这里的两处设计是刻意为之,而非偶然。全协议规则会覆盖所有端口,无论其端口字段值如何;
    #   对于策略无法读取端口的规则,则予以报告而非放行。
    package terraform.network
    
    admin_ports := {22, 3389, 3306, 5432}
    
    public_cidrs := {"0.0.0.0/0", "::/0"}
    
    # AWS 提供程序对全协议规则写入 from_port 0 和 to_port 0,这会开放所有端口,
    # 因此端口字段不能按字面解读。
    all_protocols := {"-1", "all"}
    
    # 形式 1 和 2:内联入站块,IPv4 和 IPv6。
    exposures contains exposure if {
    	some resource in input.resource_changes
    	resource.type == "aws_security_group"
    	some ingress in resource.change.after.ingress
    	some field in ["cidr_blocks", "ipv6_cidr_blocks"]
    	some cidr in object.get(ingress, field, [])
    	exposure := {
    		"address": resource.address,
    		"protocol": object.get(ingress, "protocol", ""),
    		"from_port": object.get(ingress, "from_port", null),
    		"to_port": object.get(ingress, "to_port", null),
    		"cidr": cidr,
    	}
    }
    
    # 形式 3:AWS 提供程序自 v5 起推荐的独立规则。
    exposures contains exposure if {
    	some resource in input.resource_changes
    	resource.type == "aws_vpc_security_group_ingress_rule"
    	some field in ["cidr_ipv4", "cidr_ipv6"]
    	cidr := object.get(resource.change.after, field, null)
    	is_string(cidr)
    	exposure := {
    		"address": resource.address,
    		"protocol": object.get(resource.change.after, "ip_protocol", ""),
    		"from_port": object.get(resource.change.after, "from_port", null),
    		"to_port": object.get(resource.change.after, "to_port", null),
    		"cidr": cidr,
    	}
    }
    
    # 形式 4:已弃用的独立规则,仍存在于大多数现有基础设施中。
    exposures contains exposure if {
    	some resource in input.resource_changes
    	resource.type == "aws_security_group_rule"
    	resource.change.after.type == "ingress"
    	some cidr in object.get(resource.change.after, "cidr_blocks", [])
    	exposure := {
    		"address": resource.address,
    		"protocol": object.get(resource.change.after, "protocol", ""),
    		"from_port": object.get(resource.change.after, "from_port", null),
    		"to_port": object.get(resource.change.after, "to_port", null),
    		"cidr": cidr,
    	}
    }
    
    # 规则实际覆盖的端口。当策略无法判断时,此值未定义。
    covered_ports(exposure) := [0, 65535] if {
    	exposure.protocol in all_protocols
    }
    
    covered_ports(exposure) := [exposure.from_port, exposure.to_port] if {
    	not exposure.protocol in all_protocols
    	is_number(exposure.from_port)
    	is_number(exposure.to_port)
    }
    
    deny contains msg if {
    	some exposure in exposures
    	exposure.cidr in public_cidrs
    
    	# 若端口落在 [from_port, to_port] 区间内,则认为该规则覆盖此端口。
    	range := covered_ports(exposure)
    	some port in admin_ports
    	port >= range[0]
    	port <= range[1]
    
    	msg := sprintf(
    		"%s: 入站规则将端口 %d 暴露给 %s",
    		[exposure.address, port, exposure.cidr],
    	)
    }
    
    # 对公网开放但策略无法读取端口的规则会被报告。
    # 放行此类规则相当于策略在宽松方向上进行猜测。
    deny contains msg if {
    	some exposure in exposures
    	exposure.cidr in public_cidrs
    	not covered_ports(exposure)
    
    	msg := sprintf(
    		"%s: 针对 %s 的入站规则端口无法被此策略评估(%v 至 %v)",
    		[exposure.address, exposure.cidr, exposure.from_port, exposure.to_port],
    	)
    }
    
    # 此策略已知如何检查的地址。门控利用此集合来区分
    # “无违规”与“未检查”。
    considered contains resource.address if {
    	some resource in input.resource_changes
    	resource.type in {
    		"aws_security_group",
    		"aws_vpc_security_group_ingress_rule",
    		"aws_security_group_rule",
    	}
    }
    

    从上往下读这段代码:

    1. Rego 中的规则体是合取逻辑。每一行都必须成立,而 some ... in 语句负责迭代,因此 OPA 会遍历资源、Ingress 规则、字段和端口的所有组合。

    2. 三条独立的 exposures 规则共同定义了同一个集合,这在 Rego 中称为增量定义。日后若要增加第五种形态,只需多写一个代码块,下方的逻辑完全无需改动。

    3. object.get(ingress, field, []) 在字段缺失时返回空列表,因此仅包含 IPv4 的规则在策略查找 ipv6_cidr_blocks 时不会报错。

    4. covered_ports 在此处充当安全阀门。因为“全协议”规则在计划中报告端口范围从 0 到 0,实际却打开了所有端口,所以不能字面解读端口字段;若某规则的端口为 null 或字符串,该函数将未定义,第二条 deny 规则会将其判定为违规。

    5. considered 不参与最终判定。它记录的是该策略实际涉及哪些资源,第 5 步会利用这一点,避免报告策略并未真正覆盖的通过结果。

    运行命令:

    opa test policy
    opa check --strict policy
    opa fmt --diff policy
    
    PASS: 20/20
    

    给 opa test 加上 -v 参数,即可显示每条测试的详情。

    opa check --strict 能捕获不安全变量和被遮蔽的导入,而 opa fmt --diff 在格式已符合规范时不输出任何内容。OPA 使用制表符格式化 Rego。这两者都应纳入 CI 流程,且优先级高于其他步骤。

    第二个策略关于所有权标记,因此创建 policy/tags.rego:

    # METADATA
    # title: Every managed resource carries ownership tags
    # description: |
    #   Terraform emits `tags: null` for a resource with no tags at all, so a
    #   policy that reaches into `after.tags` skips exactly the resources with
    #   the worst tagging. `tags_of` coerces that null to an empty object.
    package terraform.tags
    
    required_tags := {"owner", "cost-center", "data-classification"}
    
    # Resource types that genuinely cannot carry tags.
    untaggable := {"aws_iam_policy_attachment", "aws_route_table_association"}
    
    in_scope contains resource if {
    	some resource in input.resource_changes
    	some action in resource.change.actions
    	action in {"create", "update"}
    	not resource.type in untaggable
    }
    
    tags_of(resource) := tags if {
    	tags := resource.change.after.tags
    	is_object(tags)
    } else := {}
    
    deny contains msg if {
    	some resource in in_scope
    	some tag in required_tags
    	value := object.get(tags_of(resource), tag, "")
    	trim_space(value) == ""
    	msg := sprintf("%s: missing required tag %q", [resource.address, tag])
    }
    
    considered contains resource.address if {
    	some resource in in_scope
    }
    

    这两行代码为什么这样写,原因如下:

    1. 真正的检查逻辑是 trim_space(value) == ""。必须这样写,因为在 Rego 中只有 false 和未定义才是假值,空字符串是真值。如果只做存在性检查,owner = "" 也能通过——这正是标签策略本应杜绝的形式主义合规。

    2. tags_of 的 else := {} 分支,是为了修复我自己写出的一个 bug——直到第 9 步才被发现。

    第 4 步:对着真实计划运行

    用 TerraGoat 计划评估一个包:

    opa eval --data policy --input plan.json --format pretty 'data.terraform.network.deny'
    
    [
      "aws_security_group.web-node: ingress rule exposes port 22 to 0.0.0.0/0"
    ]
    

    策略只报告了一个违规项,对端口 80 保持沉默。同一个资源里端口 80 也是对外开放的,但公网 Web 服务器本来就该如此。如果检查把两者都标记出来,人们很快就会学会无视它。

    每个包跑一条命令没法扩展,好在 Rego 可以用一条查询聚合整个命名空间:

    opa eval --data policy --input plan.json --format pretty \
      'union({v | v := data.terraform[_].deny})'
    
    [
      "aws_security_group.web-node: 入站规则将 22 端口暴露给了 0.0.0.0/0",
      "aws_security_group.web-node: 缺少必需标签 \"cost-center\"",
      "aws_security_group.web-node: 缺少必需标签 \"data-classification\"",
      "aws_security_group.web-node: 缺少必需标签 \"owner\"",
      "aws_vpc.web_vpc: 缺少必需标签 \"cost-center\"",
      "aws_vpc.web_vpc: 缺少必需标签 \"data-classification\"",
      "aws_vpc.web_vpc: 缺少必需标签 \"owner\""
    ]
    

    {v | v := data.terraform[_].deny} 是一个推导式,用于收集 data.terraform 下所有包中的 deny 集合,union 会将它们扁平化。只需在该命名空间下放入新的策略文件即可自动生效,命令无需更改。

    步骤 5:将判定结果转化为退出码

    CI 门禁通过退出状态码进行沟通。如果将两种不同性质的失败混用同一个状态码,会导致损坏的流水线在一个月内被误判为通过。该工具使用以下规则:

    状态码 判定结果 含义
    0 pass(通过) 策略已执行,已检查资源,未发现违规
    1 fail(失败) 策略已执行,发现违规
    2 vacuous(空转) 策略已执行,但未检查到任何相关资源,结果无效
    2 broken(损坏) 工具或其输入不可用

    第四行是大多数门禁最常出错的环节。如果计划文件中没有策略关注的任何资源,报告“通过”便是在提供此次运行无法保证的安全性。它拥有独立的判定名称,且不返回退出码 0。

    创建 policy_gate.py:

    """使用 Rego 策略目录评估 Terraform 部署计划。"""
    
    import argparse
    import json
    import pathlib
    import shutil
    import subprocess
    import sys
    
    PASS, FAIL, BROKEN = 0, 1, 2
    
    DENY_QUERY = "union({v | v := data.terraform[_].deny})"
    CONSIDERED_QUERY = "union({v | v := data.terraform[_].considered})"
    
    def die(message: str) -&gt; None:
        print(f"policy-gate: {message}", file=sys.stderr)
        sys.exit(BROKEN)
    
    def load_plan(path: pathlib.Path) -&gt; dict:
        try:
            text = path.read_text()
        except OSError as exc:
            die(f"无法读取 {path}: {exc.strerror}")
        try:
            return json.loads(text)
        except json.JSONDecodeError as exc:
            die(f"{path}:{exc.lineno}:{exc.colno}: 无效的 JSON: {exc.msg}")
    
    def query(opa: str, policy_dirs: list[pathlib.Path], plan: pathlib.Path, expr: str) -&gt; list:
        command = [opa, "eval", "--input", str(plan), "--format", "raw"]
        for directory in policy_dirs:
            command += ["--data", str(directory)]
        command.append(expr)
    
        result = subprocess.run(command, capture_output=True, text=True)
        if result.returncode != 0:
            die(f"OPA 执行失败: {result.stderr.strip() or result.stdout.strip()}")
        values = json.loads(result.stdout)
        # 规则若产生非字符串值属于策略错误;对混合类型列表排序会引发无关联的 TypeError
        for value in values:
            if not isinstance(value, str):
                die(f"{expr} 产生了 {type(value).__name__} 类型值;"
                    "deny 和 considered 规则必须返回字符串")
        return values
    
    def main() -&gt; int:
        parser = argparse.ArgumentParser(description=__doc__)
        parser.add_argument("--plan", required=True, type=pathlib.Path,
                            help="来自 `terraform show -json` 的 JSON 文件")
        parser.add_argument("--policy", required=True, nargs="+", type=pathlib.Path,
                            help="一个或多个包含 .rego 文件的目录")
        parser.add_argument("--opa", default="opa", help="opa 可执行文件路径")
        args = parser.parse_args()
    
        if shutil.which(args.opa) is None:
            die(f"{args.opa} 不在 PATH 中")
        for directory in args.policy:
            if not directory.is_dir():
                die(f"{directory} 不是目录")
    
        plan = load_plan(args.plan)
        if "resource_changes" not in plan:
            die(f"{args.plan} 缺少 resource_changes 键;是否为 Terraform 计划文件?")
    
        considered = query(args.opa, args.policy, args.plan, CONSIDERED_QUERY)
        if not considered:
            # 此处报告通过会声称一种本次运行无法保证的安全性
            print(f"无效通过: 未有任何策略检查这 {len(plan['resource_changes'])} 个计划资源",
                  file=sys.stderr)
            return BROKEN
    
        violations = sorted(query(args.opa, args.policy, args.plan, DENY_QUERY))
        if violations:
            print(f"未通过: 在 {len(considered)} 个检查资源中发现 {len(violations)} 个违规项",
                  file=sys.stderr)
            for violation in violations:
                print(f"  - {violation}", file=sys.stderr)
            return FAIL
    
        print(f"通过: 检查了 {len(considered)} 个资源,无违规")
        return PASS
    
    if __name__ == "__main__":
        sys.exit(main())
    

    这个脚本做了以下几件事:

    1. argparse 把 --plan 和 --policy 标记为 required=True,且 --policy 使用 nargs="+",因此策略列表为空时会在解析阶段直接报错。

    2. load_plan 会报告 JSON 语法错误的行号和列号,因为 json.JSONDecodeError 自带 lineno 和 colno——一个只说"JSON 无效"的关卡会让别人排查一下午。

    3. 所有失败路径都通过 die 处理,统一以退出码 2 结束:缺少 opa、文件不可读、计划中没有 resource_changes 键,这些都属于工具本身的故障。

    4. considered 查询在 deny 查询之前运行。如果什么都没检查,运行结果就是 VACUOUS,根本没有机会输出通过。

    5. 违规项会排序,因此同一份计划每次运行产生的输出完全一致(逐字节相同),CI 日志之间的 diff 才有意义。

    标题为 policy-lab 的终端窗口。运行 opa test policy 显示 PASS: 20/20。用 TerraGoat 计划运行 python3 policy_gate.py 显示 FAIL: 在 2 个受检资源中共 7 处违规,列出一条将 aws_security_group.web-node 的 22 端口暴露给 0.0.0.0/0 的入站规则,以及 aws_security_group.web-node 和 aws_vpc.web_vpc 上缺失的 6 个必需标签。echo $? 返回 1。

    这份计划中的两个资源共有 7 处违规,且退出码可供 CI 直接处理。

    标题为 policy-lab 的终端窗口。对空计划运行 python3 policy_gate.py 显示 VACUOUS: 没有策略检查任何资源(计划资源数为 0),echo $? 返回 2 而不是 0。

    同一个工具面对空计划的结果——这种情况下关卡如果回答"通过"就是在撒谎。

    GitHub Actions 工作流会先测试这些策略,再用它们去做判定:

    name: policy
    
    on: [pull_request]
    
    jobs:
      policy:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v5
    
          - name: Install OPA
            run: |
              curl -L -o /usr/local/bin/opa \
                https://openpolicyagent.org/downloads/v1.20.2/opa_linux_amd64_static
              chmod +x /usr/local/bin/opa
    
          # The policies are code. Lint and test them before trusting them.
          - name: Check policy syntax
            run: opa check --strict policy
    
          - name: Verify formatting
            run: opa fmt --fail --diff policy
    
          - name: Test policies
            run: opa test policy --verbose --coverage --format json &gt; coverage.json
    
          # Only now does anything get judged.
          - name: Evaluate Terraform plan
            run: python3 policy_gate.py --plan plan.json --policy policy
    

    先让这道门禁只上报结果,持续运行两周,再正式启用拦截功能。那些看起来毫无问题的策略,往往会因你现有的基础设施结构而失效。你希望在日志里发现问题,而不是在发布被阻断后才察觉。

    Step 6: 在准入阶段强制执行

    步骤 5 中的门禁检查的是你打算部署的内容。它无法看到从某台笔记本电脑发起的 kubectl apply、供应商的 Helm 图表,或 Operator 自行创建的 Pod。针对这些场景,你需要准入控制,而 Kubernetes 现已内置该功能。

    ValidatingAdmissionPolicy 自 v1.30 起已正式发布,它在 API server 内直接评估 CEL 表达式,无需部署或维护 Webhook:

    ```html
    apiVersion: admissionregistration.k8s.io/v1
    kind: ValidatingAdmissionPolicy
    metadata:
      name: require-trusted-registry
    spec:
      failurePolicy: Fail
      matchConstraints:
        resourceRules:
          - apiGroups: [""]
            apiVersions: ["v1"]
            operations: ["CREATE", "UPDATE"]
            resources: ["pods"]
      variables:
        # A Pod has three container lists. A policy that reads only
        # spec.containers is bypassed by moving the image to an initContainer.
        - name: allImages
          expression: >-
            object.spec.containers.map(c, c.image) +
            object.spec.?initContainers.orValue([]).map(c, c.image) +
            object.spec.?ephemeralContainers.orValue([]).map(c, c.image)
      validations:
        - expression: >-
            variables.allImages.all(i, i.startsWith('registry.internal.example.com/'))
          messageExpression: >-
            'images must come from registry.internal.example.com: ' +
            variables.allImages.filter(i,
              !i.startsWith('registry.internal.example.com/')).join(', ')
          reason: Forbidden
    

    策略本身不会生效,直到有绑定将其激活。这种机制允许你先在单个命名空间进行试点:

    apiVersion: admissionregistration.k8s.io/v1
    kind: ValidatingAdmissionPolicyBinding
    metadata:
      name: require-trusted-registry-binding
    spec:
      policyName: require-trusted-registry
      validationActions: ["Deny"]
      matchResources:
        namespaceSelector:
          matchLabels:
            policy.example.com/enforce: "true"
    

    将 validationActions 设为 ["Warn", "Audit"],给一个命名空间打上标签,观察一周,然后切换为 ["Deny"] 并扩大选择器范围。

    对 initContainers 的处理很关键,因为绕过镜像来源策略最常见的方式就是把镜像移到 initContainer 中。使用 ? 可选字段语法配合 .orValue([]) 可以安全地读取可能不存在的列表,避免整个表达式报错。

    由于手头没有集群,我无法应用这两个清单文件。这些清单已对照 v1 参考架构进行检查,但从未被真实的 API 服务器接受,因此请将其视为起点,并在 Warn 模式下逐步推进(本就应该这样操作)。

    ```

    另外两个引擎在生产环境中也被广泛使用,第一个是 Kyverno。它已于 2026 年 3 月从 CNCF 毕业,被 Bloomberg、Coinbase、Deutsche Telekom、LinkedIn 和 Spotify 用于生产环境。它的策略用 YAML 编写,平台团队无需学习新语言,而且它还能处理内置策略不管的生成、镜像签名验证和清理工作。

    如果你希望用一个 Rego 代码库同时覆盖 Kubernetes、Terraform 和 CI,那么 OPA Gatekeeper 是正确的选择——本教程正是朝着这个方向构建的。Mutation 功能现在也已内置:MutatingAdmissionPolicy 在 v1.36 中转为稳定版。

    第 7 步:让模型来写策略

    写策略很繁琐,而模型最擅长处理繁琐。所以顺理成章的做法就是让模型来写。

    但有一个坑,你自己花一分钟就能验证出来——我在本步骤末尾就是这么做的:公开训练数据中的大量 Rego 代码是 Rego v0 方言,随着 2025 年 1 月 OPA 1.0 发布,这种方言已无法被解析。模型倾向于使用它见过的最常见写法,结果就是写出当前解析器拒绝的方言。

    卡拉布里亚大学团队 2025 年的一篇预印本论文 ARPaCCino 在 Terraform 案例中报告了同样的现象:在不使用工具的情况下直接要求生成 Rego,Qwen3-30B 和 GPT-4o 各自生成的 5 条策略中语法正确的都是 0 条。在 OPA 文档上增加检索也没有任何改善。但给模型一个可以运行 opa check 并读取报错的循环后,正确率分别提升到 5 条中的 4 条和 5 条。

    这些数字来自一个小型案例研究,所以重点是把握其结论方向,而非具体数字。

    解决问题的正是一个反馈循环,而且这个循环不需要任何额外成本——因为你已经在用 opa check --strict、opa fmt 和 opa test 把它搭建好了。

    搭建这条闭环的关键,在于一处反转来确保可信度:人来写测试,模型来写策略。测试用例具体且审阅成本低——你扫一眼 JSON 数据就能在三秒内确认“该拒绝”。而包含嵌套推导式的 Rego 规则阅读门槛较高,且极易误读。让人类承担审阅成本低的工作,让机器在输出可被机械化校验的地方发挥作用。

    新建 policy_forge.py 文件:

    """从一条英文规则生成 Rego 策略,只有通过工具链验证才保留。"""
    
    import argparse
    import pathlib
    import re
    import subprocess
    import sys
    import tempfile
    
    WRITTEN, REJECTED, BROKEN = 0, 1, 2
    
    SYSTEM = """你编写 Open Policy Agent 的 Rego v1(OPA 1.0+)策略。
    
    规则:
    - 每个规则体都要用 `if`,多值规则用 `contains`。
    - 不要输出 `import rego.v1`;在 OPA 1.0+ 上它是多余的。
    - 输入是 `terraform show -json` 生成的 JSON。
    - 只返回一个 ```rego 代码块,不要输出其他内容。"""
    
    def extract_rego(reply: str) -&gt; str:
        blocks = re.findall(r"```rego\n(.*?)```", reply, re.DOTALL)
        if not blocks:
            raise ValueError("model returned no rego block")
        if len(blocks) &gt; 1:
            raise ValueError(f"model returned {len(blocks)} rego blocks; expected one")
        return blocks[0]
    
    def verify(opa: str, policy: str, tests: pathlib.Path) -&gt; tuple[bool, str]:
        with tempfile.TemporaryDirectory() as tmp:
            bundle = pathlib.Path(tmp)
            # 测试文件保留原文件名;策略文件用一个不会
            # 与之冲突的名字,无论调用方把测试文件命名成什么。
            (bundle / "candidate_policy.rego").write_text(policy)
            (bundle / tests.name).write_text(tests.read_text())
    
            for command in ([opa, "check", "--strict"], [opa, "test"]):
                result = subprocess.run(command + [str(bundle)], capture_output=True, text=True)
                if result.returncode != 0:
                    return False, (result.stdout + result.stderr).strip()
        return True, "opa check and opa test both passed"
    
    def forge(rule: str, tests: pathlib.Path, ask, opa: str, attempts: int) -&gt; str:
        transcript = [{
            "role": "user",
            "content": (
                f"为这条规则编写 Rego 策略:\n\n{rule}\n\n"
                f"它必须满足以下测试:\n\n```rego\n{tests.read_text()}```"
            ),
        }]
    
        for attempt in range(1, attempts + 1):
            reply = ask(transcript)
            policy = extract_rego(reply)
            ok, output = verify(opa, policy, tests)
            headline = next(iter(output.splitlines()), "no output from the toolchain")
            print(f"attempt {attempt}: {'PASS' if ok else 'FAIL'} - {headline}", file=sys.stderr)
            if ok:
                return policy
            transcript += [
                {"role": "assistant", "content": reply},
                {"role": "user", "content": f"工具链拒绝了上面的策略:\n\n{output}\n\n请修复。"},
            ]
    
        raise RuntimeError(f"no policy survived {attempts} attempts")
    
    def claude(model: str):
        import anthropic
    
        client = anthropic.Anthropic()
    
        def ask(transcript: list[dict]) -&gt; str:
            response = client.messages.create(
                model=model,
                max_tokens=16000,
                system=[{
                    "type": "text",
                    "text": SYSTEM,
                    "cache_control": {"type": "ephemeral"},
                }],
                thinking={"type": "adaptive"},
                messages=transcript,
            )
            return "".join(b.text for b in response.content if b.type == "text")
    
        return ask
    
    def main() -&gt; int:
        parser = argparse.ArgumentParser(description=__doc__)
        parser.add_argument("--rule", required=True, help="the policy, in one English sentence")
        parser.add_argument("--tests", required=True, type=pathlib.Path,
                            help="a _test.rego file you wrote by hand")
        parser.add_argument("--out", required=True, type=pathlib.Path,
                            help="where to write the policy, only if it passes")
        parser.add_argument("--model", default="claude-opus-5")
        parser.add_argument("--attempts", type=int, default=4)
        parser.add_argument("--opa", default="opa")
        args = parser.parse_args()
    
        if args.attempts &lt; 1:
            print("policy-forge: --attempts must be at least 1", file=sys.stderr)
            return BROKEN
        if not args.tests.is_file():
            print(f"policy-forge: {args.tests} does not exist", file=sys.stderr)
            return BROKEN
    
        try:
            policy = forge(args.rule, args.tests, claude(args.model), args.opa, args.attempts)
        except RuntimeError as exc:
            print(f"policy-forge: {exc}; nothing written", file=sys.stderr)
            return REJECTED
        except ValueError as exc:
            print(f"policy-forge: {exc}", file=sys.stderr)
            return BROKEN
    
        args.out.write_text(policy)
        print(f"policy-forge: verified policy written to {args.out}", file=sys.stderr)
        return WRITTEN
    
    if __name__ == "__main__":
        sys.exit(main())
    
    
    
    export ANTHROPIC_API_KEY=...
    python3 policy_forge.py \
      --rule "No security group may expose an administrative port to the public internet." \
      --tests policy/network_test.rego \
      --out policy/generated.rego
    

    这个循环提供了以下保证:

    1. 验证以子进程方式运行:从不问模型"你的策略是否正确"。由 opa check 和 opa test 来裁决,它们的退出码是循环唯一采信的证据。

    2. 失败信息以原始工具输出返回,不做摘要:编译错误和测试失败是模型能收到的信息量最大的反馈,改写反而会丢掉最有用的部分。

    3. 不通过就不落盘:forge 要么返回经过验证的策略,要么直接抛出异常,因此不会出现"重试次数用完了,未验证的策略就进了仓库"的情况。

    我用脚本模拟的模型驱动这个循环,这样即使没有 API key 也能复现结果。三次回复分别是:一个符合 v0 语法的策略、一个看似合理但有错的 v1 策略,以及第 3 步的那个策略:

    attempt 1: FAIL - 2 errors occurred during loading:
    attempt 2: FAIL - policy/network_test.rego:63:
    attempt 3: PASS - opa check and opa test both passed
    

    第一次是 Rego v0,也就是 deny[msg] { ... } 的写法——OPA 1.0 在 2025 年 1 月发布后就不再解析这种语法了。而公开训练数据里绝大多数都是这种写法。opa check --strict 在测试之前就把它拒之门外。

    第二次是合法的 Rego v1。大多数工程师审查时都可能放行,但它在测试套件上只拿到了 8 项中的 4 项。原因是它直接把 ingress.from_port 与 admin_ports 比较而忽略了范围,并且只读取了 cidr_blocks。opa check 没有任何报错,因为代码在语法上完全没问题。

    一个语法完全正确的策略,仍然可能是错误的策略——而整个循环里唯一知道你真实意图的,就是你亲手写下的测试套件。

    修复循环不是单调递进的:每次尝试都是基于错误信息的一次全新生成,之前已经奏效的东西不会延续下来,所以第四次尝试可能丢失第三次已经具备的属性。这套设计里没有任何机制能检测到这一点,因为被检查的只有你写的那个测试套件。 给重试次数设上限,让测试套件持续扩充,并把每一条生成的策略都当作一个 pull request,必须有人批准后才能合并。

    第 8 步:治理 Agent 本身

    AI agent 也是一个行为主体,它会调用工具,因此每次工具调用都是一个必须有人做出的授权决策。 业界很快就对此达成共识:Amazon Bedrock AgentCore Policy 于 2026 年 3 月正式发布,在网关层用 Cedar 评估 agent 的工具调用。常见的开源做法是在 MCP 工具网关前面放一个 OPA sidecar。 研究也在把这条边界往前推:来自华盛顿大学(Defects4J 背后的团队)的一篇 2026 年预印本论文,Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents(Winston、Winston 和 Just),把自然语言策略编译成 SMT 约束,并用 Z3 求解器拦截不合规的调用。 这些工作指向同一个结论:写在 system prompt 里的策略不等于执行。真正的执行,是坐在调用链路上的拦截器,它能够返回"拒绝",让这次调用根本不发生。 创建 agent/authz.rego:
    # METADATA
    # title: Agent 工具调用授权
    # description: |
    #   在每次工具调用执行前评估一次。决策结果有三个值而不是两个,
    #   因为一个值得部署的 agent 有时需要做一些事情,应该由人来批准,
    #   而不是由策略直接放行。
    package agent.authz
    
    tool_grants := {
    	"support": {"search_orders", "read_customer", "issue_refund"},
    	"analytics": {"search_orders", "run_query"},
    }
    
    write_tools := {"issue_refund", "run_query"}
    
    refund_ceiling_cents := 10000
    
    # 未映射的角色、未知工具或格式错误的输入都会落到这里。
    default decision := {"effect": "deny", "reasons": ["no matching grant"]}
    
    decision := {"effect": effect_for(reasons), "reasons": reasons} if {
    	count(granted) &gt; 0
    	reasons := escalations
    }
    
    granted contains role if {
    	some role in input.agent.roles
    	input.tool in object.get(tool_grants, role, set())
    }
    
    effect_for(reasons) := "allow" if count(reasons) == 0
    
    effect_for(reasons) := "require_approval" if count(reasons) &gt; 0
    
    escalations contains reason if {
    	input.tool in write_tools
    	not input.session.human_in_loop
    	reason := sprintf("%q writes state and the session is unattended", [input.tool])
    }
    
    # 如果退款请求中读不到金额,就无法对照上限做检查,因此需要升级处理。
    # 如果这里静默放行,恰好会放过攻击者精心构造的那类调用。
    escalations contains reason if {
    	input.tool == "issue_refund"
    	not positive_amount
    	reason := "refund amount is missing, unreadable, or not positive"
    }
    
    # 负数金额就是披着退款外衣的扣款。
    positive_amount if {
    	amount := object.get(input, ["arguments", "amount_cents"], null)
    	is_number(amount)
    	amount &gt; 0
    }
    
    escalations contains reason if {
    	input.tool == "issue_refund"
    	amount := object.get(input, ["arguments", "amount_cents"], null)
    	is_number(amount)
    	amount &gt; refund_ceiling_cents
    	reason := sprintf(
    		"refund of %d cents exceeds the %d cent ceiling",
    		[amount, refund_ceiling_cents],
    	)
    }
    
    # 关键词匹配只是粗略的防护,这里主要是展示参数级规则的写法。
    # 真正涉及数据的场景需要 SQL 解析器:这种方式能拦住 `DROP TABLE`,
    # 但拦不住换一种写法的同类语句。
    destructive_sql := `(?i)\b(drop|truncate|delete|alter|grant|revoke)\b`
    
    escalations contains reason if {
    	input.tool == "run_query"
    	regex.match(destructive_sql, object.get(input, ["arguments", "statement"], ""))
    	reason := "statement contains a destructive SQL keyword"
    }
    
    
    

    决策词汇是封闭的,直接列在这里:

    Effect 调用方的行为
    allow 执行该工具
    require_approval 暂停,把原因展示给人工,批准后才执行
    deny 拒绝,且不提供人工审批通道

    策略为什么要这样设计:

    1. default decision 默认为 deny:无法识别的工具、忘记映射的角色、格式错误的输入,最终都会落到这里。默认放行的策略恰恰会在没人预料到的输入上失败敞开——而攻击者挑的正是这些输入。

    2. 三个值:非黑即白的授权只能在“拦住有用的工作”和“放行危险的工作”之间二选一,正是第三个值让高自主 agent 变得可以接受。

    3. 原因以集合返回:所有适用的原因都会被收集。凌晨三点收到审批提示时,“退款 250000 美分超过 10000 美分上限”能让人知道该怎么办;“违反策略”则不行。

    4. 参数也要检查:同样是 issue_refund,5 英镑是日常操作,2500 英镑就是大事。既然参数是 agent 自己选的,只看工具名的粒度就太粗了。

    十五个测试覆盖了整张决策表,包括角色列表为空的 agent、没有任何金额的退款,以及用换行符藏起 DROP 的查询:

    opa test agent -v
    
    PASS: 15/15
    

    启动服务,试一次调用:

    opa run --server --addr localhost:8181 agent/
    
    curl -s localhost:8181/v1/data/agent/authz/decision \
      -d '{"input":{"agent":{"roles":["support"]},"tool":"issue_refund",
           "arguments":{"amount_cents":250000},"session":{"human_in_loop":true}}}' | jq .result
    
    {
      "effect": "require_approval",
      "reasons": [
        "refund of 250000 cents exceeds the 10000 cent ceiling"
      ]
    }
    

    在有东西根据策略的答复拒绝继续执行之前,这份策略只是一纸空文。所以还需要一个客户端:

    import json
    import urllib.request
    
    OPA_URL = "http://localhost:8181/v1/data/agent/authz/decision"
    
    class PolicyDenied(Exception):
        pass
    
    class ApprovalRequired(Exception):
        pass
    
    def authorize(agent, tool, arguments, session):
        payload = json.dumps({"input": {
            "agent": agent, "tool": tool,
            "arguments": arguments, "session": session,
        }}).encode()
        req = urllib.request.Request(
            OPA_URL, data=payload, headers={"Content-Type": "application/json"}
        )
        with urllib.request.urlopen(req, timeout=2) as resp:
            body = json.load(resp)
    
        # OPA 在查询无匹配结果时会返回 200 和 {}。必须默认拒绝。
        decision = body.get("result", {"effect": "deny", "reasons": ["policy unavailable"]})
    
        if decision["effect"] == "deny":
            raise PolicyDenied("; ".join(decision["reasons"]))
        if decision["effect"] == "require_approval":
            raise ApprovalRequired("; ".join(decision["reasons"]))
        return decision
    
    search_orders    -&gt; 允许
    issue_refund     -&gt; 需要审批(退款金额 250000 美分,超过了 10000 美分的上限)
    delete_account   -&gt; 拒绝(没有匹配的授权)
    

    注意 body.get("result", ...) 这里:当查询没有匹配结果时,OPA 会返回状态码 200 和一个空的 {},所以直接写 body["result"] 会抛出 KeyError,而根据你的 agent 框架处理异常的方式,这可能导致系统默认放行(fail open)。每一层都要默认拒绝,包括解析环节。

    在你的框架的工具执行钩子中调用 authorize(),位置放在工具函数真正执行之前。总共也就十几行代码,但它能把系统提示词里那些"客气的建议"变成真正的边界。

    第 9 步:我踩过的坑

    第 2 步和第 3 步展示的是最终的策略版本。我最初尝试的是更简单的写法(也就是大多数教程止步于的版本),而那版和你在上文看到的版本之间的差距,正是这一部分最有价值的内容。

    打标签策略漏掉了最棘手的资源

    我最初天真的 in_scope 规则以 resource.change.after.tags 结尾,字面意思是"只保留带有标签的资源"。

    但它的实际行为比这更糟:Terraform 对完全没有标签的资源会输出 tags: null,而对未定义字段的查询会让规则体判定失败,于是这类资源就彻底脱离了管控范围。

    TerraGoat 计划里有两个资源:security group 带 5 个 git_* 标签、没有所有权标签,而 VPC 则一个标签都没有。

    jq -r '.resource_changes[] | "\(.address): tags=\(.change.after.tags | type)"' plan.json
    
    aws_security_group.web-node: tags=object
    aws_vpc.web_vpc: tags=null
    
    opa eval --data naive  --input plan.json --format pretty 'count(data.terraform.tags.deny)'
    opa eval --data policy --input plan.json --format pretty 'count(data.terraform.tags.deny)'
    
    3
    6
    

    naive 版漏掉的 3 条都来自那个完全没打标签的资源——策略只揪出了带部分标签的那个,却对没标签的那个视而不见。

    tags_of 加上 else := {} 分支就是解法,总共三行代码。

    网络策略只认四种写法中的一种

    我最初的朴素网络策略只读 aws_security_group 和 cidr_blocks,这也是所有教程里的写法。但 Terraform 表达同一条 ingress 规则其实有四种方式。

    ![图示标题为"Terraform 描述同一条 ingress 规则的四种方式"。图中显示四个方框。左上角是 aws_security_group.ingress[].cidr_blocks,填充淡蓝色并标注为已读取。另外三个用红色轮廓和红色斜线标注为未读取:aws_security_group.ingress[].ipv6_cidr_blocks、aws_vpc_security_group_ingress_rule.cidr_ipv4,以及已弃用的 aws_security_group_rule。图例说明:蓝色实心填充表示策略会检查此处,红色斜线表示不检查。](https://cdn.hashnode.com/uploads/covers/5f3a74bfc4d5973f55c91c8c/ab61d36d-78bf-45f7-a073-e7013216e3e1.png align="center")

    同一条规则,四种编码,而朴素策略只读了左上角那一种。

    我又规划了一个包含三个 security group 的真实配置,它们利用策略读不到的那几种写法把 SSH 对全世界开放。两个版本都在 code/ 目录里,可以复现:

    opa eval --data naive --input plan-evasion.json \
      --format pretty 'data.terraform.network.deny'
    
    []
    

    结果是:三个暴露在公网的 SSH 端口,零违规,门禁还会打印 PASS 并以退出码 0 放行。

    一个自己留着漏洞的测试

    端口范围还有第二个问题,而元凶正是我自己写的测试。我写过一个叫 test_tolerates_null_ports 的用例,断言 protocol: "-1" 的入站规则不会产生任何违规,理由是与空端口做比较时,策略不应该崩溃。

    但全协议规则会开放所有端口,AWS provider 会把它记为 from_port: 0 和 to_port: 0。范围检查把这两个值字面地理解成了"端口 0"这一个端口,结果 AWS 里最宽松的规则反而拿到了"零违规"的评价。

    对应的 plan 在 code/plan-all-protocols.json,只包含一个资源:

    jq -c '.resource_changes[] | select(.type=="aws_security_group")
           | .change.after.ingress[0] | {protocol,from_port,to_port,cidr_blocks}' \
       plan-all-protocols.json
    opa eval --data naive --input plan-all-protocols.json \
       --format pretty 'data.terraform.network.deny'
    
    {"protocol":"-1","from_port":0,"to_port":0,"cidr_blocks":["0.0.0.0/0"]}
    []
    

    所有协议、所有端口,对整个互联网敞开,却被我特意写下的一个测试认证为合规。解决办法是引入 covered_ports:把全协议规则映射到完整的端口范围,对无法解析的端口则返回 undefined,并新增一条 deny 规则来报告这种 undefined 情况。原来那个测试删掉了,取而代之的是四个新测试:

    opa eval --data policy --input plan-all-protocols.json --format pretty 'data.terraform.network.deny'
    
    [
      "aws_security_group.wide_open: ingress rule exposes port 22 to 0.0.0.0/0",
      "aws_security_group.wide_open: ingress rule exposes port 3306 to 0.0.0.0/0",
      "aws_security_group.wide_open: ingress rule exposes port 3389 to 0.0.0.0/0",
      "aws_security_group.wide_open: ingress rule exposes port 5432 to 0.0.0.0/0"
    ]
    
    原始来源: freeCodeCamp

    评论 (0)