云原生运维安全【免费下载链接】cloud-custodianRules engine for cloud security, cost optimization, and governance, DSL in yaml for policies to query, filter, and take actions on resources项目地址https://gitcode.com/gh_mirrors/cl/cloud-custodian点击查看免费下载本篇技术指南围绕 cloud-custodian GCP 提供者c7n_gcp中gcp.vertex-ai-endpoint资源的指标metrics过滤测试展开完整讲解如何借助 Terraform fixture 与辅助脚本为test_vertexai_endpoint_metrics录制真实的 Vertex AI 在线预测指标数据。读完本文你将掌握该 fixture 的完整结构与录制工作流部署 sklearn 模型、发送预测、等待prediction_count指标出现、录制后清理并理解GCPMetricsFilter底层如何按metric-key构造时间序列查询、测试断言为何使用ALIGN_SUM与REDUCE_NONE从而能够在本地为 cloud-custodian GCP 测试录制新的 flight data。一、背景为什么需要专门录制 Vertex AI Endpoint 指标cloud-custodian 的 GCP 支持中gcp.vertex-ai-endpoint资源对应 Vertex AI 在线预测端点。针对该资源c7n_gcp注册了类型为metrics的过滤器见 tools/c7n_gcp/c7n_gcp/resources/vertexai.py它继承自 GCP 通用的 GCPMetricsFilter。测试文件 tools/c7n_gcp/tests/test_vertexai.py 中的test_vertexai_endpoint_metrics通过回放replayflight data 来验证这一过滤器。问题的关键在于空的 Vertex AI Endpoint 不会产生任何在线预测指标。aiplatform.googleapis.com/prediction/online/prediction_count这类时间序列只有在端点真正部署了模型并收到预测请求后才会出现。因此为了让测试断言resources[0][c7n.metrics][metric_name][points]非空录制阶段必须先做一系列真实业务操作——这正是 vertexai_endpoint_metrics/readme.md 所描述的内容。该 fixture 的核心设计是先用 Terraform 创建空端点 GCS bucket只提供基础设施成本可控、状态可预期再用run_prediction.py完成模型上传、部署与预测等指标出现后由测试录制最后由cleanup_prediction.py撤销部署。下面逐一展开。二、Fixture 文件结构总览fixture 位于tools/c7n_gcp/tests/terraform/vertexai_endpoint_metrics/目录共 5 个文件文件作用main.tfTerraform 定义创建空 Vertex AI Endpoint 与 GCS buckettf_resources.jsonpytest-terraform记录的 apply 后资源状态供脚本读取端点与 bucket 信息run_prediction.py录制辅助脚本训练 sklearn 模型、上传 GCS、部署到端点、发送预测、等待指标cleanup_prediction.py清理脚本撤销模型部署并删除已上传模型readme.md录制工作流说明本文核心依据三、Terraform fixture 详解空端点与制品桶main.tf 内容如下provider google {} resource random_id suffix { byte_length 2 } locals { suffix ${terraform.workspace}-${random_id.suffix.hex} } resource google_vertex_ai_endpoint default { name c7n-endpoint-${local.suffix} display_name c7n-endpoint-${local.suffix} location us-central1 region us-central1 } resource google_storage_bucket artifacts { name c7n-vertex-metrics-${local.suffix} location US force_destroy true uniform_bucket_level_access true }要点解析随机后缀random_id生成 2 字节随机 hex与 Terraform workspace 名拼接成suffix保证多环境并发 apply 时资源名不冲突。默认 workspace 为default因此 tf_resources.json 中实际资源名为c7n-endpoint-default-7509、bucket 名为c7n-vertex-metrics-default-7509。端点刻意保持为空google_vertex_ai_endpoint.default不部署任何模型deployed_models为空列表见 tf_resources.json 中deployed_models: []。录制阶段的所有活性都由脚本注入terraform teardown 因此简单可靠。固定区域端点固定us-central1与测试中query: [{location: location}]的取值一致也与后续容器镜像us-docker.pkg.dev/vertex-ai/prediction/sklearn-cpu.1-3:latest的部署区域匹配。bucket 属性force_destroy true允许直接删除含对象的桶方便 teardownuniform_bucket_level_access true启用统一桶级访问控制避免 ACL 干扰。tf_resources.json由pytest-terraform生成其resources键按资源类型组织google_vertex_ai_endpoint.default、google_storage_bucket.artifacts、random_id.suffixrun_prediction.py正是从这个文件反解出端点project、name、location与 bucketname的。此外tests/conftest.py 中的pytest_terraform_modify_state钩子会对 tfstate 调用sanitize_project_name将功能测试账号信息脱敏后再写入该 JSON。四、录制前置条件按照 readme.md 的 Prerequisites 章节录制前需要满足目标 GCP 项目已配置 Application Default CredentialsADCrun_prediction.py依赖google.cloud.aiplatform、google.cloud.storage等 SDK 使用默认凭证访问项目资源。Terraform 已对该 fixture apply 完成确保tf_resources.json已生成或使用现有录制好的版本端点与 bucket 真实存在。安装 Python 依赖uv pip install google-cloud-aiplatform google-cloud-storage scikit-learn numpy2.0numpy2.0是硬性约束fixture 使用的 Vertex AI sklearn serving 容器为sklearn-cpu.1-3见run_prediction.py中serving_container_image_uri该容器内嵌的 numpy 主版本为 1.x若本地上传的model.pkl由 numpy 2.x 序列化模型加载时可能因二进制格式不兼容而失败。这与 readme 中numpy2.0matters because the Vertex AI sklearn serving container used here issklearn-cpu.1-3的说明一致。五、录制工作流断点驱动的三步曲readme 给出的录制流程针对测试 test_vertexai_endpoint_metrics核心思路是利用调试断点把真实业务动作插入到回放录制过程中间。完整步骤如下在test_vertexai_endpoint_metrics函数顶部、session_factory test.replay_flight_data(...)之前设置第一个断点在测试函数末尾设置第二个断点以便在 Terraform teardown 之前先执行清理。以record 模式启动测试cloud-custodian 测试框架通过test.record_flight_data录制 HTTP 交互C7N_FUNCTIONAL环境下LazyReplay.value为 False即走真实录制参见 tests/conftest.py。测试停在第一个断点时运行python tools/c7n_gcp/tests/terraform/vertexai_endpoint_metrics/run_prediction.py等待脚本成功结束——readme 明确提示这可能耗时2030 分钟原因见下文第六节的时间预算分析。回到测试并继续执行测试随即以真实环境为基础录制端点查询与指标查询的 HTTP 响应。测试停在第二个断点时运行python tools/c7n_gcp/tests/terraform/vertexai_endpoint_metrics/cleanup_prediction.py清理完成后再执行 Terraform teardown销毁端点和 bucket。这一断点插入模式保证了录制的 flight data 中既包含带deployedModels的端点列表响应也包含可返回时间序列点的projects.timeSeries.list响应使回放模式下测试断言len(resources) 1与points非空都能成立。六、run_prediction.py从零造出预测指标run_prediction.py 是整个录制的发动机其执行链可拆为 7 个阶段1. 复用 c7n_gcp 客户端基础设施脚本通过REPO_ROOT Path(__file__).resolve().parents[5]定位仓库根目录把tools/c7n_gcp插入sys.path直接导入from c7n_gcp.client import Session from c7n_gcp.resources.vertexai import VertexAIEndpoint这意味着它复用与 c7n 运行时完全相同的 GCP Session 与端点客户端保证脚本行为与策略执行一致例如用VertexAIEndpoint.get_location_client轮询端点部署状态。2. 读取 Terraform 状态并记录运行状态从tf_resources.json取端点与 bucket拼出标准端点名endpoint_name ( fprojects/{project_id}/locations/{location}/endpoints/{endpoint[name]} )同时维护一个run_prediction_state.jsonl状态文件append_state(**data)逐行追加 JSON记录bucket_name、endpoint_name、location、project_id后续每个阶段产生的artifact_prefix、artifact_uri、uploaded_model_name、deployed_model_id也追加进去——这份状态文件正是cleanup_prediction.py的输入。3. 训练并上传 sklearn 模型在临时目录训练一个最简单的LinearRegressionx_train [[1.0],[2.0],[3.0]]y_train [1.0,2.0,3.0]用pickleprotocol4落盘再上传到 GCS bucket 的vertexai-endpoint-metrics/{uuid}前缀下得到artifact_uri gs://bucket/prefix。4. 上传模型到 Vertex AI Model Registryuploaded_model aiplatform.Model.upload( display_nameuploaded_model_display_name, artifact_uriartifact_uri, serving_container_image_uri( us-docker.pkg.dev/vertex-ai/prediction/sklearn-cpu.1-3:latest ), syncTrue, )容器镜像选择sklearn-cpu.1-3即 readme 强调numpy2.0的原因所在。syncTrue保证阻塞至上传完成。5. 部署到端点deployed_model uploaded_model.deploy( endpointaiplatform_endpoint, deployed_model_display_namedeployed_model_display_name, machine_typen1-standard-2, min_replica_count1, max_replica_count1, syncTrue, )部署规格为单副本n1-standard-2这是整个流程中最耗时的一步创建 VM、拉起 serving 容器通常需要数分钟。随后脚本用VertexAIEndpoint.get_location_client(session, location, projects.locations.endpoints)轮询端点get接口最多 30 次、每次间隔 10 秒直到deployedModels非空并从返回中找到与deployed_model_display_name匹配的id记录为deployed_model_id。6. 发送在线预测请求for _ in range(3): prediction aiplatform_endpoint.predict(instances[[1.0]]) assert prediction.predictions time.sleep(2)连续发送 3 次预测每次间隔 2 秒。这 3 次真实调用是prediction_count时间序列数据点的直接来源。7. 轮询等待指标落盘脚本构造 Cloud Monitoring 查询验证指标aiplatform.googleapis.com/prediction/online/prediction_count是否已经可查metric_filter ( fmetric.type {metric_type} AND f( resource.labels.endpoint_id {endpoint_name.split(/)[-1]} ) )查询参数与 c7n 的GCPMetricsFilter保持一致aggregation_alignmentPeriod86400s、aggregation_perSeriesAlignerALIGN_SUM、aggregation_crossSeriesReducerREDUCE_NONE、viewFULL。脚本最多轮询 20 次、每次间隔 30 秒一旦metric_response.get(timeSeries)非空即成功退出否则断言失败。时间预算部署轮询最长 30×10s 指标轮询最长 20×30s 模型上传/部署本身的耗时叠加后即 readme 所说20-30 分钟的由来。这解释了为何必须在断点处阻塞等待而不能在测试函数内同步完成。七、cleanup_prediction.py录制后的资源回收cleanup_prediction.py 读取run_prediction_state.jsonl逐行合并为单个 dict依次执行撤销部署若存在endpoint_name与deployed_model_id调用endpoint.undeploy(deployed_model_id..., syncTrue)释放副本 VM删除模型若存在uploaded_model_name调用model.delete(syncTrue)清理 Model Registry 中的模型删除状态文件STATE_PATH.unlink()保证下次录制从干净状态开始。执行顺序有讲究必须先 undeploy 再 delete model模型正被部署时不可删除这也是它在 Terraform teardown 之前运行的原因——避免force_destroy清理时遗留孤儿资源或删除失败。八、测试侧策略、断言与 metric-key 约束录制完成后测试 test_vertexai_endpoint_metrics 在回放模式下验证以下策略policy test.load_policy( { name: vertexai-endpoint-metrics, resource: gcp.vertex-ai-endpoint, query: [{location: location}], filters: [ {type: value, key: displayName, value: endpoint_display_name}, { type: metrics, name: aiplatform.googleapis.com/prediction/online/prediction_count, aligner: ALIGN_SUM, days: 1, op: greater-than, value: 0, }, ], }, session_factorysession_factory, )关键点测试用terraform(vertexai_endpoint_metrics)装饰器注入 fixture从vertexai_endpoint_metrics.resources[google_vertex_ai_endpoint][default]取display_name与location保证策略与 fixture 一一对应指标名取常量metric_typedays: 1限定查询最近一天aligner: ALIGN_SUM将原始计数聚合成日总和op: greater-than, value: 0筛选出确有预测流量的端点断言部分验证了 c7n 指标附着的格式约定metric_name f{metric_type}.ALIGN_SUM.REDUCE_NONE assert metric_name in resources[0][c7n.metrics] assert resources[0][c7n.metrics][metric_name] is not None assert resources[0][c7n.metrics][metric_name][points]指标名.aligner.reducer正是 filters/metrics.py 中self.c7n_metric_key %s.%s.%s % (self.metric, self.aligner, self.reducer)的产物。此外test_vertexai_endpoint_metrics_invalid_metric_key 验证了VertexAIEndpointMetricsFilter.validate的约束Vertex AI Endpoint 指标只支持metric-key为resource.labels.endpoint_id传入metric.labels.deployed_model_id会抛出FilterValidationError匹配信息only supports metric-key resource.labels.endpoint_id。九、底层原理GCPMetricsFilter 如何按 endpoint_id 查指标1. metric-key 与资源名解析vertexai.py 为端点资源定义了metric_key resource.labels.endpoint_id classmethod def get_metric_resource_name(cls, resource, metric_keyNone): # Endpoint metrics are keyed by the terminal endpoint id. return resource[name].split(/)[-1]即端点完整名projects/p/locations/us-central1/endpoints/id取末段得到端点 ID再与 Monitoring 时间序列标签resource.labels.endpoint_id对齐。2. 查询构造与批处理GCPMetricsFilter.process 中days默认 14aligner默认ALIGN_NONEreducer默认REDUCE_NONEperiod-start默认auto相对当前时间滚动窗口可选start-of-day对齐 UTC 自然日通过session.client(monitoring, v3, projects.timeSeries)调用execute_query(list, ...)请求参数含aggregation_alignmentPeriod即整个查询窗口秒数、aggregation_perSeriesAligner、aggregation_crossSeriesReducer、aggregation_groupByFields、view: FULLbatch_resources第 174-197 行按metric_key ...与OR拼接资源过滤条件达到BATCH_SIZE即分片避免单个 filter 过长get_batched_query_filter第 199-213 行最终生成形如metric.type aiplatform.googleapis.com/prediction/online/prediction_count AND ( resource.labels.endpoint_id ... OR ... )的查询串。3. 匹配与附着split_by_resource用jmespath_search(metric_key, m)把返回的时间序列按 endpoint_id 建索引process_resource将完整时间序列对象写入资源字典的c7n.metrics键取出首个数据点的值执行op比较greater-than、less-than等默认less-than支持missing-value兜底无指标场景。这正是run_prediction.py第 7 阶段手工复现同一套查询参数的原因——两者共用同一语义。十、实践建议与注意事项务必先录制再回放不要直接对空端点运行测试否则timeSeries为空points断言必然失败这也正是 readme 把录制流程单独成文的动机。预留充足时间一次完整录制 2030 分钟建议在断点等待期间不要操作终端避免脚本中途失败后状态文件残留失败重跑前可手动删除run_prediction_state.jsonl。严格遵循依赖版本本地环境保持numpy2.0否则与sklearn-cpu.1-3容器不兼容。注意项目隔离录制会在目标 GCP 项目真实产生计费资源n1-standard-2副本、GCS 桶建议使用独立的测试项目并确认 ADC 账号具备aiplatform.endpoints.list、模型部署与 Monitoringprojects.timeSeries.list权限后者即GCPMetricsFilter.permissions见 filters/metrics.py。cleanup 与 teardown 顺序不可颠倒先跑cleanup_prediction.py再terraform destroy避免 bucket/端点销毁时残留已部署模型导致删除失败。通过上述 fixture 与脚本组合cloud-custodian 得以用完全真实、可复现的 Vertex AI 在线预测数据驱动gcp.vertex-ai-endpoint的指标过滤测试为后续任何依赖真实流量才能产生指标的 GCP 资源测试提供了可参考的录制范式。赞分享云原生运维安全【免费下载链接】cloud-custodianRules engine for cloud security, cost optimization, and governance, DSL in yaml for policies to query, filter, and take actions on resources项目地址https://gitcode.com/gh_mirrors/cl/cloud-custodian点击查看免费下载相关推荐Cloud Custodian GCP 指南用 YAML 策略清点、过滤与治理 Vertex AI EndpointCloud Custodian GCP 指南用 YAML 策略清点、过滤与治理 Vertex AI Endpoint 本文以 Cloud Custodian云原生运维安全cloud-custodian GCP 实战用策略盘点与治理 Vertex AI Model Garden 多发行商模型gcp.vertex-ai-publisher-modelcloud custodian GCP 实战用策略盘点与治理 Vertex AI Model Garden 多发行商模型gcp.vertex ai publ云原生运维安全Cloud Custodian 实战用实时策略强制 DMS Endpoint 启用 SSL 加密Cloud Custodian 实战用实时策略强制 DMS Endpoint 启用 SSL 加密 导读 本文基于 Cloud Custodian 官方示例文档云原生运维安全创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考