Siam Thanat Hack Co., Ltd.

What Google Secure AI Framework 2.0 is and how it secures AI agents

SAIF maps AI risks, components, and controls from data through agents. This page explains the core elements, Risk Map, SAIF 2.0 agent guidance, and what red teaming should prove before impact becomes real.

Executive Summary and Mandate Overview
Target Audience Security engineering teams, AI application developers, and cloud security architects
Mandatory Frequency Before go-live and after changes to models, data, prompts, tools, permissions, or deployment architecture
Required Scope Data, Infrastructure, Model, Application, the SAIF Agent Risk Map, Assurance, and Governance controls
Non-Compliance Risk Data poisoning, model tampering, prompt injection, sensitive-data disclosure, and agent rogue actions

Six core elements keep AI security connected to cybersecurity

The original six SAIF core elements remain foundational. The central idea is to extend proven controls into AI and add controls for risks introduced by data, models, and agents.

Infographic explaining: Six core elements keep AI security connected to cybersecurity
Visual guide: Six core elements keep AI security connected to cybersecurity
  • Expand strong security foundations by applying identity, secure-by-default infrastructure, supply-chain security, and security expertise to AI.
  • Extend detection and response by bringing model inputs, outputs, tool calls, and AI assets into threat intelligence, SOC, and incident-response workflows.
  • Automate defenses to reduce the time required to detect and respond to threats that adversaries can scale with AI.
  • Harmonize platform-level controls so policies remain consistent across model providers, platforms, applications, and internal teams.
  • Adapt controls through faster feedback loops using incidents, user feedback, red-team findings, and model changes.
  • Contextualize AI system risks across the end-to-end business process, data lineage, validation, and operational behavior.

The SAIF Risk Map uses four component areas to place controls correctly

The SAIF Map divides AI development into Data, Infrastructure, Model, and Application areas, then shows where risks are introduced, exposed, and mitigated.

  • The Data area covers sources, ingestion, processing, training data, and RAG data, with emphasis on provenance, privacy, quality, access, and integrity.
  • The Infrastructure area covers code, frameworks, compute, pipelines, registries, serving, and storage exposed to supply-chain attacks and misconfiguration.
  • The Model area covers training, tuning, evaluation, weights, input handling, and output handling that require robustness and disclosure testing.
  • The Application area covers interfaces, APIs, plugins, agents, and tools where prompt injection, excessive access, and rogue actions create real-world impact.
  • SAIF identifies 15 risk categories and maps each risk to controls owned by model creators, model consumers, or both.

SAIF 2.0 adds an Agent Risk Map and secure-by-design principles

SAIF 2.0 extends the framework to agents that plan, retain memory, invoke tools, and act for users. The risk is no longer limited to generated text. It includes authority to take action.

SAIF 2.0 diagram showing human controllers, limited power, and observable actions across agent input, reasoning, memory, tools, and action
An email-assistant example with trust labels, policy checks, protected context, least privilege, user approval, and audit logging
  • Agents need clearly defined human controllers who can approve changes, set policy, stop operation, and remain accountable for outcomes.
  • Agent powers need least privilege, contextual permissions, tool scope, credential isolation, and user confirmation for high-impact actions.
  • Agent actions and planning need observability through tool-call logs, reason codes, policy decisions, and tamper-resistant audit trails.
  • The Agent Risk Map highlights Rogue Actions, Sensitive Data Disclosure, Insecure Integrated Components, and Prompt Injection across perception, reasoning, memory, and tools.

Google AI Red Team tests systems, not only models

Google's AI Red Team combines threat intelligence, security expertise, and AI research to emulate adversaries against real products and features, then returns findings to defenders.

  • Start with the product threat model, then test models, data, applications, infrastructure, identity, supply chains, and human workflows.
  • Build attack paths such as indirect prompt injection entering RAG, steering an agent to call a tool, and sending a canary secret to a controlled endpoint.
  • Test layered controls including input validation, output sanitization, agent permissions, user approval, rate limits, egress restrictions, and detection.
  • Measure prevention, detection, time to triage, containment, and recovery so findings improve both the product and incident response.

Controls to prove before releasing AI and agents

SAIF groups controls under Data, Infrastructure, Model, Application, Assurance, and Governance. Red Teaming, Vulnerability Management, Threat Detection, and Incident Response apply across all risks.

  • Data controls should prove inventory, access, provenance, integrity, privacy, and resistance to data poisoning from source through production.
  • Model controls should prove input validation, output sanitization, adversarial testing, model access restrictions, and prompt-injection resilience.
  • Application controls should prove authentication, rate limits, agent permissions, user approval, secure rendering, tool isolation, and egress restrictions.
  • Assurance should include red teaming, vulnerability management, threat detection, incident response, and retesting after model, data, prompt, tool, or policy changes.

Requirements and Testing Scope Matrix

Summary of the referenced clauses, the testing scope they cover, and the expected evaluation cycle.

Reference Mandate Title Scope Required Testing Cycle
SAIF Elements 1 and 4 Foundational security and harmonized platform controls Assess identity, infrastructure, supply chains, model registries, secure-by-default tooling, and policy consistency Architectural review prior to production launch
SAIF Elements 2 and 5 AI telemetry detection, SOC response, and adaptive control Validate AI telemetry, SOC integration, threat detection, incident response, and control feedback Continuous monitoring and review after threat or behavior changes
SAIF Assurance: Red Teaming Adversarial testing across all SAIF risks Test data, infrastructure, models, applications, agent permissions, user approval, detection, and response Before go-live and after model, data, prompt, tool, permission, or architecture changes
SAIF 2.0 Agent Risk Map Human control, limited power, and observability Assess human controllers, least privilege, contextual permissions, tool isolation, and agent audit trails After adding tools, memory, agent workflows, or high-impact actions

Compliance Readiness Self-Assessment

Select items your organization has completed to evaluate your readiness score.

0%

Official source documents

Explore the latest SAIF Risk Map, risks, controls, and agent-security guidance on Google SAIF.

Go to the official source