Google Secure AI Framework (SAIF 2.0) คืออะไร และใช้ป้องกัน AI Agent อย่างไร
SAIF เปลี่ยนคำว่า AI Security ให้เป็นแผนที่ของ Risk, Component และ Control ตั้งแต่ Data ไปจนถึง Agent หน้านี้สรุป 6 Core Element, SAIF Risk Map, แนวทางใหม่ของ SAIF 2.0 และจุดที่ AI Red Teaming ต้องพิสูจน์ก่อนระบบสร้างผลกระทบจริง
6 Core Element วางรากฐานให้ AI Security ไม่แยกจาก Cybersecurity
SAIF รุ่นแรกกำหนด 6 Core Element ที่ยังเป็นฐานของกรอบปัจจุบัน แนวคิดสำคัญคือขยาย Control ที่พิสูจน์แล้วเข้าสู่ AI และเพิ่ม Control ใหม่สำหรับความเสี่ยงที่เกิดจาก Data, Model และ Agent

- Expand strong security foundations นำ Identity, Secure-by-default Infrastructure, Supply Chain Security และ Security Expertise มาใช้กับระบบ AI
- Extend detection and response รวม Model Input, Output, Tool Call และ AI Asset เข้ากับ Threat Intelligence, SOC และ Incident Response
- Automate defenses ใช้ Automation ลดเวลาตรวจจับและตอบสนองต่อภัยคุกคามที่ขยายตัวด้วย AI
- Harmonize platform-level controls ทำให้ Policy และ Control สอดคล้องกันระหว่าง Model Provider, Platform, Application และทีมภายใน
- Adapt controls through faster feedback loops ใช้ Incident, User Feedback, Red Team Finding และ Model Change ปรับ Control อย่างต่อเนื่อง
- Contextualize AI system risks ประเมิน End-to-end Business Process, Data Lineage, Validation และ Operational Behavior ไม่แยกโมเดลออกจากระบบจริง
SAIF Risk Map แบ่งระบบเป็น 4 Component Area เพื่อหา Control ที่ถูกจุด
SAIF Map แบ่งวงจรพัฒนา AI เป็น Data Layer, Infrastructure Layer, Model Layer และ Application Layer พร้อมแสดงตำแหน่งที่ Risk ถูกนำเข้า ถูกเปิดเผย และถูกลดผลกระทบ
- Data Layer ครอบคลุม Data Source, Ingestion, Processing, Training Data และ RAG Data โดยเน้น Provenance, Privacy, Quality, Access และ Integrity
- Infrastructure Layer ครอบคลุม Code, Framework, Compute, Pipeline, Registry, Serving และ Storage ที่อาจถูกโจมตีผ่าน Supply Chain หรือ Misconfiguration
- Model Layer ครอบคลุม Training, Tuning, Evaluation, Model Weight, Input Handling และ Output Handling ที่ต้องทดสอบ Robustness และ Disclosure
- Application Layer ครอบคลุม User Interface, API, Plugin, Agent และ Tool ซึ่งเป็นจุดที่ Prompt Injection, Excessive Access และ Rogue Action สร้างผลกระทบต่อโลกจริง
- SAIF ระบุ Risk 15 หมวดและเชื่อมแต่ละ Risk กับ Control ของ Model Creator, Model Consumer หรือทั้งสองฝ่าย
SAIF 2.0 เพิ่ม Agent Risk Map และหลัก Secure-by-design สำหรับ Agent
SAIF 2.0 ขยายกรอบไปยัง Agent ที่วางแผน ใช้ Memory เรียก Tool และดำเนินการแทนผู้ใช้ เพราะความเสี่ยงไม่ได้อยู่ที่คำตอบเพียงอย่างเดียว แต่อยู่ที่อำนาจในการกระทำ

- Agent ต้องมี Human Controller ที่ระบุชัด เพื่อกำหนดว่าใครอนุมัติ เปลี่ยน Policy หยุดระบบ และรับผิดชอบต่อผลลัพธ์
- อำนาจของ Agent ต้องถูกจำกัดด้วย Least Privilege, Contextual Permission, Tool Scope, Credential Isolation และ User Confirmation สำหรับ High-impact Action
- Action และ Planning ต้องสังเกตและตรวจสอบย้อนหลังได้ผ่าน Agent Observability, Tool Call Log, Reason Code, Policy Decision และ Tamper-resistant Audit Trail
- Agent Risk Map เน้น Rogue Action, Sensitive Data Disclosure, Insecure Integrated Component และ Prompt Injection ที่อาจไหลผ่าน Perception, Reasoning, Memory และ Tool
Google AI Red Team ทดสอบระบบ ไม่ใช่ทดสอบโมเดลเพียงจุดเดียว
แนวทางของ Google AI Red Team เน้นผสม Threat Intelligence, Security Expertise และ AI Research เพื่อจำลองผู้โจมตีต่อ Product และ Feature จริง แล้วส่งผลกลับไปยังทีมป้องกัน
- เริ่มจาก Threat Model ของ Product แล้วทดสอบทั้ง Model, Data, Application, Infrastructure, Identity, Supply Chain และ Human Workflow
- สร้าง Attack Path เช่น Indirect Prompt Injection ที่เข้าสู่ RAG แล้วชักนำ Agent ให้เรียก Tool และส่ง Canary Secret ออกไปยัง Endpoint ที่ควบคุมโดยทีมทดสอบ
- ทดสอบ Control แบบหลายชั้น ได้แก่ Input Validation, Output Sanitization, Agent Permission, User Approval, Rate Limit, Egress Control และ Detection
- วัดผลทั้ง Prevention, Detection, Time to Triage, Containment และ Recovery เพื่อให้ Finding นำไปปรับ Product และ Incident Response ได้
Control ที่ควรพิสูจน์ก่อนเปิดใช้ AI และ Agent
SAIF จัด Control เป็น Data, Infrastructure, Model, Application, Assurance และ Governance โดย Red Teaming, Vulnerability Management, Threat Detection และ Incident Response ใช้กับ Risk ทุกหมวด
- Data Control ต้องพิสูจน์ Inventory, Access, Provenance, Integrity, Privacy และการป้องกัน Data Poisoning ตั้งแต่ Source ถึง Production
- Model Control ต้องพิสูจน์ Input Validation, Output Sanitization, Adversarial Testing, Model Access และความทนทานต่อ Prompt Injection
- Application Control ต้องพิสูจน์ Authentication, Rate Limit, Agent Permission, User Approval, Secure Rendering, Tool Isolation และ Egress Restriction
- Assurance ต้องมี Red Teaming, Vulnerability Management, Threat Detection, Incident Response และ Retest เมื่อ Model, Data, Prompt, Tool หรือ Policy เปลี่ยน
ตารางเปรียบเทียบข้อกำหนดและการทดสอบความปลอดภัย
สรุปรายละเอียดประกาศ บทบัญญัติ และขอบเขตการทดสอบความปลอดภัยที่จำเป็น
| ประกาศ และข้อกำหนด | สาระสำคัญ | ขอบเขตการทดสอบ | รอบเวลาทดสอบ |
|---|---|---|---|
| SAIF Element 1 และ 4 | Security Foundation และ Platform Control | ตรวจ Identity, Infrastructure, Supply Chain, Model Registry, Secure-by-default Tooling และ Policy Consistency | ตรวจประเมินสถาปัตยกรรมก่อนนำไปใช้งานจริง |
| SAIF Element 2 และ 5 | Detection, Response และ Feedback Loop | ตรวจ AI Telemetry, SOC Integration, Threat Detection, Incident Response และการนำ Finding ไปปรับ Control | เฝ้าระวังต่อเนื่องและทบทวนเมื่อ Threat หรือ System Behavior เปลี่ยน |
| SAIF Assurance: Red Teaming | การทดสอบแบบ Adversarial ครอบคลุม Risk ทุกหมวด | ทดสอบ Data, Infrastructure, Model, Application, Agent Permission, User Approval, Detection และ Response | ก่อน Go-live และหลังเปลี่ยน Model, Data, Prompt, Tool, Permission หรือ Architecture |
| SAIF 2.0 Agent Risk Map | Human Control, Limited Power และ Observability | ประเมิน Human Controller, Least Privilege, Contextual Permission, Tool Isolation และ Agent Audit Trail | ทุกครั้งที่เพิ่ม Tool, Memory, Agent Workflow หรือ High-impact Action |
เครื่องมือประเมินความพร้อม
คลิกเลือกรายการที่องค์กรของท่านได้ดำเนินการแล้ว เพื่อคำนวณระดับความพร้อมเบื้องต้น
