AI-Native SDLC Security Engineering: Securing AI Coding Agents Against Poisoning and Sandbox Evasion
This guide uses the Anthropic AI-Native SDLC Playbook and coding-agent vendor security documentation to recommend risk-based controls. It is not a law or certification standard.
- Target Audience
- Engineering and security owners adopting AI coding agents, with controls selected for the product and environment
- Mandatory Frequency
- Check behavior and permissions on adoption or tool changes. Select tests and review depth for task risk. These are STH engineering recommendations.
- Required Scope
- Permissions, approvals, hooks, sandboxing, networking, dependencies and release evidence
- Non-Compliance Risk
- Context poisoning, credential theft, unauthorized commands and untrusted dependencies
AI-native SDLC is engineering guidance, not a regulatory standard
This guide uses the Anthropic AI-Native SDLC Playbook and coding-agent vendor security documentation to recommend risk-based controls. It is not a law or certification standard.

Select permissions, hooks, sandboxing and review for task risk. Verify actual boundaries and do not treat a single control as complete protection.
- Limit permissions — Scope files and commands
- Inspect tool calls — Hooks reduce risk, not all attacks
- Isolate execution — Sandbox and control network access
- Review before release — Tests, review and evidence
- The playbook connects planning, design, building, testing, deployment and maintenance through verifiable requirements and execution evidence.
- Context poisoning can arise when an agent ingests untrusted repository, issue or pull-request text as instructions.
- Overbroad file, command or network permissions can expose data and credentials even when the agent starts in a designated workspace.
- Verify the identity, provenance and version of proposed dependencies. A lockfile alone does not establish that a package is safe.
Limit permissions and inspect tool calls
Combine supported permissions, approvals and hooks. Distinguish organizational control recommendations from the capabilities of each product.
- Scope files, commands and network destinations to the task, and keep secrets out of agent context and runtime environments.
- Hooks can inspect parameters and deny unauthorized actions. Cover the relevant tools and test bypass paths. Hooks or AST inspection do not detect all prompt injection.
- Context fencing can label untrusted data. Delimiters and prompt sanitizers are not complete security boundaries.
- Define accountable approval or policy gates for high-impact actions, including unsandboxed execution and expanded access.
Verify sandbox and network boundaries in operation
Use mechanisms supported by the tool and operating system, and test their actual boundaries. A container or mode name alone does not prove complete host isolation.
- Limit file reads and writes, inspect mounts, symlinks and accessible credentials, and choose containers or microVMs where appropriate to risk.
- Restrict egress to destinations required by the task, including any external tools operating outside the sandbox.
- Check sandbox-disabled modes, privilege escalation paths and differences between local, cloud and supported operating systems.
- Apply appropriate resource limits and monitoring. Do not claim that sandboxing prevents every breakout.
Test, review and retain evidence before release
Select testing and review depth for change risk. CI and accountable owners should enforce the organization release policy.
- Test changed behavior and relevant negative cases. Apply SAST, SCA and secret scanning where appropriate to language, scope and risk.
- Use protected branches and release gates under organizational policy, with human review for high-impact changes. The playbook does not mandate a two-person rule for every task.
- Verify dependency identity and provenance alongside lockfiles and tests. Subresource integrity is not a universal dependency-verification mechanism.
- Retain necessary tool versions, changed files, test results and approvals. Redact secrets and set retention periods rather than recording every prompt and tool parameter indiscriminately.
Requirements and Testing Scope Matrix
Summary of the referenced clauses, the testing scope they cover, and the expected evaluation cycle.
| Reference | Mandate Title | Scope Required | Testing Cycle |
|---|---|---|---|
| Anthropic AI-Native SDLC Playbook | AI-native SDLC is engineering guidance, not a regulatory standard | This guide uses the Anthropic AI-Native SDLC Playbook and coding-agent vendor security documentation to recommend risk-based controls. It is not a law or certification standard. | According to applicable scope and service risk, not a universal testing cycle |
| Vendor permission and hook documentation | Limit permissions and inspect tool calls | Combine supported permissions, approvals and hooks. Distinguish organizational control recommendations from the capabilities of each product. | According to applicable scope and service risk, not a universal testing cycle |
| Vendor sandbox documentation | Verify sandbox and network boundaries in operation | Use mechanisms supported by the tool and operating system, and test their actual boundaries. A container or mode name alone does not prove complete host isolation. | According to applicable scope and service risk, not a universal testing cycle |
| STH release gate recommendations | Test, review and retain evidence before release | Select testing and review depth for change risk. CI and accountable owners should enforce the organization release policy. | According to applicable scope and service risk, not a universal testing cycle |
Compliance Readiness Self-Assessment
Select items your organization has completed to evaluate your readiness score.
Frequently Asked Questions (FAQ)
Key answers and practical guidance addressing common compliance questions.
How can context poisoning occur?
Context poisoning can arise when an agent ingests untrusted repository, issue or pull-request text as instructions.
Do hooks and context fencing prevent all prompt injection?
Hooks can inspect parameters and deny unauthorized actions. Cover the relevant tools and test bypass paths. Hooks or AST inspection do not detect all prompt injection. Context fencing can label untrusted data. Delimiters and prompt sanitizers are not complete security boundaries.
Does a container or sandbox remove all risk?
Use mechanisms supported by the tool and operating system, and test their actual boundaries. A container or mode name alone does not prove complete host isolation. Restrict egress to destinations required by the task, including any external tools operating outside the sandbox.
