Your agent demo always works: sandbox comes up, tools return 200, Bedrock answers in 800ms. Production night one: Fargate task killed mid-apply, tool API at 2s latency, AccessDenied on a KMS decrypt. Nobody practiced. AWS Fault Injection Service (FIS) lets you run controlled experiments — stop ECS tasks, throttle network, stress CPU, combine actions with stop conditions — against coding-agent pipelines so retries, idempotency, and kill switches are proven. Layer with Fargate Spot sandboxes, Lambda Destinations, AppConfig kill switches, and SCPs.
⚡ TL;DR: Define FIS experiment templates for (1) kill sandbox tasks mid-tool, (2) latency/packet loss toward tool endpoints, (3) IAM denial / credential stress where supported. Gate experiments with stop conditions (error-rate alarms). Run in a non-prod agent account on a schedule. Fail the CI “resilience” suite if the agent double-applies or leaves corrupt artifacts. Related: DynamoDB Transactions ledger, S3 conditional writes, Network Firewall.
What chaos means for coding agents
Traditional chaos kills web servers. Agent chaos targets the loop:
- Compute death — sandbox task/container stopped during
apply_patch - Latency — tool microservice or Bedrock path slows past timeouts
- Auth failure — IAM deny on S3/KMS/tool invoke mid-session
- Dependency brownout — queue delay, stream lag, Step Functions backlog
Success is not “system stays up.” Success is fail closed + idempotent resume: no double merges, no partial silent success, operator can kill via AppConfig.
| Experiment | Hypothesis | Pass criteria |
|---|---|---|
| Stop ECS sandbox mid-tool | Agent retries with same idempotency key | Ledger shows one apply; artifacts consistent |
| +1000ms tool latency | Timeouts trigger Destinations/DLQ | No hung sessions > SLO; user sees error |
| IAM deny on artifact bucket | Agent surfaces auth error | No fallback to public upload |
| Spot interruption | Rebalance continues job | Session resumes from EFS/workspace state |
Experiment template: kill the sandbox
// ✅ FIS template sketch: stop ECS tasks with coding-agent tag
{
"description": "Kill coding-agent sandbox tasks mid-flight",
"targets": {
"AgentSandboxes": {
"resourceType": "aws:ecs:task",
"selectionMode": "COUNT(1)",
"resourceTags": {
"workload": "coding-agent",
"role": "sandbox"
}
}
},
"actions": {
"StopSandbox": {
"actionId": "aws:ecs:stop-task",
"parameters": { "availabilityZoneStatus": "all" },
"targets": { "Tasks": "AgentSandboxes" }
}
},
"stopConditions": [{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:us-east-1:111122223333:alarm:agent-error-rate-critical"
}],
"roleArn": "arn:aws:iam::111122223333:role/FISExperimentRole",
"tags": { "workload": "coding-agent" }
}
# ✅ create + start (non-prod only)
aws fis create-experiment-template --cli-input-json file://agent-sandbox-kill.json
aws fis start-experiment --experiment-template-id EXT...
❌ Running FIS in production agent accounts on day one without stop conditions — chaos without brakes is an outage generator.
Latency and network experiments
Brownouts find timeout bugs faster than hard downs.
// ✅ action idea: aws:network:disrupt-connectivity or network latency actions
// (availability varies by resource type — pick supported actions in your Region)
{
"actions": {
"LatencyToTools": {
"actionId": "aws:ecs:task-network-latency",
"parameters": {
"delayMilliseconds": "1000",
"duration": "PT5M"
},
"targets": { "Tasks": "AgentSandboxes" }
}
}
}
Validate that:
- HTTP clients use bounded timeouts (no 15-minute hangs)
- Lambda Destinations or Pipe DLQs catch failures
- Planner marks tool result
failedand does not invent success
IAM denial drills
You may combine FIS compute actions with a pre-scripted IAM deny (attach an inline deny via a break-glass automation, or use SCP in a canary OU) to simulate credential/policy failures:
# ✅ verification probe after IAM deny injected on artifact writes
import boto3
s3 = boto3.client("s3")
def expect_fail_closed(bucket: str, key: str, body: bytes):
try:
s3.put_object(Bucket=bucket, Key=key, Body=body)
raise AssertionError("FAIL: put succeeded during IAM deny experiment")
except s3.exceptions.ClientError as e:
code = e.response["Error"]["Code"]
assert code in ("AccessDenied", "403"), code
# ✅ agent UX must surface this — ❌ silent skip of artifact upload
return "denied_ok"
Pair with KMS decrypt grants tests: revoke grant → tool must fail closed.
Idempotency is the real pass/fail
Chaos that “recovers” by applying a patch twice is a failed experiment. Instrument:
- Tool ledger with DynamoDB Transactions
- Artifact uploads with S3 If-None-Match
- Workspace state on EFS for resume
# ✅ post-experiment assertion helper
def assert_single_apply(ledger_table, tool_call_id: str):
item = ledger_table.get_item(Key={"toolCallId": tool_call_id})["Item"]
assert item["applyCount"] == 1, item
assert item["status"] in ("succeeded", "failed", "aborted")
Scheduling and CI
| Cadence | Experiment | Owner |
|---|---|---|
| Nightly | Stop 1 sandbox task during soak test | Platform |
| Weekly | Latency injection 5 minutes | Platform + agent team |
| Pre-release | Full game day + AppConfig kill | Eng + SRE |
| Per PR (canary acct) | Lightweight stop-task on PR stack | CI |
# ✅ EventBridge Scheduler → Lambda that starts FIS experiment (non-prod)
# Reuse patterns from overnight agent batches — do not cron from laptops
aws scheduler create-schedule \
--name nightly-agent-fis-sandbox-kill \
--schedule-expression "cron(0 2 * * ? *)" \
--flexible-time-window Mode=OFF \
--target file://fis-starter-target.json
Production checklist
- [ ] FIS role least-privilege; only tagged
workload=coding-agenttargets - [ ] Stop conditions bound to error-rate / customer-impact alarms
- [ ] Experiments blocked in prod by SCP or separate account until mature
- [ ] Idempotency assertions automated after each run
- [ ] AppConfig kill switch tested as a human control alongside FIS
- [ ] Results stored (S3) with experiment id + agent version
- [ ] Game day notes feed runbooks (Destinations, DLQ, resume from EFS)
- [ ] Spot interruption behavior covered if you use Fargate Spot
FAQ
Q: Is FIS the same as game days?
A: FIS automates fault injection. Game days add humans, comms, and decision practice. Use both.
Q: Can FIS inject Bedrock failures?
A: Prefer latency/timeouts on the client path and mock layers in non-prod. Do not DDoS shared model APIs. Chaos your dependencies and compute first.
Q: Will this break multi-tenant isolation tests?
A: It should prove them. If killing tenant A’s sandbox affects tenant B, that is a finding — fix blast radius before more features.
Blast-radius rules for agent FIS
Never target untagged resources. Require workload=coding-agent and fis=allowed tags on experiment targets so a mis-aimed template cannot stop payments or corporate bastions. Cap selection to COUNT(1) or PERCENT(10) until your resume logic is boring. Record every experiment in the same ticket queue as production incidents — chaos findings that are not tickets become folklore.
# ✅ tag contract for FIS-eligible sandboxes
Tags:
workload: coding-agent
role: sandbox
fis: allowed
env: nonprod
When an experiment fails its pass criteria, stop the chaos calendar until the idempotency fix ships. More experiments on a broken ledger only multiply corrupt artifacts.
Related reading
- AWS Fargate Spot: Cheap Ephemeral Sandboxes for Coding Agents
- Lambda Destinations: Route Failed Agent Tool Invokes
- AppConfig Feature Flags: Kill Switches for Agent Tools
- DynamoDB Transactions: Atomic Tool-Ledger Writes
If you have not killed a sandbox on purpose, you do not know whether your coding agent fails closed. FIS makes that practice routine — with stop conditions, not hope.
Last updated on October 1, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- Agentic Git Workflows: Atomic Commits From Noisy LLM Diffs
- Python String Methods: Every str Method With Real Production Examples
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
Newly added
- AWS Fault Injection Service: Chaos-Test Coding-Agent Pipelines (Sandbox Kill, Latency, IAM Denials)
- Amazon VPC Lattice: Service-to-Service Auth for Coding-Agent Tool Microservices
- Amazon Bedrock Intelligent Prompt Routing: Auto-Route Coding-Agent Calls Across Models for Cost and Latency
- AWS CloudFormation Hooks: Block Unsafe Infra Coding Agents Propose Before It Lands
- Amazon EventBridge Pipes: Wire DynamoDB Streams / SQS to Coding-Agent Tool Runners Without Glue Lambdas
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.