AWS Fault Injection Service: Chaos-Test Coding-Agent Pipelines (Sandbox Kill, Latency, IAM Denials)

2 views

Your agent demo always works: sandbox comes up, tools return 200, Bedrock answers in 800ms. Production night one: Fargate task killed mid-apply, tool API at 2s latency, AccessDenied on a KMS decrypt. Nobody practiced. AWS Fault Injection Service (FIS) lets you run controlled experiments — stop ECS tasks, throttle network, stress CPU, combine actions with stop conditions — against coding-agent pipelines so retries, idempotency, and kill switches are proven. Layer with Fargate Spot sandboxes, Lambda Destinations, AppConfig kill switches, and SCPs.

⚡ TL;DR: Define FIS experiment templates for (1) kill sandbox tasks mid-tool, (2) latency/packet loss toward tool endpoints, (3) IAM denial / credential stress where supported. Gate experiments with stop conditions (error-rate alarms). Run in a non-prod agent account on a schedule. Fail the CI “resilience” suite if the agent double-applies or leaves corrupt artifacts. Related: DynamoDB Transactions ledger, S3 conditional writes, Network Firewall.

What chaos means for coding agents

Traditional chaos kills web servers. Agent chaos targets the loop:

  1. Compute death — sandbox task/container stopped during apply_patch
  2. Latency — tool microservice or Bedrock path slows past timeouts
  3. Auth failure — IAM deny on S3/KMS/tool invoke mid-session
  4. Dependency brownout — queue delay, stream lag, Step Functions backlog

Success is not “system stays up.” Success is fail closed + idempotent resume: no double merges, no partial silent success, operator can kill via AppConfig.

Experiment Hypothesis Pass criteria
Stop ECS sandbox mid-tool Agent retries with same idempotency key Ledger shows one apply; artifacts consistent
+1000ms tool latency Timeouts trigger Destinations/DLQ No hung sessions > SLO; user sees error
IAM deny on artifact bucket Agent surfaces auth error No fallback to public upload
Spot interruption Rebalance continues job Session resumes from EFS/workspace state

Experiment template: kill the sandbox

json
// ✅ FIS template sketch: stop ECS tasks with coding-agent tag
{
  "description": "Kill coding-agent sandbox tasks mid-flight",
  "targets": {
    "AgentSandboxes": {
      "resourceType": "aws:ecs:task",
      "selectionMode": "COUNT(1)",
      "resourceTags": {
        "workload": "coding-agent",
        "role": "sandbox"
      }
    }
  },
  "actions": {
    "StopSandbox": {
      "actionId": "aws:ecs:stop-task",
      "parameters": { "availabilityZoneStatus": "all" },
      "targets": { "Tasks": "AgentSandboxes" }
    }
  },
  "stopConditions": [{
    "source": "aws:cloudwatch:alarm",
    "value": "arn:aws:cloudwatch:us-east-1:111122223333:alarm:agent-error-rate-critical"
  }],
  "roleArn": "arn:aws:iam::111122223333:role/FISExperimentRole",
  "tags": { "workload": "coding-agent" }
}
bash
# ✅ create + start (non-prod only)
aws fis create-experiment-template --cli-input-json file://agent-sandbox-kill.json
aws fis start-experiment --experiment-template-id EXT...

❌ Running FIS in production agent accounts on day one without stop conditions — chaos without brakes is an outage generator.

Latency and network experiments

Brownouts find timeout bugs faster than hard downs.

json
// ✅ action idea: aws:network:disrupt-connectivity or network latency actions
// (availability varies by resource type — pick supported actions in your Region)
{
  "actions": {
    "LatencyToTools": {
      "actionId": "aws:ecs:task-network-latency",
      "parameters": {
        "delayMilliseconds": "1000",
        "duration": "PT5M"
      },
      "targets": { "Tasks": "AgentSandboxes" }
    }
  }
}

Validate that:

  • HTTP clients use bounded timeouts (no 15-minute hangs)
  • Lambda Destinations or Pipe DLQs catch failures
  • Planner marks tool result failed and does not invent success

IAM denial drills

You may combine FIS compute actions with a pre-scripted IAM deny (attach an inline deny via a break-glass automation, or use SCP in a canary OU) to simulate credential/policy failures:

python
# ✅ verification probe after IAM deny injected on artifact writes
import boto3

s3 = boto3.client("s3")

def expect_fail_closed(bucket: str, key: str, body: bytes):
    try:
        s3.put_object(Bucket=bucket, Key=key, Body=body)
        raise AssertionError("FAIL: put succeeded during IAM deny experiment")
    except s3.exceptions.ClientError as e:
        code = e.response["Error"]["Code"]
        assert code in ("AccessDenied", "403"), code
        # ✅ agent UX must surface this — ❌ silent skip of artifact upload
        return "denied_ok"

Pair with KMS decrypt grants tests: revoke grant → tool must fail closed.

Idempotency is the real pass/fail

Chaos that “recovers” by applying a patch twice is a failed experiment. Instrument:

python
# ✅ post-experiment assertion helper
def assert_single_apply(ledger_table, tool_call_id: str):
    item = ledger_table.get_item(Key={"toolCallId": tool_call_id})["Item"]
    assert item["applyCount"] == 1, item
    assert item["status"] in ("succeeded", "failed", "aborted")

Scheduling and CI

Cadence Experiment Owner
Nightly Stop 1 sandbox task during soak test Platform
Weekly Latency injection 5 minutes Platform + agent team
Pre-release Full game day + AppConfig kill Eng + SRE
Per PR (canary acct) Lightweight stop-task on PR stack CI
bash
# ✅ EventBridge Scheduler → Lambda that starts FIS experiment (non-prod)
# Reuse patterns from overnight agent batches — do not cron from laptops
aws scheduler create-schedule \
  --name nightly-agent-fis-sandbox-kill \
  --schedule-expression "cron(0 2 * * ? *)" \
  --flexible-time-window Mode=OFF \
  --target file://fis-starter-target.json

Production checklist

  • [ ] FIS role least-privilege; only tagged workload=coding-agent targets
  • [ ] Stop conditions bound to error-rate / customer-impact alarms
  • [ ] Experiments blocked in prod by SCP or separate account until mature
  • [ ] Idempotency assertions automated after each run
  • [ ] AppConfig kill switch tested as a human control alongside FIS
  • [ ] Results stored (S3) with experiment id + agent version
  • [ ] Game day notes feed runbooks (Destinations, DLQ, resume from EFS)
  • [ ] Spot interruption behavior covered if you use Fargate Spot

FAQ

Q: Is FIS the same as game days?
A: FIS automates fault injection. Game days add humans, comms, and decision practice. Use both.

Q: Can FIS inject Bedrock failures?
A: Prefer latency/timeouts on the client path and mock layers in non-prod. Do not DDoS shared model APIs. Chaos your dependencies and compute first.

Q: Will this break multi-tenant isolation tests?
A: It should prove them. If killing tenant A’s sandbox affects tenant B, that is a finding — fix blast radius before more features.

Blast-radius rules for agent FIS

Never target untagged resources. Require workload=coding-agent and fis=allowed tags on experiment targets so a mis-aimed template cannot stop payments or corporate bastions. Cap selection to COUNT(1) or PERCENT(10) until your resume logic is boring. Record every experiment in the same ticket queue as production incidents — chaos findings that are not tickets become folklore.

yaml
# ✅ tag contract for FIS-eligible sandboxes
Tags:
  workload: coding-agent
  role: sandbox
  fis: allowed
  env: nonprod

When an experiment fails its pass criteria, stop the chaos calendar until the idempotency fix ships. More experiments on a broken ledger only multiply corrupt artifacts.

Related reading

If you have not killed a sandbox on purpose, you do not know whether your coding agent fails closed. FIS makes that practice routine — with stop conditions, not hope.

Last updated on October 1, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.