Amazon S3 Express One Zone: Sub-Millisecond Scratch for Coding-Agent Tool Artifacts

2 views

A coding-agent turn is a storm of tiny files: eslint JSON, tsc incremental caches, patch hunks, pytest JUnit, AST dumps. Parking those on general-purpose S3 Standard works until P99 open latency and request charges show up in tool timeouts. Amazon S3 Express One Zone (directory buckets) is purpose-built for high-throughput, low-latency objects in a single Availability Zone — ideal ephemeral scratch colocated with your Fargate/EC2 sandbox. Distinct from S3 Object Lock for immutable artifacts (WORM compliance) and EFS shared workspaces (POSIX multi-turn mounts): Express One Zone is hot object scratch, not a filesystem and not an audit vault.

⚡ TL;DR: Create a directory bucket in the same AZ as agent sandboxes, use session-based auth / bucket-style APIs for high QPS PUT/GET, TTL or explicit delete after the turn, and promote only keepers to Standard / Object Lock buckets. Related: S3 Conditional Writes, ElastiCache Redis scratchpads, Fargate Spot sandboxes, Macie on artifact buckets.

Where scratch actually lives in an agent stack

Layer Store Lifetime Example
Token / tiny JSON Redis / MemoryDB Seconds–minutes Tool-result cache
Hot binary blobs S3 Express One Zone Seconds–hours Build cache, lint XML, diffs
Shared checkout EFS Hours–days Multi-turn repo mount
Release artifacts S3 Standard + Object Lock Months+ Promoted patches, audit bundles

❌ Putting every ephemeral lint log into Object Lock Governance mode “for safety” — you will pay forever and cannot clean up.

Create a directory bucket (same AZ as sandboxes)

bash
# ✅ directory bucket name MUST include AZ id suffix pattern — follow current AWS naming rules
# Example shape: coding-agent-scratch--use1-az4--x-s3
aws s3api create-bucket \
  --bucket coding-agent-scratch--use1-az4--x-s3 \
  --create-bucket-configuration \
    'Location={Type=AvailabilityZone,Name=use1-az4},Bucket={DataRedundancy=SingleAvailabilityZone,Type=Directory}' \
  --region us-east-1

Colocate Fargate tasks / EC2 sandboxes in that AZ (or accept cross-AZ latency). For Spot fleets (Fargate Spot), pin capacity providers / subnet to the Express AZ for the hot path.

Agent tool I/O pattern

python
# ✅ boto3 — write mid-turn artifacts with short keys under run_id/
import boto3, os, json

s3 = boto3.client("s3")
BUCKET = os.environ["SCRATCH_DIRECTORY_BUCKET"]  # ...--x-s3

def put_scratch(run_id: str, name: str, body: bytes, content_type: str = "application/octet-stream"):
    key = f"{run_id}/{name}"
    # Conditional write when you need single-writer semantics:
    # see If-None-Match patterns on Standard; verify Express support for your SDK version
    s3.put_object(Bucket=BUCKET, Key=key, Body=body, ContentType=content_type)
    return {"bucket": BUCKET, "key": key}

def get_scratch(run_id: str, name: str) -> bytes:
    key = f"{run_id}/{name}"
    return s3.get_object(Bucket=BUCKET, Key=key)["Body"].read()

def finish_turn(run_id: str, promote_keys: list[str], durable_bucket: str):
    # ✅ promote keepers to durable bucket; delete the rest
    for name in promote_keys:
        key = f"{run_id}/{name}"
        s3.copy_object(
            Bucket=durable_bucket,
            Key=f"runs/{run_id}/{name}",
            CopySource={"Bucket": BUCKET, "Key": key},
        )
    # list + delete scratch prefix (paginate in production)
    token = None
    while True:
        kwargs = {"Bucket": BUCKET, "Prefix": f"{run_id}/"}
        if token:
            kwargs["ContinuationToken"] = token
        page = s3.list_objects_v2(**kwargs)
        objs = [{"Key": o["Key"]} for o in page.get("Contents", [])]
        if objs:
            s3.delete_objects(Bucket=BUCKET, Delete={"Objects": objs})
        if not page.get("IsTruncated"):
            break
        token = page.get("NextContinuationToken")
bash
# ✅ CLI smoke: many small objects should feel snappier than Standard in-AZ
for i in $(seq 1 100); do
  echo "hunk $i" | aws s3 cp - s3://coding-agent-scratch--use1-az4--x-s3/run-demo/hunk-$i.txt
done

Express vs Redis vs EFS — pick deliberately

Need Prefer Why
Sub-kb JSON tool memo ElastiCache / MemoryDB Microsecond RAM; no object API
100KB–100MB tool blobs at high QPS S3 Express Cheap-ish durable-ish scratch without POSIX
Editors / compilers needing POSIX EFS open, locks, shared mounts
Immutable audit package S3 + Object Lock Compliance, not speed

❌ Mounting Express as if it were EFS. It is still an object API — great for put_object/get_object, not for gcc reading a tree of includes unless you sync to local disk first.

Security for scratch that still holds secrets

Scratch is ephemeral, not public:

  • Block public access; VPC gateway / interface endpoints where applicable
  • Encrypt with KMS CMKs; grants for agent roles only (KMS decrypt grants)
  • Scan promoted durable buckets with Macie — scratch may contain .env copies agents should never promote
  • Prefix by tenantId/runId; deny cross-tenant s3:GetObject with IAM conditions
json
{
  "Effect": "Allow",
  "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
  "Resource": "arn:aws:s3:::coding-agent-scratch--use1-az4--x-s3/$${aws:PrincipalTag/tenant}/*",
  "Condition": {
    "StringEquals": {
      "aws:PrincipalTag/workload": "coding-agent"
    }
  }
}

Failure modes agents hit

Issue Symptom Mitigation
AZ outage Scratch unreachable Fail turn soft; durable path on Standard multi-AZ; do not block human review on scratch
Forgotten deletes Bill creep Lifecycle-like sweeper Lambda on prefixes older than N hours
Cross-AZ sandbox Latency back to “meh” Pin compute AZ = Express AZ
Promoting secrets Leak to durable bucket Deny promote of *.env / scan before copy
Treating as WORM Wrong compliance story Use Object Lock bucket for audit keepers
bash
# ✅ sweeper sketch — delete scratch older than 6h (run hourly)
aws lambda invoke --function-name scratch-sweeper /tmp/out.json

Production checklist

  • [ ] Directory bucket in the same AZ as hot sandboxes
  • [ ] Key layout {tenant}/{runId}/... with IAM conditions
  • [ ] Explicit promote-then-delete (or sweeper) — no infinite scratch retention
  • [ ] Durable / Object Lock buckets only for keepers
  • [ ] Redis for tiny memos; Express for blobs; EFS for POSIX
  • [ ] KMS + no public access; Macie on durable side
  • [ ] Measure tool P99 GET/PUT vs Standard baseline before celebrating
  • [ ] Chaos: kill AZ path and ensure agent fails gracefully (FIS)
  • [ ] Cost tags workload=coding-agent, tier=scratch-express

FAQ

Q: Is Express One Zone cheaper than Standard?
A: Storage and request pricing differ — for high QPS small objects, Express often wins on latency and can win on request economics; for cold rarely-touched blobs, Standard (or IA) is usually better. Benchmark your agent’s object size histogram.

Q: Can I use Object Lock on Express?
A: Object Lock is the durability/compliance tool for general purpose buckets. Keep Express for scratch; lock the promoted audit copies.

Q: Why not only local NVMe on the sandbox?
A: Local disk is fastest — until the task dies, another tool pod needs the artifact, or you want a progress UI to download a hunk. Express is the shared hot object tier between processes.

Related reading

Keep scratch hot, keep audits durable, and stop paying Standard latency for files that live ten minutes.

Last updated on October 3, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.