AWS Batch: Overnight Parallel Coding-Agent Jobs Without Babysitting EC2

2 views

You have 200 open PRs and a security scan that should finish before standup. Spinning EC2 by hand is babysitting. One fat Lambda dies at 15 minutes. AWS Batch gives you compute environments, job queues, and job definitions so overnight coding-agent fleets run in parallel with retries and Spot — then go to zero. Distinct from Fargate Spot sandboxes (interactive short tasks), EventBridge Scheduler overnight batches (cron/flexible time triggers — not the compute plane), CodeBuild sandboxes (buildspec CI shape), and FIS chaos (failure injection). Also distinct from our earlier Batch long-running jobs post (single long jobs vs parallel array fleets): here we design queue + array jobs for PR-review / repo-scan fan-out.

⚡ TL;DR: Define a Batch job definition that runs your agent container, attach a managed compute environment (Fargate or EC2/Spot), submit array jobs keyed by PR/repo, trigger the fan-out with EventBridge Scheduler, and cap spend with Budgets. Related: EventBridge Scheduler, Fargate Spot sandboxes, Step Functions agent graphs, Budgets + Cost Anomaly.

Why parallel overnight work hates EC2 babysitting

Coding-agent overnight failure modes:

  1. Serial crawl — one agent walks every repo until morning; half unfinished
  2. Lambda ceiling — 15-minute timeout kills deep PR reviews mid-tool-loop
  3. Orphan EC2 — someone left m5.4xlarge on after the job; finance pages you
  4. No retry semantics — Spot reclaim loses work with no Batch-style attempt history
Runner Best for Weak for
Lambda Seconds–minutes tools Hour-long multi-repo sweeps
CodeBuild CI-shaped buildspec Huge fan-out arrays + Spot mix
Fargate Spot (ad-hoc) Interactive sandboxes Queue fairness / array indexing
Batch Parallel jobs, retries, CE scaling Ultra-low-latency interactive chat

❌ Cron on a single bastion that sshs out to start containers “when it remembers.”

Architecture: CE + queue + array jobs

EventBridge Scheduler (02:00 IST)
  → Lambda/SFn enumerates open PRs → submits Batch array job
      Job queue: coding-agent-overnight
      Compute env: Fargate Spot + on-demand fallback
      Job def: coding-agent-pr-review:7
  → each array child: AGENT_PR_ID = AWS_BATCH_JOB_ARRAY_INDEX mapped
  → artifacts → S3; summary → SNS/Slack
bash
# ✅ managed Fargate compute environment (illustrative)
aws batch create-compute-environment \
  --compute-environment-name coding-agent-overnight-fargate \
  --type MANAGED \
  --state ENABLED \
  --compute-resources type=FARGATE,maxvCpus=256,subnets=subnet-aaa,subnet-bbb,securityGroupIds=sg-agent

aws batch create-job-queue \
  --job-queue-name coding-agent-overnight \
  --state ENABLED \
  --priority 10 \
  --compute-environment-order order=1,computeEnvironment=coding-agent-overnight-fargate

Put sensitive model keys in Secrets Manager / SSM; the job role gets bedrock:InvokeModel scoped to your region — pair with Network Firewall egress so agents cannot phone home to random IPs.

Job definition for a coding-agent container

json
{
  "jobDefinitionName": "coding-agent-pr-review",
  "type": "container",
  "platformCapabilities": ["FARGATE"],
  "containerProperties": {
    "image": "111122223333.dkr.ecr.us-east-1.amazonaws.com/coding-agent:1.8.0",
    "resourceRequirements": [
      {"type": "VCPU", "value": "2"},
      {"type": "MEMORY", "value": "4096"}
    ],
    "executionRoleArn": "arn:aws:iam::111122223333:role/batch-agent-execution",
    "jobRoleArn": "arn:aws:iam::111122223333:role/batch-agent-job",
    "networkConfiguration": {"assignPublicIp": "DISABLED"},
    "environment": [
      {"name": "AGENT_MODE", "value": "pr_review"},
      {"name": "PROMPT_VERSION", "value": "3"}
    ],
    "secrets": [
      {"name": "GITHUB_TOKEN", "valueFrom": "arn:aws:secretsmanager:us-east-1:111122223333:secret:agent/github"}
    ],
    "logConfiguration": {
      "logDriver": "awslogs",
      "options": {
        "awslogs-group": "/aws/batch/coding-agent",
        "awslogs-region": "us-east-1",
        "awslogs-stream-prefix": "pr-review"
      }
    }
  },
  "retryStrategy": {"attempts": 2},
  "timeout": {"attemptDurationSeconds": 7200}
}
python
# ✅ entrypoint reads array index → PR list object in S3
import os, json, boto3

s3 = boto3.client("s3")
idx = int(os.environ["AWS_BATCH_JOB_ARRAY_INDEX"])
manifest = json.loads(
    s3.get_object(Bucket="agent-overnight", Key="manifests/2026-10-04.json")["Body"].read()
)
pr = manifest["prs"][idx]
# run agent tool loop for this PR only
# ❌ for pr in all_prs: ... inside one job — defeats parallelism

Submit parallel array jobs (and don’t melt the API)

python
# ✅ submitter Lambda after building manifest
import boto3

batch = boto3.client("batch")

def submit_overnight(manifest_key: str, n: int):
    # Cap fan-out; huge arrays need chunking + queue fair-share
    assert 1 <= n <= 500
    return batch.submit_job(
        jobName="pr-review-2026-10-04",
        jobQueue="coding-agent-overnight",
        jobDefinition="coding-agent-pr-review:7",
        arrayProperties={"size": n},
        containerOverrides={
            "environment": [
                {"name": "MANIFEST_KEY", "value": manifest_key},
            ]
        },
    )
Pattern When Watch-out
Array job Homogeneous PR/repo units Manifest length must match array size
Many single jobs Heterogeneous CPU/memory API throttle; use queues
Step Functions Map → Batch Need per-item branching Extra orchestration cost

Trigger at 02:00 with EventBridge Scheduler — Scheduler wakes the submitter; Batch owns parallelism. Orchestrate multi-step “scan → cluster → open issues” with Step Functions when a single array child is not enough.

Spot, retries, and idempotency

Overnight fleets should prefer Spot/Fargate Spot with on-demand fallback in the compute environment. Agent work must be idempotent per PR:

  • Write review comments with a deterministic external id
  • Checkpoint tool ledger in DynamoDB before posting
  • Treat attempt 2 as resume-from-checkpoint, not double-comment
python
# ✅ idempotent comment key
comment_key = f"batch-review:{pr_number}:{prompt_version}:{commit_sha}"
# if exists in DynamoDB → skip post; still exit 0

Chaos-test reclaim and dependency failures with FIS in staging — not on Friday prod overnight.

Cost and blast-radius controls

Control How
maxvCpus on CE Hard concurrency ceiling
Job timeout Kill wedged tool loops
Budgets / anomaly Cap agent spend
SCPs Deny unexpected regions/services
Private subnets No public IP on containers

❌ Unbounded maxvCpus + Bedrock on-demand with no budget alarm — your “helpful overnight review” becomes a finance incident.

Scratch artifacts (lint logs, hunks) can land in S3 Express One Zone when hot; durable reviews go to standard S3 + your forge.

Production checklist

  • [ ] Managed CE sized with maxvCpus; Spot preferred + on-demand fallback
  • [ ] Job def pins image digest; secrets via Secrets Manager
  • [ ] Array jobs + S3 manifest; index-bounded
  • [ ] Retry + timeout; idempotent outputs per PR/commit
  • [ ] Scheduler triggers submitter; Batch owns fan-out
  • [ ] Logs to CloudWatch; traces with ADOT on tool hops
  • [ ] Budgets + tags workload=coding-agent, batch=overnight
  • [ ] Distinct from CodeBuild / Fargate interactive / long-single-job Batch post
  • [ ] Staging FIS for Spot reclaim / dependency fault

FAQ

Q: Batch Fargate vs EC2 compute environment?
A: Fargate for less patch ops and simpler networking; EC2 when you need GPUs, custom AMIs, or cheaper sustained Spot for huge fleets. Most PR-review agents are fine on Fargate.

Q: Why not only Step Functions Map with Lambda?
A: Map+Lambda is great under 15 minutes. Multi-hour repo scans, large checkouts, and Dockerized toolchains fit Batch better. Compose: SFn for workflow, Batch for heavy children.

Q: How is this different from the other Batch overnight post?
A: That post focuses on surviving long single jobs past Lambda limits. This one is parallelism: queues, array jobs, and fleet-scale overnight PR/repo sweeps.

Related reading

Submit the array. Sleep. Read the reviews at standup — not the EC2 console at 2am.

Last updated on October 4, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.