Amazon Bedrock Provisioned Throughput: Reserved Capacity for Coding-Agent Latency SLOs

2 views

Your coding agent looks fast in a quiet sandbox and then melts when twenty engineers hit “Review PR” at 10:05. On-demand Bedrock shares a regional pool — throttles and queueing show up as tool-step timeouts, not polite 429s in the UI. Amazon Bedrock Provisioned Throughput lets you purchase model units for supported foundation models / custom models so inference capacity is reserved. Pair with Intelligent Prompt Routing for cheap vs hard prompts and Model Evaluation before you lock a model into a commitment.

⚡ TL;DR: Measure tokens/minute and P95 latency under peak concurrent agent sessions; buy Provisioned Throughput (or PT with commitment) for the hot model IDs; point production modelId at the provisioned ARN; keep on-demand as overflow with explicit degrade UX; alarm on InvocationThrottles and latency. Related: Budgets + Cost Anomaly, Prompt Management, AppConfig kill switches.

When on-demand stops being “simpler”

Signal Meaning Action
ThrottlingException bursts at standup Pool contention PT for primary model
P95 TTFT > SLO while error rate OK Queueing / capacity PT + smaller maxTokens on overflow
CI agent fleet spikes hourly Predictable load shape PT sized to CI concurrency
Spiky hobby traffic Unpredictable Stay on-demand + WAF/rate limits

❌ Buying a year of PT for an experimental model you have not evaluated — commitments amplify bad model picks.

Size from agent telemetry, not vibes

bash
# ✅ pull Bedrock invocation metrics for your agent model (CloudWatch)
aws cloudwatch get-metric-statistics \
  --namespace AWS/Bedrock \
  --metric-name Invocations \
  --dimensions Name=ModelId,Value=anthropic.claude-3-5-sonnet-20241022-v2:0 \
  --start-time 2026-09-25T00:00:00Z \
  --end-time 2026-10-02T00:00:00Z \
  --period 3600 \
  --statistics Sum

aws cloudwatch get-metric-statistics \
  --namespace AWS/Bedrock \
  --metric-name InvocationThrottles \
  --dimensions Name=ModelId,Value=anthropic.claude-3-5-sonnet-20241022-v2:0 \
  --start-time 2026-09-25T00:00:00Z \
  --end-time 2026-10-02T00:00:00Z \
  --period 3600 \
  --statistics Sum

Convert peak concurrent sessions × avg tokens/request × safety factor into model units using the current Bedrock PT calculator for that model. Leave headroom for tool-heavy multi-hop graphs (Step Functions).

Purchase and wire the provisioned model ARN

bash
# ✅ create Provisioned Throughput (example shape — confirm model support & commitment in your region)
aws bedrock create-provisioned-model-throughput \
  --provisioned-model-name coding-agent-sonnet-pt \
  --model-id anthropic.claude-3-5-sonnet-20241022-v2:0 \
  --model-units 2 \
  --commitment-duration OneMonth \
  --tags Key=workload,Value=coding-agent Key=slo,Value=p95-ttft-2s

# ✅ describe until ACTIVE, then use provisionedModelArn as modelId in Converse
aws bedrock list-provisioned-model-throughputs \
  --query 'provisionedModelSummaries[?provisionedModelName==`coding-agent-sonnet-pt`]'
typescript
// ✅ production client: prefer PT ARN; overflow to on-demand with degrade flag
import { BedrockRuntimeClient, ConverseCommand } from "@aws-sdk/client-bedrock-runtime";

const PT_ARN = process.env.BEDROCK_PT_ARN!; // provisioned model ARN
const OD_ID = "anthropic.claude-3-5-sonnet-20241022-v2:0";
const client = new BedrockRuntimeClient({});

export async function converseAgent(messages: any[], allowOverflow: boolean) {
  try {
    return await client.send(
      new ConverseCommand({
        modelId: PT_ARN,
        messages,
        inferenceConfig: { maxTokens: 4096, temperature: 0.2 },
      })
    );
  } catch (e: any) {
    if (!allowOverflow || e.name !== "ThrottlingException") throw e;
    // ✅ explicit overflow path — UI should show "degraded capacity"
    return client.send(
      new ConverseCommand({
        modelId: OD_ID,
        messages,
        inferenceConfig: { maxTokens: 2048, temperature: 0.2 },
      })
    );
  }
}

Combine with routing — do not confuse the two

  • Prompt Routing chooses which model family/tier serves a prompt.
  • Provisioned Throughput reserves capacity for a specific model (or custom model) you already trust.

A common production pattern: route easy prompts to a cheap on-demand Haiku-class model; send hard PR reviews to a PT-backed Sonnet/Opus-class ARN. Version system prompts with Prompt Management so PT capacity is not wasted on prompt regressions.

Cost controls that belong next to PT

json
// ✅ Budgets filter tag for PT + on-demand agent spend
{
  "Tags": {
    "Key": "workload",
    "Values": ["coding-agent"]
  }
}

❌ Assuming PT removes the need for retries — still use bounded exponential backoff and circuit breakers.

Production checklist

  • [ ] Peak concurrency and tokens/min measured from real agent traffic (7–14 days)
  • [ ] Model frozen via evaluation scores before commitment
  • [ ] modelId in prod points at provisioned ARN; config-driven, not hard-coded in ten services
  • [ ] Overflow policy documented (fail closed vs degrade to on-demand)
  • [ ] CloudWatch alarms: throttles, P95 latency, 5xx from agent BFF
  • [ ] Cost Budgets on workload=coding-agent including PT
  • [ ] Commitment calendar: renewal / scale-down review before term ends
  • [ ] Chaos: simulate throttle and prove UX (FIS for dependent infra)

FAQ

Q: Does PT work with every Bedrock model?
A: No — support and unit definitions are model- and region-specific. Check the Bedrock console / docs for your target model before promising an SLO.

Q: PT vs provisioned concurrency on Lambda?
A: Different layers. Lambda PC warms compute. Bedrock PT reserves model inference capacity. Coding agents usually need both tuned.

Q: Can I PT a custom fine-tune?
A: Often yes for custom models imported/fine-tuned into Bedrock — validate the ARN type and minimum units for your case.

Related reading

Reserve capacity for the model that protects your SLO, route the easy work elsewhere, and treat PT as an engineering commitment — not a credit-card reflex.

Last updated on October 2, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.