Your coding agent looks fast in a quiet sandbox and then melts when twenty engineers hit “Review PR” at 10:05. On-demand Bedrock shares a regional pool — throttles and queueing show up as tool-step timeouts, not polite 429s in the UI. Amazon Bedrock Provisioned Throughput lets you purchase model units for supported foundation models / custom models so inference capacity is reserved. Pair with Intelligent Prompt Routing for cheap vs hard prompts and Model Evaluation before you lock a model into a commitment.
⚡ TL;DR: Measure tokens/minute and P95 latency under peak concurrent agent sessions; buy Provisioned Throughput (or PT with commitment) for the hot model IDs; point production
modelIdat the provisioned ARN; keep on-demand as overflow with explicit degrade UX; alarm onInvocationThrottlesand latency. Related: Budgets + Cost Anomaly, Prompt Management, AppConfig kill switches.
When on-demand stops being “simpler”
| Signal | Meaning | Action |
|---|---|---|
ThrottlingException bursts at standup |
Pool contention | PT for primary model |
| P95 TTFT > SLO while error rate OK | Queueing / capacity | PT + smaller maxTokens on overflow |
| CI agent fleet spikes hourly | Predictable load shape | PT sized to CI concurrency |
| Spiky hobby traffic | Unpredictable | Stay on-demand + WAF/rate limits |
❌ Buying a year of PT for an experimental model you have not evaluated — commitments amplify bad model picks.
Size from agent telemetry, not vibes
# ✅ pull Bedrock invocation metrics for your agent model (CloudWatch)
aws cloudwatch get-metric-statistics \
--namespace AWS/Bedrock \
--metric-name Invocations \
--dimensions Name=ModelId,Value=anthropic.claude-3-5-sonnet-20241022-v2:0 \
--start-time 2026-09-25T00:00:00Z \
--end-time 2026-10-02T00:00:00Z \
--period 3600 \
--statistics Sum
aws cloudwatch get-metric-statistics \
--namespace AWS/Bedrock \
--metric-name InvocationThrottles \
--dimensions Name=ModelId,Value=anthropic.claude-3-5-sonnet-20241022-v2:0 \
--start-time 2026-09-25T00:00:00Z \
--end-time 2026-10-02T00:00:00Z \
--period 3600 \
--statistics Sum
Convert peak concurrent sessions × avg tokens/request × safety factor into model units using the current Bedrock PT calculator for that model. Leave headroom for tool-heavy multi-hop graphs (Step Functions).
Purchase and wire the provisioned model ARN
# ✅ create Provisioned Throughput (example shape — confirm model support & commitment in your region)
aws bedrock create-provisioned-model-throughput \
--provisioned-model-name coding-agent-sonnet-pt \
--model-id anthropic.claude-3-5-sonnet-20241022-v2:0 \
--model-units 2 \
--commitment-duration OneMonth \
--tags Key=workload,Value=coding-agent Key=slo,Value=p95-ttft-2s
# ✅ describe until ACTIVE, then use provisionedModelArn as modelId in Converse
aws bedrock list-provisioned-model-throughputs \
--query 'provisionedModelSummaries[?provisionedModelName==`coding-agent-sonnet-pt`]'
// ✅ production client: prefer PT ARN; overflow to on-demand with degrade flag
import { BedrockRuntimeClient, ConverseCommand } from "@aws-sdk/client-bedrock-runtime";
const PT_ARN = process.env.BEDROCK_PT_ARN!; // provisioned model ARN
const OD_ID = "anthropic.claude-3-5-sonnet-20241022-v2:0";
const client = new BedrockRuntimeClient({});
export async function converseAgent(messages: any[], allowOverflow: boolean) {
try {
return await client.send(
new ConverseCommand({
modelId: PT_ARN,
messages,
inferenceConfig: { maxTokens: 4096, temperature: 0.2 },
})
);
} catch (e: any) {
if (!allowOverflow || e.name !== "ThrottlingException") throw e;
// ✅ explicit overflow path — UI should show "degraded capacity"
return client.send(
new ConverseCommand({
modelId: OD_ID,
messages,
inferenceConfig: { maxTokens: 2048, temperature: 0.2 },
})
);
}
}
Combine with routing — do not confuse the two
- Prompt Routing chooses which model family/tier serves a prompt.
- Provisioned Throughput reserves capacity for a specific model (or custom model) you already trust.
A common production pattern: route easy prompts to a cheap on-demand Haiku-class model; send hard PR reviews to a PT-backed Sonnet/Opus-class ARN. Version system prompts with Prompt Management so PT capacity is not wasted on prompt regressions.
Cost controls that belong next to PT
// ✅ Budgets filter tag for PT + on-demand agent spend
{
"Tags": {
"Key": "workload",
"Values": ["coding-agent"]
}
}
- Tag every PT resource and the IAM roles that call it
- Cap monthly with Budgets + Anomaly Detection
- Kill switch in AppConfig to force read-only / smaller model
❌ Assuming PT removes the need for retries — still use bounded exponential backoff and circuit breakers.
Production checklist
- [ ] Peak concurrency and tokens/min measured from real agent traffic (7–14 days)
- [ ] Model frozen via evaluation scores before commitment
- [ ]
modelIdin prod points at provisioned ARN; config-driven, not hard-coded in ten services - [ ] Overflow policy documented (fail closed vs degrade to on-demand)
- [ ] CloudWatch alarms: throttles, P95 latency, 5xx from agent BFF
- [ ] Cost Budgets on
workload=coding-agentincluding PT - [ ] Commitment calendar: renewal / scale-down review before term ends
- [ ] Chaos: simulate throttle and prove UX (FIS for dependent infra)
FAQ
Q: Does PT work with every Bedrock model?
A: No — support and unit definitions are model- and region-specific. Check the Bedrock console / docs for your target model before promising an SLO.
Q: PT vs provisioned concurrency on Lambda?
A: Different layers. Lambda PC warms compute. Bedrock PT reserves model inference capacity. Coding agents usually need both tuned.
Q: Can I PT a custom fine-tune?
A: Often yes for custom models imported/fine-tuned into Bedrock — validate the ARN type and minimum units for your case.
Related reading
- Amazon Bedrock Intelligent Prompt Routing
- Amazon Bedrock Model Evaluation
- AWS Budgets + Cost Anomaly: Cap Coding-Agent Spend
- Amazon Bedrock Prompt Management
Reserve capacity for the model that protects your SLO, route the easy work elsewhere, and treat PT as an engineering commitment — not a credit-card reflex.
Last updated on October 2, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
- Java Virtual Threads vs Traditional Threads: What Nobody Tells You
- Agentic Git Workflows: Atomic Commits From Noisy LLM Diffs
Newly added
- AWS IAM Access Analyzer: Find Over-Privileged Coding-Agent Roles Before They Leak
- Amazon CloudWatch Application Signals: SLOs and Traces for Multi-Hop Coding-Agent Tools
- AWS Config Conformance Packs: Continuous Compliance Guards for Coding-Agent Accounts
- Amazon Bedrock Provisioned Throughput: Reserved Capacity for Coding-Agent Latency SLOs
- AWS Lambda Response Streaming: Stream Coding-Agent Tokens Without Buffering Full Completions
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.