Your agent uses Claude Sonnet for “rename this variable” and for “redesign the consensus protocol.” That is burning money. Hand-built routers (embeddings + rules + a sidecar classifier) become a second product. Amazon Bedrock Intelligent Prompt Routing sits in front of a family of models and chooses which one handles each prompt based on complexity signals — so trivial tool-planning steps stay cheap and hard multi-file refactors climb the model ladder. Combine with Prompt Management, Model Evaluation, and Budgets + Cost Anomaly.
⚡ TL;DR: Create a Bedrock prompt router (intelligent routing) over an approved model family for coding-agent inference. Point
InvokeModel/ Converse at the router ARN, not a single model ID. Keep system prompts in Prompt Management; score quality with Model Evaluation before widening traffic. Watch CloudWatch for routed-model distribution and fallback rates. Cap spend with Budgets. Related: Prompt Flows, ApplyGuardrail, PrivateLink Bedrock.
The cost curve agents ignore
Coding-agent loops are mostly easy steps:
- Format tool JSON
- Summarize a 40-line diff
- Choose between 3 tools
- Write a unit test for a pure function
Interspersed with hard steps:
- Cross-service IAM redesign
- Subtle race in distributed locking
- Multi-file API migration with breaking changes
Flat “always Opus/Sonnet” pricing means you pay frontier rates for JSON formatting. Intelligent Prompt Routing exists to collapse that waste without you maintaining a brittle if/else on prompt length.
| Approach | Pros | Cons |
|---|---|---|
| Single premium model | Simple ops | Expensive on easy steps |
| Hand router (rules/ML) | Full control | You own drift + eval |
| Bedrock Intelligent Prompt Routing | Managed choice inside family | Family/region constraints; still need eval |
| Prompt Flows static branches | Explicit graphs | Not per-prompt difficulty |
How routing fits the agent stack
User / orchestrator
→ Guardrail (ApplyGuardrail)
→ Intelligent Prompt Router → Model A (fast/cheap) or Model B (strong)
→ Tool runner / parser
Routing is not a replacement for:
- Prompt Management (versioned system prompts)
- Model Evaluation (offline/online quality gates)
- Your tool-selection policy (Verified Permissions / Cedar)
It only picks which foundation model serves this invoke.
# ✅ create an intelligent prompt router (CLI shape — pin to current Bedrock API)
aws bedrock create-prompt-router \
--prompt-router-name coding-agent-default \
--models modelArn=arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-haiku-4-5-20251001-v1:0 \
modelArn=arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-5-20250929-v1:0 \
--routing-criteria responseQualityDifference=0.5 \
--fallback-model modelArn=arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-5-20250929-v1:0 \
--tags key=workload,value=coding-agent
(Exact model IDs and CLI flags move with Bedrock releases — resolve current router-capable families in your Region before production.)
# ✅ invoke via router ARN instead of a fixed modelId
import boto3
br = boto3.client("bedrock-runtime")
ROUTER_ARN = "arn:aws:bedrock:us-east-1:111122223333:prompt-router/coding-agent-default"
def agent_complete(messages: list[dict], max_tokens: int = 2048) -> dict:
# ✅ Converse against the router — Bedrock picks the member model
resp = br.converse(
modelId=ROUTER_ARN,
messages=messages,
inferenceConfig={"maxTokens": max_tokens, "temperature": 0.2},
)
# Optional: log which model served (response metadata / trace when available)
return resp
❌ Hardcoding modelId="anthropic.claude-…" in every tool planner after you paid for a router — traffic never shifts.
When to force a model (escape hatches)
Routers optimize average cost/quality. Some steps must pin:
- Final customer-facing patch explanation (compliance tone)
- Security-sensitive code review
- Eval harness baselines (reproducibility)
# ✅ pin premium model for security review; router for everyday planning
SECURITY_MODEL = "anthropic.claude-sonnet-4-5-20250929-v1:0"
def complete(messages, *, task: str):
model = SECURITY_MODEL if task == "security_review" else ROUTER_ARN
return br.converse(modelId=model, messages=messages, inferenceConfig={"maxTokens": 4096})
Use Prompt Management so the security-review prompt version is pinned independently of the router.
Evaluate before you widen traffic
Routing without evaluation is vibes with extra steps. Before migrating 100% of agent traffic:
- Build a golden set: easy renames, medium refactors, hard design prompts
- Run Bedrock Model Evaluation on single-model baselines vs router
- Track task success (tests pass / human accept), not only BLEU-ish metrics
- Canary 10% → 50% → 100% with Budgets alarms
# ✅ cheap online shadow metric: log router choice + outcome
def log_route(tenant: str, task: str, router_meta: dict, success: bool):
print({
"tenant": tenant,
"task": task,
"routed_model": router_meta.get("model"),
"success": success,
"workload": "coding-agent",
})
# ship to CloudWatch EMF / ADOT — see ADOT multi-hop traces post
Cost controls that still apply
Intelligent routing reduces waste; it does not stop a runaway agent loop.
- Budgets + Cost Anomaly on Bedrock model spend
- Max tokens per step + max steps per session
- AppConfig kill switches to force Haiku-only during incidents
- PrivateLink so routing traffic never hairpins the public internet
# ✅ budget example: Bedrock coding-agent tag
aws budgets create-budget --account-id 111122223333 --budget file://bedrock-agent-budget.json
Production checklist
- [ ] Router covers only approved models your legal/security team cleared
- [ ] Agent code uses router ARN by default; pin list documented
- [ ] Model Evaluation compared router vs best single model on golden set
- [ ] CloudWatch dashboards: invokes by member model, error rate, P95 latency
- [ ] Fallback model set; alarm on fallback spike
- [ ] Guardrails applied before/after route (ApplyGuardrail)
- [ ] Prompt versions managed separately from routing
- [ ] Cost anomaly alert + AppConfig “router_off → haiku_only” switch tested
FAQ
Q: Is this the same as cross-Region inference profiles?
A: No. Inference profiles help with capacity/Region. Intelligent Prompt Routing chooses among models for quality/cost. You may combine both.
Q: Can I route between Claude and Llama arbitrarily?
A: Routers are constrained to supported families/combinations AWS documents for the feature. Do not assume any pair works — check current Bedrock docs for your Region.
Q: Will routing hurt coding quality?
A: It can if you skip evaluation. Gate rollout on task-success metrics; keep pin paths for high-risk tasks.
Operational rollout playbook
Week 1: shadow mode — log what the router would pick while still calling Sonnet for all traffic (if your stack supports dual-invoke; otherwise canary 5% on internal tenants only). Week 2: router for tool_plan and summarize_diff tasks only; pin security_review and architecture. Week 3: expand to general chat if Model Evaluation deltas stay within your SLA (e.g. ≤2% absolute drop in “tests pass after agent patch”). Rollback is a one-line modelId change plus AppConfig flag — practice it.
// ✅ feature-flagged router usage
const modelId = flags.usePromptRouter
? process.env.BEDROCK_ROUTER_ARN!
: process.env.BEDROCK_PINNED_MODEL!;
await bedrock.converse({ modelId, messages, inferenceConfig: { maxTokens: 2048 } });
Track $/successful task, not only $/1k tokens. Routing wins when successful tasks get cheaper without acceptance-rate collapse.
Related reading
- Amazon Bedrock Prompt Management: Versioned System Prompts
- Amazon Bedrock Model Evaluation: Score Coding-Agent Outputs
- AWS Budgets + Cost Anomaly: Cap Coding-Agent Spend
- Bedrock ApplyGuardrail API: Pre/Post Filters for Tool I/O
Route the easy steps downmarket, pin the hard ones, and prove quality with evaluation — that is how coding-agent fleets stay fast without lighting money on fire.
Last updated on October 1, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- Agentic Git Workflows: Atomic Commits From Noisy LLM Diffs
- Python String Methods: Every str Method With Real Production Examples
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
Newly added
- AWS Fault Injection Service: Chaos-Test Coding-Agent Pipelines (Sandbox Kill, Latency, IAM Denials)
- Amazon VPC Lattice: Service-to-Service Auth for Coding-Agent Tool Microservices
- Amazon Bedrock Intelligent Prompt Routing: Auto-Route Coding-Agent Calls Across Models for Cost and Latency
- AWS CloudFormation Hooks: Block Unsafe Infra Coding Agents Propose Before It Lands
- Amazon EventBridge Pipes: Wire DynamoDB Streams / SQS to Coding-Agent Tool Runners Without Glue Lambdas
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.