When a coding-agent session feels “slow,” the blame ping-pongs between Bedrock, the tool Lambda, EFS, and “the frontend.” Amazon CloudWatch Application Signals (built on OpenTelemetry / CloudWatch agent instrumentation) discovers services, charts golden signals, and lets you attach SLOs so a sandbox P99 regression pages you before Twitter does. Use it alongside FIS chaos tests and Lambda Destinations — Signals tells you what breached; those tell you how you fail closed.
⚡ TL;DR: Instrument agent BFF + tool runtimes with the CloudWatch agent / ADOT; enable Application Signals; define SLOs on latency and availability per tool service; link traces for slow multi-hop sessions; alarm on burn rate, not only hard error counts. Related: Step Functions agent graphs, VPC Lattice tool auth, App Runner agent APIs, EventBridge Pipes runners.
What to model as “services”
Treat each hop as a service Application Signals can see:
agent-bff— HTTP API that starts sessionsplanner— Bedrock-facing Lambda / containertool-lint,tool-sandbox,tool-rag— discrete tool microservicessession-store— dependencies (DynamoDB/MemoryDB) as observed dependencies
| Golden signal | Agent meaning |
|---|---|
| Latency | Time to first useful token / tool result |
| Faults | 5xx, tool crashes, Bedrock 5xx mapped |
| Errors | Business failures (lint failed) — keep separate from faults |
| Traffic | Sessions started / tools invoked per minute |
❌ One mega-service name coding-agent for everything — you lose which tool burned the SLO.
Enable Application Signals on the agent fleet
# ✅ example: enable Application Signals in the account/region (console or APIs evolve — verify)
aws application-signals list-services --region us-east-1
# ✅ CloudWatch agent config snippet for host/container metrics + traces (OTLP)
# ✅ cwagent-config excerpt — traces + application signals oriented
traces:
traces_collected:
otlp:
grpc_endpoint: 0.0.0.0:4317
http_endpoint: 0.0.0.0:4318
// ✅ Node tool service: OpenTelemetry SDK exporting to ADOT / CW agent
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";
import { Resource } from "@opentelemetry/resources";
import { SemanticResourceAttributes } from "@opentelemetry/semantic-conventions";
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: "tool-sandbox",
"workload": "coding-agent",
}),
traceExporter: new OTLPTraceExporter({ url: "http://cloudwatch-agent:4317" }),
});
sdk.start();
Propagate traceparent / AWS X-Ray headers across VPC Lattice hops and Step Functions task tokens where applicable so one session is one trace.
Define SLOs that match product promises
// ✅ conceptual SLO: 99% of sandbox tool invocations < 8s over 28 days
{
"service": "tool-sandbox",
"sloName": "sandbox-latency-p99-8s",
"goal": 99.0,
"periodDays": 28,
"indicator": {
"type": "LATENCY",
"thresholdMs": 8000,
"operation": "POST /run"
}
}
Practical agent SLOs:
- Session start availability — BFF 5xx < 0.1%
- Planner TTFT — first token < 2s at P95 (pairs with Bedrock PT when you publish that path)
- Sandbox success — non-infra failures excluded from fault budget
- Multi-hop budget — total graph time < 60s P95 for “quick review” mode
Trace a slow multi-hop review
When Application Signals flags a latency burn:
- Open the service operation → exemplars / traces
- Find spans:
bff → planner → tool-rag → tool-sandbox - Check dependency: EFS cold read? DynamoDB throttle? Bedrock retry loop?
- Confirm fail-closed behavior with Destinations / DLQs
# ✅ Logs Insights — correlate sessionId once you have the trace id
fields @timestamp, @message
| filter sessionId = "sess_01JEXAMPLE"
| sort @timestamp asc
| limit 100
Alarms: burn rate over raw error count
# ✅ alarm on elevated fault rate for tool-sandbox (wire to SLO burn in console)
aws cloudwatch put-metric-alarm \
--alarm-name coding-agent-sandbox-fault-rate \
--namespace AWS/ApplicationSignals \
--metric-name Fault \
--dimensions Name=Service,Value=tool-sandbox Name=Operation,Value="POST /run" \
--statistic Sum --period 60 --threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--evaluation-periods 3
Tune namespaces/dimensions to what Application Signals emits in your account — names evolve; prefer console-created alarms exported to IaC after the first green dashboard.
Production checklist
- [ ] Every tool microservice has a distinct
service.name - [ ] OTLP export path authenticated / private (no public collectors)
- [ ] SLOs defined for BFF, planner, top 3 tools
- [ ] Trace propagation across Lattice / HTTP / async boundaries documented
- [ ] Dashboards linked from on-call runbook
- [ ] Chaos tests (FIS) prove SLO burn alerts fire
- [ ] Cost: sampling rules so 100% of tiny heartbeats do not bankrupt traces
- [ ] Tag
workload=coding-agenton instrumented resources
FAQ
Q: Do we still need X-Ray console?
A: Application Signals sits on the same telemetry spine. Use Signals for SLO/service maps; drop to traces for forensics — do not maintain three competing dashboards.
Q: How does this differ from “just ADOT”?
A: ADOT is the pipeline. Application Signals is the product UX for services + SLOs on top of that telemetry.
Q: Async tool jobs via SQS/Pipes?
A: Instrument producers/consumers; use links / baggage for sessionId. Pure fire-and-forget without correlation IDs will look like random spikes.
Related reading
- AWS Fault Injection Service: Chaos-Test Coding-Agent Pipelines
- Step Functions: Orchestrate Multi-Step Coding-Agent Graphs
- Amazon EventBridge Pipes: Tool Runners
- Lambda Destinations: Failed Agent Tool Invokes
Instrument the hops, attach SLOs to the tools customers feel, and debug with traces — not with six engineers grepping six log groups.
Last updated on October 2, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
- Java Virtual Threads vs Traditional Threads: What Nobody Tells You
- Distributed Locks Reality Check: When Redis Redlock Is the Wrong Tool
Newly added
- Amazon Bedrock Custom Model Import: Host Fine-Tuned Coding Models Inside Your Account
- AWS Systems Manager Session Manager: Audited Break-Glass Shell into Coding-Agent Sandboxes
- Amazon S3 Express One Zone: Sub-Millisecond Scratch for Coding-Agent Tool Artifacts
- AWS AppSync GraphQL Subscriptions: Push Live Coding-Agent Progress Without Polling
- Amazon Neptune: Graph Memory for Code Dependency Reasoning in Coding Agents
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.