Amazon CloudWatch Application Signals: SLOs and Traces for Multi-Hop Coding-Agent Tools

2 views

When a coding-agent session feels “slow,” the blame ping-pongs between Bedrock, the tool Lambda, EFS, and “the frontend.” Amazon CloudWatch Application Signals (built on OpenTelemetry / CloudWatch agent instrumentation) discovers services, charts golden signals, and lets you attach SLOs so a sandbox P99 regression pages you before Twitter does. Use it alongside FIS chaos tests and Lambda Destinations — Signals tells you what breached; those tell you how you fail closed.

⚡ TL;DR: Instrument agent BFF + tool runtimes with the CloudWatch agent / ADOT; enable Application Signals; define SLOs on latency and availability per tool service; link traces for slow multi-hop sessions; alarm on burn rate, not only hard error counts. Related: Step Functions agent graphs, VPC Lattice tool auth, App Runner agent APIs, EventBridge Pipes runners.

What to model as “services”

Treat each hop as a service Application Signals can see:

  1. agent-bff — HTTP API that starts sessions
  2. planner — Bedrock-facing Lambda / container
  3. tool-lint, tool-sandbox, tool-rag — discrete tool microservices
  4. session-store — dependencies (DynamoDB/MemoryDB) as observed dependencies
Golden signal Agent meaning
Latency Time to first useful token / tool result
Faults 5xx, tool crashes, Bedrock 5xx mapped
Errors Business failures (lint failed) — keep separate from faults
Traffic Sessions started / tools invoked per minute

❌ One mega-service name coding-agent for everything — you lose which tool burned the SLO.

Enable Application Signals on the agent fleet

bash
# ✅ example: enable Application Signals in the account/region (console or APIs evolve — verify)
aws application-signals list-services --region us-east-1

# ✅ CloudWatch agent config snippet for host/container metrics + traces (OTLP)
yaml
# ✅ cwagent-config excerpt — traces + application signals oriented
traces:
  traces_collected:
    otlp:
      grpc_endpoint: 0.0.0.0:4317
      http_endpoint: 0.0.0.0:4318
typescript
// ✅ Node tool service: OpenTelemetry SDK exporting to ADOT / CW agent
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";
import { Resource } from "@opentelemetry/resources";
import { SemanticResourceAttributes } from "@opentelemetry/semantic-conventions";

const sdk = new NodeSDK({
  resource: new Resource({
    [SemanticResourceAttributes.SERVICE_NAME]: "tool-sandbox",
    "workload": "coding-agent",
  }),
  traceExporter: new OTLPTraceExporter({ url: "http://cloudwatch-agent:4317" }),
});
sdk.start();

Propagate traceparent / AWS X-Ray headers across VPC Lattice hops and Step Functions task tokens where applicable so one session is one trace.

Define SLOs that match product promises

json
// ✅ conceptual SLO: 99% of sandbox tool invocations < 8s over 28 days
{
  "service": "tool-sandbox",
  "sloName": "sandbox-latency-p99-8s",
  "goal": 99.0,
  "periodDays": 28,
  "indicator": {
    "type": "LATENCY",
    "thresholdMs": 8000,
    "operation": "POST /run"
  }
}

Practical agent SLOs:

  • Session start availability — BFF 5xx < 0.1%
  • Planner TTFT — first token < 2s at P95 (pairs with Bedrock PT when you publish that path)
  • Sandbox success — non-infra failures excluded from fault budget
  • Multi-hop budget — total graph time < 60s P95 for “quick review” mode

Trace a slow multi-hop review

When Application Signals flags a latency burn:

  1. Open the service operation → exemplars / traces
  2. Find spans: bff → planner → tool-rag → tool-sandbox
  3. Check dependency: EFS cold read? DynamoDB throttle? Bedrock retry loop?
  4. Confirm fail-closed behavior with Destinations / DLQs
bash
# ✅ Logs Insights — correlate sessionId once you have the trace id
fields @timestamp, @message
| filter sessionId = "sess_01JEXAMPLE"
| sort @timestamp asc
| limit 100

Alarms: burn rate over raw error count

bash
# ✅ alarm on elevated fault rate for tool-sandbox (wire to SLO burn in console)
aws cloudwatch put-metric-alarm \
  --alarm-name coding-agent-sandbox-fault-rate \
  --namespace AWS/ApplicationSignals \
  --metric-name Fault \
  --dimensions Name=Service,Value=tool-sandbox Name=Operation,Value="POST /run" \
  --statistic Sum --period 60 --threshold 5 \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --evaluation-periods 3

Tune namespaces/dimensions to what Application Signals emits in your account — names evolve; prefer console-created alarms exported to IaC after the first green dashboard.

Production checklist

  • [ ] Every tool microservice has a distinct service.name
  • [ ] OTLP export path authenticated / private (no public collectors)
  • [ ] SLOs defined for BFF, planner, top 3 tools
  • [ ] Trace propagation across Lattice / HTTP / async boundaries documented
  • [ ] Dashboards linked from on-call runbook
  • [ ] Chaos tests (FIS) prove SLO burn alerts fire
  • [ ] Cost: sampling rules so 100% of tiny heartbeats do not bankrupt traces
  • [ ] Tag workload=coding-agent on instrumented resources

FAQ

Q: Do we still need X-Ray console?
A: Application Signals sits on the same telemetry spine. Use Signals for SLO/service maps; drop to traces for forensics — do not maintain three competing dashboards.

Q: How does this differ from “just ADOT”?
A: ADOT is the pipeline. Application Signals is the product UX for services + SLOs on top of that telemetry.

Q: Async tool jobs via SQS/Pipes?
A: Instrument producers/consumers; use links / baggage for sessionId. Pure fire-and-forget without correlation IDs will look like random spikes.

Related reading

Instrument the hops, attach SLOs to the tools customers feel, and debug with traces — not with six engineers grepping six log groups.

Last updated on October 2, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.