If your coding-agent API waits for the entire Bedrock (or OpenAI-compatible) completion before writing HTTP bytes, users stare at a spinner while you pay for idle connection time — and a 40k-token patch review can hit Lambda’s 6 MB synchronous response wall. AWS Lambda response streaming lets Node.js and Python runtimes flush chunks via a Function URL (InvokeMode=RESPONSE_STREAM) or InvokeWithResponseStream, so first-token latency matches the model instead of “model + buffer + serialize.” Distinct from API Gateway WebSockets for multi-turn agents (bidirectional sessions) and App Runner HTTP APIs (long-lived containers): streaming is the sync request path that still looks like one HTTP response.
⚡ TL;DR: Enable response streaming on a Function URL (or use the invoke stream API), write an
awslambda.streamifyResponse/LambdaResponseStreamhandler that pipes Bedrock token events to the stream, set generousHttpResponseTimeouton the URL, and cap max stream duration so runaway agents cannot hold sockets forever. Measure TTFB vs buffered baseline. Related: Lambda SnapStart, Recursive Loop Protection, Lambda Destinations, Bedrock Intelligent Prompt Routing.
Why buffering kills agent UX
Coding-agent responses are long and incremental:
- Planner emits a plan outline (first 200 tokens matter for trust)
- Tool progress events (
lint_started,diff_ready) should paint in the UI mid-flight - Final patches are large — buffering them doubles peak memory and risks 413/timeout
| Path | First byte | Max payload | Good for |
|---|---|---|---|
Buffered Invoke / API GW proxy |
After full completion | ~6 MB sync | Short JSON tool results |
| Lambda response streaming (Function URL) | As you write() |
Stream limits (see docs; much larger) | Token + progress SSE-like chunks |
| WebSocket API | Push anytime | Session-shaped | Multi-turn bidirectional chat |
| App Runner / ECS | Framework streaming | Container limits | Always-on agent gateways |
❌ Returning { "completion": "...." } after await bedrock.converse(...) finishes — users wait for the slowest token.
Enable streaming on the Function URL
# ✅ create / update Function URL for RESPONSE_STREAM
aws lambda create-function-url-config \
--function-name coding-agent-stream \
--auth-type AWS_IAM \
--invoke-mode RESPONSE_STREAM \
--cors '{"AllowOrigins":["https://app.example.com"],"AllowMethods":["POST"],"AllowHeaders":["content-type","authorization"]}'
# ✅ raise the HTTP response timeout (seconds) so long reviews can finish
aws lambda update-function-url-config \
--function-name coding-agent-stream \
--invoke-mode RESPONSE_STREAM
Wire IAM so only your BFF / API Gateway can invoke. Public unauthenticated Function URLs for agents are an invitation to burn Bedrock quota — rate-limit at the edge with API Gateway + WAF if you front them.
Node.js: streamifyResponse + Bedrock tokens
// ✅ handler.js — stream tokens as NDJSON lines (easy for browsers / CLIs)
const { BedrockRuntimeClient, ConverseStreamCommand } = require("@aws-sdk/client-bedrock-runtime");
const client = new BedrockRuntimeClient({});
exports.handler = awslambda.streamifyResponse(async (event, responseStream, _context) => {
const httpStream = awslambda.HttpResponseStream.from(responseStream, {
statusCode: 200,
headers: {
"Content-Type": "application/x-ndjson",
"Cache-Control": "no-cache",
},
});
let body;
try {
body = JSON.parse(event.body || "{}");
} catch {
httpStream.write(JSON.stringify({ type: "error", message: "invalid_json" }) + "\n");
httpStream.end();
return;
}
const cmd = new ConverseStreamCommand({
modelId: body.modelId || "anthropic.claude-3-5-sonnet-20241022-v2:0",
messages: body.messages,
inferenceConfig: { maxTokens: body.maxTokens || 4096, temperature: 0.2 },
});
try {
const resp = await client.send(cmd);
for await (const ev of resp.stream) {
if (ev.contentBlockDelta?.delta?.text) {
// ✅ flush early — do not accumulate
httpStream.write(
JSON.stringify({ type: "token", text: ev.contentBlockDelta.delta.text }) + "\n"
);
}
if (ev.messageStop) {
httpStream.write(JSON.stringify({ type: "stop", reason: ev.messageStop.stopReason }) + "\n");
}
}
} catch (err) {
// ❌ swallowing errors leaves the client hanging with an open stream
httpStream.write(JSON.stringify({ type: "error", message: String(err.name || err) }) + "\n");
} finally {
httpStream.end();
}
});
# ✅ Python: Lambda response streaming with a text IO wrapper
import json
from awslambdaric.lambda_response_stream import LambdaResponseStream # runtime provides stream APIs
# Prefer the documented decorator / stream writer for your runtime version.
# Pattern: open stream → write chunks → close. Never buffer the full completion.
InvokeWithResponseStream from a trusted BFF
When the browser must not talk to Lambda directly, your Nest/Express BFF can pipe the stream:
// ✅ AWS SDK v3 — pipe InvokeWithResponseStream to the client
import {
LambdaClient,
InvokeWithResponseStreamCommand,
} from "@aws-sdk/client-lambda";
const lambda = new LambdaClient({});
export async function proxyAgentStream(res: NodeJS.WritableStream, payload: unknown) {
const out = await lambda.send(
new InvokeWithResponseStreamCommand({
FunctionName: "coding-agent-stream",
Payload: Buffer.from(JSON.stringify(payload)),
})
);
for await (const event of out.EventStream!) {
if (event.PayloadChunk?.Payload) {
res.write(event.PayloadChunk.Payload);
}
if (event.InvokeComplete?.ErrorCode) {
// surface failure; pair with Destinations for async paths
throw new Error(event.InvokeComplete.ErrorCode);
}
}
}
Limits and failure modes agents hit
| Issue | Symptom | Mitigation |
|---|---|---|
| Idle timeout | Stream cuts mid-diff | Heartbeat NDJSON every N seconds; raise URL timeout |
| Client disconnect | Lambda keeps calling Bedrock | Check context.getRemainingTimeInMillis(); abort Bedrock stream |
| Dual billing | Tokens + Lambda-GB-seconds | Cap maxTokens; kill switch via AppConfig |
| Recursive fan-out | Stream handler invokes more Lambdas | Recursive Loop Protection |
| Cold start before first token | Slow TTFB | SnapStart on Java; provisioned concurrency on Node if needed |
# ✅ alarm: function errors during streaming reviews
aws cloudwatch put-metric-alarm \
--alarm-name coding-agent-stream-errors \
--namespace AWS/Lambda \
--metric-name Errors \
--dimensions Name=FunctionName,Value=coding-agent-stream \
--statistic Sum --period 60 --threshold 3 \
--comparison-operator GreaterThanOrEqualToThreshold \
--evaluation-periods 1
Production checklist
- [ ] Function URL (or ALB/Lambda) uses
RESPONSE_STREAM, not buffered default - [ ] Auth is IAM / JWT at BFF — not open
NONEauth on a Bedrock-backed function - [ ] Handler writes tokens incrementally; no full-string concatenate before
end() - [ ] Content-Type is
application/x-ndjsonortext/event-streamwith a documented client parser - [ ] Max tokens + AppConfig kill switch bound runaway cost
- [ ] SnapStart / concurrency tuned so TTFB is model-bound, not init-bound
- [ ] Errors written as a final NDJSON error event; metrics + Destinations for async siblings
- [ ] Load test: 100 concurrent streams; watch Bedrock throttle and Lambda concurrency
- [ ] Document ownership tag
workload=coding-agent
FAQ
Q: Can API Gateway REST/HTTP APIs do Lambda response streaming?
A: Function URLs and InvokeWithResponseStream are the first-class paths. For API Gateway, many teams stream from a container (App Runner/ECS) or terminate WebSockets instead — verify current API GW support for your region before betting the design.
Q: Is this a replacement for WebSockets?
A: No. Streaming is one response, many chunks. WebSockets win for multi-turn push, cancel, and server-initiated tool events after the HTTP request ended.
Q: How do tool calls fit?
A: Emit {type:"tool_call",...} NDJSON lines, run the tool in-process or via Step Functions (agent graphs), then continue token streaming — or switch the long graph off-stream and push progress on a bus.
Related reading
- AWS Lambda SnapStart: Cut Coding-Agent Cold Starts
- Lambda Destinations: Route Failed Agent Tool Invokes
- Amazon Bedrock Intelligent Prompt Routing
- API Gateway + WAF: Rate-Limit Public Coding-Agent Endpoints
Stream the tokens, bound the cost, and stop making senior engineers wait for a buffer that only exists because the handler returned a string.
Last updated on October 2, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
- Java Virtual Threads vs Traditional Threads: What Nobody Tells You
- Agentic Git Workflows: Atomic Commits From Noisy LLM Diffs
Newly added
- AWS IAM Access Analyzer: Find Over-Privileged Coding-Agent Roles Before They Leak
- Amazon CloudWatch Application Signals: SLOs and Traces for Multi-Hop Coding-Agent Tools
- AWS Config Conformance Packs: Continuous Compliance Guards for Coding-Agent Accounts
- Amazon Bedrock Provisioned Throughput: Reserved Capacity for Coding-Agent Latency SLOs
- AWS Lambda Response Streaming: Stream Coding-Agent Tokens Without Buffering Full Completions
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.