AWS Lambda Response Streaming: Stream Coding-Agent Tokens Without Buffering Full Completions

2 views

If your coding-agent API waits for the entire Bedrock (or OpenAI-compatible) completion before writing HTTP bytes, users stare at a spinner while you pay for idle connection time — and a 40k-token patch review can hit Lambda’s 6 MB synchronous response wall. AWS Lambda response streaming lets Node.js and Python runtimes flush chunks via a Function URL (InvokeMode=RESPONSE_STREAM) or InvokeWithResponseStream, so first-token latency matches the model instead of “model + buffer + serialize.” Distinct from API Gateway WebSockets for multi-turn agents (bidirectional sessions) and App Runner HTTP APIs (long-lived containers): streaming is the sync request path that still looks like one HTTP response.

⚡ TL;DR: Enable response streaming on a Function URL (or use the invoke stream API), write an awslambda.streamifyResponse / LambdaResponseStream handler that pipes Bedrock token events to the stream, set generous HttpResponseTimeout on the URL, and cap max stream duration so runaway agents cannot hold sockets forever. Measure TTFB vs buffered baseline. Related: Lambda SnapStart, Recursive Loop Protection, Lambda Destinations, Bedrock Intelligent Prompt Routing.

Why buffering kills agent UX

Coding-agent responses are long and incremental:

  1. Planner emits a plan outline (first 200 tokens matter for trust)
  2. Tool progress events (lint_started, diff_ready) should paint in the UI mid-flight
  3. Final patches are large — buffering them doubles peak memory and risks 413/timeout
Path First byte Max payload Good for
Buffered Invoke / API GW proxy After full completion ~6 MB sync Short JSON tool results
Lambda response streaming (Function URL) As you write() Stream limits (see docs; much larger) Token + progress SSE-like chunks
WebSocket API Push anytime Session-shaped Multi-turn bidirectional chat
App Runner / ECS Framework streaming Container limits Always-on agent gateways

❌ Returning { "completion": "...." } after await bedrock.converse(...) finishes — users wait for the slowest token.

Enable streaming on the Function URL

bash
# ✅ create / update Function URL for RESPONSE_STREAM
aws lambda create-function-url-config \
  --function-name coding-agent-stream \
  --auth-type AWS_IAM \
  --invoke-mode RESPONSE_STREAM \
  --cors '{"AllowOrigins":["https://app.example.com"],"AllowMethods":["POST"],"AllowHeaders":["content-type","authorization"]}'

# ✅ raise the HTTP response timeout (seconds) so long reviews can finish
aws lambda update-function-url-config \
  --function-name coding-agent-stream \
  --invoke-mode RESPONSE_STREAM

Wire IAM so only your BFF / API Gateway can invoke. Public unauthenticated Function URLs for agents are an invitation to burn Bedrock quota — rate-limit at the edge with API Gateway + WAF if you front them.

Node.js: streamifyResponse + Bedrock tokens

javascript
// ✅ handler.js — stream tokens as NDJSON lines (easy for browsers / CLIs)
const { BedrockRuntimeClient, ConverseStreamCommand } = require("@aws-sdk/client-bedrock-runtime");
const client = new BedrockRuntimeClient({});

exports.handler = awslambda.streamifyResponse(async (event, responseStream, _context) => {
  const httpStream = awslambda.HttpResponseStream.from(responseStream, {
    statusCode: 200,
    headers: {
      "Content-Type": "application/x-ndjson",
      "Cache-Control": "no-cache",
    },
  });

  let body;
  try {
    body = JSON.parse(event.body || "{}");
  } catch {
    httpStream.write(JSON.stringify({ type: "error", message: "invalid_json" }) + "\n");
    httpStream.end();
    return;
  }

  const cmd = new ConverseStreamCommand({
    modelId: body.modelId || "anthropic.claude-3-5-sonnet-20241022-v2:0",
    messages: body.messages,
    inferenceConfig: { maxTokens: body.maxTokens || 4096, temperature: 0.2 },
  });

  try {
    const resp = await client.send(cmd);
    for await (const ev of resp.stream) {
      if (ev.contentBlockDelta?.delta?.text) {
        // ✅ flush early — do not accumulate
        httpStream.write(
          JSON.stringify({ type: "token", text: ev.contentBlockDelta.delta.text }) + "\n"
        );
      }
      if (ev.messageStop) {
        httpStream.write(JSON.stringify({ type: "stop", reason: ev.messageStop.stopReason }) + "\n");
      }
    }
  } catch (err) {
    // ❌ swallowing errors leaves the client hanging with an open stream
    httpStream.write(JSON.stringify({ type: "error", message: String(err.name || err) }) + "\n");
  } finally {
    httpStream.end();
  }
});
python
# ✅ Python: Lambda response streaming with a text IO wrapper
import json
from awslambdaric.lambda_response_stream import LambdaResponseStream  # runtime provides stream APIs

# Prefer the documented decorator / stream writer for your runtime version.
# Pattern: open stream → write chunks → close. Never buffer the full completion.

InvokeWithResponseStream from a trusted BFF

When the browser must not talk to Lambda directly, your Nest/Express BFF can pipe the stream:

typescript
// ✅ AWS SDK v3 — pipe InvokeWithResponseStream to the client
import {
  LambdaClient,
  InvokeWithResponseStreamCommand,
} from "@aws-sdk/client-lambda";

const lambda = new LambdaClient({});

export async function proxyAgentStream(res: NodeJS.WritableStream, payload: unknown) {
  const out = await lambda.send(
    new InvokeWithResponseStreamCommand({
      FunctionName: "coding-agent-stream",
      Payload: Buffer.from(JSON.stringify(payload)),
    })
  );
  for await (const event of out.EventStream!) {
    if (event.PayloadChunk?.Payload) {
      res.write(event.PayloadChunk.Payload);
    }
    if (event.InvokeComplete?.ErrorCode) {
      // surface failure; pair with Destinations for async paths
      throw new Error(event.InvokeComplete.ErrorCode);
    }
  }
}

Limits and failure modes agents hit

Issue Symptom Mitigation
Idle timeout Stream cuts mid-diff Heartbeat NDJSON every N seconds; raise URL timeout
Client disconnect Lambda keeps calling Bedrock Check context.getRemainingTimeInMillis(); abort Bedrock stream
Dual billing Tokens + Lambda-GB-seconds Cap maxTokens; kill switch via AppConfig
Recursive fan-out Stream handler invokes more Lambdas Recursive Loop Protection
Cold start before first token Slow TTFB SnapStart on Java; provisioned concurrency on Node if needed
bash
# ✅ alarm: function errors during streaming reviews
aws cloudwatch put-metric-alarm \
  --alarm-name coding-agent-stream-errors \
  --namespace AWS/Lambda \
  --metric-name Errors \
  --dimensions Name=FunctionName,Value=coding-agent-stream \
  --statistic Sum --period 60 --threshold 3 \
  --comparison-operator GreaterThanOrEqualToThreshold \
  --evaluation-periods 1

Production checklist

  • [ ] Function URL (or ALB/Lambda) uses RESPONSE_STREAM, not buffered default
  • [ ] Auth is IAM / JWT at BFF — not open NONE auth on a Bedrock-backed function
  • [ ] Handler writes tokens incrementally; no full-string concatenate before end()
  • [ ] Content-Type is application/x-ndjson or text/event-stream with a documented client parser
  • [ ] Max tokens + AppConfig kill switch bound runaway cost
  • [ ] SnapStart / concurrency tuned so TTFB is model-bound, not init-bound
  • [ ] Errors written as a final NDJSON error event; metrics + Destinations for async siblings
  • [ ] Load test: 100 concurrent streams; watch Bedrock throttle and Lambda concurrency
  • [ ] Document ownership tag workload=coding-agent

FAQ

Q: Can API Gateway REST/HTTP APIs do Lambda response streaming?
A: Function URLs and InvokeWithResponseStream are the first-class paths. For API Gateway, many teams stream from a container (App Runner/ECS) or terminate WebSockets instead — verify current API GW support for your region before betting the design.

Q: Is this a replacement for WebSockets?
A: No. Streaming is one response, many chunks. WebSockets win for multi-turn push, cancel, and server-initiated tool events after the HTTP request ended.

Q: How do tool calls fit?
A: Emit {type:"tool_call",...} NDJSON lines, run the tool in-process or via Step Functions (agent graphs), then continue token streaming — or switch the long graph off-stream and push progress on a bus.

Related reading

Stream the tokens, bound the cost, and stop making senior engineers wait for a buffer that only exists because the handler returned a string.

Last updated on October 2, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.