Amazon Bedrock Prompt Management: Version and A/B Test Coding-Agent System Prompts

2 views

Your coding agent’s “personality” is a 2k-token system prompt someone edited at 11pm. It lives in a Lambda env var, a YAML in the monorepo, and three Slack threads — none of which the runtime can pin, roll back, or A/B. Amazon Bedrock Prompt Management treats system and tool prompts as versioned assets with variants you invoke by identifier — so a bad edit is a version pin change, not an emergency redeploy. Distinct from Prompt Caching pitfalls (reuse cached prefix tokens), Intelligent Prompt Routing (auto-pick a model), Model Evaluation (score outputs), and Custom Model Import (bring weights): here we version the prompt text coding agents use for system + tool instructions.

⚡ TL;DR: Create a Prompt Management prompt for each coding-agent role (planner, patcher, reviewer), publish versions, attach A/B variants for experiments, invoke by prompt ARN + version from Converse, and gate promotions with eval harnesses. Related: Intelligent Prompt Routing, Model Evaluation, Provisioned Throughput, Verified Permissions.

Why git-only prompts fail agents in production

Coding-agent prompt failure modes that scream “you needed Prompt Management”:

  1. Silent drift — staging and prod run different system prompts; nobody noticed until tool-call JSON broke
  2. Rollback theater — “revert the commit” still leaves warm Lambdas / Step Functions with old env vars for minutes
  3. A/B by hope — you flip a feature flag that swaps two strings in code; metrics never tag which prompt version won
  4. Tool-spec sprawl — system prompt says “use run_tests” while toolConfig renamed it to pytest_run last week
Approach Strength Weak for
Hardcoded string in code Simple PR review Runtime version pin / A/B
SSM / Secrets string Centralized No first-class variants / Bedrock integration
Prompt Caching Cheaper repeated prefixes Does not version or A/B the text
Prompt Management Versioned assets + variants + invoke by ARN Not a substitute for eval harnesses

❌ Editing the system prompt in a Notion doc and asking ops to “paste it into Lambda when they get a chance.”

Create a versioned coding-agent system prompt

Model each agent role as its own prompt asset. Keep variables for repo, language, and policy packs — not for the entire instruction dump.

bash
# ✅ create prompt asset (illustrative CLI shape — use console/API for your region)
aws bedrock-agent create-prompt \
  --name coding-agent-patcher-system \
  --description "System prompt for patcher role: edit files, run tests, never force-push" \
  --default-variant default \
  --variants '[
    {
      "name": "default",
      "templateType": "TEXT",
      "templateConfiguration": {
        "text": {
          "text": "You are a coding-agent patcher for {{repo}}. Languages: {{langs}}. Rules: (1) never force-push (2) run tests before claiming done (3) output tool calls only via provided tools. Banned APIs: {{banned}}.",
          "inputVariables": [
            {"name": "repo"},
            {"name": "langs"},
            {"name": "banned"}
          ]
        }
      },
      "modelId": "anthropic.claude-sonnet-4-20250514-v1:0"
    }
  ]'

Publish an immutable version after review:

bash
# ✅ createPromptVersion after peer review + smoke eval
aws bedrock-agent create-prompt-version \
  --prompt-identifier arn:aws:bedrock:us-east-1:111122223333:prompt/ABCDEF \
  --description "v3: add banned APIs + require pytest before done"

Pin version 3 in prod config (SSM Parameter Store path /prod/coding-agent/patcher/prompt-version). Staging can float to DRAFT or a newer version.

Invoke from Converse / agent runtime by prompt identifier

Wire the agent host to resolve the prompt ARN + version, fill variables, then call Converse (or your Bedrock Agents action group). Do not paste the full system text into every request from app code.

python
# ✅ resolve versioned prompt, then Converse with toolConfig
import boto3, json

bedrock = boto3.client("bedrock-runtime")
agent = boto3.client("bedrock-agent")

PROMPT_ID = "ABCDEF"
PROMPT_VERSION = "3"  # from SSM — not hardcoded forever

def render_system(repo: str, langs: str, banned: str) -> str:
    # GetPrompt / GetPromptVersion returns template; fill variables server-side
    # (API names vary by SDK version — keep a thin adapter)
    p = agent.get_prompt(promptIdentifier=PROMPT_ID, promptVersion=PROMPT_VERSION)
    text = p["variants"][0]["templateConfiguration"]["text"]["text"]
    return (
        text.replace("{{repo}}", repo)
        .replace("{{langs}}", langs)
        .replace("{{banned}}", banned)
    )

TOOL_CONFIG = {
    "tools": [
        {
            "toolSpec": {
                "name": "apply_patch",
                "description": "Apply a unified diff to the workspace",
                "inputSchema": {"json": {"type": "object", "properties": {"diff": {"type": "string"}}, "required": ["diff"]}},
            }
        },
        {
            "toolSpec": {
                "name": "pytest_run",
                "description": "Run pytest subset; required before claiming done",
                "inputSchema": {"json": {"type": "object", "properties": {"path": {"type": "string"}}, "required": ["path"]}},
            }
        },
    ]
}

def patch_turn(user_msg: str, repo: str):
    system = render_system(repo, "python,typescript", "subprocess.call,os.system")
    # ❌ stuffing a different ad-hoc system string per request
    return bedrock.converse(
        modelId="anthropic.claude-sonnet-4-20250514-v1:0",
        system=[{"text": system}],
        messages=[{"role": "user", "content": [{"text": user_msg}]}],
        toolConfig=TOOL_CONFIG,
    )

Pair with Verified Permissions so tool names in the prompt match Cedar-allowed actions. Trace which prompt version ran with ADOT attributes prompt.id / prompt.version.

A/B variants without forking the agent binary

Prompt Management variants let you ship two instruction styles (terse vs checklist-heavy) under one prompt identifier. Route traffic in your orchestrator — not by duplicating Lambdas.

python
# ✅ sticky A/B: hash session_id → variant; log for Model Evaluation join
import hashlib

def pick_variant(session_id: str) -> str:
    h = int(hashlib.sha256(session_id.encode()).hexdigest(), 16)
    return "terse" if h % 100 < 20 else "default"  # 20% terse

def system_for_session(session_id: str, repo: str) -> tuple[str, str]:
    variant = pick_variant(session_id)
    # load variant text from GetPrompt; fill vars
    text = render_system_variant(variant, repo)
    return text, variant
Experiment Control Treatment Success metric
Tool reminder style default checklist terse “call pytest_run” % turns with pytest before “done”
Diff format unified diff only structured JSON patch apply_patch success rate
Ban list verbosity short list list + examples policy-violation rate

Feed outcomes into Model Evaluation datasets tagged with prompt_version + variant. Promote winners by creating a new prompt version that adopts the winning variant text — then pin prod to that version.

Promotion pipeline (don’t hot-edit prod)

PR reviews prompt text in repo (source of truth for humans)
  → CI creates/updates DRAFT in Prompt Management
  → smoke eval suite (golden PRs / tool-call schema checks)
  → createPromptVersion
  → staging pin floats
  → canary 5% sessions on new version
  → prod pin via SSM/AppConfig

Use AppConfig kill switches to flip the pin back to the previous version in seconds if tool-call quality collapses — without rewriting Prompt Management history.

❌ Publishing a new prompt version and pointing 100% of prod at it with no canary and no eval gate.

Prompt Management vs the Bedrock neighbors

Capability Job Not for
Prompt Management Version + variant prompt assets Caching token prefixes
Prompt Caching Cheaper repeated prefixes A/B lifecycle
Intelligent Prompt Routing Pick model by prompt Editing system text
Model Evaluation Score outputs / compare models Storing the prompt
Custom Model Import Host fine-tuned weights Prompt strings
Provisioned Throughput Reserved model capacity Prompt versioning

Compose them: Prompt Management pins the text → Routing picks the model → Caching saves tokens on the stable prefix → Evaluation proves the change.

Failure modes and ops

Issue Symptom Mitigation
Variable mismatch {{banned}} left literal in model input Schema test in CI; fail deploy if unresolved
Version skew Two Step Functions map states use different pins Single SSM path per role; deny inline overrides
Variant leak Metrics untagged Always emit prompt.version + variant spans
Over-long system Context bloat Keep role prompts lean; put house style in Knowledge Bases
Stale DRAFT in prod Unreviewed text live IAM deny GetPrompt without version in prod roles
bash
# ✅ IAM: prod agent role may GetPrompt only with version resource pattern
# arn:aws:bedrock:region:acct:prompt/ID:3  — not DRAFT

Cap spend when experiments multiply with Budgets + Cost Anomaly.

Production checklist

  • [ ] One Prompt Management asset per agent role (planner / patcher / reviewer)
  • [ ] Variables for repo/langs/policy — not the whole instruction novel
  • [ ] Immutable versions; prod pins version number via SSM/AppConfig
  • [ ] A/B variants sticky per session; metrics tagged
  • [ ] Eval gate before createPromptVersion; canary before 100%
  • [ ] Tool names in prompt match toolConfig + Cedar policies
  • [ ] Trace prompt.id / prompt.version / variant on every turn
  • [ ] Document distinct use of Caching / Routing / Evaluation / Import
  • [ ] Kill-switch pin rollback tested quarterly

FAQ

Q: Can Prompt Management replace our prompt git repo?
A: Keep git as the human review surface; treat Prompt Management as the runtime catalog. CI syncs reviewed text → DRAFT → version. Don’t invent prompts only in the console.

Q: Prompt Management vs Agents “prompts” in Bedrock Agents?
A: Bedrock Agents have their own instruction fields. For custom Converse/tool loops (most coding-agent platforms), Prompt Management is the portable versioned store. Align names so humans aren’t hunting two sources of truth.

Q: Do variants work with Prompt Caching?
A: Yes — but each variant is a different prefix. Expect separate cache warmups; don’t A/B five variants at 1% each or you pay cold-prefix tax everywhere.

Related reading

Version the prompt. Pin the version. A/B the variant. Let the model write code — not invent which instructions it was supposed to follow.

Last updated on October 4, 2026

Deep-dive PDF

Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.

Already subscribed? or open the subscribe page.


Discover more from CheatCoders

Subscribe to get the latest posts sent to your email.

Comments

No comments yet. Why don’t you start the discussion?

Leave a comment

No account needed. Name and email are optional.