Self-hosting a fine-tuned CodeLlama / Mistral / Llama variant means CUDA drivers, autoscaling guesswork, and a pager when the H100 pool evaporates. Amazon Bedrock Custom Model Import (where available for your architecture) lets you import compatible model weights into Bedrock so coding agents call them through familiar InvokeModel / Converse surfaces — with IAM, PrivateLink, and Guardrails — instead of babysitting inference servers. Distinct from Provisioned Throughput (reserved capacity on Bedrock-hosted models), Model Evaluation (scoring outputs), Intelligent Prompt Routing (picking among models), and Prompt Management (versioned prompts): Custom Model Import is about your weights, Bedrock’s inference path.
⚡ TL;DR: Package weights in a supported format, upload to S3, create a Bedrock model import job, get a model ARN, point agent
modelIdat it, evaluate before promote, then optionally buy Provisioned Throughput on the imported model for latency SLOs. Related: PrivateLink for Bedrock, ApplyGuardrail, Budgets + Cost Anomaly.
When import beats self-host (and when it does not)
| Situation | Prefer |
|---|---|
| Custom weights, want IAM + Converse + Guardrails | Custom Model Import |
| Only need lower latency on Anthropic/Amazon foundation models | Provisioned Throughput |
| Choosing cheap vs smart model per prompt | Intelligent Prompt Routing |
| Measuring which checkpoint wins on your eval set | Model Evaluation |
| Exotic architectures / unsupported formats | Self-host on SageMaker / EKS GPUs |
❌ Assuming every Hugging Face checkpoint imports cleanly — check Bedrock’s supported architectures and formats for your region before promising leadership a date.
Package and import
# ✅ weights in supported layout (example: SafeTensors / model artifacts per docs)
# Upload to an encrypted bucket in the same region as Bedrock import
aws s3 sync ./export/coding-ft-v3/ s3://bedrock-model-imports/coding-ft-v3/ \
--sse aws:kms --sse-kms-key-id arn:aws:kms:us-east-1:123456789012:key/YOUR_KEY
# Create import job (CLI shape — verify flag names for your CLI version)
aws bedrock create-model-import-job \
--job-name coding-ft-v3-$(date +%Y%m%d) \
--imported-model-name coding-ft-v3 \
--role-arn arn:aws:iam::123456789012:role/BedrockModelImportRole \
--model-data-source '{
"s3DataSource": {
"s3Uri": "s3://bedrock-model-imports/coding-ft-v3/"
}
}'
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:ListBucket"],
"Resource": [
"arn:aws:s3:::bedrock-model-imports",
"arn:aws:s3:::bedrock-model-imports/*"
]
},
{
"Effect": "Allow",
"Action": ["kms:Decrypt"],
"Resource": "arn:aws:kms:us-east-1:123456789012:key/YOUR_KEY"
}
]
}
Trust policy must allow bedrock.amazonaws.com to assume the import role. Tag the S3 prefix data-class=model-weights and deny public ACLs.
Point coding agents at the imported model
# ✅ same Converse path your agent already uses — swap modelId
import boto3, json
bedrock = boto3.client("bedrock-runtime")
IMPORTED = "arn:aws:bedrock:us-east-1:123456789012:imported-model/coding-ft-v3"
def agent_turn(messages, tools):
resp = bedrock.converse(
modelId=IMPORTED, # ❌ hardcoding a foundation model ID after you paid for fine-tune
messages=messages,
toolConfig={"tools": tools} if tools else None,
inferenceConfig={"maxTokens": 4096, "temperature": 0.2},
)
return resp
Gate traffic with AppConfig kill switches: model_route=imported_v3|foundation_fallback. Keep a foundation-model fallback when import jobs fail or the custom model throttles.
Evaluate before you bet production
# ✅ run Model Evaluation jobs against your coding golden set before cutover
# Compare imported vs Claude/Sonnet (or your prior baseline) on:
# - compile success rate
# - unit tests passed
# - security lint regressions
# - latency P95
Use Bedrock Model Evaluation for automated scoring, then a human review sample for “house style” that metrics miss. Only then wire Intelligent Prompt Routing to include the imported ARN as a candidate — routing without eval is how you ship a polite regressor.
Latency and capacity after import
Import ≠ infinite capacity. For standup demos and CI review bursts:
- Import model → validate quality
- Optionally purchase Provisioned Throughput on the imported model (guide)
- Cap spend with Budgets + Anomaly Detection
- Private inference path via PrivateLink
# ✅ Guardrails still apply — filter tool I/O around custom models too
from botocore.client import BaseClient
def guarded_converse(client: BaseClient, **kwargs):
# ApplyGuardrail pre/post — see dedicated post
return client.converse(**kwargs)
Failure modes
| Issue | Symptom | Mitigation |
|---|---|---|
| Unsupported arch | Import job fails | Check docs; convert/export to supported format; SageMaker fallback |
| Silent quality drop | Agent “works” but tests fail | Eval harness gates; canary % traffic |
| Cost surprise | Bill spike after cutover | Budgets; per-tenant token metrics (EMF) |
| Prompt mismatch | Imported model ignores old system prompt | Revisit Prompt Management versions for the new tokenizer/chat template |
| Region skew | Weights in us-west-2, agents in us-east-1 | Import in the agent region; do not cross-region every token |
Production checklist
- [ ] Confirm architecture/format support in your Bedrock region
- [ ] Encrypted S3 source; import role least privilege
- [ ] Model ARN wired via config flag — not hardcoded in five Lambdas
- [ ] Eval + canary before 100% traffic
- [ ] Guardrails / ApplyGuardrail on tool I/O
- [ ] PrivateLink for egress-locked sandboxes
- [ ] Provisioned Throughput if P95 SLOs demand it
- [ ] Budgets + anomaly alerts on the imported model
- [ ] Fallback foundation model documented and tested
- [ ] Ownership tags
workload=coding-agent,model=imported-ft-vN
FAQ
Q: Is Custom Model Import the same as Bedrock fine-tuning jobs?
A: No. Fine-tuning jobs create adapted models inside Bedrock’s training flow. Custom Model Import brings weights you already trained elsewhere (when formats match). Many teams fine-tune outside, then import.
Q: Can I use Continual Pretraining / LoRA adapters?
A: Only if the resulting artifacts match what Import accepts. Merge adapters into a supported full/compatible checkpoint when required — do not assume raw LoRA folders import.
Q: Do I still need Prompt Management?
A: Yes. Imported models often need different chat templates and stop sequences. Version prompts per modelId or you will debug “the model got dumber” when you only changed routing.
Related reading
- Amazon Bedrock Provisioned Throughput: Reserved Capacity
- Amazon Bedrock Intelligent Prompt Routing
- Amazon Bedrock Model Evaluation: Score Coding-Agent Outputs
- Amazon Bedrock Prompt Management: Versioned System Prompts
Import the weights. Keep the agent API. Stop renting a second career as a GPU SRE unless you truly need to.
Last updated on October 3, 2026
Most viewed
- Python Decorators Explained: From Simple Wrappers to Production Patterns
- AI Agent Frameworks in 2025: LangGraph vs CrewAI vs AutoGen vs Raw API
- REST API Design Best Practices: The Patterns That Make APIs a Joy to Use
- Java Virtual Threads vs Traditional Threads: What Nobody Tells You
- Distributed Locks Reality Check: When Redis Redlock Is the Wrong Tool
Newly added
- Amazon Bedrock Custom Model Import: Host Fine-Tuned Coding Models Inside Your Account
- AWS Systems Manager Session Manager: Audited Break-Glass Shell into Coding-Agent Sandboxes
- Amazon S3 Express One Zone: Sub-Millisecond Scratch for Coding-Agent Tool Artifacts
- AWS AppSync GraphQL Subscriptions: Push Live Coding-Agent Progress Without Polling
- Amazon Neptune: Graph Memory for Code Dependency Reasoning in Coding Agents
Deep-dive PDF
Get the expanded guide for this post — extra diagrams-style checklists, failure modes, and a production walkthrough. Free when you subscribe to CheatCoders.
Already subscribed? or open the subscribe page.
Discover more from CheatCoders
Subscribe to get the latest posts sent to your email.