The AI Architecture Bootcamp is a free, hands-on course for software engineers who want to build AI systems that hold up in production, mostly on AWS. It is organized as 10 modules, each with a project. Every module is one guide with tested code, diagrams, troubleshooting and an FAQ, and every claim is checked against the official documentation.
How to use this bootcamp
- Work through the modules in order; each builds on the previous one.
- Run the code as you read. Examples run locally; AWS calls are marked, and their request shapes are checked against the current SDK.
- Finish each module with its project, then use the troubleshooting table when something breaks in your own system.
Available modules
Module 1: LLM Foundations for Coding Systems
Days 1 to 10. A documentation assistant that answers from your runbooks, cites what it used, refuses without evidence, and has an eval gate.
Covers: Tokens and context budgets, the KV cache, embeddings, vector indexes, a debuggable RAG pipeline, chunking, prompt contracts, structured outputs, evals.
Prerequisites: Python and basic AWS familiarity. Reading time: about 30 minutes, plus time to run the code.
Module 2: Agents, Tools and Production Guardrails
Days 11 to 20. An agent loop with typed tools, budgets, memory tiers, guardrails and approval gates, ending in a diff-scoped PR review bot.
Covers: Tool contracts, planning patterns, memory, Bedrock Guardrails, token cost curves, honest streaming, human approvals, safe logging, online evaluation.
Prerequisites: Module 1. Reading time: about 32 minutes, plus time to run the code.
Module 3: Production RAG for Code
Days 21 to 30. Multi-tenant code search on Bedrock Knowledge Bases with hybrid retrieval, reranking, tenant isolation and checked citations.
Covers: Hybrid search with RRF, rerankers, metadata filters, safe query rewriting, call-graph expansion, citations, re-indexing on push, multimodal inputs, a failure taxonomy.
Prerequisites: Modules 1 and 2. Reading time: about 32 minutes, plus time to run the code.
Coming soon
These modules are in preparation and will be linked here when they are published.
- Module 4: Serving AI on AWS: Lambda, Bedrock, Step Functions, EventBridge (Days 31 to 40)
- Module 5: Fine-Tuning and Model Customization: When It Beats Prompting (Days 41 to 50)
- Module 6: Multi-Agent Systems Without Chaos (Days 51 to 60)
- Module 7: Security for AI Coding Agents (Days 61 to 70)
- Module 8: Cost, Latency and Scale for AI Workloads (Days 71 to 80)
- Module 9: Memory, Multimodal and Domain Copilots (Days 81 to 90)
- Module 10: AI Architecture Leadership and Capstone (Days 91 to 100)
Deep dives
Standalone guides that go further on topics the modules touch:
- Bedrock Knowledge Bases: RAG over a monorepo
- Bedrock Agents: tool use, memory and guardrails
- Bedrock Prompt Management for system prompts
- Bedrock prompt caching and batch inference
- Lambda sandboxes for coding-agent tools
- PrivateLink for Bedrock
- Secrets Manager rotation for agent tool credentials
- Serverless API on AWS: from first endpoint to production
Lessons for upcoming modules
Until each module above is published, its individual day lessons stay available here.
Module 4: Serving AI on AWS: Lambda, Bedrock, Step Functions, EventBridge
- Day 31: Lambda + Bedrock: Sync, Stream, and Batch
- Day 32: Bedrock Converse API Tool Choice in Production
- Day 33: Agents on Step Functions, Not Infinite Loops
- Day 34: EventBridge as the Agent's Async Backbone
- Day 35: Idempotent Tool Calls Against DynamoDB
- Day 36: VPC Endpoints and Secret Hygiene for AI
- Day 37: Provisioned Throughput vs On-Demand Cliffs
- Day 38: Multi-Region Model Failover
- Day 39: Observability: Traces Across Prompt, Retrieval, Tools
- Day 40: Project: On-Call Copilot That Cannot Mutate Prod
Module 5: Fine-Tuning and Model Customization: When It Beats Prompting
- Day 41: When Fine-Tunes Beat Prompting
- Day 42: Preference Data Without Vanity Accept Rates
- Day 43: Distillation for Edge and Lambda
- Day 44: Continued Pretrain vs Instruct Tune vs Adapters
- Day 45: Catastrophic Forgetting Checks
- Day 46: Synthetic Data That Doesn't Clone GitHub Noise
- Day 47: Serving Custom Models Behind the Same API
- Day 48: License and Training-Data Risk
- Day 49: Eval for Fine-Tunes: Blind A/B on Real Tickets
- Day 50: Project: Internal SDK Autocomplete Model
Module 6: Multi-Agent Systems Without Chaos
- Day 51: Multi-Agent Roles: Planner, Implementer, Critic
- Day 52: Blackboard vs Message Bus Orchestration
- Day 53: Avoiding Agent Ping-Pong
- Day 54: Specialist Tools per Agent
- Day 55: Swarm Failure Modes
- Day 56: Supervisor Patterns on Bedrock Agents
- Day 57: Coding Agents That Only Emit Intents
- Day 58: Spec-First: OpenAPI Remains Source of Truth
- Day 59: Multi-Agent Eval: Task Suites, Not Vibes
- Day 60: Project: Migration Strangler Assistant
Module 7: Security for AI Coding Agents
- Day 61: Prompt Injection in Jira, Email, and RAG
- Day 62: SSRF and Tool Allowlists
- Day 63: IAM for Agents: Roles, Not God Keys
- Day 64: Secret-Aware Context Filters
- Day 65: PII Redaction Before Embeddings
- Day 66: Audit Trails Humans Can Replay
- Day 67: Supply Chain: Model and Image Provenance
- Day 68: Abuse and Cost Attacks
- Day 69: Policy as Code for Shell and Apply
- Day 70: Project: Secure Coding-Agent Sandbox on ECS Fargate
Module 8: Cost, Latency and Scale for AI Workloads
- Day 71: Token Budgets per PR and per Engineer
- Day 72: Prompt Caching Pitfalls
- Day 73: Draft-Then-Verify Routing
- Day 74: Batch Inference Overnight
- Day 75: Autoscaling Retrieval and Embed Jobs
- Day 76: p99 of Agents: Queueing, Not Just Model Latency
- Day 77: Caching RAG Answers Safely
- Day 78: Load Shedding When the Model Is Sick
- Day 79: FinOps Dashboards Engineers Open
- Day 80: Project: Platform AI Gateway
Module 9: Memory, Multimodal and Domain Copilots
- Day 81: Long-Term Memory Without Becoming Creepy
- Day 82: Voice and Realtime Agents
- Day 83: Multimodal Coding: Screenshots of Broken UIs
- Day 84: SQL and Warehouse Copilots
- Day 85: Time-Series and Anomaly Copilots
- Day 86: Search and Ranking Copilots
- Day 87: Workflow Mining From Logs
- Day 88: Simulation and Digital Twins for Tools
- Day 89: On-Device and Edge Completions
- Day 90: Project: Incident Timeline Summarizer
Module 10: AI Architecture Leadership and Capstone
- Day 91: Architecture Review: Draw the Box Diagram First
- Day 92: Threat Model an AI Feature
- Day 93: SLOs for AI Features
- Day 94: Change Management for Prompts
- Day 95: Incident Response When the Agent Goes Wrong
- Day 96: Hiring and Team Topology for AI Platform
- Day 97: Build vs Buy the Agent Runtime
- Day 98: Capstone Spec: Pick One Production Problem
- Day 99: Capstone Build Week Checklist
- Day 100: Capstone Ship: Demo, Postmortem, Next 90 Days
Watch: 100 Days of AI on YouTube
Short videos from CheatCoders, one AI engineering topic per day.
Last updated October 10, 2026