Skip to content

Topic

Agentic AI on AWS

Building agents on AWS: Bedrock AgentCore, Strands, MCP, RAG, tools, and memory.

16 posts.

Field Notes: A Gateway policy can't see the user if your agent calls with its own token

A re:Post question asked how to stop an AgentCore agent acting beyond the user who asked. I built a demo to find out. A Cedar policy does nothing if the agent calls with its own token. Two gotchas, and the real limit: a stale token claim.

Introducing the AI Architect Roadmap

Eight rungs mapping what shipping production Agentic AI on AWS actually takes, marked honestly as covered, strong, flagship or gap rather than filled in to look finished.

Not a Python tutorial: the patterns that bite in production agents

The senior-lens Python patterns that actually bite in a production agent loop: why boto3 blocking the event loop is the first one, and four more that follow from it.

You probably don't need to fine-tune

A post framed fine-tuning as a 2026 interview trap. Here is the AWS version: the ladder before you touch weights, the four job types and three serving paths Bedrock forks 'fine-tune' into, and what each one costs, priced today.

Your LLM security diagram defends the wrong layer

The LLM security diagram you have seen a dozen times is a threat map. Read as a defence it makes you patch every box at the layer the attacker controls. The fix is one deterministic boundary the diagram leaves out, in code the model never touches.

The LLM is not a security boundary

Designing a production agent over sensitive data: no control makes the flow hole-free. You rank the layers, assume each one leaks, and stack them so no single hole reaches the data. Here is the code that does it.

Field Notes: The AgentCore Memory write that returns success and reads back empty

AgentCore long-term memory has a read-after-write gotcha the docs skip: a direct BatchCreateMemoryRecords write returns 201 and stays unsearchable for 15 to 30 seconds. Measured, with the two-tier model that explains it.

Part 2: The MCP Server: Turning ADRs and Incidents into a Queryable Org-Knowledge Surface

The agent doesn't read your wiki. It calls four tools that pull frontmatter-filtered chunks out of a Bedrock Knowledge Base. The contract, the code, and the small decisions between an agent that reads your docs and one that knows your org.

Part 3: Wiring It Into AWS DevOps Agent: AgentSpace, register-service, and the IAM Trust Policy That Ate My Afternoon

The MCP server is done. Now plug it into AWS DevOps Agent: three CDK stacks, the AgentSpace and register-service flow, the composite-principal trust policy you will get wrong first try, and an OIDC gotcha that broke my blog deploy for a month.

Part 1: Intent vs State. How AWS DevOps Agent Closes the Gap Between What Your System Is and What You Decided It Should Be

When something breaks at 3am, you look at logs, metrics, traces. You don't go and re-read the ADR your team wrote in January. AWS DevOps Agent does. Here's why that changes the first hour of an incident.

Part 6: Cost & Performance for Bedrock AgentCore: Prompt Caching, Model Selection, and CloudWatch Alarms

Real cost breakdown of running an AgentCore agent: prompt caching savings, when to use Nova Pro vs Claude Sonnet, PriceClass_100, idle timeouts, and how to set alarms before your bill surprises you.

Part 5: CI/CD for Bedrock AgentCore with GitHub Actions and AWS OIDC (No Stored Credentials)

How to build a complete CI/CD pipeline for AgentCore using GitHub Actions OIDC: no stored AWS keys, dual-tag ECR strategy, automated Runtime updates, and multi-environment promotion.

Part 4: Running Your AgentCore Agent Locally with Docker (The Right Way)

How to build and run your AgentCore container locally with real AWS credentials, the correct linux/amd64 platform flag, the .env.local pattern, and how to test with curl.

Part 3: Building the AI Agent with Strands Agents SDK, Prompt Caching, and AgentCore Memory

How to build the Python agent that runs inside AgentCore: Strands SDK setup, prompt caching that cuts costs by 90%, dual-model strategy, tool definitions, and AgentCore Memory integration.

Part 2: CDK Infrastructure for Amazon Bedrock AgentCore (And Every Gotcha You'll Hit)

A complete CDK v2 TypeScript stack for Bedrock AgentCore, with inline comments for every deployment trap: naming constraints, ECR bootstrap, missing L1 constructs, VPC endpoint conflicts, and more.

Part 1: Why I Chose Amazon Bedrock AgentCore (And What Lambda Gets Wrong for AI Agents)

Before writing a single line of agent code, I spent a week figuring out where to run it. Here's the architecture decision that changed everything, and the Lambda limitations that forced my hand.