Field Notes
Diagnostic posts from real production AI engagements: what the dashboard didn't catch, what the load test missed, what the CloudWatch metric actually said. Evidence-led, customer sanitised, written so a peer running the same stack can apply the lesson on Monday.
The harness is one integer column
The fix for a poison-pill row re-billing a model every sixty seconds forever was not a rewrite. It was one INT column, capped at five, incremented atomically in the database. The fix itself shipped with a gap, and that gap is the actual lesson.
You probably don't need to fine-tune
A post framed fine-tuning as a 2026 interview trap. Here is the AWS version: the ladder before you touch weights, the four job types and three serving paths Bedrock forks 'fine-tune' into, and what each one costs, priced today.
Your quality alert needs 32 samples
Part 3 left me a label-free signal that responds to a real regression. It does not come with a threshold. Here is where the line actually goes, what a false page costs, and why the reflex answer computes to a negative number.
Your golden dataset is too easy
I could not detect a deleted guardrail with an LLM judge, a judge-free assertion, or six label-free signals. Three instruments, one null. The instrument was never the problem: shorten the source and the same gate goes from p = 1.000 to p = 0.0020.
Your regression gate needs a power calculation
I deliberately broke my summariser's prompt, then failed to detect it two ways: with an LLM judge over a golden dataset, and with a judge-free deterministic assertion. Removing the judge changed nothing. Here is the calculation that would have told me first.
A clean pass rate is not calibration
I built an LLM-as-judge eval on my own blog and got a suspiciously perfect 16/16. Here's the three-round test I ran before trusting that number: single-variable corruption, and a self-consistency check the research says most teams skip.
Your LLM security diagram defends the wrong layer
The LLM security diagram you have seen a dozen times is a threat map. Read as a defence it makes you patch every box at the layer the attacker controls. The fix is one deterministic boundary the diagram leaves out, in code the model never touches.
The LLM is not a security boundary
Designing a production agent over sensitive data: no control makes the flow hole-free. You rank the layers, assume each one leaks, and stack them so no single hole reaches the data. Here is the code that does it.
Field Notes: The AgentCore Memory write that returns success and reads back empty
AgentCore long-term memory has a read-after-write gotcha the docs skip: a direct BatchCreateMemoryRecords write returns 201 and stays unsearchable for 15 to 30 seconds. Measured, with the two-tier model that explains it.
Field Notes: Turning prompt caching on for a production Bedrock workload
Strands' BedrockModel ships with prompt caching off. Two kwargs turn it on, one per-model gotcha catches you, and a 10-turn driver measures 99.9% and 99.8% hit ratios against an 8,156-token production system prefix. The usage block proves it in seconds.
Field Notes: Three things I learned diagnosing a production Bedrock workload
Three findings from a real customer engagement on AWS Bedrock: what a load test was actually doing, why p95 latency was 45 seconds, and the prompt-caching default that costs every team money. Plus the three CloudWatch metrics that catch all three.