Skip to content
Field Notes

Field Notes

Diagnostic posts from real production AI engagements: what the dashboard didn't catch, what the load test missed, what the CloudWatch metric actually said. Evidence-led, customer sanitised, written so a peer running the same stack can apply the lesson on Monday.

The harness is one integer column

The fix for a poison-pill row re-billing a model every sixty seconds forever was not a rewrite. It was one INT column, capped at five, incremented atomically in the database. The fix itself shipped with a gap, and that gap is the actual lesson.

#aws#observability#reliability#agents#finops

You probably don't need to fine-tune

A post framed fine-tuning as a 2026 interview trap. Here is the AWS version: the ladder before you touch weights, the four job types and three serving paths Bedrock forks 'fine-tune' into, and what each one costs, priced today.

#aws#bedrock#fine-tuning#lora#finops

Your quality alert needs 32 samples

Part 3 left me a label-free signal that responds to a real regression. It does not come with a threshold. Here is where the line actually goes, what a false page costs, and why the reflex answer computes to a negative number.

#aws#bedrock#evals#observability#monitoring

Your golden dataset is too easy

I could not detect a deleted guardrail with an LLM judge, a judge-free assertion, or six label-free signals. Three instruments, one null. The instrument was never the problem: shorten the source and the same gate goes from p = 1.000 to p = 0.0020.

#aws#bedrock#evals#llm-as-judge#observability

Your regression gate needs a power calculation

I deliberately broke my summariser's prompt, then failed to detect it two ways: with an LLM judge over a golden dataset, and with a judge-free deterministic assertion. Removing the judge changed nothing. Here is the calculation that would have told me first.

#aws#bedrock#evals#llm-as-judge#regression-testing

A clean pass rate is not calibration

I built an LLM-as-judge eval on my own blog and got a suspiciously perfect 16/16. Here's the three-round test I ran before trusting that number: single-variable corruption, and a self-consistency check the research says most teams skip.

#aws#bedrock#evals#llm-as-judge#observability

Your LLM security diagram defends the wrong layer

The LLM security diagram you have seen a dozen times is a threat map. Read as a defence it makes you patch every box at the layer the attacker controls. The fix is one deterministic boundary the diagram leaves out, in code the model never touches.

#genai#security#agents#prompt-injection#architecture

The LLM is not a security boundary

Designing a production agent over sensitive data: no control makes the flow hole-free. You rank the layers, assume each one leaks, and stack them so no single hole reaches the data. Here is the code that does it.

#genai#security#agents#prompt-injection#architecture

Field Notes: The AgentCore Memory write that returns success and reads back empty

AgentCore long-term memory has a read-after-write gotcha the docs skip: a direct BatchCreateMemoryRecords write returns 201 and stays unsearchable for 15 to 30 seconds. Measured, with the two-tier model that explains it.

#aws#bedrock#agentcore#memory#agents

Field Notes: Turning prompt caching on for a production Bedrock workload

Strands' BedrockModel ships with prompt caching off. Two kwargs turn it on, one per-model gotcha catches you, and a 10-turn driver measures 99.9% and 99.8% hit ratios against an 8,156-token production system prefix. The usage block proves it in seconds.

#aws#bedrock#agentcore#finops#strands

Field Notes: Three things I learned diagnosing a production Bedrock workload

Three findings from a real customer engagement on AWS Bedrock: what a load test was actually doing, why p95 latency was 45 seconds, and the prompt-caching default that costs every team money. Plus the three CloudWatch metrics that catch all three.

#aws#bedrock#performance#observability#finops