Skip to content

Blog · Tag

#genai

9 posts tagged #genai.

The ML you need to operate LLMs, not train them

You do not need to train a transformer to run one well. What tokens, context budget and sampling knobs actually mean, verified against the current Bedrock Converse API, including a temperature/top_p gotcha that silently breaks on newer Claude models.

Introducing the AI Architect Roadmap

Eight rungs mapping what shipping production Agentic AI on AWS actually takes, marked honestly as covered, strong, flagship or gap rather than filled in to look finished.

You probably don't need to fine-tune

A post framed fine-tuning as a 2026 interview trap. Here is the AWS version: the ladder before you touch weights, the four job types and three serving paths Bedrock forks 'fine-tune' into, and what each one costs, priced today.

Your quality alert needs 32 samples

Part 3 left me a label-free signal that responds to a real regression. It does not come with a threshold. Here is where the line actually goes, what a false page costs, and why the reflex answer computes to a negative number.

Your golden dataset is too easy

I could not detect a deleted guardrail with an LLM judge, a judge-free assertion, or six label-free signals. Three instruments, one null. The instrument was never the problem: shorten the source and the same gate goes from p = 1.000 to p = 0.0020.

Your regression gate needs a power calculation

I deliberately broke my summariser's prompt, then failed to detect it two ways: with an LLM judge over a golden dataset, and with a judge-free deterministic assertion. Removing the judge changed nothing. Here is the calculation that would have told me first.

A clean pass rate is not calibration

I built an LLM-as-judge eval on my own blog and got a suspiciously perfect 16/16. Here's the three-round test I ran before trusting that number: single-variable corruption, and a self-consistency check the research says most teams skip.

Your LLM security diagram defends the wrong layer

The LLM security diagram you have seen a dozen times is a threat map. Read as a defence it makes you patch every box at the layer the attacker controls. The fix is one deterministic boundary the diagram leaves out, in code the model never touches.

The LLM is not a security boundary

Designing a production agent over sensitive data: no control makes the flow hole-free. You rank the layers, assume each one leaks, and stack them so no single hole reaches the data. Here is the code that does it.