Skip to content

Blog · Tag

#finops

5 posts tagged #finops.

The harness is one integer column

The fix for a poison-pill row re-billing a model every sixty seconds forever was not a rewrite. It was one INT column, capped at five, incremented atomically in the database. The fix itself shipped with a gap, and that gap is the actual lesson.

You probably don't need to fine-tune

A post framed fine-tuning as a 2026 interview trap. Here is the AWS version: the ladder before you touch weights, the four job types and three serving paths Bedrock forks 'fine-tune' into, and what each one costs, priced today.

Every dashboard was green while the agent burned six figures a year

The most expensive AI agent failures don't throw an error, they hide. One ran at a six-figure-a-year rate for days while every dashboard stayed green, because the signals that catch it are the ones nobody watches. The two instruments you are missing.

Field Notes: Turning prompt caching on for a production Bedrock workload

Strands' BedrockModel ships with prompt caching off. Two kwargs turn it on, one per-model gotcha catches you, and a 10-turn driver measures 99.9% and 99.8% hit ratios against an 8,156-token production system prefix. The usage block proves it in seconds.

Field Notes: Three things I learned diagnosing a production Bedrock workload

Three findings from a real customer engagement on AWS Bedrock: what a load test was actually doing, why p95 latency was 45 seconds, and the prompt-caching default that costs every team money. Plus the three CloudWatch metrics that catch all three.