Run AI agents in production without burning money or trust.
I'm Alexey. I write about the unglamorous side of agents: cost control, reliability, and tooling — including on-chain agents. Every post ships a small, runnable tool you can paste and run with no API key, plus the real output.
Recent posts
-
The best config in your bake-off didn't win. Selection did.
Best-of-K eval selection bias: the top config's observed pass rate is biased up. It shows up even when all K configs are equal, and it grows with K.
-
Zero failures isn't zero risk: the rule of three for evals
Rule of three for evals: zero failures in N runs is a count, not a rate. With 0 failures in 100 runs, the 95% upper bound on the true rate is 2.95%.
-
Your A/B eval is paired. Your stat test probably isn't.
Two prompts on one eval set are paired data. The two-proportion SE my harness printed said 'collect more'; McNemar said it was already decided.
-
A Spend Cap That Stops Counting Is Already Fail-Open
A spend cap that stops counting is already fail-open. When the cost oracle goes quiet, the question is whether the ledger keeps moving. Five strategies.
Building a tool for AI-agent developers?
This is a small, niche blog for developers building and operating AI agents. If your product serves them, you can sponsor a post or a placement — always disclosed, always something I can actually run. The tools here stay free to read.
