WRITING
Our Production Model Finished Fifth
9 min readFor months our WhatsApp sales agent answered customers on a model nobody had ever benchmarked. When I finally tested that slot against ten alternatives it came fifth — and my own harness lied to me four times on the way to finding out.
/AI/EVALUATION/LLMOPSAn AI That Can Describe an Action Cannot Necessarily Perform It
6 min readA proactive follow-up told a customer their payment link was on the way. No order existed and no link existed. What that incident changed about how I let model output reach people.
/AI/RELIABILITY/POSTMORTEMAI Self-Improvement Needs New Data, Not Just Better Loops
5 min readClosed-loop self-improvement may be impossible for models trained mostly on their own outputs. That is a boundary condition, not the end of AI progress.
/AI/MACHINE LEARNING/DATA