An LLM feature that works in a notebook is the easy 20%. The questions I find more interesting come after it: does the retrieval surface the right evidence on inputs nobody curated, can the system tell you when it doesn't know, what does a request cost, and how would you notice if it quietly got worse?
That shows up throughout this work. A semantic cache and a cross-encoder make a clinical retrieval system fast and accurate enough to trust. A multi-agent planner checks every price against the tool payload it came from and does its arithmetic in Python, because a model that mis-adds a budget is worse than no budget at all. A coding agent has to produce an observed failing test and a passing containerized run before it's allowed to call itself done.
I also report what didn't work. One project retracts its own earlier retrieval finding once the ablation was run properly — the corrected result is indistinguishable from zero, and it sits in the README next to the wins. Another measures a near-doubling of retrieval recall and says plainly, in the same table, that answer F1 barely moved. A result you'd bury when it disagrees with you isn't a measurement.
I'm comfortable across the whole path — retrieval and indexing, fine-tuning, the serving layer, the eval harness, and the tracing and cost accounting that keep it honest — on AWS, GCP and Azure.