A model that does well in validation is only the start. The questions I find more interesting come after it: does this hold on data it hasn't seen, does it change a decision anyone is actually making, and is it worth what it costs to run?
That shows up throughout this work. A forecasting model is judged by a newsvendor cost simulation and a pre-registered A/B test rather than by WMAPE — and the gate said HOLD twice before it said SHIP. A segmentation model is treated as something with a lifecycle, with drift thresholds calibrated against the data's own month-to-month variation instead of textbook constants, which cut false retrain alerts from 21 windows out of 22 down to the four real regime changes.
I also report what didn't work. Two full iterations went into a model that was more accurate and still lost money, and both are written up next to the version that finally earned its launch. A result you'd bury when it disagrees with you isn't a measurement.
I'm comfortable across the whole workflow — framing the problem, moving the data with Spark, dbt and Airflow, building and tuning the model, wiring the serving layer, and setting up the monitoring that catches drift — on AWS, GCP and Azure.