Where we startedThe catalog was useful, the recommendation area was not

The product had a large catalog and a static set of suggestions. Most people saw roughly the same items. The area existed, but it behaved like merchandising space rather than a product that could learn from the person using it.

We did not have labeled training data or time for a multi-quarter machine-learning program. What we did have was behavior: views, searches, repeated interest, abandoned actions and categories a user consistently ignored. I asked the team to build the smallest complete system that could use those signals and prove whether relevance improved.

The product callImprove the ranking before redesigning the cards

One proposal was to redesign the module. Better cards and a carousel would make it look new, but the same generic list would still sit underneath. We kept the interface mostly stable so we could isolate whether the ranking itself made a difference.

We also chose one outcome before development began. It was not a click. The recommendation existed to help a user complete a meaningful action, so that was what we measured. This kept product, data and engineering aligned when several attractive secondary metrics moved in different directions.

The first releaseA ranking we could explain and measure

The first ranker used a small set of understandable inputs: recency and depth of engagement, declared interests and category affinity. Every impression carried its position and eventual outcome. We released to a slice of traffic and kept a control group.

An explainable score was the right 0 to 1 choice. The team could answer why an item appeared, debug surprising results and change one input at a time. We also reserved some space for exploration so the system could learn about new interests instead of repeating what it already knew.

I held the launch until the instrumentation and control were ready. That added work to the first release, but it protected the only clean chance we had to learn whether the product was actually better.

From release to productThe learning loop mattered more than version one

The action metric improved against the live control. More importantly, the team now had a repeatable loop: collect signals, rank, measure, inspect and adjust. Product discussions moved from opinions about what should be promoted to evidence about what helped users act.

Scaling required different work from the initial launch. We had to monitor signal quality, keep ranking changes reversible, handle new users and new catalog items, and prevent one team from becoming the permanent owner of every tuning decision. The system became a product when other teams could operate it safely, not when the first experiment succeeded.

What I learned

I would keep measurement inside the definition of the first release. A dashboard can come later. A missing baseline cannot.

I would treat cold start as a first-class path sooner. New users and new items have the least data and often matter most. We handled them, but too late in the first design.

The first version did not need to be clever. It needed to be measurable, explainable and easy for the team to improve.