I'm Tony, a Machine Learning Engineer at Spotify in London. I build models and production ML systems that learn from user behaviour and optimise user journeys. I also write about recommender systems and foundation models for structured data.
My work spans modelling, experimentation and the engineering needed to put ML into production. At Spotify, I work on personalising subscription experiences, connecting predictions to decisions and evaluating their impact through experiments.
I'm particularly interested in how models learn useful representations of people and structured data. Recently, I've been exploring generative recommenders, tabular and relational foundation models, and how these approaches connect to practical decision-making.
I studied Mathematics at the Chinese University of Hong Kong and Statistics at the University of Warwick, where my dissertation explored vision-language models for chest X-ray report generation.
Away from work, I enjoy singing, badminton and playing music.
Spotify
Leading the development of personalised subscription grace-period decisioning, from heuristic rules towards predictive and causal ML. Building supporting inference services and tooling for policy evaluation and experimentation.
Trainline
Worked on customer lifetime value prediction and contextual bandits for personalised content, supporting conversion optimisation and profitable growth.
Guidehouse
Worked on digital twins for gas distribution networks, using property graphs and autoencoder-based synthetic data generation.
Dept. of Medicine, University of Hong Kong
Worked on convolutional neural networks for orthopaedic medical-image classification.
Tailoring product experiences to individual customers using behavioural signals and machine learning.
Models and pipelines that support acquisition, conversion, and lifetime value across the customer journey.
Designing and analysing A/B tests so product and business teams can make confident decisions.
Connecting predictions to actions — turning model outputs into policies, thresholds, and business rules.
Estimating real effects of interventions when randomised experiments aren't feasible or affordable.
Relational foundation models and other enterprise FMs that predict directly from business data — and how to extend them from zero-shot predictions to zero-shot actions.
Thoughts on machine learning, experimentation, growth, and decision-making.
How Spotify's GLIDE turns an LLM into a personalised recommender — semantic IDs, soft prompts, and why users don't need their own codes.
A look at how semantic IDs compress item embeddings into tokens, why simple residual K-Means can outperform RQ-VAE, and what generative recommenders do with them.
I benchmarked zero-shot and continued-pretrained Relational Transformers against XGBoost and RelGT on two RelBench tasks — 18 hours on a 48GB Mac. One task was competitive; the other collapsed.
Relational foundation models learn directly from the relational data businesses already have — no hand-built feature pipelines. Here's why that shifts the ML stack.
TabFMs and Kumo's RFMs are quietly automating feature engineering and model training. What's left for MLEs? Policy and decision-making.