← All projects Graduate project · Modeling

Music Popularity Predictor

STT 810 · Michigan State · 2022
R Regression Feature engineering Data viz

How much of a song's popularity is actually written into its audio features — and how much is just noise?

This was the final project for STT 810 — Mathematical Statistics for Data Scientists, a course in the M.S. Data Science program taught by Dr. Paul Speaker. Working with a partner, Chris Grandy, I set out to predict a track's popularity from a Spotify dataset of roughly 20,000 songs and 15 numerical and categorical features.

Histogram showing the distribution of song popularity, roughly bell-shaped and centered in the 50s to 60s
The target. Popularity is roughly bell-shaped and centered in the 50s–60s — most songs are middling, with far fewer true hits or flops.

The approach

We treated it as a supervised learning problem, built entirely in R. The pipeline ran from exploratory analysis and feature engineering through visualization and modeling, validated on a 70/30 train–test split. Along the way we leaned on the standard R toolkit for this kind of work:

  • Wrangling & stats: dplyr, tidyverse, MASS, caTools
  • Correlation & visualization: ggplot2, corrplot, Plotly, patchwork

The final model used a small, deliberate set of predictors — energy, key, and speechiness — which the exploratory work flagged as the features with any real relationship to popularity.

What we found

The honest headline: popularity in this dataset was close to random. Most features carried little signal, and even the best predictors explained only a modest slice of the variance. Rather than force a more complex model to chase noise, we kept the model simple and let the data set the ceiling.

Correlation matrix of all audio features; the song popularity row shows only faint correlations with every feature
The story is the top row. Song popularity barely correlates with any single feature — while the audio features correlate plenty with each other (energy with loudness; acousticness against energy). Green is positive, plum is negative; bigger, deeper circles mean a stronger relationship.
Two-dimensional density of song popularity against energy, showing a broad hotspot around energy 0.7 and popularity 60
Even the stronger features stay soft. Popular songs cluster around energy ≈ 0.7 — but the hotspot is wide and flat. High energy nudges popularity; it never determines it.
Observed versus predicted song popularity; predicted values form a narrow band near the mean while observed values scatter across the full range
The punchline. The model's predictions (teal) barely leave the mean — its safest bet is "about average" — while the true values (coral) scatter across the whole range. When the features don't carry the signal, this is what an honest model looks like.

That turned out to be the real lesson of the project — not every question has a strong enough signal to build a genuinely predictive model from, and recognizing that early is its own kind of result. It's a lesson that has stuck with me well beyond the course: interrogate whether the relationship exists before investing in the model that assumes it does.

Data

The dataset came from a public Spotify song popularity collection on Kaggle.