Skip to content

guide

How movie recommendation algorithms work: the four main types

Leads the technical side. The Filmatic app was his idea, and he created the algorithm and the framework the app runs on.

13 min read

Every movie recommendation you have ever been shown came out of one of a small number of algorithms. The names sound technical, but the ideas underneath are simple, and knowing which one you are looking at explains most of what feels right or wrong about a recommendation.

This guide covers the four main types of movie recommendation algorithm: user-based collaborative filtering, item-based collaborative filtering, matrix factorisation and content-based filtering. It also explains how real systems combine them. We wrote it for anyone who has wondered why an app suggested a particular movie, and for anyone building one.

What a recommendation algorithm actually predicts

Start with a table. Every user is a row, every movie is a column, and each cell holds that person’s rating of that movie. A recommendation algorithm fills in the empty cells. It predicts the rating you would give a movie you have not seen, then shows you the movies with the highest predictions.

The table is almost entirely empty. In their 2009 paper on the subject, Yehuda Koren, Robert Bell and Chris Volinsky describe explicit ratings as a sparse matrix, because any one person has rated only a small share of the movies available. Every algorithm below is a different strategy for guessing well from very little.

Ratings are not the only input. The same paper separates explicit feedback, meaning stars, thumbs and likes, from implicit feedback, meaning what people actually do: what they watch, search for, finish or abandon. Implicit data is far denser, because everyone generates it whether or not they ever rate anything. It is also noisier, because watching something is not the same as liking it.

Collaborative filtering: people who agree with you

Collaborative filtering ignores what a movie is about. It only looks at behaviour. If you and another person have rated many of the same movies the same way, a movie they loved and you have not seen is a good bet for you.

Koren and his co-authors credit the developers of Tapestry, an early recommender, with coining the term. The version most people would recognise arrived in 1994, when Paul Resnick, John Riedl and colleagues presented GroupLens at the CSCW conference. GroupLens recommended Usenet news articles by letting readers share ratings and predicting each reader’s interest from people who rated like them.

Researchers call that approach user-based collaborative filtering. Its strength is that it finds connections no description would ever contain. Two movies can share nothing on paper and still appeal to the same people for reasons nobody wrote down.

Its weakness is scale and churn. Tastes shift, new users arrive every day, and comparing every person with every other person gets expensive quickly.

Item-based collaborative filtering: movies rated alike

In January 2003, Greg Linden, Brent Smith and Jeremy York of Amazon published a short paper in IEEE Internet Computing that flipped the comparison around. Instead of finding customers similar to you, their system found items similar to each other, where “similar” meant bought or rated together by the same people. Their argument was that comparing items rather than customers scales to very large data sets.

For movies, item-based collaborative filtering works like this. Koren, Bell and Volinsky use Saving Private Ryan as the example. Its neighbours are the movies that tend to get similar ratings from the same people, which might be war movies, Spielberg movies and Tom Hanks movies. To predict your rating for it, the system looks at how you rated those neighbours.

This is the algorithm behind most “because you watched” rows. It is stable, because relationships between movies change more slowly than people do, and it is easy to explain, because every recommendation points back to something you rated.

Matrix factorisation: taste axes nobody labelled

The neighbourhood methods above compare things directly. Matrix factorisation takes a different route, and it is the approach the Koren paper made famous.

It assumes that a handful of hidden factors explain most ratings. Every movie gets a short list of numbers describing how much it has of each factor, and every user gets a matching list describing how much they care about each one. A predicted rating is how well the two lists line up. The paper describes characterising users and movies on 20 to 100 such factors, all learned from the ratings alone.

Nobody tells the algorithm what the factors mean. According to the paper, some turn out to be obvious, such as comedy against drama, amount of action or orientation to children. Some are vaguer, such as depth of character development or quirkiness. Nobody can interpret some of them at all.

The paper’s most striking example comes from Netflix data. One learned axis ran from lowbrow comedies and horror aimed at a male or adolescent audience at one end, to drama or comedy with serious undertones and strong female leads at the other. Another separated independent, critically acclaimed and quirky movies from mainstream formulaic ones. No human labelled either axis. The algorithm found them because they explained how people rated.

Latent factor models are also why “movies like this” can surprise you in a good way. Two movies can sit close together on the hidden axes without sharing a genre, a director or a decade.

The Netflix Prize, and what it taught the field

Matrix factorisation went mainstream because of a public competition. On 2 October 2006 Netflix released about 100 million anonymised movie ratings and challenged anyone to beat its own recommendation algorithm, Cinematch, by 10 percent on prediction error.

It took almost three years. On 21 September 2009 a merged team called BellKor’s Pragmatic Chaos won. Another team, The Ensemble, matched their result, and BellKor’s Pragmatic Chaos won because they submitted 20 minutes earlier. The winning solution was a blend of many models rather than one clever algorithm.

Two lessons came out of it, and both matter more than the leaderboard.

Accuracy is not the product. In April 2012, Xavier Amatriain and Justin Basilico wrote on the Netflix Tech Blog that Netflix had evaluated the new methods offline, and that the extra accuracy did not seem to justify the engineering effort needed to run them in production. A small gain in prediction error rarely survives the cost of running it for millions of people.

Ratings identify people. In 2008, Arvind Narayanan and Vitaly Shmatikov showed that they could match the “anonymous” Netflix Prize records to named people by comparing them with public ratings on IMDb. Netflix cancelled a planned second contest in March 2010. For anyone building a recommender, a person’s ratings are personal data, whatever the file calls them.

Content-based filtering: describing the movie itself

Content-based filtering does the opposite of collaborative filtering. It ignores other users and describes each movie by its own attributes: genre, cast, crew, era, keywords, plot summary, or the text critics and audiences have written about it. It then recommends movies whose description resembles the ones you liked.

The Koren paper uses Pandora’s Music Genome Project as the classic example from music, where trained analysts score every song on hundreds of characteristics. Modern systems usually skip the analysts. They turn text about each movie into embeddings, long lists of numbers that place similar movies near each other, and measure similarity as distance.

Its great advantage is the cold start problem. The same paper notes that collaborative filtering is generally more accurate, but it cannot handle new products and new users, and that content filtering is superior in exactly that case. A movie released this morning has no ratings, but it already has a synopsis, a cast and a director. We cover the new-user side of this in why recommendation apps ask you to rate things first.

Its weakness is sameness. A system that only recommends what resembles your history keeps you inside it. Pure similarity also has a geometric quirk, where a few titles sit close to almost everything, which we explain in why every “movies like this” list shows the same movies.

The four types side by side

AlgorithmWhat it learns fromStrongest atWeakest at
User-based collaborative filteringRatings from people like youUnexpected links between moviesScale, and new users
Item-based collaborative filteringMovies rated alike by the same peopleStable, explainable “because you watched” picksBrand new movies
Matrix factorisationHidden factors inferred from all ratingsAccuracy when ratings are plentifulCold start, and explaining itself
Content-based filteringThe movie’s own attributes and textNew movies, niche moviesVariety, and surprise

Hybrids: how real systems combine them

Almost nobody ships just one of these. In 2002, Robin Burke published a survey of hybrid recommender systems in User Modeling and User-Adapted Interaction, cataloguing the ways designers combine methods: weighting their scores together, switching between them depending on how much data exists, or feeding the output of one into another.

At very large scale the usual shape is a two-stage pipeline. Paul Covington, Jay Adams and Emre Sargin described YouTube’s version at the ACM RecSys conference in 2016. A candidate generation model narrows an enormous catalogue down to a small set of possibilities, and a separate ranking model orders those candidates using many more signals. Collaborative and content methods often feed the first stage, and the second stage decides what you actually see.

The part the algorithm does not decide

Every recommendation algorithm above predicts something. Which something is a business decision, not a technical one. A system tuned to maximise viewing hours behaves differently from one tuned to predict what you will rate highly, even with identical code underneath.

A streaming service also recommends only from its own catalogue. That is a hard limit no algorithm can remove, and it is the subject of why Netflix’s recommendations feel worse every year.

Where Filmatic fits

Filmatic is a hybrid with a content-based core. It describes every movie and show using the text written about it, builds a taste profile from your own swipes, and ranks candidates against that profile. It then filters to what is streaming on the services you have in your country.

We chose that shape because it handles brand new titles from the day they appear, and because a content-based core can say why it picked something. The full walkthrough, including what the approach is bad at, is in how the Filmatic recommendation algorithm works. If you want to compare the apps built on these ideas rather than the ideas themselves, see the best movie recommendation apps.

The short version

  • Collaborative filtering recommends what people like you enjoyed. It is accurate with lots of data and blind to anything new.
  • Item-based collaborative filtering recommends movies rated alike. It powers most “because you watched” rows.
  • Matrix factorisation learns hidden taste axes from ratings. It won the Netflix Prize and is hard to explain.
  • Content-based filtering matches movies on what they are. It handles new titles and tends towards sameness.
  • Real systems are hybrids, usually a fast candidate stage followed by a careful ranking stage.

Questions

What is the best algorithm for movie recommendations?

There is no single best one. Collaborative filtering, and matrix factorisation in particular, tends to be the most accurate when a system has plenty of ratings, because it learns from how many people behave. Content-based filtering is better for new movies and new users, because it only needs a description of the movie. Most production recommenders are hybrids that use both and then rank the combined candidates.

What is the difference between collaborative filtering and content-based filtering?

Collaborative filtering looks only at behaviour. It recommends a movie because people whose ratings resemble yours liked it, and it never needs to know what the movie is about. Content-based filtering looks only at the movie. It describes each title by its genre, cast, crew or the text written about it, and recommends titles that resemble ones you already liked.

How does matrix factorisation work in a recommender system?

It treats all ratings as one large table of users against movies, with most cells empty, and learns a short list of hidden factors for every user and every movie. A predicted rating is how well the two lists line up. The factors are never labelled, but in Netflix data some of them turned out to track recognisable things, such as lowbrow comedy against serious drama.

How does Netflix recommend movies?

Netflix does not publish its current system in detail. What is public is its history. It ran the Netflix Prize from 2006 to 2009 to improve how well it predicted ratings, and in 2012 its engineers wrote that the extra accuracy from the newest methods did not justify the engineering effort of running them. It can also only recommend titles in its own catalogue in your country.

Why do new movies and new users get poor recommendations?

Because collaborative filtering learns from ratings, and a new movie or a new user has almost none. Researchers call this the cold start problem. Content-based methods avoid it for new movies, since a description exists from day one, and apps avoid it for new users by asking for a few ratings before making their first pick.

Keep reading

Everything optional starts switched off. Nothing in a category loads until you allow it.