Skip to content

guide

Why Netflix thumbnails are different for everyone

Leads the technical side. The Filmatic app was his idea, and he created the algorithm and the framework the app runs on.

12 min read

Open Netflix on your phone and on a friend’s, and the same movie can wear two different faces. One of you gets a close-up of the lead. The other gets a side character, a car, or a couple about to kiss. Neither image is a mistake. Netflix chose each one, for each of you, on purpose.

Netflix calls this artwork personalisation, and the algorithm behind it is a contextual bandit. Netflix has described the system in two public engineering posts, in 2016 and 2017. This guide walks through how it works, why Netflix built it, and where it goes wrong.

Why the thumbnail matters so much

In May 2016, Gopal Krishnan of Netflix wrote that if the service does not capture a member’s attention within 90 seconds, that member will likely lose interest and do something else. The same post says Netflix’s studies found that members look at the artwork first, and only then decide whether to read the title, the synopsis or anything else.

So the thumbnail is not decoration. It is the first, and often the only, piece of evidence a viewer sees before deciding whether a title deserves a click. A recommendation algorithm can pick exactly the right movie and still lose you with the wrong picture.

Step one: one best image for everyone

Netflix’s first approach, described in that 2016 post, did not personalise at all. It tried to find the single best image for each title.

One early test used The Short Game, a movie about grade school children competing at golf. The default artwork did not make it obvious the movie was about kids, so Netflix made variants and showed each one to a different group of members. Some variants widened the audience.

Netflix then scaled the idea up with an explore and exploit test. During the explore phase it showed every candidate image to a sample of members and measured each image’s take rate, the number of plays divided by the number of times the image appeared on screen. During the exploit phase it showed the winning image to everyone else.

Two findings from that work still shape the system. Longer tests showed that simply rotating the artwork every so often was not as good as finding a genuinely better image. And the winning images tended to share traits, such as expressive faces that convey the tone of the title, with winners that could differ from one region to another.

Step two: the best image for you

One best image still ignores the fact that people want different things from the same movie. In December 2017, Ashok Chandrashekar, Fernando Amat, Justin Basilico and Tony Jebara published Artwork Personalization at Netflix, which describes picking the best image per member.

Their example is Good Will Hunting. Someone who watches a lot of romantic movies might respond to artwork of Matt Damon and Minnie Driver. Someone who watches a lot of comedies might respond to Robin Williams. The same logic works for cast: a viewer who watches a lot of Uma Thurman might see Pulp Fiction with her on it, and a John Travolta fan might see him instead.

The post is careful to say that Netflix does not write rules like these by hand. It lets the data decide which signals matter.

How a contextual bandit works

A conventional way to improve a recommender is slow. Collect a batch of data, train a model, then run an A/B test against the current system for weeks. The Netflix post points out the cost of that: every member in the losing group spends the whole test with the worse experience. In the language of the field, that cost is regret.

A contextual bandit learns while it serves, which cuts that waiting. The name comes from the “multi-armed bandit”, a classic problem about choosing between slot machines whose payouts you do not know. The contextual version adds information about the situation, and here the situation is the viewer.

For each viewer and each title, the system does three things.

  1. Read the context. The 2017 post lists the kinds of signal it can use: the titles a member has played, their genres, the member’s past interaction with this title, their country, their language preferences, their device, the time of day and the day of the week.
  2. Pick an image. A model predicts, for each candidate image, the probability that this viewer will play the title and actually watch it. Netflix says a title typically has up to a few dozen candidates. The system shows the image with the highest prediction.
  3. Explore a little. Some of the time, the system deliberately shows a different image, chosen with controlled randomness. Without that, it would only ever learn about images it already favours. Netflix names simple schemes like epsilon-greedy, where a small fixed share of choices is random, and adaptive schemes that explore more when the model is less certain.

The system logs every choice, including the random ones, along with whether the viewer played the title. That log becomes the next round of training data.

Nobody invented contextual bandits for movies. One of the best known papers in the area, by Lihong Li, Wei Chu, John Langford and Robert Schapire at WWW 2010, applied the approach to news articles on the Yahoo! front page and introduced an algorithm called LinUCB, which the Netflix post lists as one of its options.

Testing a new algorithm without showing it to anyone

The logged randomness has a second use. Because Netflix knows exactly how likely each image was to be shown, it can estimate how a new selection algorithm would have performed, without deploying it.

Researchers call the technique replay, and it comes from a 2011 paper by Li, Chu, Langford and Xuanhui Wang. Take the historical sessions where the random choice happened to match what the new algorithm would have picked, and measure the take rate across just those sessions. That gives a fair offline estimate.

Netflix wrote that the contextual bandits beat both random selection and the one-best-image approach on replay, that a live A/B test then confirmed a significant lift, and that the system was rolled out to all members. One detail from the results is telling. Personalisation helped most when the viewer had never interacted with the title before, which is exactly when a picture has to do the most work.

Why personalised artwork is hard to get right

The 2017 post is unusually candid about the problems, and they are the interesting part.

Only one image can win. A normal recommendation row shows many titles, so the system learns from which one you pick. A thumbnail is a single choice. If you play the title, you can only have played it after seeing the image you were shown, so the system has to work out whether the image made the difference or whether you would have watched anyway.

Changing the image has a cost. A fresh image can make you reconsider a title. It can also stop you recognising something you meant to watch. Netflix wrote that it controls exploration so that a viewer’s artwork does not change too often, partly for that reason and partly so it can tell which image deserves credit.

The page is the unit, not the image. A bold close-up stands out in a row of muted images. A whole row of bold close-ups stands out from nothing.

Clickbait wins clicks. An image that misrepresents a title can still earn plays. Netflix says it measures the quality of the viewing that follows, not just the play, so that an image that gets people to start and then abandon a title does not win. It also says the candidate pool has to be representative of the title in the first place.

It runs at enormous scale. The post gives a peak of over 20 million image requests per second, all of which need an answer quickly.

When personalisation feels like targeting

The representativeness problem became public in October 2018. As Joe Berkowitz reported in Fast Company, the writer Stacia L. Brown posted that her artwork for the movie Like Father showed two Black supporting actors who barely appear in the movie, rather than its leads, Kristen Bell and Kelsey Grammer. Other viewers reported similar examples.

Netflix said it does not ask members for their race, gender or ethnicity, and that it uses viewing history to personalise. Both things can be true at once. An algorithm that only sees viewing history can still learn patterns that line up with who someone is, because what people watch is not independent of who they are.

The lesson for anyone building a recommender is that the algorithm optimises within the pool it is given. Whether a side character on the thumbnail fairly represents the movie is an editorial judgement, and no take rate can make it.

What this means when you browse

  • The thumbnail is a prediction about you. It shows the part of the title the system expects you to respond to, not necessarily what the title is mostly about.
  • Read past the image. If a title catches your eye, check the synopsis and the cast before assuming it is the movie the picture suggests.
  • A changed image is not a new title. If artwork for something you skipped looks different, the system may be exploring, or your history may have changed what it predicts.
  • The picture chooses how, not what. Artwork personalisation decides how a title is shown to you. Which titles reach your homepage at all is a separate system with its own goals, covered in why Netflix’s recommendations feel worse every year.

For the algorithms that pick the titles in the first place, from collaborative filtering to matrix factorisation, see how movie recommendation algorithms work. Filmatic takes a plainer route to the same problem of persuading you a title is worth your evening. Every pick comes with a sentence on why it suits you, described in how the Filmatic recommendation algorithm works.

Questions

Why does Netflix show different thumbnails to different people?

Because Netflix personalises the artwork as well as the recommendations. Each title has several candidate images, and an algorithm predicts which one a particular viewer is most likely to play, using signals such as what they have watched before. Netflix described the system publicly in 2017, using Good Will Hunting as the example, where a viewer of romances might see Matt Damon and Minnie Driver and a viewer of comedies might see Robin Williams.

What is a contextual bandit?

It is a type of online learning algorithm that makes a choice, sees the result, and updates itself as it goes, rather than waiting for a full experiment to finish. The context is what the algorithm knows about the situation, such as the viewer. It mostly picks the option it predicts is best, and occasionally picks something else at random so it keeps learning. That trade is called explore and exploit.

Why did the thumbnail for a show I already know suddenly change?

Netflix chooses personalised artwork per viewer, and the choice can change as the system learns, either because your viewing history shifted or because it is still exploring which image works. Netflix wrote in 2017 that changing artwork between sessions is a known trade off, since a new image can prompt a second look but can also make a title harder to find again, and that it limits how often selections change.

Does Netflix use race to choose thumbnails?

Netflix says it does not. In October 2018, after viewers reported artwork for Like Father that featured minor Black cast members rather than the leads, Netflix said it does not ask members for their race, gender or ethnicity and uses viewing history to personalise. Viewing history can still correlate with those things, which is why the episode is a standard example of how personalisation can feel targeted.

Is personalised thumbnail artwork just clickbait?

It can drift that way, and Netflix's own engineers named the risk. Their 2017 write up says they judge an image by the quality of the viewing it leads to, not just the click, so an image that gets plays that are quickly abandoned does not win. The harder problem is whether every image in the pool fairly represents the title, which is an editorial choice the algorithm cannot make.

Keep reading

Everything optional starts switched off. Nothing in a category loads until you allow it.