I spent a year testing thumbnails as carefully as an individual channel allows, which is to say not very carefully, but more carefully than most people bother with.
What I mostly learned is how difficult it is to learn anything, and that is worth writing down because the confident advice on this subject is not supported by anything I could reproduce.
Why individual testing is hard
The core problem is that you cannot run a clean experiment on your own channel.
Two videos are never comparable. Different topics, different titles, different days, different traffic mix, different competition on the day.
Testing two thumbnails on the same video sequentially is better but still confounded, because a video's performance changes with age regardless of what you do to it.
The built-in test features that some platforms provide are the only genuinely clean method, and even those need a decent volume of impressions to distinguish a real difference from noise. On a small channel most tests will not reach that.
Which means most people concluding that a thumbnail style works are looking at a handful of videos and a lot of variance.
The things that held up
Two, and only two, produced differences I saw repeatedly across enough videos to believe them.
Legibility at small size. Whatever is on the thumbnail has to read at the size it is actually displayed, which on a phone is very small. Most of the thumbnails I made in my first two years failed this and I never noticed because I was looking at them full size while making them.
The test is to shrink the image to about a centimetre wide and look at it. If you cannot tell what it is, nothing else about it matters.
Contrast with the surroundings. A thumbnail sits in a grid of other thumbnails, mostly bright, mostly saturated, mostly with a face and large text. Something visually different from that grid gets looked at.
Which is a moving target, because what is different depends on what everyone else is doing. The advice to use bright saturated colours was probably right when most thumbnails were dull, and is now advice to blend in.
The things that did not hold up
Faces. The received wisdom is that faces perform better, and across my videos I could not find a consistent effect. Some of my best performers had faces and some did not, and the split looked random.
I suspect the real finding underneath this is that a face is a reliable way to get a clear focal point, and the focal point is what matters rather than the face specifically.
Exaggerated expressions. I tried this properly, felt ridiculous doing it, and the results were within noise. On my channel, with my audience, it did nothing.
Text quantity. Advice varies between no text and short text and I could not distinguish them. What did seem to matter was whether the text duplicated the title, which is a wasted opportunity either way.
Arrows and circles. No detectable effect for me.
The thing that mattered more than any of it
The relationship between the thumbnail and the title.
My best performers were the ones where the thumbnail and title did different jobs. The title states the subject, the thumbnail shows something that raises a question the title does not answer, or vice versa.
My worst were the ones where both said the same thing, which gives someone no reason to click because they already have the whole proposition.
This is not a thumbnail finding exactly. It is a packaging finding, and it suggests that testing thumbnails independently of titles is testing half of a unit.
Click-through rate is a trap on its own
Worth stating clearly. It is possible to improve click-through rate and make the channel worse.
A thumbnail that overstates what the video contains gets more clicks and worse retention, and the retention signal appears to matter more downstream. I did this once with a thumbnail that implied something more dramatic than the content and the video had my best click-through rate of that quarter and my worst watch time.
The metric I now look at is click-through and retention together. A thumbnail change that raises one and lowers the other is not an improvement.
What I actually do now
Make the thumbnail before making the video, quite often. If I cannot express the idea as a small image and a short title, the idea may not be as clear as I think.
Check legibility at real size, always.
Look at what is currently in the grid for my subject and deliberately do something other than that.
Make sure the title and image are not saying the same thing.
And stop there. I spent a lot of the test year making four variants of everything and the variance between them was smaller than the variance between videos, which means the effort was going into the wrong place.
The uncomfortable conclusion
Thumbnails matter, and beyond a competent baseline the returns to more effort are small and hard to detect.
The gap between a bad thumbnail and a decent one is large. The gap between a decent one and an optimised one, at least on a channel of my size, was lost in the noise.
Which suggests getting to competent quickly and then spending the effort somewhere with better returns, which for me turned out to be the first two minutes of the video.