Aggregate review scores are widely used as a summary of critical opinion, and the conversion from written criticism to a single figure loses precisely the information that matters most.

The conversion happens before the averaging

Many reviews do not carry a score at all, so aggregators assign one by interpreting the text, which turns a judgment expressed in prose into a number chosen by a third party.

That step is not neutral. A review expressing admiration alongside serious reservations can plausibly be scored several different ways, and the choice determines its contribution.

Some aggregators simplify further by sorting reviews into positive and negative, which discards magnitude entirely and treats a mild recommendation identically to an enthusiastic one.

Averages erase the shape of the response

A work that half the critics considered remarkable and half considered a failure produces the same middling figure as one that everybody found unobjectionable and dull.

Those are opposite outcomes for a reader deciding what to watch, since the divisive work may be exactly what someone wants and the mediocre one is unlikely to be.

The number changes behaviour on both sides

Scores affect what audiences choose, so distributors treat them as commercial assets rather than as commentary, and marketing is built around the figure once it exists.

Some outlets adjust their scoring practices in response, and the compression of ratings towards the upper end of a scale is a predictable consequence of scores having consequences.

Critics who write carefully about mixed responses find their work reduced to whichever side of a threshold the aggregator placed them on.

Which critics are counted is itself a judgment

Aggregators decide which publications qualify for inclusion, and that decision shapes the result before any review is read.

A narrow pool produces a consensus figure that may reflect a particular set of tastes rather than critical opinion in general, and the composition is rarely visible to readers.

What the number is actually good for

Aggregates are reliable at the extremes, where near-unanimous responses are genuinely informative and the averaging destroys little.

In the middle they are close to useless, and the useful information sits in the individual reviews the score was built from.

Reading two critics whose tastes are known produces a better prediction than any aggregate, which is why the score works best as an entry point rather than as a verdict.