Star Ratings and Review Counts: What a Small Sample Can and Cannot Tell You

What a single review is worth

Two televisions in this catalog carry a star rating of 1.0. The Samsung DU6900 series fifty inch set shows 1.0 stars from one rating, and the Sharp AQUOS seventy five inch commercial set shows the same. Two other items carry 5.0 stars from one rating, including the LG 5.0 cubic foot front load washer and the Roku 2 XD streaming player.

All four of those numbers carry the same amount of information about the product, which is close to none. One rating tells you one person had one experience. It does not distinguish a defective unit from a defective delivery from a buyer who wanted a different model, and it does not distinguish a genuinely good machine from a happy first customer.

The instinct to read 1.0 as a warning and 5.0 as an endorsement is the single most common mistake in listing reading, and it is worth breaking deliberately. A rating is an average, and an average of one is just the observation.

Where an average starts to carry weight

There is no exact threshold, but the arithmetic gives a useful feel for it. With four ratings, one additional review moves the average by up to a full star. With ten, it moves it by up to half a star. With a hundred, a single review shifts it by hundredths and the number has genuinely settled.

So read small counts as ranges rather than values. The Hisense 32A4NR thirty two inch set shows 5.0 stars from four ratings. That is compatible with a very good product and equally compatible with an average one whose first four buyers were pleased. The Bosch SHE53B75UC dishwasher shows 3.9 from two ratings, which is one four star and one four star, or one five and one three. Neither of those numbers should move a decision by itself.

Somewhere above roughly thirty ratings the average becomes worth quoting, and above a few hundred it becomes stable enough to compare against a rival product with a similar count. Below thirty, treat the star figure as a placeholder and go and find the evidence somewhere else.

Why appliance listings have so few reviews

A low count on a major appliance is normal and is not a signal about the product. People buy a refrigerator once a decade and a television once every five or six years, so a listing accumulates ratings at a fraction of the rate a kitchen gadget does.

Related articles you may like:  Package Dimensions Versus Product Dimensions, and the Measurement That Catches People Out

The GE GNE27JYMFS french door refrigerator is a good illustration. It has been listed since 2020 and carries 37 ratings at 3.6 stars. Thirty seven ratings across five years is not a neglected product, it is the normal review velocity for a large appliance that most buyers order through a dealer rather than online. The 3.6 is worth noticing, but it rests on a sample small enough that a handful of delivery damage reports would produce it.

Contrast that with the two Zojirushi rice cookers in this catalog. The NS-ZCC10 listing carries 7 ratings while the induction model in the same lineup carries over four thousand. These are both long established machines from the same brand. The difference is not quality, it is which listing became the one buyers land on.

What a very large review count actually tells you

Large counts measure volume, not merit. The two highest counts in this catalog belong to a cold brew pitcher and a television wall bracket, not to any appliance. The Takeya cold brew coffee maker carries over sixty seven thousand ratings. The Mounting Dream MD2380 full motion mount carries over forty three thousand at 4.8 stars.

Both of those are inexpensive, frequently replaced items that a very large number of households buy. The count is telling you the product has been sold at enormous scale and survived it. That is genuinely useful, because at forty thousand ratings a serious design fault would show in the distribution. What it is not telling you is that the item is better than a competitor with four hundred ratings and the same average.

The comparison that works is between products of similar type and similar age. Comparing a bracket with forty thousand ratings against a dishwasher with two is not a comparison at all.

Ratings that belong to a different product

Merged variations

Amazon merges color, size and configuration variants into a single review pool. A rating on a 43 inch set counts toward the 75 inch set, and a rating on the black version counts toward the white. That is usually reasonable and occasionally not, because the large panel and the small panel in a television lineup often use different backlight designs, and a stand mixer bowl size can change the motor.

The practical step is to filter reviews to the variant you are buying, which is an option on the listing page, and see whether the average holds. If a lineup shows 4.5 overall and 3.6 for the size you want, the 4.5 was never about your unit.

Related articles you may like:  International Versions and Gray Imports: Warranty, Voltage and Support

Renewed and duplicate listings

Renewed items are separate listings with separate pools, so a strong average under a new machine does not carry over. This catalog also contains several cases where the same product exists twice under different listings, such as the pair of Ninja SP101 ovens in different colors, which carry 165 and 129 ratings respectively rather than one combined figure. If you find a low count, it is worth checking whether a second listing for the same model has the reviews.

Read the shape, not the mean

Once a count is high enough to be meaningful, the average is the least informative thing about it. Two products can both average 4.2 and mean opposite things. One made mostly of fours is a competent machine nobody loves. One made of a large block of fives with a hard tail of ones is usually a reliability story, and the ones will tell you what fails and how soon.

So open the one star and two star reviews first and look for repetition. A dozen different complaints is noise. Eight people describing the same seal leaking at month four is a specification. Then sort by most recent, because manufacturers make running changes and a fault reported three years ago may have been fixed, or a good product may have been cost reduced since.

Sales rank is a different measurement again

Listings carry a best sellers rank, and it is frequently mistaken for a quality score. It measures recent sales velocity within a category and nothing else. The GE refrigerator above is ranked 2,094th in Appliances and 366th within refrigerators while holding 3.6 stars from a small sample. Those two facts are not in tension, because rank is about how many people bought it and stars are about what those people thought.

Rank is useful for one thing: it tells you whether a product is currently in wide distribution, which is a decent proxy for whether spare parts and accessories will still be available in a few years. Treat it as an availability signal, not a recommendation.

A workable way to use ratings

Use the count first and the average second. Under about thirty ratings, ignore the star figure entirely and judge the product on its specifications, its brand’s service record and the manufacturer’s own documentation. Between thirty and a few hundred, treat the average as a rough band rather than a number. Above that, read the distribution and the recent reviews rather than the headline.

And keep the comparison honest. Compare a washer against other washers of similar age and a television against other sets in the QLED TVs group, not against an accessory that half the country owns.