10 October 2026

4.8 Stars From 200 Reviews or 4.4 From 20,000? How We Weigh Them

Two listings sit side by side. One is rated 4.8 stars by 200 buyers. The other is rated 4.4 by 20,000. Which is the safer purchase?

Most people feel the answer is "the second one, probably" without being able to say by how much. Our ranking has to put a number on it, so this post shows the number.

The two signals

Every listing we rank gets a score built from five signals. Two of them come from buyers:

  • Star rating is worth 40% of the score. A 5.0 earns all 40 points, a 4.0 earns 32, and so on in a straight line.
  • Review volume is worth 20%. It is not a straight line. It grows with the logarithm of the review count and reaches its maximum at 1,000 reviews.

That second rule is the interesting one. Here is what it gives a listing for its review count alone, out of 20 points:

Reviews Points out of 20
3 4.0
10 6.9
100 13.4
200 15.4
1,000 20.0
20,000 20.0

Going from 10 reviews to 100 nearly doubles the points. Going from 1,000 to 20,000 adds nothing. That is deliberate: the hundredth review tells you far more than the ten-thousandth does. Once a thousand people have rated something, a few thousand more rarely move the average.

Working the example

Add the two signals together for each listing:

Listing Rating points (of 40) Review points (of 20) Total (of 60)
4.8 stars, 200 reviews 38.4 15.4 53.8
4.4 stars, 20,000 reviews 35.2 20.0 55.2

The 4.4 wins, by 1.4 points out of 60. That is a narrow margin, and it should be. A 0.4-star gap is large; it is nearly cancelled by the fact that the lower rating has a hundred times the evidence behind it.

Change the example slightly and the result flips. A 4.8 from 1,000 reviews scores 58.4 and beats the 4.4 comfortably, because it no longer has an evidence gap to make up. A 5.0 from 3 reviews scores 44.0 and loses to both: a perfect rating from three people is close to no information.

What the other 40% does

Buyer signals are 60% of the score. The rest is price (25%), whether the listing is in stock and how recently we checked (10%), and a small fixed weight for the store (5%). So in a real ranking the 1.4-point margin above is easily overturned by price. If the 4.8-star listing is meaningfully cheaper, it will come out on top, and the product page will list the lower price among its reasons.

What this cannot tell you

These limits are real and worth stating.

We do not read the reviews. The score uses the star average and the count that the store publishes. It does not know what the reviews say, and it cannot tell a product with a known defect from one without, if both average 4.4.

We cannot detect manipulated reviews. If a rating has been inflated, our score inherits the inflation. A large review count makes manipulation harder but not impossible.

Counts are sometimes shared. Stores often pool reviews across colors and sizes of one product, so a rarely bought variant can display the review count of the popular one.

The cap is a choice. Stopping at 1,000 reviews means we treat 1,000 and 50,000 as equally trustworthy. You could argue for a higher ceiling. We chose a point past which, in our judgment, more reviews stop changing the picture.

A ranking like this is good at one thing: sorting many listings quickly using evidence that would take you an hour to collect by hand. It is not a substitute for reading the negative reviews of the one you are about to buy. Do that part yourself. It takes five minutes and it is where the specific problems show up.

The full formula, including the stock-freshness thresholds, is on the How we pick products page.

Written by Ajay Kanani, who built PickBestProducts. See how we rank products.