Are "User Reviews" Reliable?

Nowadays, almost every e-commerce site (Amazon, Taobao, etc.) offers a "user review" feature, intended to let buyers determine the quality of a product. On the surface, this seems to give the public a sense of fairness and openness — but is that really the case? An article titled "Are User Reviews Reliable?" in this year's eighth issue of Scientific American (Chinese edition) discusses how relying solely on "user reviews" to judge a product can be unfair. Now a trial begins: the plaintiff is "User Reviews," the defendant is the article from Scientific American, and the judge is mathematics.

Screenshot of Taobao user reviewsScreenshot of Taobao user reviews

The trial begins... "User Reviews" insists that what it shows reflects reality, while Scientific American argues there's something amiss. What will the verdict be? more

Let's assume the following scenario:

1. Ratings are divided into three levels, denoted 1, 2, and 3, with 3 being the best.
2. Readers can be divided into buyers and non-buyers, and we assume each group makes up half. As is the case on most current websites, only buyers have the right to leave a review.
3. Buyers can be split into two categories (each accounting for 50%):
1) Those who bought the book after "understanding" it (here "understanding" means having some knowledge of the book's actual content before purchase, not merely browsing the table of contents on the website). In this case, most readers tend to like the book, so we assume this group rates it 3.
2) Those who bought the book without understanding it. In this case, readers might really like it, feel indifferent, or dislike it, with all three outcomes equally likely.**
4. Since non-buyers have no reviewing rights on the website, we do not further subdivide this group.

Under these assumptions, suppose 12 people review a book, giving ratings of 1, 2, and 3 in equal numbers. The average score for the book should then be 2. This is the result obtained through an actual, objective survey — not through "user reviews."

Now suppose that among these 12 people, only 6 actually purchased the book. For the 3 people who are "buyer type 1," their total score is 3×3=9. For the other 3 people who are "buyer type 2," their total score is 1×1+1×2+1×3=6. The combined total is 15, so the average is 15÷6=2.5, higher than the actual value of 2.

The discussion above only assumes a very simple scenario, but it is still representative. First, it's unlikely that buyers and non-buyers really split evenly; for most books, buyers are probably fewer than non-buyers, so this assumption is already quite conservative (in other words, favorable to "User Reviews" winning the case). Second, considering real-world situations — if you've ever shopped online (especially bought books), you'll know that "user reviews" typically happen right after receiving the goods: if the item isn't damaged, people casually give the seller a positive review (since, to us, this feels like a trivial matter), rather than leaving a review only after careful "evaluation" over time. So the two buyer categories in our thought experiment are likewise fairly conservative estimates (again favoring "User Reviews"), since in reality buyers' ratings tend to skew even higher — and some websites even default to a positive review if the buyer doesn't leave one at all.

And yet, even after stacking the deck in favor of "User Reviews" with all these "favorable pieces of evidence," the final verdict remains: "User Reviews" loses the case. Although 2.5 doesn't seem all that much bigger than 2, once we scale this up and factor in real-world conditions, we find that the deviation from the truth ends up being quite substantial. So perhaps an old piece of advice still holds true for consumers: shop carefully!

English translation of a post from 科学空间 | Scientific Spaces by 苏剑林. Original: https://kexue.fm/archives/894
Translated automatically with claude-sonnet-5; all equations are reproduced verbatim from the source. Copyright remains with the original author.