Methodology v1.0 · public

How a fake review checker should work: in the open.

Every HonestStars grade is the sum of eight public signals. The weights, the grade bands and the rules for when we refuse to grade are all on this page. If the method changes, the version number changes and the change is logged below.

What a grade means

A grade describes how trustworthy a listing's reviews look, on a 0 to 100 scale. It is a statistical judgement about patterns in review data. It is never a statement about the seller, the brand, or the product itself, and it is not a claim that any individual review is untrue.

The eight signals

Weights add up to 100. Each signal produces a sub-score; the grade is the weighted sum, then the ceiling below is applied. The three signals that moved the score most are the three reasons you see.

Signals and weights in methodology v1.0
SignalWeightWhat we look at
S1 Review velocity spikes20%We compare how many reviews arrive each week with the product’s own history. A sudden burst, such as a large share of 5★ reviews landing within days, lowers confidence in the grade.
S2 Rating shape15%Real products tend to collect a spread of opinions. A rating distribution with a wall of 5★, a clump of 1★ and almost nothing between is unusual, and we weigh it.
S3 Verified-purchase ratio10%We look at the share of high-star reviews that are marked as verified purchases compared with the rest of the listing.
S4 Copied phrasing15%We look for groups of reviews that share long runs of the same wording. We show matches so you can open and read them yourself.
S5 AI-written likelihood10%Some reviews read as machine-generated. This is a likelihood, never a certainty, so it carries a modest weight.
S6 Reviewer depth10%We look at how much history the reviewers behind the 5★ ratings have, for example profiles with a single review.
S7 Stars vs words10%A 5★ rating next to text that describes a defect is a mismatch worth noting.
S8 Listing hygiene10%Reviews that talk about a different item than the one on the page can come from merged or reused listings.

The severe-signal ceiling

Every grade starts as a weighted sum of up to eight checks, so a single weak check can only move the score by its own weight. That would let a listing with one glaring problem, such as sixty near-identical reviews, still look "mostly genuine". So we add a ceiling: when a check has enough reviews to judge and the evidence is strong (for example a burst week holding a quarter of all reviews, or dozens of copied texts), the score is capped at C; two such checks, or one plus another concern, cap it at D; three cap it at F. The ceiling never raises a score, is never triggered by our low-trust AI-text check on its own, and when it applies we say so in the first reason, in plain words. It grades the reviews we could read on that day, not the seller.

Grade ceilings when the severe-signal rule applies
When it appliesHighest possible scoreWorst grade
1 severe check64C
2 severe checks, or 1 plus another concern54D
3 or more39F

Only S1, S4, S6, S7, S8 can trigger the ceiling. Rating shape, verified ratio and AI-written text never do, because they have honest explanations or carry less trust.

Grade bands

A grade always shows a letter and a word, never colour alone.

Grade bands
GradeScoreWord
A85–100Genuine
B70–84Mostly genuine
C55–69Mixed
D40–54Doubtful
F0–39Unreliable
Not graded< 5 reviewsNot graded

Confidence and when we refuse to grade

  • Fewer than 5 reviews: not graded. We won't guess.
  • 5 to 19 reviews: graded, but the grade is a rough guide. Confidence is always low, a single odd review counts for less than it would in a bigger sample, and the best possible grade is B. We say how many reviews we could read, and the first reason says why the grade stops at B.
  • Fewer than 50 reviews read: the grade is shown with low confidence and says how many reviews we could read.
  • Missing data: some pages, such as independent Shopify stores, lack dates, verified flags or reviewer profiles. We re-weight over the signals we can compute and lower confidence. If fewer than three signals can be computed, we don't grade.
  • Every verdict shows the sample size and the method version it was graded with.

Adjusted rating

Alongside the grade we show the listed rating and an adjusted rating with suspect reviews removed, for example "4.6★ listed → 4.5★ real". The suspect reviews are the ones flagged by the signals above, and you can open them on the store to read them yourself.

What we never do

  • We never call a seller or product "fake" or a "scam". We grade reviews.
  • We never let a brand, seller or advertiser pay to change a grade.
  • We never use affiliate links or earn from purchases.

Known limits

No automated method is perfect. Signals can fire on honest listings, for example after a legitimate viral moment, and can miss careful manipulation. That is why we show the reasons and the evidence, so you can judge for yourself. In particular:

  • Brushed orders can pass. When someone really buys the item and then posts a glowing review, the account is verified, aged and the dates look natural. Only a few signals can catch this, so an A does not mean there are no paid reviews.
  • Slow, managed campaigns can pass the timing check, and reviews that are paraphrased by a person or a tool can pass the similarity check.
  • We read at most about 250 reviews, even when a listing has thousands. On most stores these are the newest ones, so a grade leans toward recent activity.
  • Text signals work in English only. Listings in other languages are graded on the rating, timing, reviewer and verification signals, with lower coverage.
  • Small listings are a rough guide. With 5 to 19 reviews, chance can make a listing look better or worse than it is, so we measure each check against a 20-review sample, show low confidence and stop the grade at B.
  • Generic pages are often low confidence. Where a page gives us only a few of the signals, we say so and show Low confidence at best.
  • Detecting AI-written reviews is unreliable. That signal is weak by design, so it carries a modest weight and can never trigger the severe-signal ceiling on its own.
  • A grade describes the reviews we could read on a given date. It is not a statement about the seller.

If you believe a grade contains a factual error, use the grade corrections form.

Version history

  • v1.0: initial public method with eight signals and the severe-signal ceiling.