Categories

AutomotiveBeauty & Personal CareElectronicsGrocery & Gourmet FoodHealth & HouseholdHome & KitchenIndustrial & ScientificMusical InstrumentsOffice ProductsPet SuppliesSports & OutdoorsTools & Home Improvement

View all categories

Methodology:
Trufl Score

The Trufl Score is a holistic quality rating for products on a scale from 40 to 100. It combines four independent signals – customer feedback (50 %), brand reputation (20 %), bestseller rank (20 %), and technical features (10 %) – into a comparable rating within the same product category.

What does the Trufl Score measure?

The Trufl Score answers one central question: "How good is this product compared with other products in the same category?"

Unlike simple star ratings, which can be misleading across different product categories, the Trufl Score ensures a fair comparison: a highly rated puzzle is not compared with a security camera, but only with other products in its category.

The score is built from four independent rating pillars, each capturing a different perspective on product quality. Combining multiple signals balances distortions from individual data sources and produces a robust overall rating.

The four pillars of the rating

Each pillar captures a distinct aspect of product quality. The weighting reflects the reliability and relevance of each data source.

Customer feedback 50 %

Reviews and ratings from real buyers – the most direct indicator of actual product quality in everyday use.

Brand reputation 20 %

The brand’s quality history within the category, plus public perception on social media.

Bestseller rank 20 %

Market validation through real purchase decisions – the “wisdom of the crowd” as an objective reality check.

Technical features 10 %

Objective analysis of category-specific product attributes such as material quality, feature set, and build quality.

Pillar 1: Customer feedback

Customer feedback is the core of the Trufl Score and has the strongest influence on the overall result. It captures real buyer experience and is composed of multiple signals for a nuanced picture.

Multi-signal approach

Instead of relying only on the average star rating, the Trufl Score analyses several dimensions of customer feedback:

  • Overall ratings: The average star rating from platforms such as Amazon and Google Shopping, normalised to a common scale. Ratings from other platforms are adjusted by category because they tend to be higher on average than on Amazon. This signal has the greatest weight within the customer-feedback pillar.
  • Review sentiment: Beyond star counts alone, we analyse what customers actually write. Only product-related experience is included – feedback about delivery, packaging, or seller service is filtered out so logistics issues do not distort the score.
  • Price perception: How do customers rate value for money? Products that buyers mostly see as overpriced receive a small deduction.
  • Malfunction reports: Systematic quality issues matter most: when many customers independently report the same defect, that feeds into a score deduction. Isolated one-off complaints are not overweighted.

Correction for other-platform ratings

Ratings from platforms other than Amazon (e.g. Google Shopping) are not taken at face value. For each category we calculate a platform offset from overlapping products and correct for systematic rating differences. That keeps comparisons fair – regardless of where the rating originated.

Expert-review bonus

Verified professional test results can further strengthen the customer-feedback pillar: products that score especially well in independent expert reviews receive a moderate bonus of up to 4.5 % within this pillar.

Statistical normalisation

Different product categories naturally have very different rating levels. The Trufl Score corrects those distortions for customer feedback, brand reputation, and bestseller rank so only products within the same category are compared fairly. Technical feature analysis is the exception: it is deliberately not normalised relative to the category, but scored on absolute capability.

Pillar 2: Brand reputation

Brand reputation captures a manufacturer’s quality history within a given product category. It has two components:

Review-based brand score

How have all of a brand’s products performed within the category? Brands that consistently deliver high quality receive a higher score. Products with more reviews are weighted more heavily – a single niche product influences the brand score less than a bestseller with thousands of reviews.

Social sentiment analysis

We also analyse public perception of the brand on social platforms such as Reddit and Twitter. We capture not only tone (positive vs. negative), but also the quality and relevance of discussions. Product-related criticism counts more than general complaints about shipping or a website.

Brand-position bonus

Leading brands in a category receive an additional bonus based on their overall performance within that product group.

Pillar 3: Bestseller rank

Bestseller rank reflects the “wisdom of the crowd”: which products do consumers actually buy most often? The data foundation is the Amazon Sales Rank.

1

Collection

Amazon Sales Rank is collected for every product in the category

2

Cleanup

Misclassified products and accessories are filtered out of the bestseller list

3

Percentile calculation

Rank is converted into a percentile relative to the cleaned category size

4

Score integration

The result feeds into the overall score as market validation

Bestseller proxy for other-platform products

Not every product has an Amazon Sales Rank. For products sold mainly on other platforms, we calculate a bestseller proxy from rating activity on that platform. The value is adjusted by category and weighted conservatively so market relevance outside Amazon is still represented fairly.

Category cleanup

Amazon bestseller lists often include misclassified products (e.g. computer games in a baby-monitor category). Those outliers are removed systematically so only relevant products influence the rating.

This signal provides an objective reality check: products that have proven themselves in the market receive corresponding recognition in the score. A category winner among 10,000 competitors receives the highest possible value.

Pillar 4: Technical feature analysis

Every product category has specific quality attributes that can be assessed objectively. Relevant properties are extracted with AI from technical datasheets and product descriptions, then scored:

1

AI-assisted feature extraction

For each product category, relevant product attributes are identified automatically from technical data and descriptions and stored in the database.

2

Weighting

The identified attributes are ranked and weighted by their relevance to customers and product quality.

3

Scoring

Each product is scored on its attributes to identify the product with the strongest features in the category.

Absolute rather than relative scoring

Unlike the other pillars, technical features are not scored via relative category normalisation. Actual capability is used instead – so simple products are not artificially pushed up or down.

Fairness principle

If ratings are not yet available for individual attributes, a neutral baseline is used. Products are not penalised simply because data is missing for some aspects.

Data quality and minimum requirements

To keep the Trufl Score meaningful, strict quality requirements apply to the underlying data:

Bayesian smoothing

Products with few reviews could receive extreme scores by chance alone. The Trufl Score uses statistical smoothing that gently pulls new-product scores toward the category average. As more data accumulates, the score is driven increasingly by real feedback.

Minimum data basis

A product must have at least one real signal (e.g. customer reviews or bestseller rank) to receive a Trufl Score. Products with no data basis are excluded from scoring – preventing empty records from distorting the overall distribution.

Scoring and recognition

The Trufl Score is calculated on a scale from 40 to 100 points. The weighted results of all four pillars are combined for each product and normalised within the category. We then apply an S-curve (sigmoid) transform so similar top products are easier to tell apart.

40
55
70
85
100

The higher a product’s Trufl Score, the better it performs versus other products in the same category. The rating is always relative to the category – a score of 85 for headphones is not directly comparable with a score of 85 for suitcases.

The rating pillars at a glance

Customer feedback

Analysis of product ratings, product-focused review sentiment, price perception, and malfunction reports – including platform correction and an optional expert-review bonus.

Brand reputation

Review-based brand quality history within the category, plus social-sentiment analysis from relevant discussions on platforms such as Reddit and Twitter.

Bestseller rank

Market validation via Amazon Sales Rank and a conservative bestseller proxy for products on other platforms.

Technical features

AI-assisted extraction and objective scoring of category-specific product attributes based on absolute capability.

Studies and ratings are conducted independently and published editorially. The Trufl Score is calculated fully from data, without influence from manufacturers or retailers. Positively rated providers may, after the study is complete, obtain a licence to use test seals for advertising. Licensing has no influence on the methodology or the results.