How we test and score

We would rather tell you exactly what we did than imply we did more. This page is the method behind every guide and review on Meln.

What we check ourselves

Pricing, on the vendor's own page, on the date shown at the top of each article. Free-tier limits, by signing up where that is possible. Feature availability per plan. Cancellation and refund terms, from the vendor's published policy. Affiliate program terms, because they often reveal how a company thinks about its own marketing (several humanizer companies, for example, forbid affiliates from calling the product a "bypass"). Company registration where it is public. Where we could not verify something, we say so in the text rather than filling the gap.

What we take from others, and how we label it

Accuracy and performance figures come from independent sources: academic benchmarks such as RAID, published comparative studies such as Scribbr's detector test, peer-reviewed work such as the Patterns study on non-native English writers, regulator notices, court records, and third-party reviewers who published their method. Every figure is attributed. When a vendor's own number is the only one available, we label it a claim. When two sources disagree, we show both. When a figure is repeated widely but cannot be traced to a primary source, we leave it out and say that we did.

What we read from users

Trustpilot, G2, Capterra, app stores and the relevant communities, read by theme rather than by star average. We care about the last twelve months more than the lifetime figure, we note when a company replies to negative reviews and when it does not, and we treat billing complaints as part of the product.

The scoring rubric

Standalone reviews carry a score out of five built from six category scores, which are shown on the page. The categories differ by product type (a detector is scored on false positives, a companion app on memory, a trading bot on security record) but the weighting principle is constant: we weight the dimension most likely to hurt a buyer more heavily than the dimension most likely to impress one. For detectors that means false positives over raw detection. For companion apps it means real monthly cost over feature count. For trading bots it means custody and breach history over bot variety.

ScoreWhat it means
4.5 to 5Buy it if you are in the target group; we found no significant problem
4 to 4.4Recommended for the target group with a caveat you should read
3.5 to 3.9Good at its core job, with a real weakness in price, trust or quality
3 to 3.4Usable, but a competitor probably serves you better
Under 3We would not buy it, and we say why

Buyer's guides

Comparison guides rank tools against each other on the same evidence, with a quick-picks list at the top labelled by use case and, where we think it is warranted, a tool we would skip. We include that skip pick even when the tool pays a commission. The ranking is the editor's judgement based on the factors above; it is not a formula and we do not pretend it is.

Re-checking

Prices in these categories change constantly. We re-verify pricing on each page on a rolling schedule and update the "pricing checked" date when we do. If a product changes materially (an acquisition, a breach, a pricing restructure), we update the review and note the change.

What we do not do

We do not run secret benchmarks and report a single pass rate as if it were definitive; the variance between runs in these categories makes that misleading. We do not accept payment for placement. We do not let affiliate relationships affect rankings. We do not fabricate hands-on anecdotes. And we do not make recommendations about academic integrity, financial decisions or personal relationships; we review software and we are clear about the limits of that.