How We Test Resale Price-Check Apps
Every ranking on this site comes out of the same question: when a reseller points a phone at an unlabeled item at a garage sale, which tool gives the truest answer, fastest? That's harder to measure than it sounds — "accuracy" for a pricing tool has at least five moving parts, and a tool can be excellent at one and useless at another. This page lays out exactly what we measure, how we score it, and where our testing has limits, so you can judge our rankings for yourself.
The five things we measure
A resale price-check tool has one job — turn a photo of an item into a trustworthy number — but that job breaks into five distinct skills. We score each one separately, because the tool that wins on speed is rarely the one that wins on depth, and a reseller deciding buy-or-pass in ten seconds needs different things than someone valuing an estate.
- Identification accuracy. Did it correctly work out what the item is — brand, model, and the specifics that move the price — from the photo alone? Everything downstream is worthless if this step is wrong, because a confident price for the wrong item is worse than no price at all.
- Real sold prices, not asking prices. Does the number come from listings that actually sold, or from what sellers are hoping to get? This is the single biggest divider between tools, and we explain it in its own section below.
- Condition awareness. A mint example and a scuffed one are different items at the same title. Does the tool account for condition, or does it hand you one blended median and let you overpay for the beat-up version?
- Speed at the table. A price that arrives after you've already had to decide is a price you didn't use. We measure how long the full answer takes on a phone, on real conditions, not a demo.
- Honesty on the misses. Every tool fails on some items — a one-of-a-kind piece, a bad photo, a thin market. The tools we rank highest are the ones that say so instead of inventing a confident number from three stray listings.
What "identification accuracy" actually means — a worked example
Most "we tested it" claims are a reviewer eyeballing a few scans and forming an impression. That's not repeatable, and it quietly favors whichever tool the reviewer already likes. The fix is a fixed set of items whose correct identity is written down in advance — a human verifies the true brand, the true core identity, and the defining specifics of each item before any tool sees it. Then every tool is graded against that same answer key, and the same photo scores the same way every time.
Here is a concrete version of that, from MarketplaceIQ's own internal evaluation (disclosed). The set is 17 items with human-verified identities — a deliberate mix of the things that break photo identification: costume jewelry with tiny hallmarks, a Bakelite bangle, branded apparel, a plush toy, kitchenware. The interesting result isn't the headline score; it's what happened when a single recognition step was removed.
MarketplaceIQ identifies an item with three independent recognition passes that have to agree — a text-and-logo reader that pulls the maker's marks and printed model numbers off the photo, a reverse-image search that finds visually matching products, and a reasoning step that cross-checks the two and resolves conflicts. To test how much that cross-checking matters, the evaluation re-ran the same 17 items with one of those passes switched off:
| Identification score | All three passes | One pass removed |
|---|---|---|
| Correct brand | 100% | 79% |
| Correct core identity | 100% | 86% |
Internal MarketplaceIQ evaluation, 17 human-verified items, photos included. "Core identity" = the right item even if a brand nuance is missed. Dropping one recognition pass cost 21 points of brand accuracy — the items that flipped were the ones where a small printed mark, not the overall shape, was the whole answer.
That 21-point drop is the case for weighting cross-checked, multi-signal identification heavily in our rankings. A tool leaning on a single recognition method can look fine on obvious items and quietly fail on exactly the pieces where identification matters most — the marked, the branded, the easily-confused-with-a-reproduction. It's also why we don't take a tool's own "confidence" score at face value; we check whether it got the item right against the answer key, not whether it felt sure.
Why we only trust sold prices
This is the criterion that reorders every ranking, so it gets said plainly. An asking price is what a seller hopes to get. A sold price is what a buyer actually paid. Active-listing pages are full of items priced with heroic optimism that will sit unsold for months, so any tool that averages active listings systematically overstates what an item is worth — and a reseller who trusts that number overpays at the table.
So a tool only gets credit from us for pricing built on real, completed sold-listing data. And even then we read the median as a floor for the common version of an item, not a verdict on the specific piece in hand — because a category median is the average of the forgettable examples, and the money is usually in telling the forgettable one from the score. Tools that surface the individual sold comps (so you can see the spread and find the ones that match your item's condition) rank above tools that hand you one blended number with no way to check it. If you want the reseller-facing version of this argument, our guide to checking eBay sold prices walks through it with examples.
How we check that the price matches the item
Getting the identity right and pulling sold data are two steps, and the join between them is its own failure point: a tool can identify an item perfectly and then search for the wrong comps. We spot-check that join on a batch of real items — does the tool's search actually return listings for this item, and does it broaden sensibly when the exact match is too thin to price rather than returning an empty shrug or, worse, a confident number off three unrelated listings? A tool that reliably finds 50+ genuine comps for a common item but knows to widen the net (and to say the data is thin) on an obscure one is doing the hard part right. This is where "honesty on the misses" and "real sold prices" meet.
Where our testing has limits
A methodology page that only lists strengths isn't a methodology page. Here's what our testing does not settle:
- Sample sizes are modest. A 17-item accuracy set and batches of a few dozen priced items are enough to expose real differences — like a 21-point accuracy swing — but they are not a census. We treat them as strong signal, not proof to three decimal places, and we date them.
- We publish our own tool's internal numbers. Where a figure comes from MarketplaceIQ's evaluation rather than a competitor's, we can describe the method but we can't independently run the competitor's internals the same way. We say which is which, every time.
- Markets move. Sold-price data has a rolling window, categories heat up and cool off, and tools ship changes. A ranking is a snapshot; the dated ones on this site link to always-current sources where they exist.
- "Best" is per job. Our whole premise is that there's rarely one overall winner — the fastest field app isn't the deepest archive. We score the skills separately and match each tool to the moment it's best at.
How this turns into a ranking
Put together, the five criteria explain why our guides read the way they do. A tool wins the field-speed pick by being fast and condition-aware; it wins the "what's it worth" pick by being accurate on identification for a non-expert; it loses the multi-platform row if it's eBay-anchored, and it loses the collectibles-history row to a deep archive. No single tool sweeps every row, which is exactly why our three-way comparison declares a winner per row instead of an overall champion, and why the best-apps guide ranks by job. If you want to see the criteria applied end to end, start there.
See the method in action
The accuracy and sold-price criteria above are exactly what MarketplaceIQ is built around — triple-checked identification and real eBay sold comps from one photo. Try it free and judge the read yourself. Free tier plus a 14-day Plus trial, no credit card.
Scan an item to check its value →