E-commerce
Find duplicate products in a catalog file
Upload a product list as CSV or XLSX and get exact and near-duplicate groups with the reason for each match. Nothing is merged; you decide in a review table.
or drop it here
- CSV / XLSX
- Up to 50 rows
- 64,000 extracted characters maximum
SKU-101 | Blue Mug | 350 ml | EAN 000111 SKU-219 | Mug Blue | 350 ml | EAN 000111
- Candidate A: SKU-101
- Locked identifier: EAN 000111
Duplicate-candidate workbook, match-evidence CSV and human decision table
One clear job, from source to download
- 1
Add the source
Supported formats and limits are visible before the upload.
- 2
Confirm the settings
Review the exact source, options, units and access before processing.
- 3
Inspect and download
Check the preview and warnings, then unlock the complete package.
Finding duplicate products before a catalog cleanup
Your product list goes in
You upload a product list as CSV or XLSX. One run takes up to 50 rows and 64,000 extracted characters. The matcher looks at the identifiers and attributes already in your file, such as SKU, title, size and barcode. A shared EAN or GTIN counts as strong evidence but not as proof, because catalogs do reuse those codes by mistake. A file over the limit is rejected before processing rather than checked in part. The job costs 3 AI credits and needs an account.
Paired rows and a decision table
You get a workbook of duplicate candidates, a match-evidence CSV and a decision table. Each candidate shows both source rows in full, so you judge the pair on the real values and not on a score. The evidence CSV records which fields agreed and which did not. Every candidate ID is kept. Nothing is merged, deleted or rewritten, and your source file is never written to.
Sensitivity and variant locks
Strict keeps candidates close to exact identifier and attribute agreement. Balanced tolerates more wording variation in titles, so you see more pairs and reject more of them. Name the column that holds variant identity, such as size or colour, and rows that differ there are never proposed as the same product. Which rows to merge stays your decision, and you make it in the table, one pair at a time.
Questions before you run it
Will it merge or delete rows in my catalog?
No. The output is a decision table for a person to work through. Every candidate ID is preserved and your source file is never written to.
How do I stop it treating size or colour variants as duplicates?
Name the column that holds variant identity and those values are locked, so two rows differing there are not proposed as the same product.
What is the difference between Strict and Balanced sensitivity?
Strict keeps candidates close to exact identifier and attribute agreement. Balanced tolerates more wording variation in titles, which surfaces more pairs and also more that you will reject.
Does it use barcodes such as EAN or GTIN?
A shared identifier is strong evidence and is shown in the paired values for each candidate. It is not treated as proof on its own, because catalogs do reuse those codes by mistake.
Why does every candidate show both rows in full?
So the decision rests on the source values rather than on a score. The match-evidence CSV records which fields agreed and which did not.