CSV Cleaner & Validator
Clean and Validate CSV or TSV Files: find broken CSV rows before they break your import.
Find structural and value errors, approve safe cleanup and retain a line-by-line issue trail.
- Maximum
- 50 MB · 100,000 rows · 250 columns
- File handling
- Draft 2 hours · result 24 hours
Larger files need Pro. Failed jobs are never counted.
- Prepared example is free
- 3 free jobs a day
- Cancel anytime
- Rows checked
- 18,420
- Issues
- 37
- Safe changes
- 12
Formats, limits & file handling
- Works with
- CSV, TSV
- Limit
- 50 MB · 100,000 rows · 250 columns
- You receive
- Clean CSV, protected CSV, XLSX profile and issues report
- File handling
- Draft 2 hours · result 24 hours
Upload one CSV or TSV, see the detected format immediately and get a useful baseline check without building a schema first. Approve every content-changing cleanup.
THREE CLEAR STEPS
From your files to a usable result.
- 01Add your input
Limits and supported formats are visible before you begin.
- 02Review the result
Check the preview, findings or artwork before you download.
- 03Download
Use your free job, plan or credits. Failed jobs are never counted.
Interactive synthetic data check
See the exact row before approving a fix.
The sample stays in this browser. Toggle safe changes and filter the prepared findings.
Data quality without guesswork
Start with a useful automatic check
A CSV import often fails before anyone knows whether the problem is the delimiter, quoting, headers or a value deep in the file. WoluTools first recognises UTF-8, UTF-8 with BOM or Windows-1252, then detects comma, semicolon, tab or pipe structure across logical records and shows the result before processing. A preview makes the chosen header row and columns visible. The baseline audit needs no schema and checks malformed structure, inconsistent field counts, empty or duplicate headers, blank rows, outside whitespace, duplicate records, control characters and text that spreadsheet software might execute.
Files are processed as streams up to 100,000 rows and 250 columns, with a 1 MB limit for each field, instead of loading the complete table into memory. A structurally ambiguous row is blocked and retained in the issue trail with its exact source row; the worker never guesses where a missing quote or delimiter belongs.
Content changes require a deliberate choice
Technical output normalisation—UTF-8, consistent line endings and RFC-style CSV quoting—does not change field meaning. Removing outer whitespace, blank rows or duplicate records does, so each action appears in the final confirmation before the job starts. Deduplication also needs a confirmed key and a choice to keep the first or last occurrence.
Optional column rules cover required values, text, integer, decimal, date, Boolean, email and URL types, minimum and maximum values, length, uniqueness, allowed values and a bounded regular expression. Dates and decimal separators stay explicit. The tool does not infer an uncertain regional date, create a missing value or transform a leading-zero identifier into a number.
Keep a verifiable trail after cleanup
The main cleaned CSV contains only the transformations you approved. A separate spreadsheet review copy prefixes recognized formula starters to reduce execution risk; no CSV mitigation is universal, so review it and keep it separate from machine imports. The ZIP adds an issue CSV, XLSX profile, reusable validation schema, README and a manifest with source hash, settings, counts and a SHA-256 inventory of every artifact. Completion and billing happen only after the ZIP passes CRC, exact-member, byte-size and hash verification.
Before cleaning data
CSV Cleaner & Validator FAQ
Specific answers about schema-free checks, safe changes, formulas and output.
Can I run a useful check without creating a schema?
Yes. The baseline check detects delimiter and encoding, malformed structure, uneven fields, empty or duplicate headers, empty rows, whitespace, duplicate rows, control characters and spreadsheet-formula risks.
Which CSV formats are accepted?
Version 2 accepts CSV and TSV up to 50 MB, 100,000 data rows, 250 columns and 1 MB per field. UTF-8, UTF-8 with BOM and Windows-1252 are supported; comma, semicolon, tab and pipe are detected across logical records.
Does the cleaner guess dates or missing values?
No. It never invents missing content and never converts an ambiguous date. Date, decimal and Boolean normalisation requires a visible column rule and a confirmed cleanup choice.
What happens to a row with the wrong number of fields?
The row is blocked and recorded with its source row number. The tool does not guess where a missing delimiter or quotation mark belongs.
How are spreadsheet formulas handled?
The ordinary cleaned CSV preserves approved text. A separate spreadsheet review copy prefixes recognized formula starters to reduce execution risk. No CSV mitigation is universal, so review this copy and keep it separate from machine imports. Issue exports use the same protection.
Can duplicate rows be removed?
Yes, only after you enable deduplication and confirm the key plus whether the first or last occurrence is kept. Without that confirmation, duplicates remain findings.
What is included in the download?
The verified ZIP contains cleaned.csv, a risk-reduced spreadsheet review copy, issues.csv, data-profile.xlsx, validation-schema.json, README.txt, cleaning-manifest.json and blocked-rows.csv when structural rows were excluded. The manifest inventories every artifact by byte size and SHA-256.
Is this a Google Merchant feed validator?
No. It checks general delimited-data quality. Google-specific product requirements remain in the separate Google Merchant Feed Validator.
Available now
Inspect the issue trail before sharing a file.
The synthetic demo is open. Your own files run here without an account: 3 free jobs a day, or up to 200 a day with a Pro plan.