YouTube Caption Cleaner & Sync
Fix SRT subtitle timing online
Move every cue earlier or later by the same amount, and find cues that overlap, end before they start or run out of order. SRT, VTT and SBV. The spoken words stay as they are.
1 SRT, VTT or SBV file up to 10 MB and 50,000 cues
Shift all cues
- + moves captions later, − moves them earlier
- Seconds × 1000: 1.2 s late = −1200
- −600,000 to +600,000 ms, in 10 ms steps
Add the matching video or audio next to the caption file. The shift field opens in the review step, so you can listen before every cue moves. The tool can also suggest the offset from the audio.
- Free: 3 jobs a day, no account
- Video or audio up to 180 min
- Failed jobs are never counted
Pro up to 200 jobs a day, €12.99/month or €89.99/year. See plans
What a +640 ms shift does
Example cues from a fictional file. Every start and end time moves by the same amount; the text is not touched.
| Cue | Before | After |
|---|---|---|
| 1 | 00:00:01,200 00:00:03,900 | 00:00:01,840 00:00:04,540 |
| 2 | 00:00:04,100 00:00:06,750 | 00:00:04,740 00:00:07,390 |
| 3 | 00:00:07,000 00:00:09,480 | 00:00:07,640 00:00:10,120 |
You download the shifted file as SRT, VTT or SBV, in any combination. Optional reports
Analysis draft kept 2 hours, result 24 hours. File security
How it works
- 1Add the caption file
SRT, WebVTT or SBV. For a shift, add the matching video or audio as well.
- 2Enter the offset
Choose Use a custom global shift and type the milliseconds. Tick other fixes only if you want them.
- 3Download
The shifted caption file in the formats you picked.
Optional extras: cleaning checks and reports Overlap, reading-speed and duplicate fixes, plus CSV, HTML and JSON files in the ZIP
Also in the ZIP
- Clean SRT, VTT or SBVAny combination of the three formats
- caption-issues.csvOriginal and final timing per cue
- caption-quality-report.htmlRemaining warnings explained
- processing-manifest.jsonSource hash, settings and versions
Cleaning fixes you can tick in the review step: trim overlaps, split long cues, merge adjacent duplicates, remove invalid markup, normalise encoding and spacing.
- Cues checked
- 486
- Timing issues
- 14
- Suggested shift
- +640 ms
Open the prepared example Fictional 486-cue file, runs in your browser, nothing is uploaded
Interactive synthetic caption review
See the exact cue before changing its timing.
This fictional subtitle file stays in the browser. Switch evidence and repairs to inspect a prepared result; no caption or media is uploaded.
Caption-only evidence shown. No wording is changed.
How the check works What is checked, what media adds, what is never changed
Caption repair without invented words
Start with the file that YouTube actually imports
Subtitle problems are often mechanical rather than linguistic: a timestamp ends before it starts, two cues overlap, an export uses broken encoding, or a long line becomes unreadable on a small screen. WoluTools parses SRT, WebVTT and SBV locally, preserves the original wording and builds a cue-level issue list. The baseline check works without the matching video and reports invalid timing, cue order, negative values, empty or duplicate cues, unsupported markup, long lines, short flashes and reading speed.
Every finding identifies the cue, original time range and reason. A severity label always includes text and an icon, so the review does not depend on colour. Large files are bounded at 10 MB and 50,000 cues, and a malformed structure is blocked rather than guessed into a different sentence.
Add media only when timing evidence matters
An optional video or audio file lets the worker verify total duration and detect non-silent audio activity with FFmpeg. It can reveal captions outside the media, a consistent global offset or a small bounded drift. This is deterministic activity evidence, not speech recognition: the tool does not know which words were spoken, cannot judge meaning and never claims frame-perfect lip sync.
A proposed shift or stretch remains a suggestion until the user listens and confirms it. Global offsets are bounded to ten minutes and drift scale to five percent in either direction. The report keeps the original measurement, suggestion and confirmed adjustment together, making a wrong choice visible instead of silently rewriting every cue.
Approve structural changes one by one
Safe normalization can standardize UTF-8, line endings, numbering and outer whitespace. Separate choices control invalid-markup removal, overlap trimming, long-cue splitting and adjacent duplicate merging. Splitting reuses the existing words; merging removes repeated adjacent text. Neither action invents punctuation, corrects spelling, translates captions or fills missing dialogue.
The final ZIP can include any confirmed combination of clean SRT, WebVTT and SBV files. A formula-protected issues CSV records original and final timing, the accessible HTML report explains remaining warnings, and the manifest binds source hashes, settings, rules and engine versions. The source is removed after processing and encrypted results expire after 24 hours.
Before repairing captions
YouTube Caption Cleaner & Sync FAQ
Exact answers about wording, timing evidence, formats, exports and deletion.
Which subtitle files are supported?
Version 1 accepts one SRT, WebVTT or SBV text file up to 10 MB and 50,000 cues. UTF-8 is preferred; Windows-1252 can be converted to UTF-8 with the conversion recorded.
Can I check captions without uploading a video?
Yes. Caption-only mode checks syntax, encoding, cue order, overlaps, duration, line length, reading speed, markup, duplicates and empty cues.
How does optional synchronisation work?
A video or audio file lets the worker compare cue boundaries with media duration and locally detected non-silent audio activity. It may suggest a bounded global shift or drift scale, but the user must listen and confirm it.
Does WoluTools transcribe or rewrite the spoken words?
No. It does not transcribe, translate or invent wording. Normalisation affects encoding, whitespace, numbering and syntax. Splitting or merging text blocks requires explicit confirmation.
Which timing changes require confirmation?
Trimming overlaps, splitting long cues, merging adjacent duplicates, shifting every cue and stretching timing drift are all visible choices bound to the reviewed source hash.
What reading-speed limits are used?
The report marks more than 20 characters per second for review and more than 30 as a blocker. These are transparent workflow thresholds, not a guarantee of accessibility or comprehension.
What is included in the download?
The ZIP contains the selected clean SRT, VTT and/or SBV files, a protected cue-level issues CSV, accessible HTML quality report and a manifest with hashes, settings, versions and remaining findings.
When are files deleted?
Encrypted analysis drafts and source parts expire after two hours and are removed after completion, failure or cancellation. The encrypted result expires after 24 hours or immediately when deleted.