WoluTools

YouTube Caption Cleaner & Sync

Fix SRT subtitle timing online

Move every cue earlier or later by the same amount, and find cues that overlap, end before they start or run out of order. SRT, VTT and SBV. The spoken words stay as they are.

Choose caption file

1 SRT, VTT or SBV file up to 10 MB and 50,000 cues

Shift all cues

  • + moves captions later, − moves them earlier
  • Seconds × 1000: 1.2 s late = −1200
  • −600,000 to +600,000 ms, in 10 ms steps

Add the matching video or audio next to the caption file. The shift field opens in the review step, so you can listen before every cue moves. The tool can also suggest the offset from the audio.

  • Free: 3 jobs a day, no account
  • Video or audio up to 180 min
  • Failed jobs are never counted

Pro up to 200 jobs a day, €12.99/month or €89.99/year. See plans

What a +640 ms shift does

Example cues from a fictional file. Every start and end time moves by the same amount; the text is not touched.

Cue times before and after a +640 ms shift
CueBeforeAfter
100:00:01,200
00:00:03,900
00:00:01,840
00:00:04,540
200:00:04,100
00:00:06,750
00:00:04,740
00:00:07,390
300:00:07,000
00:00:09,480
00:00:07,640
00:00:10,120

You download the shifted file as SRT, VTT or SBV, in any combination. Optional reports

Analysis draft kept 2 hours, result 24 hours. File security

How it works

  1. 1
    Add the caption file

    SRT, WebVTT or SBV. For a shift, add the matching video or audio as well.

  2. 2
    Enter the offset

    Choose Use a custom global shift and type the milliseconds. Tick other fixes only if you want them.

  3. 3
    Download

    The shifted caption file in the formats you picked.

Optional extras: cleaning checks and reports Overlap, reading-speed and duplicate fixes, plus CSV, HTML and JSON files in the ZIP

Also in the ZIP

  • Clean SRT, VTT or SBVAny combination of the three formats
  • caption-issues.csvOriginal and final timing per cue
  • caption-quality-report.htmlRemaining warnings explained
  • processing-manifest.jsonSource hash, settings and versions

Cleaning fixes you can tick in the review step: trim overlaps, split long cues, merge adjacent duplicates, remove invalid markup, normalise encoding and spacing.

EXAMPLE RESULTSample data
Cues checked · 486
Cue overlap4
Reading speed above 20 cps7
Confirmed global shift+640 ms
Clean SRT · VTT · issues report
Sample report for a 486-cue fileEvery finding names the cue, its time range and the reason
Cues checked
486
Timing issues
14
Suggested shift
+640 ms
Open the prepared example Fictional 486-cue file, runs in your browser, nothing is uploaded

Interactive synthetic caption review

See the exact cue before changing its timing.

This fictional subtitle file stays in the browser. Switch evidence and repairs to inspect a prepared result; no caption or media is uploaded.

✓
No public uploadPrepared 486-cue example
Cues checked486
Timing errors4
Reading warnings7
Suggested shift—
CueTimingCaption wordingStatus
Prepared caption packageclean-captions.srt · clean-captions.vttcaption-issues.csv · caption-quality-report.htmlprocessing-manifest.json

Caption-only evidence shown. No wording is changed.

How the check works What is checked, what media adds, what is never changed

Caption repair without invented words

Start with the file that YouTube actually imports

Subtitle problems are often mechanical rather than linguistic: a timestamp ends before it starts, two cues overlap, an export uses broken encoding, or a long line becomes unreadable on a small screen. WoluTools parses SRT, WebVTT and SBV locally, preserves the original wording and builds a cue-level issue list. The baseline check works without the matching video and reports invalid timing, cue order, negative values, empty or duplicate cues, unsupported markup, long lines, short flashes and reading speed.

Every finding identifies the cue, original time range and reason. A severity label always includes text and an icon, so the review does not depend on colour. Large files are bounded at 10 MB and 50,000 cues, and a malformed structure is blocked rather than guessed into a different sentence.

Add media only when timing evidence matters

An optional video or audio file lets the worker verify total duration and detect non-silent audio activity with FFmpeg. It can reveal captions outside the media, a consistent global offset or a small bounded drift. This is deterministic activity evidence, not speech recognition: the tool does not know which words were spoken, cannot judge meaning and never claims frame-perfect lip sync.

A proposed shift or stretch remains a suggestion until the user listens and confirms it. Global offsets are bounded to ten minutes and drift scale to five percent in either direction. The report keeps the original measurement, suggestion and confirmed adjustment together, making a wrong choice visible instead of silently rewriting every cue.

Approve structural changes one by one

Safe normalization can standardize UTF-8, line endings, numbering and outer whitespace. Separate choices control invalid-markup removal, overlap trimming, long-cue splitting and adjacent duplicate merging. Splitting reuses the existing words; merging removes repeated adjacent text. Neither action invents punctuation, corrects spelling, translates captions or fills missing dialogue.

The final ZIP can include any confirmed combination of clean SRT, WebVTT and SBV files. A formula-protected issues CSV records original and final timing, the accessible HTML report explains remaining warnings, and the manifest binds source hashes, settings, rules and engine versions. The source is removed after processing and encrypted results expire after 24 hours.

CreatorsFix an exported subtitle file before a scheduled upload.
EditorsHand off clean captions with exact timing evidence and no unexplained rewrite.
EducatorsFind unreadably fast or overlong cues across lectures and lessons.
AgenciesReview client captions consistently without sending media to an AI service.

Before repairing captions

YouTube Caption Cleaner & Sync FAQ

Exact answers about wording, timing evidence, formats, exports and deletion.

Which subtitle files are supported?

Version 1 accepts one SRT, WebVTT or SBV text file up to 10 MB and 50,000 cues. UTF-8 is preferred; Windows-1252 can be converted to UTF-8 with the conversion recorded.

Can I check captions without uploading a video?

Yes. Caption-only mode checks syntax, encoding, cue order, overlaps, duration, line length, reading speed, markup, duplicates and empty cues.

How does optional synchronisation work?

A video or audio file lets the worker compare cue boundaries with media duration and locally detected non-silent audio activity. It may suggest a bounded global shift or drift scale, but the user must listen and confirm it.

Does WoluTools transcribe or rewrite the spoken words?

No. It does not transcribe, translate or invent wording. Normalisation affects encoding, whitespace, numbering and syntax. Splitting or merging text blocks requires explicit confirmation.

Which timing changes require confirmation?

Trimming overlaps, splitting long cues, merging adjacent duplicates, shifting every cue and stretching timing drift are all visible choices bound to the reviewed source hash.

What reading-speed limits are used?

The report marks more than 20 characters per second for review and more than 30 as a blocker. These are transparent workflow thresholds, not a guarantee of accessibility or comprehension.

What is included in the download?

The ZIP contains the selected clean SRT, VTT and/or SBV files, a protected cue-level issues CSV, accessible HTML quality report and a manifest with hashes, settings, versions and remaining findings.

When are files deleted?

Encrypted analysis drafts and source parts expire after two hours and are removed after completion, failure or cancellation. The encrypted result expires after 24 hours or immediately when deleted.