What this converter does
The parser reads sequence numbers, start and end timestamps, and caption text. It can show a clean script or preserve timing labels, previews each cue with a device voice, and estimates whether a line is likely to overrun its subtitle window before local WAV generation.
Why timing conflicts need review
The conflict check estimates spoken duration from word count and selected speed; it is a warning, not a measured render. Generated speech has its own pauses and pronunciation. Shorten flagged cues or generate scene-sized clips, then align them against the original timeline in an editor.
Privacy, limits, and licensing
Parsing and text inference run in the current browser. Generation is limited to 1,000 cleaned characters, with a visible cancel control between segments. Model files require a first download. Verify rights to the script, selected voice, model, and distribution before publishing.
Questions people ask
Does this tool upload my content?
No. Tool input is processed in the current browser. Required code or model files can be downloaded from the providers disclosed in the privacy policy.
Do I need an account?
No account or subscription is required. Browser and device limits still apply.
Can I use the result commercially?
This website does not grant rights to text, recordings, device voices, models, or third-party assets. Check every applicable license for the intended use.