AI Video Caption Generator

Upload a short video and get real, correctly-timed TikTok/Reels-style captions burned in by a self-hosted Whisper AI model.

Caption Generator

Configure & run

Real Self-Hosted AI

Drag & Drop Video Here

or click to browse (MP4, WebM, MOV -- up to 90s, ~200MB)

Captions are burned in using the spoken language's own script. Selecting Urdu specifically uses a larger, more accurate Whisper model for it (the same dual-model setup as our Voice Typing tool).

Preview

Upload a short video to begin...

Download

The .srt file has the same real timestamps as the burned-in captions -- import it into your own editor if you'd rather style the captions yourself.

Real Self-Hosted Whisper AI + ffmpeg

About the AI Video Caption Generator

Real Self-Hosted AI (Whisper + ffmpeg/libass) • Processed on Our Server, Deleted Immediately, Never Shared With a Third Party

1. What This Tool Actually Does

When you click "Generate Captions," a Caption Style Studio popup opens first, letting you pick a font (Classic, Montserrat, Bebas Neue, Inter, or Roboto -- real font files installed on our server), a text color (white, yellow, cyan, or pink), and an effect (a thin classic outline, a bolder "impact" outline with a drop shadow, or a solid highlight background box, TikTok/Instagram-caption-box style). These choices are genuinely applied server-side, not cosmetic -- confirming the popup sends them to our server, which builds a real libass force_style string from your selection before burning anything in. Only after you confirm does the real pipeline run: your video is uploaded to our own server, ffmpeg extracts the audio track, a self-hosted Whisper speech-to-text model (the same one that powers our Voice Typing tool) transcribes it with real word-level timestamps, those words are grouped into short, TikTok/Reels-style caption bursts instead of long full-sentence blocks, and ffmpeg's libass-backed subtitle renderer burns those captions directly onto the video frames using your chosen style. This is a genuine transcription-and-render pipeline running on our own hardware, not a call to a paid third-party captioning API. The uploaded video is deleted from our server immediately after processing and never shared with any third party.

Being honest about the style picker's preview: the popup shows a live CSS approximation of your chosen font/color/effect so you have a rough idea before generating, but it is only an approximation -- the real caption look comes from ffmpeg/libass rendering on our server, which can differ subtly from the CSS preview (most noticeably the highlight box's exact opacity). If you skip the popup's choices entirely, the defaults reproduce this tool's original look exactly (white text, thin black outline, no box).

Being honest about limits: this runs on the server's CPU, not a GPU, so processing takes real measured seconds -- in our own testing, a ~16 second clip finished in roughly 8-14 seconds end-to-end, a ~28 second clip in roughly 20-22 seconds, and a ~52 second clip in roughly 45-50 seconds (audio extraction is near-instant; transcription is fast; the video re-encode for burning in the captions is the slowest step and scales with video length). Accuracy depends on the same real-world factors any speech-to-text engine faces -- background noise, accents, overlapping speakers, and unclear audio all reduce accuracy, and the auto-detect language feature can occasionally misidentify very short or ambiguous audio. Uploads are capped at 90 seconds and roughly 200MB to keep processing time reasonable on this server.

2. How to Use It

  1. Upload your video: drag and drop an MP4, WebM, or MOV clip (up to 90 seconds and roughly 200MB), or click "Browse Video."
  2. Pick a spoken language or leave it on Auto-Detect, then click "Generate Captions".
  3. Choose your caption style in the Caption Style Studio popup that opens -- font, color, and effect -- then click "Confirm & Generate Captions". Only then does the real upload begin. A loading overlay shows real elapsed time while audio extraction, transcription, and the caption burn-in run on our server.
  4. Preview and download: watch the result in the player, then download the captioned video and/or the standalone .srt subtitle file.

3. Why Short Caption Bursts Instead of Full Sentences

Whisper's own transcription segments are often a full sentence or more, which wraps into three or four lines and covers a large part of a vertical frame -- not what real TikTok/Reels-style captions look like. Instead, this tool uses Whisper's real per-word timestamps to regroup the transcript into short bursts of a few words each, synced closely to when they're actually spoken, which is both more readable and closer to the familiar short-form caption style.

4. Frequently Asked Questions

Caption Style Studio

Pick a font, color, and effect -- these are burned into the real video by ffmpeg, not just a cosmetic preview.

YOUR CAPTION TEXT

Approximate CSS preview -- the real burned-in video is rendered server-side by ffmpeg/libass and may look subtly different (especially the box effect's exact opacity).

Copied to clipboard!