Subtitle generator (SRT/VTT)

Auto-generate captions with free AI.

Drop a file here, or click to choose

Audio or video · runs a free AI model in your browser (first run downloads ~40MB)

Private — files are processed in your browser and never uploaded.

Most people watch videos on mute, so captions are no longer optional — they boost watch time, accessibility and reach. But paid transcription services add up fast, especially across a lot of content.

This free subtitle generator transcribes your video or audio and produces ready-to-use SRT and VTT caption files — using an AI speech-to-text model that runs entirely in your browser. There is no API cost, no upload and no sign-up.

How to use it

  1. 1

    Add a video or audio file

    Drop in your clip. The first run downloads a small AI model (about 40MB) to your browser; after that it is cached.

  2. 2

    It transcribes on your device

    The speech-to-text model runs locally, breaking the audio into timed caption chunks — nothing is sent to a server.

  3. 3

    Download SRT or VTT

    Review the captions and download an SRT or VTT file to upload alongside your video or import into an editor.

Why use this tool

Free AI transcription

A Whisper-based model runs in your browser, so there are no per-minute API fees no matter how much you transcribe.

SRT and VTT output

Get standard subtitle files that YouTube, editors and players accept everywhere.

Completely private

Your audio never leaves your device — the model does all the work locally.

Why captions are worth it

A large share of social video is watched without sound, and captions keep those viewers engaged instead of scrolling away. They also make your content accessible to deaf and hard-of-hearing viewers, help non-native speakers follow along, and give platforms text to understand and rank your video by. Uploading an SRT file (rather than relying on auto-captions) means you control the wording, spelling of names, and timing.

On-device AI and its trade-offs

This tool uses a compact Whisper model compiled to run in the browser. That is what makes it free and private — the transcription happens on your own hardware with no server involved. The trade-off is speed and absolute accuracy: a small model is quick to load but less precise than a large paid one, and long files take longer on slower devices. It is excellent for short-form clips and a great starting draft for longer videos; always give the captions a quick read and fix any names or terms before publishing.

Private by design — nothing is uploaded

Unlike most online converters that send your files to a remote server, this tool does all of its work directly inside your web browser. Your image, video or audio never leaves your device, there is nothing to delete afterwards, and there is no account to create. That makes it faster (no upload or download round-trip on top of the processing) and far safer for anything personal or client-confidential. When you close the tab, everything is gone.

Tips & best practices

  • Add captions to every clip — most social video is watched on mute, and captions lift watch time and reach.
  • Upload your own SRT rather than relying on auto-captions so you control spelling, names and timing.
  • Give the generated captions a quick read to fix names and specialist terms before you publish.
  • Use the VTT file for web players and the SRT file for YouTube and most editors.
  • For long videos, transcribe in shorter sections to keep the in-browser model fast and responsive.
  • Clear audio produces far better captions, so start from the cleanest recording you have.

Frequently asked questions

Is the transcription really free?+

Yes. The AI model runs in your browser, so there are no API fees or per-minute charges no matter how many files you transcribe.

What file formats do I get?+

You can download SRT and VTT subtitle files, which work with YouTube, video editors and web players.

Is my audio uploaded to a server?+

No. The speech-to-text model runs entirely on your device, so your file never leaves your browser.

How accurate is it?+

It uses a compact Whisper model that is very good for clear speech and short clips. For long or noisy audio, review and lightly edit the captions before publishing.

Why does the first run take a moment?+

The first time, your browser downloads the AI model (about 40MB). After that it is cached, so subsequent runs start quickly.

About our free creator tools

This is one of a growing set of free tools from Faceless Post, built for creators who publish consistently across YouTube, TikTok, Instagram, Facebook, X and LinkedIn. Every tool is free to use, needs no account, and adds no watermark. Most run entirely inside your browser, so your images, video and audio are processed on your own device and never uploaded to a server — there is nothing stored on our side and nothing to delete. That makes them fast, safe for client and personal files, and usable even on flaky connections once the page has loaded.

We build these because the day-to-day of content creation is full of small friction — converting a file, trimming a clip, resizing an image, cleaning metadata, grabbing a thumbnail. Handling those in one place, for free, saves you from juggling sketchy upload sites. And when you are ready to stop doing it all by hand, Faceless Post can schedule, caption and auto-publish your content to every platform from a single calendar. Explore the full free tools library or start free to automate your whole posting workflow.