← All tools
📋

Video & Audio Summarizer: TL;DR, key points and chapters

Drop a video or audio file — a lecture, podcast, meeting or talk — and get a short TL;DR, the key points, chapters with timestamps and the full transcript with clickable times. The speech model runs in your browser, so the file never leaves your device.

Free · No sign-up
📚 Guide: write YouTube chapters from a transcriptOpen →

How to use Video & Audio Summarizer

  1. 1. Drop, choose or paste a video or audio file (MP4, MOV, WebM, MP3, M4A, WAV…). It opens on your device — nothing is uploaded.
  2. 2. Pick the speech model (Tiny is fastest, Base is more accurate) and leave the language on Auto-detect or choose it. The estimate shows how long it will take on your device.
  3. 3. Press Transcribe & summarize. Whisper runs in your browser, then you get a TL;DR, key points and chapters — click any timestamp to play from there, and rename chapters if you like.
  4. 4. Copy the YouTube chapters block into your video description, or download the transcript (.txt), the summary (.md) or subtitles (.srt / .vtt).

Frequently asked questions

Is my video or audio uploaded?

No. The file is decoded and transcribed inside your browser; only the speech model (80–300 MB, once) is downloaded. For the summary, Chrome's on-device AI is used when available; otherwise the transcript text (not the recording) is sent to our free cloud AI to write it and is not stored. Choose Basic summary only to keep the transcript on your device too.

How long does it take?

It depends on the device and the model. On a recent laptop's processor, Whisper Tiny transcribes about 10 times faster than real time (a 30-minute talk in about 3 minutes) and Base about 6 times; older laptops and phones are slower, WebGPU is usually faster. You see an estimate before you start and the time left while it runs.

How are the chapters made?

The transcript is split where its vocabulary changes: the words said in the minute or so before each point are compared with the words after it (TF-IDF similarity), and clear drops become chapter starts. Each chapter is titled by its most typical sentence, or by an AI headline (Chrome's on-device AI or our free cloud AI). You can rename them.

Are the YouTube chapters ready to paste?

Yes, when the video has at least three topic sections: the block starts at 00:00, has at least three chapters and each is at least 10 seconds long, as YouTube requires. If there are fewer topics, choose 'More chapters' or add your own.

Which summary do I get?

In desktop Chrome 138 or newer with its built-in AI (Gemini Nano, on device) you get an AI-written TL;DR, key points and chapter titles in English, Spanish or Japanese. Elsewhere, and for other languages, our free cloud AI writes them from the transcript text. With Basic summary only (or if the cloud AI is busy) you get the transcript's most central sentences. The label always says which engine was used.

How long can the file be?

On a computer, recordings up to about two hours and 2 GB usually work; on a phone, about 30 minutes with the Tiny or Base model. The audio is kept in memory while it is transcribed, so very long files can run out of memory.

🎬 Browse all video & audio guides and tools →