Transcription
Every recording and import gets a word-level transcript with exact timestamps, produced locally by Whisper — no per-minute transcription bill. The transcript is the editing surface for everything below.
Transcripts, filler-word detection, and silence removal — Descript-style, built in.
Every recording and import gets a word-level transcript with exact timestamps, produced locally by Whisper — no per-minute transcription bill. The transcript is the editing surface for everything below.
Click words in the transcript to cut them from the video — the audio and scenes close up around the cut. Trim a whole imported clip or work scene by scene, and preview playback skips your cuts live so you hear the result before applying.
Transcription is tuned to keep disfluencies rather than hide them, then the trim view highlights every umm, uh, and false start as a cuttable token. One click highlights them; one more cuts them all.
Silences are detected across the take and removed in one action — the fastest way to turn a rambling recording into a tight Short.
Cutting words out of a video is really a text operation, so ReelMint makes it one: click words in the transcript, they leave the video, and the audio closes up around the cut. Transcripts are produced locally by Whisper with word-level timestamps — no per-minute transcription bill — and the pipeline deliberately keeps the umms rather than hiding them, so the trim view can flag every one as a cuttable token.
That's how a four-minute ramble becomes a tight Short in minutes: cut the fillers in one action, remove the silences in another, preview with your cuts applied live, and render only what's left.
Transcription, word-level trim, filler-word cuts, and silence removal are in the free tier. No feature gates.
Start free