Clip anything by prompt
Type what you want — “every time they mention pricing”, “the funniest 30 seconds” — and get clips that match, plus an optional time range.
vs. Opus: prompt + explicit time-range filter, on every plan.
Drop in a podcast, webinar or stream. Latch finds the moments, scores them with a virality score you can actually read, reframes to follow the speaker, adds animated captions in 30+ languages, pulls B-roll, and schedules the posts.
No credit card. Runs on your machine or our cloud.
Live preview is a CSS illustration — real output uses your video.
Built for podcasters, educators, agencies and streamers who post every day and want to know why a clip was picked.
*On a single consumer GPU with faster-whisper large-v3; CPU-only is slower.
Three steps. Every one of them is inspectable — no black boxes between your video and the clips.
Paste a YouTube, podcast or Zoom link — or upload an MP4 up to several hours long. We extract audio, normalize it, and transcribe with word timestamps and speaker labels.
ingest → transcribe
Chapters are detected from topic shifts and scene cuts. Candidate windows snap to sentence boundaries, then each is scored on hook, flow, value, trend, emotion and clarity — with written reasons.
chapters → candidates → score
Tweak captions, swap the aspect, toggle jump-cuts, apply your brand kit. Download MP4 + SRT, hand the timeline to Premiere or Final Cut, or schedule straight to your channels.
plan → render → publish
Twelve things Latch does out of the box. Each card notes how it differs from the closed alternative.
Type what you want — “every time they mention pricing”, “the funniest 30 seconds” — and get clips that match, plus an optional time range.
vs. Opus: prompt + explicit time-range filter, on every plan.
Six sub-scores (hook, flow, value, trend, emotion, clarity) and plain-English reasons for each pick. Heuristic first, LLM-refined when you add a key.
vs. Opus: a single 0–99 number with no breakdown.
Face and speech tracking keep the active speaker centred when going 16:9 → 9:16. Two speakers? Auto split-screen.
vs. Opus: tracking without automatic split-screen.
Six presets — Pop, Karaoke, Box, Bounce, Minimal, Outline — keyword highlight, emoji, per-platform safe zones. Word-level timing.
vs. Opus: 30+ Whisper languages instead of 20+, safe zones built in.
“Um”, “like”, dead air — detected from the transcript and waveform, turned into jump-cuts you can toggle one by one.
vs. Opus: every cut is listed and reversible, not baked in.
Keyword-driven stock footage from Pexels, or your own local library, dropped in fullscreen, picture-in-picture or top-third.
vs. Opus: B-roll on all plans, including self-host, and your own library.
Each clip ships with a title, a hook line, a description and hashtags tuned to the transcript — editable before you post.
vs. Opus: generated with the same model you bring, not a locked one.
Logo placement and scale, primary and accent colours, intro/outro clips, watermark text and a progress bar. Save as many as you need.
vs. Opus: unlimited kits on self-host; no plan gating.
9:16, 1:1, 16:9 and 4:5 from the same plan, up to 4K, with loudness normalisation and per-platform encode presets.
vs. Opus: 4K and every aspect on every plan.
Hand the whole cut to Premiere Pro (XML), Final Cut (FCPXML) or any NLE via EDL — including reframe keyframes and cuts.
vs. Opus: XML export is Pro-only there; here it's standard.
Queue clips to YouTube, TikTok, Instagram, LinkedIn and X. Each platform's length, caption and aspect rules are validated before you hit save.
vs. Opus: validation errors are surfaced in the UI, not at publish time.
One `pip install`, a CLI, a REST API with SSE progress. Bring your own Whisper and LLM keys. No per-minute credits — ever.
vs. Opus: API is Business-tier; ours is MIT-licensed and included.
Captions come from word-level timestamps, not sentence guesses. Pick a preset, set your highlight colour, and Latch keeps text out of the TikTok, Reels and Shorts UI zones automatically.
Bold, outlined, the current word scales up and turns your highlight colour.
Face detection plus who-is-talking from the transcript drive a smooth crop path. When two people trade lines, Latch switches to a stacked split-screen instead of whip-panning.
16:9 source · 2 speakers detected
9:16 · auto split
Every candidate gets six sub-scores and a list of the concrete reasons behind them. Heuristics run on every clip; when you add an LLM key, the reasons get sharper and the score is marked “hybrid”.
Candidate 03 · 0:41
“Why does nobody talk about this?”
Why this scored 84 · hybrid
Pick a clip, a platform and a time. Latch checks length, aspect ratio, caption limits and hashtag count against each platform's current rules before it saves — and tells you exactly what to fix.
A straight feature comparison. Where they are ahead, we say so.
† Opus Clip facts per Opus public pricing and feature pages, September 2026. Plans and limits change; check their site for current details. Opus Clip is a trademark of its owner; Latch is not affiliated.
No per-minute credits. Cloud plans are not live yet — join the waitlist and we'll email you when they open.
MIT-licensed. Your GPU, your keys, no limits.
For one channel that posts every day.
For creators and small teams who want it all.
For agencies running many channels.
Cloud prices shown are planned launch prices in USD and may change before launch. Source minutes reset monthly and do not roll over.
Candidates always snap to sentence boundaries, so you never get a clip that starts mid-word. Ranking is heuristic by default and improves noticeably when you add an LLM key. In our internal test set of podcast episodes, the top-3 picks matched a human editor's picks roughly two times out of three — good enough to save hours, not good enough to skip a 30-second review.
Transcription and captions run on Whisper, which handles 30+ languages well (English, Spanish, Portuguese, German, French, Italian, Dutch, Japanese, Korean, Mandarin, Hindi and more). Titles, hooks and descriptions are generated in the transcript's language when an LLM key is present.
ffmpeg and Python 3.11+. Any modern CPU works; a GPU with 8 GB+ VRAM makes Whisper large-v3 roughly 10× faster. Face tracking uses OpenCV/MediaPipe and is optional — without it, reframe falls back to center or static modes.
Self-hosted: nothing leaves your machine except calls you configure (your LLM provider, your stock-footage provider). Cloud (when it launches): media is stored encrypted, processed in an isolated worker, and deleted 30 days after the project is deleted. We don't train on your content.
Six sub-scores from 0 to 100 — hook (how the first seconds grab), flow (self-contained, no dangling references), value (concrete takeaways), trend (topic momentum), emotion (tonal shifts, laughter, emphasis) and clarity (speech rate, filler density). Each score comes with written reasons that cite timestamps, so you can check them. The total is a weighted blend you can tune.
Yes. Export a Premiere Pro XML, a Final Cut Pro FCPXML or an EDL for the whole project. Cuts, reframe keyframes and caption timing are included where the format supports them.
Never on self-host. Cloud Creator and up are also watermark-free. There is no free cloud tier planned, precisely so we never have to watermark.
The engine is open source (MIT) and runs on your own hardware with your own API keys, so there are no per-minute credits. The score is explained rather than a single number, every cut is listed and reversible, timeline export works on every plan, and the REST API is included. Opus has a longer track record and a hosted product today; our cloud plans are still on a waitlist.
Paste a link and see ranked candidates in minutes — or clone the repo and run the whole thing on your own machine tonight.