How it works
Semantic video transcript search
ClipIt runs semantic search over the transcript of your recording. You ask in ordinary language; matches come back with the quote, why it fit, and the time range on the source video — not a bare keyword hit list.
Transcript search that still plays the video
A searchable transcript is only useful if you can hear the moment it points to. ClipIt keeps playback, quote, and rationale together so “Ctrl-F on the text” never becomes the whole product.
The transcript is generated when the recording becomes ready — the same captions Mux already makes for playback. ClipIt does not ask you to upload a separate text file, and it does not leave you jumping between a document and a player hoping the timecodes still match.
You can download the source as .srt at any time. Searching it here is for finding meaning; the file is for editors who need the words elsewhere.
Keyword find vs. asking a question
| Keyword / Ctrl-F | ClipIt question | |
|---|---|---|
| Input | Exact words you hope were said | The idea you need answered |
| Misses | Synonyms, paraphrases, “they never said X” | Fewer — meaning can still match |
| Output | Offset in text | Range + evidence + optional clip |
Exhaustive vs. one answer
Some jobs need every mention of a topic; others need the single clearest explanation. ClipIt supports both shapes of ask — see the recording studio’s question kinds — and still routes each result through the same evidence review.
Exhaustive lists are easy to over-trust. “Every time they mention pricing” will also catch a joke, a correction, and the host repeating the guest. That is why each hit still has a quote and a range: coverage is not the same as a publishable clip.
Related: AI video search and search video by asking a question.
Speakers, once you name them
Premium captions can label voices as speaker_1, speaker_2, and so on. You name them by listening to a sample, not by guessing from the waveform. After that, questions like “what did the host ask about refunds?” have someone to attach to. Until the voices are named, speaker search is not a promise — it is a later unlock.
The labels from the caption job are not real names. The product will not invent “Alex” from a voiceprint. If you skip naming, exhaustive asks still work — they just cannot filter by who said it.
A worked search
Forty-five minutes of office hours. A student writes “when did they explain the midterm curve?” You are not looking for the word “curve” — they may have said “how the exam will be scored.” Ask that. The first hit is the syllabus recap; the quote makes that obvious. The second is the actual explanation. You play fifteen seconds, save the range, and send the unlisted finding. No MP4 required for that reply.
Later you need every mention of office-hour times for the course page. That is an exhaustive ask on the same file. Some hits will be asides. Coverage is the point; you still pick which ranges become clips.
What it costs to search a transcript
Captions are part of ingest. A 45-minute lecture spends 45 source minutes and arrives with a transcript you can search and download. Free is 120 source minutes total, one-time. Paid plans renew the source allowance monthly; see Pricing. Asking “every mention of photosynthesis” does not spend more source minutes than asking once.
Translating those captions into another language is a separate, optional meter on paid plans. Searching the original transcript does not require it.
What this is not
It is not a research repository. There is no participant database, no tag taxonomy across a study, and no project folder that holds fifty interviews under one query. If that is the job, Dovetail is built for it; ClipIt for researchers is the honest narrower case — one interview, one question, one citable range.
It is also not a live meeting bot. Nothing joins your Zoom. You upload a file you already have. If you need a bot in the call, that is a different product — see ClipIt vs Grain.
Questions
How do I find every time a topic is mentioned in a recording?
Ask for exhaustive mentions of that topic after the recording is ready. Review each hit against the source before you save clips — exhaustive lists are for coverage, not automatic publishing.
Can I search across my whole library?
Not yet. Each question runs against one recording. Cross-library search is a later product, because it changes retrieval cost and has to stay free of source-minute charges. Today the cheap move is to ask the file you already have.
Does the transcript cost extra?
No. Captions are part of ingest. You paid source minutes for the file; the text is included. Translation of those captions is a separate, optional step on paid plans.
Can I upload my own transcript instead?
Not today. Search uses the captions generated at ingest so the timecodes stay attached to the picture. A sidecar file you made elsewhere is useful in an editor; it is not what ClipIt retrieves against.