Transcription & speakers
Every word on the record, with a name against each voice.
Most of what is said in your footage is unfindable. You remember someone explaining the thing, but not which file it was in or how far through. Orcah transcribes every video as it indexes, splits the audio into speaker turns, and lets you put a name to each voice.
Getting started with transcription
- 1.
Point Orcah Studio at a folder
Transcription runs as part of the same indexing pass as scenes, faces and objects. There is no separate step and no per-minute bill.
- 2.
Name the voices once
Diarization separates the speakers on its own, and hands you each one as an unnamed voice. Label it once and every other clip that voice appears in picks up the name.
- 3.
Search by what was said
Quote a line, or part of one. Results come back as scenes with their timecode, so you land on the moment rather than the file.
What it can do
- Word-level timing: every word carries its own timestamp, so turns land on the exact second.
- Speaker separation: the audio is split into turns by voice before anything is named.
- Named voices: label a speaker once and the name follows them across the whole library.
- Follows playback: the transcript scrolls in step with the video, and a click on any line jumps the player there.
- Searchable by person: combine a name and a topic — find anyone talking about a subject, in any clip.
- Runs on-device: the audio never leaves your Mac, which keeps NDA footage and unreleased cuts where they belong.
Questions worth asking
- What happens to speakers it can't identify?
- They arrive as unnamed voices rather than being guessed at. You simply see an unlabelled speaker until you give it a name.
- Can I search for a phrase and a person at the same time?
- Yes — that is the point of naming voices. Once a speaker is labelled you can ask for the moments where that person talks about a given subject, and the agent will combine the speaker filter with the transcript search for you.
- Is the audio uploaded anywhere to be transcribed?
- No. The models download once, then transcription and diarization both run on your Mac.
Find the moment. Not the file.
Download immediately, models included. Your footage stays on your Mac.