Coming soon
Automatic speaker detection (diarization)
Cast will work out who is speaking when from the audio alone and label the transcript by speaker, so a two-host or interview episode arrives already split — no clicking through paragraphs to mark who said what.
Today you label speakers yourself: select a paragraph, pick a name, and the label follows that person through cuts and rearrangements. That takes a few clicks per handover and never guesses wrong, but on an hour-long interview it is the most tedious part of the job.
Speaker detection would run a diarization model during transcription. You would still be able to rename or correct a speaker, and your corrections would take priority over the model everywhere the labels are used — chapters, captions, the exported transcript.
It is built for interviews and co-hosted shows recorded on one track. If each guest already has their own file, the split is already known and this feature adds nothing.
Last updated