Edit a multicam podcast with one shared microphone
Most tools need a separate audio track per person. CastCut works out who is talking from a single mixed recording, then cuts to them.
Why one microphone is the hard case
Camera cutting is easy when every person has their own track: whichever track is loud is the person talking. That is how most automatic editors work, and it is why they ask for a lav on everyone or a multitrack recorder.
Real rooms are not like that. One shotgun on a boom, a table mic between two chairs, a field recorder catching the whole conversation: all of it is one signal with several people in it. Loudness tells you nothing about who spoke.
What CastCut does instead
It separates the voices inside the single recording. Silero voice activity detection finds the speech, then overlapping 1.5 second windows get a WeSpeaker voice fingerprint, and those fingerprints are grouped into people. Speaker changes snap to the nearest breath or pause, so the cuts land where a human editor would put them.
Then it works out who is who on screen. Apple Vision counts the faces on each camera, and lip movement is sampled while each voice talks, so a voice is matched to the face whose mouth moves with it. A camera showing one face is a close-up; a camera showing two or more is the wide shot.
From there the edit is ordinary: one person talking goes to their camera, or a punch-in on the wide shot framed on their face if they do not have one. Two people talking, or a quick back-and-forth, goes wide.
What you need
- One audio recording that covers the whole conversation. A recorder, a mixer, or even a camera's own sound.
- At least one camera that kept rolling. Two or more gives more to cut between, but a single wide camera works.
- No clapper and no matching timecode. Files are lined up by the sound itself.
Where it is honest about not knowing
If a voice cannot be matched to a face, the edit says so on the Multicam page and stays on the wide shot for that person rather than cutting to the wrong camera. You can set who each face is yourself from the Media page, and what you set survives the next analysis.
Questions
Can you edit a multicam podcast with only one microphone?
Yes. CastCut separates the voices inside a single mixed recording using a local voice fingerprint model, then matches each voice to a face on camera by sampling lip movement. You get camera cuts driven by who is talking, from one shared mic, with no per-person audio track.
Do I need a lavalier or a separate track for each person?
No. Separate tracks work and are slightly more reliable, but they are not required. A single shotgun, a table mic or a field recorder picking up the whole room is enough.
What about two people talking at the same time?
Sustained cross-talk is detected and the edit holds the wide shot while it lasts, instead of flicking between close-ups.
Does it work if the people sound similar?
Usually. Voices are grouped by fingerprint rather than by pitch, and the count is fixed to the number of faces the camera shows, which stops two people being merged into one. Where it cannot tell, it says so instead of guessing.
How many people can share one microphone?
Three or four work well. Everyone on one wide camera also works: each person gets their own close-up, cropped from the wide, in landscape and in vertical.
Try it on your next episode
2 free exports, everything unlocked. No account, no email, no card.