Setups by Scenario
Pick your meeting below and copy the settings. Each section tells you what to turn on, what to leave alone, and one thing that trips people up.
Remote calls: 1:1s, sales calls, client check-ins
What makes these hard: two voices, usually one on your speakers. If Clearminutes only captures your microphone, half the conversation is missing.
Settings:
- Microphone: on (default)
- System audio: on, pointed at the device the call plays through
- Model: large-v3-turbo
- Diarization: on (Pro; the model downloads from Settings → Transcription) — gives you Speaker 1 / Speaker 2 labels you can rename
- Echo Suppression (Pro): on, in Settings → Recording ("Suppress Echo / Speaker Bleed") if you use speakers instead of headphones
On free, without diarization, a two-person call comes back as one flat block of text. Still accurate. Just not split by speaker.
The one trip-up: on a Mac, the most common reason the remote side never makes it in is the Screen Recording permission (System Settings → Privacy & Security → Screen Recording). Grant it, restart Clearminutes, test again. Screen-sharing apps that route audio through a virtual device can do the same damage; if a transcript comes back with only your side, check the system-audio device in Settings → Audio before anything else.
Interviews and research: when you quote people later
Accuracy is the whole point. A wrong name in a published quote costs more than twenty extra minutes of processing.
Settings:
- Model: large-v3-turbo for most; large-v3 when the interview is on the record
- Language: set it manually in Settings → Transcription → Language if you know it; auto-detect is good but not psychic
- Diarization: on (Pro) if multiple people speak
Then use a custom prompt (Pro) with your template: "Extract quotes verbatim, keep speaker names, list follow-up questions." Same interview, same template, every week.
Ten standups a week
Speed beats everything. A 15-minute standup on Parakeet comes back in about two minutes.
Settings:
- Model: Parakeet (Apple Silicon, English; ~670 MB download in Settings → Transcription)
- Auto-summarize: on (Pro), with a short prompt like "Action items only, one line per person"
- Custom template: keep it to three lines. Standup summaries don't need prose.
This is the setup that makes you stop noticing transcription time exists.
Lectures, talks, long solo recordings
Two rules for long recordings: pick a fast model, and control the audio up front.
Settings:
- Model: Parakeet (single speaker, English) or large-v3-turbo (technical vocabulary, multiple languages)
- Record in the Clearminutes app itself, or use Audio Import (Pro) if you captured the lecture on a recorder or phone
- Summaries shine here: a 90-minute lecture needs a structure, not a wall of text
Solo, one voice, good acoustics: this is Parakeet's best case.
Mixed rooms: people in the room and on the call
The hardest setup, and the one that most rewards getting audio right before recording.
- Mic: captures the room
- System audio: captures the call
- Model: large-v3-turbo (overlapping voices are the hardest case)
- Diarization: on (Pro), then rename the Speaker 1 / Speaker 2 labels to real names in the transcript
- Headphones all round if you can get them. Open speakers mean your mic hears remote speech a second time, which muddies speaker labels. Echo Suppression (Pro), in Settings → Recording, trims the doubled copy of the audio if open speakers are unavoidable.
Quiet speakers and soft-spoken colleagues
Someone speaks just above a whisper and the transcript comes back with gaps. A few levers, cheapest first:
- Microphone Gain (free): a slider in Settings → Recording. Nudge it up for the quiet talker.
- System audio levels: Clearminutes already normalizes quiet incoming levels, so soft remote voices don't vanish on their own.
- VAD Sensitivity (Pro): lower it (slider left toward "Low") in Settings → Recording. The default is tuned for normal conversation; a just-above-whisper voice can sit under it. Lowering the threshold tells the speech detector to keep more borderline audio.
Non-native speakers and accented speech
- Model matters more here than anywhere else: large-v3 or large-v3-turbo, never small
- Parakeet is tuned for English meetings; fine for accented English, weak for switching languages mid-meeting
- Set the language explicitly instead of relying on auto-detect
Meetings you recorded elsewhere
Recorded on a phone dictaphone, or a client's Zoom recording file? Audio Import (Pro) takes the file and runs the transcription engine of your choice. Nothing about the import path forces a different model.
If the notes don't appear
Transcription is unlimited on free and the recording always saves, even when the transcript doesn't. The app tells you which step failed: "No audio detected" or "No speech detected" means the speech detector found nothing (audio too quiet, or VAD filtered it), and "System audio wasn't captured" means the remote side never made it in (Screen Recording permission on macOS, or the wrong output device). Fixing it and re-running the transcript uses retranscription, which is a Pro feature. See Transcription Problems.
Not covered here?
Tell us what meeting broke and why, through your support ticket or the client portal. Gaps like that become new sections.