Skip to content

Speakers

Quendian groups speech by voice and labels each speaker s1, s2, s3, and so on. Where those labels come from depends on the transcription provider you’re using:

  • AssemblyAI, Deepgram, and Quendian Cloud return speaker labels as part of transcription itself. No extra setup; labels are always there.
  • Whisper, gpt-4o-transcribe, and other transcription-only providers don’t return speaker labels. Quendian can produce them on-device via an opt-in toggle in Preferences → Speakers → “Identify speakers on-device” — see Local diarization below.

Either way, the labels look the same in the transcript and scale with how many distinct voices the system detected.

If you don’t want speaker labels — for example, you’re transcribing a solo voice memo, or you’re using a BYOK provider and find on-device labeling noisier than helpful — leave the toggle off. The transcript renders as plain text without speaker captions.

When you’re using a transcription provider that doesn’t return speaker labels — Whisper, gpt-4o-transcribe, Groq Whisper, or a custom OpenAI-compatible endpoint — Quendian can run an on-device pass after each recording to group speech by voice. This is opt-in via Preferences → Speakers → “Identify speakers on-device”.

The pass runs entirely on your Mac. After a session stops, Quendian feeds the audio through a small speech model that clusters segments by acoustic similarity and assigns each cluster a label (s1, s2, etc.). Your audio doesn’t leave your machine for this step.

Local diarization works well when the recording has 2-3 distinct voices in a quiet room. It can mislabel speech when voices sound similar — same gender, similar accent, similar pitch — or when there’s significant background noise. It’s best understood as a useful approximation, not ground truth.

For meetings where transcript quality matters more than BYOK cost, the natively-diarizing providers (AssemblyAI, Deepgram, or Quendian Cloud) produce more reliable speaker labels — they use server-side models trained on far more meeting audio than fits on a Mac.

Turn this on if you want speakers with cheaper BYOK transcription and you’re recording the kind of audio it handles well (small group, decent mic, low background noise). Leave it off if you only care about the words, or if you’ve tried it and noisy labels distract more than help.

The toggle applies to new sessions only. Existing recordings aren’t re-labeled retroactively.

The auto-assigned labels (s1, s2, etc.) are placeholders. You can rename any speaker directly in the session view:

  • Per-session rename — click a speaker label in the transcript to rename it for that session only. Other sessions are unaffected.
  • Global rename — rename a speaker globally to apply the new name everywhere that speaker has been identified across all sessions.

Each speaker is assigned a color swatch for visual distinction in the transcript. You can change any speaker’s color using the color picker, which offers 7 preset swatches plus an Auto option.

The Auto option lets Quendian assign a color automatically. The special “You” identity (your own voice) is assigned the iris color from the app’s color scheme.

The session view includes a talk-share visualization showing how speaking time was distributed across speakers in that session. This is a proportional breakdown — helpful for reviewing meeting balance or identifying who dominated a call.