Files
WhispAssist/docs/adr/0005-diarization-sherpa-onnx.md
T

5.0 KiB

ADR-0005 — Speaker diarization: sherpa-onnx

  • Status: Accepted
  • Date: 2026-06-30
  • Context source: Design doc §"Speaker Diarization and Naming"

Context

WA must label transcript segments by speaker (Speaker 1/2…), allow naming during and after a meeting, support merging over-split speakers, and feed calendar participant names into the naming UI — all offline.

Decision

Use sherpa-onnx offline speaker diarization: a pyannote segmentation model + a speaker-embedding extractor (e.g. 3D-Speaker ERes2Net) + clustering, all ONNX and fully local. Wrap it behind a diarization::Diarizer trait. Diarization runs as post-processing over the recorded audio (audio is the source of truth, ADR-0006), producing speaker-ID-tagged spans that are aligned to whisper segments by timestamp overlap.

During live recording, show provisional speaker turns from segmentation (cheap), then refine labels in the post-meeting pass; this keeps latency low while improving final accuracy.

Consequences

  • Positive: purpose-built, offline, ONNX (shares the runtime with the NPU transcription path); supports the segmentation+embedding+clustering pipeline the design assumes; models are downloadable and swappable.
  • Negative: separate models to download/manage (model-management UI, FR-MODEL-1); C/C++ FFI to integrate; clustering may over/under-split → the merge workflow (FR-SPK-3) is required, not optional.
  • Speaker IDs (S1, S2, …) are internal and stable per meeting; name mappings live in the DB and are applied at render/export time, never destructively rewritten onto segments.

Phase 3 refinement — per-stream "You" attribution (FR-SPK, 2026-07-14)

The single-pass-over-the-whole-recording model above blind-clusters a summed mono signal (mic + loopback), then guesses which cluster is the user via a voiceprint match. In practice that clustering is unreliable in both directions on the summed signal — it over-split a 2-speaker call into 83 clusters at one threshold and merged two clearly distinct voices into one at another — and the voiceprint can only label a cluster, never create the separation.

So when the microphone is enabled, WA no longer relies on clustering to find the user:

  • A live mic-activity timeline (audio::MicActivity) records, per 100 ms of audio.wav frame time, when the mic was speech-level — captured in the loopback writer, the one place the mic and loopback exist separately in the recording's own timebase. Those ranges become "You" spans directly (no clustering, no embedding).
  • sherpa clustering is then run over the recording with the "You" ranges masked out (zeroed), so it only ever splits the far side into Speaker N. The mic can never pollute or merge into a far-side cluster.
  • The timeline is persisted as mic_activity.json (retained meetings) so reprocess_transcript re-attributes identically without the live capture.

The original blind-clustering + voiceprint pass remains the fallback for mic-off recordings and imports (no timeline). Naming stays uniform (build_name_map: "You", then Speaker 2…). Segment IDs and the names-in-DB rule are unchanged; only the source of the spans changes.

Phase 3.5 refinement — dual-channel capture (supersedes Phase 3's timeline, 2026-07-14)

The Phase 3 timeline above reconstructed the mic/far-side split after the fact from a summed-mono audio.wav plus a mic_activity.json sidecar, masking the mic ranges before clustering. That was reliable at stop but fragile on reprocess (re-aligning a sidecar against a mono mix — it collapsed to a single speaker) and carried a parallel persistence path.

Decision: when the mic is enabled, record audio.wav as stereo with the streams separated — left = microphone ("You"), right = system/loopback ("Speaker") instead of summing them. The separation is then intrinsic to the recording:

  • Diarization runs on the right channel only → Speaker N; "You" comes from left-channel voice activity (vad_spans). The mic is never clustered, by construction.
  • Reprocess recomputes both from the file — no sidecar, no masking; stop and reprocess agree.
  • Transcription still downmixes (L+R) to the same summed mono; playback folds to mono and bundle export folds to dual-mono so shared/played audio is normal.
  • A meetings.audio_layout flag (summed | split, carried in the bundle manifest) distinguishes recordings; existing summed recordings keep the blind-clustering + voiceprint path. On a stereo render endpoint this is size-neutral (the file was already 2ch with the mic summed into both).

This retires MicActivity / mic_activity.json / masked diarization from Phase 3 (see the retirement note at attribute_split in commands.rs); the capture-silence handling, diarize_samples, and build_name_map naming are retained. Segment IDs and names-in-DB unchanged.

Revisit if

A single model gives joint ASR + diarization with better accuracy, or whisper.cpp gains production diarization.