Files
WhispAssist/docs/adr/0005-diarization-sherpa-onnx.md
T

1.7 KiB

ADR-0005 — Speaker diarization: sherpa-onnx

  • Status: Accepted
  • Date: 2026-06-30
  • Context source: Design doc §"Speaker Diarization and Naming"

Context

WA must label transcript segments by speaker (Speaker 1/2…), allow naming during and after a meeting, support merging over-split speakers, and feed calendar participant names into the naming UI — all offline.

Decision

Use sherpa-onnx offline speaker diarization: a pyannote segmentation model + a speaker-embedding extractor (e.g. 3D-Speaker ERes2Net) + clustering, all ONNX and fully local. Wrap it behind a diarization::Diarizer trait. Diarization runs as post-processing over the recorded audio (audio is the source of truth, ADR-0006), producing speaker-ID-tagged spans that are aligned to whisper segments by timestamp overlap.

During live recording, show provisional speaker turns from segmentation (cheap), then refine labels in the post-meeting pass; this keeps latency low while improving final accuracy.

Consequences

  • Positive: purpose-built, offline, ONNX (shares the runtime with the NPU transcription path); supports the segmentation+embedding+clustering pipeline the design assumes; models are downloadable and swappable.
  • Negative: separate models to download/manage (model-management UI, FR-MODEL-1); C/C++ FFI to integrate; clustering may over/under-split → the merge workflow (FR-SPK-3) is required, not optional.
  • Speaker IDs (S1, S2, …) are internal and stable per meeting; name mappings live in the DB and are applied at render/export time, never destructively rewritten onto segments.

Revisit if

A single model gives joint ASR + diarization with better accuracy, or whisper.cpp gains production diarization.