# Fix plan: speaker diarization accuracy ("You", live labels, 83-speaker explosion) Status: planned 2026-07-13 (v0.5.2). Investigation confirmed in code; see memanto memories `a9e5f29e` (diagnosis) and the entry referencing this file (fix order). Execute phases **in order** — each phase is independently shippable and verifiable, and later phases assume earlier ones landed. ## Symptoms (user report, all reproduced against code) 1. Live recording shows `S1` for every line; the user's own speech should show "You". 2. Post-stop diarization of a 2-person conversation produced **83** speakers. 3. Re-transcribing (English Medium) collapsed every segment to speaker "You". ## Root causes (verified anchors) | # | Cause | Anchor | |---|-------|--------| | A | Live segments always carry the pre-diarization `"S1"` placeholder | `src-tauri/src/transcription/mod.rs:222`, emitted at `src-tauri/src/commands.rs:480` | | B | The 15s provisional diarization tick relabels only the backend buffer and emits `diarization://updated`, which **no frontend code listens to** (`src/lib/api.ts` has no listener); live view renders raw labels with no name map (`src/lib/views/TranscriptNotes.svelte:527`) | `src-tauri/src/commands.rs:509-567` | | C | "You" (mic voiceprint match) runs only in `stop_recording`, never live | `src-tauri/src/commands.rs:666-695`, `src-tauri/src/diarization/voiceprint.rs` | | D | `reprocess_transcript` never diarizes: fresh segments all default to `"S1"`, then the meeting's **stale** name map (`S1 → "You"` from the original run) is reused, so the whole transcript renders "You" | `src-tauri/src/commands.rs:1646-1742` | | E | Clustering over the **summed mic+loopback mono** WAV with `num_clusters: -1, threshold: 0.5` over-clusters badly (overlapped speech → mixed embeddings; short pyannote chunks → unstable ERes2Net embeddings) | `src-tauri/src/diarization/mod.rs:68-75` | --- ## Phase 0 — Tame the cluster explosion (do first; everything else is useless at 83 clusters) Smallest possible change, lands before Phase 1 so re-diarization doesn't reproduce the 83-speaker mess. 1. In `SherpaDiarizer::new` (`src-tauri/src/diarization/mod.rs:73`) raise `threshold` from `0.5` to `0.7` as a code constant with a comment naming the tuning evidence (83 clusters for 2 speakers at 0.5 on mixed mono audio). sherpa fast-clustering semantics: larger threshold → fewer clusters. Do **not** add a settings knob yet. 2. In `segment_to_span`'s caller (`SherpaDiarizer::diarize`), drop spans shorter than 700 ms before returning — sub-second chunks carry unstable embeddings and only cause label churn in `assign_by_overlap` (which already keeps a segment's prior label when no span overlaps). 3. Unit test: spans under the minimum are filtered; existing overlap tests still pass. 4. **Verify empirically**: re-run diarization over the user's retained 2-person recording (the 0.5.2 test meeting) and confirm the cluster count lands near 2–4, not 83. If 0.7 still over-clusters, try 0.8 before considering per-meeting configurability. Acceptance: 2-person retained recording diarizes to ≤4 speakers. ## Phase 1 — `reprocess_transcript` re-diarizes + persistent mic voiceprint (fixes "everything is You") 1. **Persist the voiceprint at stop.** In `stop_recording` (`src-tauri/src/commands.rs` after the transcription worker join), when `session.mic_voice_sample` exists **and** `session.retention` is on, write the mic sample as 16 kHz mono `voiceprint.wav` next to `audio.wav` in `meeting_dir(&meeting_id)`. ADR-0009 gate: it is retained audio of the user's voice, so it lives and dies with `audio.wav` — every code path that deletes `audio.wav` (retention off at finalize, delete-recording command) must delete `voiceprint.wav` too. Update `docs/03-data-model.md` (meeting dir file list). 2. **Re-diarize on reprocess.** In `reprocess_transcript` (`src-tauri/src/commands.rs:1646`), after `transcribe_file`: build the diarizer via `diarizer_from_installed_models` (inside `spawn_blocking`, same as `stop_recording`), `diarize(&wav_path)`, then `assign(&mut segments, &spans)`. Missing models → skip gracefully (same degradation as live). 3. **Rebuild the name map instead of reusing the stale one.** Old labels are meaningless after re-clustering. If `voiceprint.wav` exists, run `voiceprint::match_mic_speaker` against the new spans to get a fresh `You`/`Speaker N` map; otherwise use an empty map (raw `S1…` labels). Pass `speaker_infos_from_segments(&segments, &new_names)` to `finalize_meeting` instead of `meeting.speakers` (`src-tauri/src/commands.rs:1713`), and persist the new names via the store the same way `stop_recording` does. Deliberate policy: user-typed names from the original run are dropped on reprocess because they key to dead labels — note this in the command's doc comment. 4. `import_media` reuses this path (`commands.rs:1825` "same batch path"); confirm imported meetings get diarized speakers too (they have no voiceprint — expect raw labels). 5. Tests: unit test the fresh-name-map policy (stale map not reused); manual: re-transcribe the 0.5.2 meeting with English Medium → segments show distinct speakers, the user's own lines show "You". Acceptance: re-transcription yields per-speaker labels again, with "You" on the mic speaker when a voiceprint exists; never a single-speaker collapse. ## Phase 2 — Live labels reach the UI, including live "You" (fixes "S1 for everyone") Backend (`src-tauri/src/commands.rs`, the 15s tick at 509–567): 1. After `diarizer.assign` in the tick, **re-emit every committed segment whose speaker changed** via the existing `transcript://segment` event (ids are stable; the frontend store already replaces by id — `src/lib/stores/recording.svelte.ts:49-53`). No new event needed for relabeling. 2. In the same tick, run `voiceprint::match_mic_speaker` with `session.mic_voice_sample.samples()` against the fresh spans (clusters re-shuffle every tick, so match every tick; candidate audio is already capped at 10 s per cluster). Merge results into `session.speaker_names` **without overwriting user-set names** — same already-named guard as the post-stop pass (`commands.rs:676-683`). The tick's existing `diarization://updated` emit then carries the "You" display name. 3. Frontend: add an `onDiarizationUpdated` listener to `src/lib/api.ts` (payload `{ meetingId, speakers: SpeakerInfo[] }`, already documented in `docs/04-api-contracts.md:188`). Recording store gains a `speakers` state updated by it; live rendering at `TranscriptNotes.svelte:527` passes that list to `speakerName` (the finalized path at :344 already does this). 4. Docs: note in `docs/04-api-contracts.md` that `transcript://segment` may re-emit a committed segment with an updated `speaker` (replace-by-id contract). 5. Tests: frontend store test — a re-emitted segment with the same id replaces the old one; manual: during a live 2-person call, own speech flips to "You" within ~15–30 s. Acceptance: during recording, labels differentiate speakers and the mic speaker shows "You" while talking (within one tick), not just after stop. ## Phase 3 — Per-stream attribution (accuracy end-game, larger change, needs design sign-off) Blind clustering of a summed mono signal is the ceiling on accuracy. WA knows which samples are mic before `spawn_mixer` sums them (`commands.rs:342-371`, `src-tauri/src/audio/mod.rs` MicBridge/mixer). Design sketch — do NOT start without reviewing ADR-0005 and getting sign-off, since it changes the diarization contract: - Record a coarse mic-activity timeline during capture (e.g. per-100 ms mic-RMS-dominant flags, negligible memory). - At stop: mic-dominant ranges become "You" spans directly; run sherpa only over the remaining (far-side) ranges to split the *other* participants; merge span lists before `assign_by_overlap`. - Voiceprint match remains as fallback for meetings without the timeline (imports, old recordings). - Update ADR-0005 (or add a new ADR) + `docs/02-architecture.md` data flow. Acceptance: 2-person call yields exactly "You" + 1 speaker with near-zero attribution errors on non-overlapping speech; overlapping speech attributes to the dominant stream. --- ## Cross-cutting rules for the executing agent - CLAUDE.md applies in full: `cargo fmt` + `cargo clippy -- -D warnings`, no `unwrap()`/`expect()` on user-reachable paths, conventional commits referencing FR-SPK-*, commit each file right after finishing it (one commit per file), docs updated in the same change when contracts move. - Memory discipline: `memanto agent activate whispassist` first; `memanto remember` every decision/tuning result (especially the empirical threshold from Phase 0) with full metadata; record phase completion so the next agent knows where to resume. - No new egress, no new settings defaults ON; `voiceprint.wav` is local retained audio under the existing ADR-0009 consent/retention gate.