Files
WhispAssist/docs/plans/2026-07-13-diarization-phase3-per-stream-attribution.md
T

244 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Phase 3 design sketch: per-stream speaker attribution (FR-SPK)
Status: **signed off & in progress**, 2026-07-13. Deepens Phase 3 of
[`2026-07-13-diarization-speaker-accuracy.md`](./2026-07-13-diarization-speaker-accuracy.md).
The §§1–8 below are the original sketch (ADR-0005 review + alternatives); the **finalized
design** the user signed off on is at the top and supersedes the RMS-dominance heuristic.
---
## Finalized design (signed off 2026-07-13)
Simpler than the sketch's dominance heuristic, because the user records on **headphones** so the
mic never hears the far side (no bleed):
> **The microphone stream is "You" by definition. Diarization runs on the loopback (speaker)
> audio only, never the mic. Any loopback speaker is "Speaker N".**
Mechanism, all in **`audio.wav` frame time** (so You-spans and sherpa-spans share one clock):
1. **Mic-activity timeline** — captured in the loopback writer, the one place mic and loopback
exist separately in `audio.wav`'s timebase (`capture_loop`, where `bridge.pull()` is folded
in). Per 100 ms, flag whether the pulled mic chunk was speech-level (reuse `VOICE_ENERGY_FLOOR`).
Held in a shared `MicActivity` (sibling to `VoiceSample`); persisted for reprocess.
2. **At stop / reprocess:**
- Collapse active windows → **"You" spans** (dropping sub-`MIN_SPAN_MS` blips).
- **Mask** the You ranges out of `audio.wav` (zero them) and run sherpa on the remainder → far
side splits into **"Speaker N"** spans. Clustering never sees the mic, so the mic can't
pollute or merge with a far-side cluster.
- Merge span lists; `assign_by_overlap` unchanged.
3. **Voiceprint (Phase 1/2) becomes the fallback** for meetings with no timeline (imports, older
recordings). Live-tick voiceprint "You" is retired once the timeline drives labels.
Timebase note: with the capture-silence fix (FR-CAP-7) `audio.wav` is now continuous, so the
frame-indexed timeline aligns cleanly with what sherpa reads back via `read_wav_mono_16k`.
Implementation status (all landed 2026-07-13/14): **(a)** `MicActivity` struct + capture wiring —
done; **(b)** stop-path attribution + masked diarize (`phase3_attribute`) — done; **(c)** persist
`mic_activity.json` + reprocess uses it (imports fall back, no timeline) — done; **(d)** ADR-0005 +
`docs/02`/`03` amended — done. **Not yet done:** empirical user test; and the *live* 15s tick still
uses the Phase 2 voiceprint, not the timeline (a later polish so live labels match the stop result).
---
_Original sketch (ADR-0005 review, risks, alternatives) follows._
Phases 0–2 shipped: cluster explosion tamed (0.7 + 700 ms floor), reprocess re-diarizes with a
persistent `voiceprint.wav`, and live "You" reaches the UI. Those all still treat diarization as
**blind clustering of a summed mono signal**, then guess which cluster is the mic. Phase 3
removes the guess for the mic side by using information WA throws away today.
---
## 1. ADR-0005 review — what it commits us to
ADR-0005 (Accepted, 2026-06-30) decides: sherpa-onnx offline diarization (pyannote segmentation
+ ERes2Net embedding + clustering), behind the `diarization::Diarizer` trait, run as a
**post-processing pass over recorded audio**, aligned to whisper segments by **timestamp
overlap**. Speaker IDs (`S1`…) are internal/stable; names map in the DB, applied at render time,
**never rewritten onto segments**.
What Phase 3 must respect vs. what it may change:
- **Keeps (contract-level):** the `Diarizer` trait, post-pass-over-audio model, `SpeakerSpan`
→ segment alignment by overlap (`assign_by_overlap`), stable internal labels, names-in-DB.
Phase 3 produces *more* `SpeakerSpan`s from a better source; it does not rewrite segments or
move naming.
- **Bends (needs the ADR noted/updated):** ADR-0005 assumes one clustering pass over "the
recorded audio." Phase 3 introduces a **second, non-clustered source of spans** (the mic
activity timeline) and restricts clustering to the **far-side** audio. That is new enough to
warrant an ADR amendment or a short ADR-0005a, because a future reader will otherwise expect
all spans to come from sherpa.
- **ADR's "Revisit if"** (joint ASR+diarization model, or whisper.cpp diarization) is unrelated
— Phase 3 is orthogonal and does not trigger it.
Conclusion: Phase 3 fits inside the trait and the overlap-alignment contract. The only doc debt
is recording that "mic-dominant spans bypass clustering," which is a genuine deviation from
ADR-0005's single-source assumption. **Recommend: amend ADR-0005** (Consequences section) rather
than a new ADR — same decision, refined.
---
## 2. The asset we currently discard
When the mic is enabled (FR-CAP-7), `start_recording` (`commands.rs:342`) wires two independent
16 kHz-mono streams into `spawn_mixer` (`audio/mod.rs:1020`). The `Mixer` (`audio/mod.rs:979`)
holds `loopback` and `mic` **time-aligned in the same buffer** and sums them sample-for-sample
into the transcript stream — then the per-stream identity is **gone**. Diarization later reads
`audio.wav` (the summed signal) and has to reverse-engineer which cluster was the mic. That
reversal is the entire reason `voiceprint.rs` exists.
Two facts make this cheap to exploit:
1. **Per-stream RMS already exists.** `audio_level(mono, mic)` (`audio/mod.rs:522`) runs on every
WASAPI chunk for *both* directions (`is_loopback` at :564, emitted at :678). We already know,
per ~10 ms chunk, how loud each side is. We just don't persist it against time.
2. **The two streams are already aligned at the mixer**, in the **same timebase the transcript
segments derive from** (the streaming worker consumes the mixer's summed output; segment
timestamps are cumulative-samples-fed / 16). So a timeline indexed by cumulative mixer-output
samples shares the segments' clock — tighter than sherpa spans, which come from `audio.wav`.
---
## 3. Core idea
> The mic stream is a **known** speaker. Don't cluster to find it — record when it's dominant,
> call those spans "You", and run sherpa only on the **far-side** audio to split the *other*
> participants.
Pipeline at stop (mic enabled + retained):
```
mixer ─┬─► loopback samples ──► [far-side WAV] ──► sherpa cluster ──► Speaker 2..N spans
│ │
└─► mic vs loopback RMS per 100 ms ──► mic-dominant ranges ──► "You" spans
│
merge span lists ──► assign_by_overlap ──► segments
```
The mic-dominant spans need no embedding, no clustering, no voiceprint — they are attribution by
construction. Voiceprint (Phases 1–2) stays as the **fallback** for meetings without a timeline
(imports, pre-Phase-3 recordings, mic-disabled meetings).
---
## 4. The two real risks (why this needs sign-off, not just a ticket)
### 4a. Acoustic bleed — the mic is not purely "You"
Without echo cancellation (WA has none — fully local, minimal), a user on **speakers** (not
headphones) has the far side playing into the room and back into the mic. Naive "mic has signal →
You" would attribute the far side to the user whenever the far side is loud.
Mitigation (this is why the parent plan says *mic-**dominant**,* not *mic-active*): a range is
"You" only when **mic RMS meaningfully exceeds loopback RMS** in that window (e.g.
`mic_rms > k · loopback_rms` with `k ≈ 2`, plus an absolute mic-VAD floor so silence isn't
"You"). When both are comparably loud → treat as far-side/ambiguous, let sherpa/overlap decide.
This degrades gracefully: headphone users get near-perfect mic isolation; speaker users get
"You" only on clear self-speech and fall back to the old behavior during overlap. **`k` is a
calibration knob, not a constant to bury** — real rooms differ.
### 4b. Timebase — three clocks, currently reconciled by tolerance
- **Segments** are in *mixer-output* time (streaming worker over the summed stream).
- **sherpa spans** are in *`audio.wav`* time (loopback thread's byte-accurate WAV, mic bridged in
at native rate then resampled to 16 kHz on read).
- The **new mic-timeline** would be in *mixer-output* time.
Today `assign_by_overlap` already bridges segments↔sherpa-spans across the first two clocks and
tolerates the small skew. Phase 3 merges a *third* source. The mic-timeline is actually the
**closest** to segment time (shared origin), so aligning You-spans to segments is *more* reliable
than sherpa's. The residual risk is only that You-spans (mixer time) and Speaker-2..N spans
(`audio.wav` time) coexist in one merged list with a small relative skew at the boundaries. In
practice that costs at most a fraction of a `MIN_SPAN_MS` at each turn boundary — below segment
granularity. **Verification item, not a blocker.** If it proves visible, the clean fix is to run
the far-side sherpa pass over the mixer's *loopback lane* buffered to a WAV rather than over
`audio.wav`, putting everything on one clock; heavier, so deferred until measured.
---
## 5. Concrete design
**Capture (live, near-zero cost).** Add a mic-activity accumulator the mixer writes to. In
`Mixer::drain_ready` (or the `mixer_loop`), for each emitted window compute `mic_rms` and
`loopback_rms` over the aligned overlap (samples are already in hand there) and append a coarse
flag per 100 ms:
```rust
// audio/mod.rs — new, alongside Mixer
struct MicActivity { flags: Vec<bool>, /* one per 100 ms of mixer output */ }
// ponytail: Vec<bool> at 10 Hz = 10 bytes/s ≈ 36 KB/hour. No ring buffer, no config.
```
Expose it via the same `Arc` sharing pattern `VoiceSample` uses so `stop_recording` can read it.
Reuse `audio_level`'s RMS math — do **not** add a second RMS implementation.
**Persist (retained meetings).** At stop, alongside `voiceprint.wav`, write the timeline as
`mic_activity.json` (or pack into an existing artifact) so **reprocess** can reuse it — same
ADR-0009 retention/consent gate as `voiceprint.wav`, same lives-and-dies-with-`audio.wav` rule
(delete it everywhere `audio.wav` is deleted). Tiny, plaintext-or-sealed to match.
**Attribute (stop + reprocess).**
1. Collapse the flag timeline into `You` `SpeakerSpan`s (merge adjacent flags, drop runs
`< MIN_SPAN_MS`, reuse Phase 0's floor).
2. Build a **far-side-only** 16 kHz signal: zero out (or excise) mic-dominant ranges from the
loopback lane, run the existing `SherpaDiarizer::diarize` over it → `Speaker 2..N` spans
(label numbering starts at 2, mirroring `voiceprint::build_name_map`).
3. Merge the two span lists; `assign_by_overlap` unchanged.
4. Fallback chain: no timeline → today's Phase-1/2 voiceprint path; no models → raw `S1`.
**Contract/doc changes.**
- Amend **ADR-0005** Consequences: "mic-dominant spans are attributed directly from a per-stream
activity timeline and bypass clustering; sherpa runs on the far-side residual only."
- `docs/02-architecture.md` data-flow: note the mixer emits a mic-activity side-channel.
- `docs/03-data-model.md`: add `mic_activity.json` to the meeting-dir list (like `voiceprint.wav`).
- No new egress; no new default-on setting; `k` threshold lives in code with a calibration comment.
**Tests.**
- Unit: flag-run → span collapse (boundaries, sub-`MIN_SPAN_MS` drop) — pure, table-driven.
- Unit: dominance classifier (`mic`, `loop` RMS pairs → You / far / ambiguous), incl. the bleed
case (both loud → not You).
- Manual: 2-person call on **headphones** → exactly "You" + 1 speaker, near-zero errors; repeat
on **speakers** → verify bleed doesn't mislabel far-side as You.
---
## 6. Alternatives considered (and why not)
- **Do nothing / stop at Phase 2.** Legitimate. Phases 0–2 already fix the reported bugs
(83-cluster explosion, reprocess collapse, live "You"). Phase 3 is an **accuracy ceiling
raise**, not a bug fix. *If 2-person calls are the dominant case and Phase 2's voiceprint "You"
tests well in the field, Phase 3 may not clear the cost/benefit bar — measure Phase 2 first.*
This is the ponytail-honest recommendation: **verify Phase 2 empirically before building 3.**
- **"Mic active → You" (no dominance test).** Simpler, but wrong under speaker bleed (4a). The
dominance test is the minimum that survives real rooms.
- **Full source separation / AEC.** Much larger, new deps, defeats "minimal + local." The
dominance heuristic gets ~90 % of the benefit for ~5 % of the effort.
- **Better voiceprint only** (bigger sample, re-match every tick — already done in Phase 2).
Still bottlenecked on sherpa's ability to *form* a clean mic cluster in the summed signal;
Phase 3's point is to stop relying on that.
---
## 7. Open questions for sign-off
1. **Build it now, or verify Phase 2 first?** Recommendation: run the empirical 2-person
re-diarization (outstanding for Phases 0–2) and a Phase-2 live "You" check **before**
committing to Phase 3. If Phase 2 accuracy is acceptable, Phase 3 becomes optional.
2. **Dominance threshold `k` and mic-VAD floor** — accept a code-constant default with a
calibration comment (proposed), or expose a hidden setting? Recommendation: constant first.
3. **Far-side signal construction** — zero-out mic-dominant ranges in the loopback lane (simple,
preserves timeline) vs. excise-and-concat (shorter audio, shifts timestamps → needs remap).
Recommendation: **zero-out**, keeps one timebase.
4. **ADR: amend 0005 vs. new ADR-0005a?** Recommendation: amend.
5. **Scope of `mic_activity.json`** — persist for reprocess (proposed) or compute live-only and
accept that reprocess falls back to voiceprint? Persisting is cheap and keeps reprocess at
parity.
## 8. Rough effort
Backend-only until the far-side pass; no frontend change (labels already flow via Phase 2). Est.
~1–1.5 days: mixer side-channel + accumulator (½ day), span collapse + far-side pass + merge
(½ day), persistence + reprocess wiring + docs/tests (½ day). Contained, reversible, behind the
existing mic-enabled + retention gates.