Files
WhispAssist/docs/03-data-model.md
T

425 lines
21 KiB
Markdown

# 03 — Data Model
## On-disk layout
Default root: `%LOCALAPPDATA%\WhispAssist\` (user-configurable, FR-STORE-2).
```
%LOCALAPPDATA%\WhispAssist\
├── wa.db # SQLite index (relations, metadata, FTS)
├── settings.json # app settings (theme, provider, storage, presets)
├── models\ # downloaded models (whisper, diarization)
│ ├── whisper-base.q5.bin
│ ├── seg-pyannote-3.0.onnx
│ └── spk-eres2net.onnx
└── meetings\
└── <meeting_id>\ # one folder per meeting (uuid)
├── audio.wav # canonical recording — present ONLY if "Record" was on (ADR-0009). Mic on → stereo split (L=mic/"You", R=loopback/"Speaker", `audio_layout=split`); mic off → summed. Playback folds to mono; export folds to dual-mono (ADR-0005 Phase 3.5)
├── voiceprint.wav # mic-only 16kHz sample for "You" re-matching on reprocess of a SUMMED recording (FR-SPK); lives/dies with audio.wav
├── transcript.json # canonical transcript (segments+speakers+timings)
├── manual_notes.json # raw user-authored notes captured live during recording
├── notes.md # the final notes document: manual notes + transcript, merged at finalize
├── summary.json # LLM summary, decisions, action items (if generated)
└── briefs/ # feature briefs distilled from this meeting (ADR-0011), if any
└── <brief_id>.json # agent-ready spec served via the MCP `get_feature_brief` tool
```
A **bundle export** (`export_meeting` / `bulk_export_meetings` with `format: "bundle"`, FR-STORE-4)
copies a meeting's `audio.wav`, `transcript.json`, `notes.md`, and `summary.json` (all decrypted)
into a destination folder plus a `meeting.json` manifest (the `MeetingBundle`: title, timestamps,
duration, language/backend/model, tags, speakers, and confirmed action items). `import_meeting_bundle`
reconstructs each such folder under a fresh meeting id — the portable format for moving recordings
between computers.
Rule: while a meeting is in progress a working WAV is the source of truth for crash recovery. On
finalize, it is **kept** as `audio.wav` if "Record this meeting" was on, or **deleted** if not
(FR-REC-1/4) — deletion happens only after `transcript.json` is finalized. `transcript.json` and
`summary.json` are **derived** and regenerable (regenerable only while the audio still exists —
i.e. for recorded meetings).
`notes.md` is **generated once, at finalize**, by merging `manual_notes.json` (freeform notes
typed live during the recording, plus any per-moment annotations — see below) with the rendered,
speaker-tagged transcript (`notes::MarkdownNotes::merge`). After that it is the user's own
document, freely editable via `update_notes` exactly like before this changed — nothing
re-renders or overwrites it afterward. In particular, renaming or merging a speaker after finalize
updates the `speakers` table and the live UI display, but does **not** retroactively rewrite text
already baked into `notes.md` (same as any other manual edit isn't retroactively touched either —
this was a pre-existing clobber bug this redesign also fixes: renaming a speaker used to silently
overwrite the whole file). A crash-recovery finalize (T2.8) and a post-finalize batch
re-transcription (T3.8) both re-render `notes.md` from scratch and so both re-read
`manual_notes.json` from disk to fold the same manual notes back in.
### `manual_notes.json`
```jsonc
{
"schema": 1,
"freeform_md": "string — the user's running notes, typed live in the Notes pane while recording",
"segment_notes": [
{ "anchor_ms": 12345, "text": "string", "created_at": 1735000000, "updated_at": 1735000010 }
]
}
```
`segment_notes[].anchor_ms` is a timestamp into the recording (a clicked transcript segment's
`start_ms`), not a segment id — a later batch re-transcription can renumber/regenerate segment
ids, but never moves the moment in time a note was attached to. At merge time, each note is placed
right after whichever transcript paragraph's time span contains its `anchor_ms`; a note whose
anchor doesn't land inside any paragraph surfaces under an "Other notes" section instead of being
silently dropped. Written to disk on every edit via `update_live_notes`/`set_segment_note`
(`04-api-contracts.md`) — live-session only, same write-through-for-crash-safety spirit as
`transcript.json` accumulating during recording.
## SQLite schema (`wa.db`)
```sql
-- A meeting is the top-level record.
CREATE TABLE meetings (
id TEXT PRIMARY KEY, -- uuid v4
title TEXT NOT NULL DEFAULT 'Untitled meeting',
started_at INTEGER NOT NULL, -- unix epoch seconds
ended_at INTEGER, -- null while recording
duration_secs INTEGER, -- finalized on stop
folder_path TEXT NOT NULL, -- absolute path to meeting folder
audio_path TEXT, -- audio.wav; NULL if not retained (ADR-0009)
recorded INTEGER NOT NULL DEFAULT 0, -- 1 = audio retained as .wav, 0 = transcript-only
status TEXT NOT NULL, -- recording|transcribing|ready|recovering|error
language TEXT, -- detected/selected language code
backend_used TEXT, -- npu|nvidia|amd|intel|cpu
model_used TEXT, -- e.g. whisper-base
audio_layout TEXT, -- split|summed|NULL; how audio.wav's channels are laid out (0009, ADR-0005 Phase 3.5)
calendar_event_id TEXT, -- FK -> calendar_events.id (nullable)
template_id TEXT, -- note-template id (T8.1, FR-NOTE-5); catalog is
-- a built-in Rust list (notes::templates), not a table
created_at INTEGER NOT NULL,
updated_at INTEGER NOT NULL
);
-- Internal speakers detected per meeting (S1, S2, …) and their assigned names.
CREATE TABLE speakers (
id TEXT PRIMARY KEY, -- uuid
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
label TEXT NOT NULL, -- 'S1','S2',… (internal, stable per meeting)
display_name TEXT, -- user/participant name (nullable)
participant_id TEXT REFERENCES participants(id), -- if mapped to a calendar attendee
color TEXT, -- UI color hint
merged_into TEXT, -- non-null: this label folds into another label's
-- row at render/export time (T4.5, FR-SPK-3); segment
-- speaker IDs in storage are never rewritten (FR-SPK-5)
UNIQUE(meeting_id, label)
);
-- People known from calendar/.pst; reused across meetings for continuity.
CREATE TABLE participants (
id TEXT PRIMARY KEY, -- uuid
name TEXT NOT NULL,
email TEXT, -- nullable
UNIQUE(name, email)
);
CREATE TABLE meeting_participants ( -- attendee list per meeting
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
participant_id TEXT NOT NULL REFERENCES participants(id),
role TEXT, -- organizer|required|optional
PRIMARY KEY (meeting_id, participant_id)
);
-- Calendar events imported from .pst (or future sources).
CREATE TABLE calendar_events (
id TEXT PRIMARY KEY, -- uuid (stable from source uid if available)
source TEXT NOT NULL, -- pst|graph|ics
subject TEXT,
organizer TEXT,
starts_at INTEGER,
ends_at INTEGER,
description TEXT,
raw_uid TEXT -- source's own id for dedup
);
-- Re-importing the same event (matched by source+raw_uid) updates it in place (T6.2).
CREATE UNIQUE INDEX idx_calendar_events_source_uid ON calendar_events(source, raw_uid)
WHERE raw_uid IS NOT NULL;
-- Attendees per imported calendar event, populated at import time (T6.1) —
-- distinct from meeting_participants, which links a *recording* to people
-- once a meeting is attached to an event (T6.3/T6.6).
CREATE TABLE calendar_event_participants (
calendar_event_id TEXT NOT NULL REFERENCES calendar_events(id) ON DELETE CASCADE,
participant_id TEXT NOT NULL REFERENCES participants(id),
role TEXT, -- organizer|required|optional
PRIMARY KEY (calendar_event_id, participant_id)
);
CREATE TABLE tags (
id TEXT PRIMARY KEY,
name TEXT NOT NULL UNIQUE
);
CREATE TABLE meeting_tags (
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
tag_id TEXT NOT NULL REFERENCES tags(id) ON DELETE CASCADE,
PRIMARY KEY (meeting_id, tag_id)
);
-- Action items parsed from LLM output, then user-confirmed.
CREATE TABLE action_items (
id TEXT PRIMARY KEY,
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
text TEXT NOT NULL,
owner TEXT, -- assignee name (nullable)
due_at INTEGER, -- nullable
confirmed INTEGER NOT NULL DEFAULT 0, -- 0=suggested, 1=confirmed by user
reminder_set INTEGER NOT NULL DEFAULT 0,
created_at INTEGER NOT NULL
);
-- Feature briefs distilled from a meeting for coding-agent handoff (ADR-0011). The JSON body
-- lives in the meeting's briefs/ folder; this table indexes it for the MCP tools.
CREATE TABLE feature_briefs (
id TEXT PRIMARY KEY, -- uuid
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
title TEXT NOT NULL,
target_repo TEXT, -- optional repo hint for the agent
path TEXT NOT NULL, -- briefs/<id>.json
exposed INTEGER NOT NULL DEFAULT 0, -- visible to the MCP server? (scope control, FR-MCP-3)
created_at INTEGER NOT NULL
);
-- Audit of what an MCP client (agent) read (FR-MCP-5).
CREATE TABLE mcp_access_log (
id TEXT PRIMARY KEY,
at INTEGER NOT NULL,
tool TEXT NOT NULL, -- e.g. get_feature_brief
meeting_id TEXT, -- subject, if any
client TEXT -- client-reported name, if provided
);
-- Configured sync/upload destinations (ADR-0010). Credentials are NOT stored here —
-- only a reference into the OS credential store (FR-SYNC-6).
CREATE TABLE sync_targets (
id TEXT PRIMARY KEY, -- uuid
name TEXT NOT NULL, -- user label, e.g. "Home Nextcloud"
kind TEXT NOT NULL, -- webdav|onedrive|dropbox|box
provider_hint TEXT, -- nextcloud|owncloud|cloudreve|seafile|synology|generic
base_url TEXT, -- WebDAV URL (kind=webdav)
remote_base_path TEXT NOT NULL DEFAULT '/WhispAssist',
username TEXT, -- WebDAV username (secret is in credential store)
credential_ref TEXT NOT NULL, -- key into OS credential store
enabled INTEGER NOT NULL DEFAULT 0, -- off by default (FR-SYNC-1)
upload_transcript INTEGER NOT NULL DEFAULT 1,
upload_notes INTEGER NOT NULL DEFAULT 1,
upload_summary INTEGER NOT NULL DEFAULT 1,
upload_recording INTEGER NOT NULL DEFAULT 0, -- only meaningful if a meeting is recorded
trigger_on_finalize INTEGER NOT NULL DEFAULT 1,
allow_plaintext_lan INTEGER NOT NULL DEFAULT 0, -- FR-SYNC-7
encrypt_before_upload INTEGER NOT NULL DEFAULT 0, -- FR-SYNC-10
created_at INTEGER NOT NULL
);
-- One upload job per (artifact, target). The durable queue (FR-SYNC-5).
CREATE TABLE sync_jobs (
id TEXT PRIMARY KEY, -- uuid
target_id TEXT NOT NULL REFERENCES sync_targets(id) ON DELETE CASCADE,
meeting_id TEXT NOT NULL REFERENCES meetings(id) ON DELETE CASCADE,
artifact TEXT NOT NULL, -- transcript|notes|summary|recording
local_path TEXT NOT NULL,
remote_path TEXT NOT NULL,
sha256 TEXT, -- skip-if-unchanged idempotency
status TEXT NOT NULL, -- pending|uploading|done|failed|skipped
attempts INTEGER NOT NULL DEFAULT 0,
last_error TEXT,
next_attempt_at INTEGER, -- backoff schedule
bytes_total INTEGER,
bytes_sent INTEGER NOT NULL DEFAULT 0,
updated_at INTEGER NOT NULL,
UNIQUE(target_id, meeting_id, artifact)
);
-- Full-text search over transcript + notes (FR-SEARCH-1).
CREATE VIRTUAL TABLE meeting_fts USING fts5(
meeting_id UNINDEXED,
title,
transcript_text,
notes_text,
tokenize = 'porter unicode61'
);
CREATE INDEX idx_meetings_started ON meetings(started_at DESC);
CREATE INDEX idx_speakers_meeting ON speakers(meeting_id);
CREATE INDEX idx_action_meeting ON action_items(meeting_id);
CREATE INDEX idx_syncjobs_status ON sync_jobs(status, next_attempt_at);
CREATE INDEX idx_briefs_meeting ON feature_briefs(meeting_id);
```
Schema migrations are versioned (e.g. `sqlx::migrate!`), each migration numbered and forward-only.
## `transcript.json` (canonical transcript)
```jsonc
{
"schema": 1,
"meeting_id": "f1c2…",
"language": "en",
"model": "whisper-base",
"backend": "nvidia",
"segments": [
{
"id": 0,
"start_ms": 1240, // offset from recording start
"end_ms": 4880,
"speaker": "S1", // internal speaker id; name resolved at render time
"text": "Let's start with the roadmap.",
"confidence": 0.91, // optional
"interim": false, // true while streaming, false once finalized
},
],
"speakers": [
{ "label": "S1", "display_name": "Alex" },
{ "label": "S2", "display_name": null },
],
}
```
Names are stored as a mapping and applied at render/export — segments keep the internal `speaker`
label so re-diarization or renaming never requires rewriting every segment (FR-SPK-5).
## `summary.json`
```jsonc
{
"schema": 1,
"generated_at": 1751299200,
"provider": "ollama",
"model": "llama3",
"summary_md": "## Summary\n…",
"decisions": ["Adopt Tauri for the shell"],
"action_items": [
// Same shape as the `action_items` table row (ActionItem) minus the row
// never having existed yet: `id`/`due_at` are null until the user
// reviews and confirms a drafted item (FR-LLM-3), at which point
// `confirm_action_items` creates the real row.
{
"id": null,
"text": "Send the API contract draft",
"owner": "Jordan",
"due_at": null,
"confirmed": false,
},
],
}
```
## `briefs/<brief_id>.json` (feature brief — ADR-0011)
Agent-ready spec the MCP `get_feature_brief` tool returns. Designed to drop straight into a coding
agent's context. Written by `create_feature_brief` (M1); one file per brief under the meeting's
`briefs/` folder, **sealed at rest with the vault** when unlocked (T8.8), exactly like
`summary.json`. The `feature_briefs` table indexes it (id, meeting_id, title, target_repo, path,
exposed) for `list_feature_briefs` and the MCP scope check — the file is the source of truth; the
row is the index. The IPC `FeatureBrief` type (`04-api-contracts.md`) is the subset returned to the
UI / MCP: everything below **except** the `schema`/provenance envelope (`generated_at`, `provider`,
`model`, `source`).
```jsonc
{
"schema": 1,
"id": "b7a1…",
"meeting_id": "f1c2…",
"generated_at": 1751299200, // envelope: which model distilled this, and when
"provider": "ollama",
"model": "llama3",
"title": "Bulk CSV export for the reporting view",
"problem": "Customer can't get their data out for offline analysis.",
"desired_outcome": "One-click CSV export of the current filtered report.",
"acceptance_criteria": [
"Export button on the report toolbar",
"Respects active filters and column order",
"Streams large exports without blocking the UI",
],
"target_repo": "acme/reporting-web", // optional hint the user supplies at create time
"context_excerpts": [
// verbatim transcript quotes that ground the request — NOT model paraphrase;
// the builder selects them from the real transcript (speaker = resolved display name)
{ "speaker": "Customer", "text": "We really need to pull this into our own spreadsheets." },
],
"source": { "meeting_title": "Acme quarterly sync", "at": 1751299200 },
}
```
Field presence: `title`, `problem`, `desired_outcome` are always strings (the builder falls back to
the meeting title / `""` on a sparse model reply); `acceptance_criteria` and `context_excerpts` may
be empty arrays. Every `context_excerpts[].text` is a verbatim substring of a real transcript
segment (the M1 grounding invariant, asserted by the golden-transcript test).
## `settings.json`
```jsonc
{
"theme": "system", // system|light|dark
"storage_root": "%LOCALAPPDATA%\\WhispAssist",
"retention": { "max_age_days": 90, "max_size_gb": 10 },
"transcription": { "model": "base", "language": "auto", "batch_reprocess": false },
"hardware": { "preferred_backend": "auto", "low_overhead": false },
"recording": {
"default_record": false, // "Record this meeting" default — OFF (FR-REC-1)
"consent_acknowledged": false, // set true after the one-time consent notice (FR-REC-2)
},
"llm": {
"provider": "ollama", // ollama|custom|anthropic|openai|off (ADR-0007/0011; "openai" not yet wired)
"endpoint": "http://localhost:11434",
"model": "llama3",
"stream": true,
// API keys for hosted providers (anthropic|openai) live in the OS credential store, not here.
"hosted_ai_acknowledged": false, // one-time "data leaves your device" notice ack (T10.3, ADR-0011)
},
// Sync target rows live in wa.db (sync_targets); secrets live in the OS credential store.
// settings.json only holds the global default. No credentials here (FR-SYNC-6).
"sync": { "enabled": false }, // master off switch — OFF by default (FR-SYNC-1)
// Local MCP server (ADR-0011). OFF by default; inbound loopback only; token in credential store.
"mcp": {
"enabled": false, // FR-MCP-1
"transport": "http", // http (127.0.0.1/mcp) | stdio
"port": 4849,
"expose": "selected", // none | selected | all (scope control, FR-MCP-3)
"expose_recordings": false, // never serve .wav unless explicitly true (FR-MCP-3)
},
"privacy": { "encrypt_at_rest": false },
// Launch WhispAssist at login (NFR-RES-4). Opt-in, OFF by default. Toggling this
// via `set_auto_start` also writes a per-user `HKCU\...\Run` entry (no admin);
// startup reconciles the OS entry to this flag (e.g. after a reinstall).
"auto_start": false,
// Closing the window hides WhispAssist to the system tray (keep running in background) instead of
// quitting; ON by default. Tray "Quit" is the real exit. Enforced in the Rust on_window_event handler.
"close_to_tray": true,
// Optional MS Graph calendar source (M4.4, T8.9, FR-CAL-6). Opt-in, explicit consent via OAuth
// PKCE — OFF by default. `credential_ref` points into the OS credential store; the token itself
// is never written here (same invariant as sync credentials, FR-SYNC-6).
"calendar": { "graph_enabled": false, "graph_credential_ref": null },
}
```
## First-run deploy seeding (`wa-defaults.ini`)
For enterprise mass-deployment, the **first** launch on a machine (before `settings.json` exists)
optionally seeds its defaults from an admin-supplied INI. First file found wins:
1. `%PROGRAMDATA%\WhispAssist\wa-defaults.ini` — machine-wide (GPO / SCCM / Intune file copy).
2. `<install dir>\wa-defaults.ini` — the bundled template (shipped fully commented → no-op).
Keys are flat `key = value` matching `settings.json` field names (bools/ints coerced), plus the
special `auto_download_model = true` which fetches the configured `whisper_model` in the background.
**No secrets** — any key containing `key`/`token`/`secret`/`credential`/`password` is ignored; those
live only in the OS credential store. After first run the file is never read again. See
`docs/enterprise-deployment.md` and `src-tauri/src/deploy.rs`.
## Retention & recovery semantics
- **Retention** (FR-STORE-2): a background job deletes whole meeting folders + rows once a meeting
exceeds the age/size policy, oldest first; never runs while a meeting is `recording`/`transcribing`.
- **Recovery** (FR-REL-1): on startup, any meeting whose **working WAV still exists** but whose
`transcript.json` is missing/partial is marked `recovering`; the user can re-run transcription from
the audio. (A cleanly finalized non-recorded meeting has no audio and needs no recovery; an
interrupted one still has its working WAV, so it is always recoverable.)
- **Sync** (FR-SYNC-5): on startup and after each finalize, the `sync` service scans `sync_jobs` for
`pending`/`failed` jobs whose `next_attempt_at` has passed and resumes them with backoff. Deleting
a meeting cascades to its `sync_jobs` (local rows only; already-uploaded remote copies are left to
the user/server).