- Nouveau moteur « Microsoft Edge » (voix neuronales gratuites, sans clé) : Henri / Denise
par défaut selon le genre, liste des voix fr-FR / fr-CA / fr-CH / fr-BE, WebSocket signé
(Sec-MS-GEC) dans le processus principal, MP3 24 kHz. Moteur par défaut des nouvelles
installations.
- Supertonic 3 ajouté au catalogue local (31 langues, 5 voix masculines + 5 féminines,
44 kHz, 129 Mo) : langue transmise au worker, genres des voix vérifiés par mesure de F0.
- Kokoro déclassé pour le français (une seule voix féminine, accent) ; le choix
masculin/féminin bascule sur le meilleur modèle installé et conserve le genre quand on
change de modèle ; Google Translate signalé comme voix féminine uniquement.
- Deux préréglages JARVIS : en ligne (Edge Henri) et hors ligne (Supertonic 3).
- Traitement du micro par Chromium (écho, bruit, gain) débrayable pour de meilleures
transcriptions au casque.
- Tests : protocole Edge (jeton, SSML, trames), sélection de voix ; README et feuille de
route (mesures Supertonic, Parakeet vs Qwen3-ASR).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MMpgFriwxiBgurUVb21oCE
Voix
- Timbre « JARVIS » (Web Audio) : hauteur légèrement abaissée, chaleur dans
les basses, présence, compression douce, courte réverbération d'intercom.
Activé par défaut, réglable dans Paramètres → Voix.
- Préréglage « Voix JARVIS » en un clic : moteur local, voix masculine
française (Piper Tom téléchargé automatiquement, Kokoro n'ayant pas de voix
française masculine), débit calme. Affichage de la voix active.
- Boutons Masculine / Féminine désormais lisibles (style du sélecteur ajouté).
Reconnaissance
- Parakeet TDT 0.6B v3 (NVIDIA NeMo, int8, 25 langues européennes dont le
français) ajouté au catalogue et recommandé : plus précis et plus rapide que
Whisper sur processeur, ponctuation incluse.
- 0,4 s de silence ajoutées avant et après chaque énoncé avant la
reconnaissance (syllabes coupées, hallucinations de Whisper sur les clips
courts).
Écoute permanente
- Les modules AudioWorklet sont livrés en fichiers statiques (public/worklets)
chargés depuis l'application : en version installée, la CSP stricte
(script-src 'self') refusait les URL blob et l'écoute permanente échouait
avec « Unable to load a worklet's module ». Repli blob conservé pour le dev.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
- Liaison Hermes : l'échec de l'API des crons (/api/jobs absente ou refusée)
ne fait plus passer la liaison en « dégradé » ; l'onglet Crons affiche la
raison. Les états « healthy / ready / ok » sont reconnus et un état
« degraded » détaille les contrôles en échec.
- Chat completions : une réponse sans fragment n'affiche plus « … » ; le corps
est relu (JSON non streamé, erreur renvoyée en HTTP 200) et sinon l'erreur
explicite « Réponse vide de Hermes » est affichée avec le début du corps.
- Voix : transcriptions parasites de Whisper (« (cliquant) », « *Claire* »,
« [Musique] », génériques de sous-titres) ignorées au lieu d'être envoyées.
- Préférence de voix masculine / féminine (Paramètres → Voix), appliquée à
tous les moteurs : Piper Tom / Pierre en local (téléchargement automatique),
onyx / nova en API OpenAI, Paul / Hortense en voix Windows. Genre des voix
du catalogue déduit des libellés.
Tests : 42 tests unitaires (santé, récupération de réponse, filtre de bruit,
préférence de voix) ; e2e Electron inchangé et vert.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
- Catalog: 3.3 MB zipformer keyword-spotting model (kws-en).
- shared/keywords: BPE encoding of wake phrases (SentencePiece table for common words,
greedy longest-match fallback over the model vocabulary) and keywords file builder.
- Worker: KeywordSpotter stream fed with 16-bit PCM, detections pushed as unsolicited
messages; engine derives the keywords file, maps sensitivity to threshold/score,
forwards detections to the renderer and re-arms after a worker restart.
- Renderer: WakeListener keeps one microphone stream, batches 256 ms frames to the
spotter and captures the command on the same stream after detection (pre-roll, VAD),
then resumes spotting; the wake word also interrupts speech. Settings: wake mode
(off / always-on / transcript filter), keyword, sensitivity, status and one-click
model download; HUD caption shows the active keyword.
- Validated: detection in the worker (fork) and through the real Electron IPC path on
Kokoro audio, no false positive on an English recording; 23 unit tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
Electron main
- utility-process channel: sherpa output copied into V8 buffers (Electron rejects external
buffers), audio exchanged as base64; worker restart on timeout, identity-safe exit handling,
deterministic native unload by restart
- webhook: loopback by default without secret, UTF-8-safe body assembly, clean restart, port validation
- files IPC: realpath-based root check, openPath allow-list (reveal-only outside document folders),
Windows reserved names; store: debounced async atomic writes with backup of corrupt files;
logger: streaming writes with rotation; log message size cap
- HTTP stream proxy: socket released on idle timeout / renderer destroyed, id validation
- model manager: inactivity timeout, retrying rm/rename (Windows locks), engine stopped before
replacing a model; WAV decoder handles float/24-bit; IPC payload validation
- window: opaque rounded window on Windows, navigation lock-down, visibility events;
Ctrl+Alt+Escape instead of the Task Manager shortcut; single-instance guard; EVEFLOW_USER_DATA
Voice pipeline
- abort semantics (SendHandle.aborted), abort before the stream opens, session id prefixes per transport
- hands-free re-arm after replies without speech, start/stop race, no chime on auto re-arm,
no silence shipped to STT (400 ms pre-roll), no transcription of empty manual stops
- TTS: bounded prefetch, cancellable segments, non-interrupting notices, volume applied at play time
- SSE CRLF split, usage in chat completions, finish_reason length, phonetic regex hoisted
HUD
- core renderer: no canvas shadows, cached colours, reusable spectrum buffer, theme read on change,
30 fps idle, stops when the window is hidden; ping flashes on send / tool / speech
- deltas coalesced per animation frame; stable auto-scroll; narrow selectors everywhere
- bundled fonts (offline), reduce-motion fix, error toasts, ops drawer below 1180 px,
interim transcript and first-token latency in the core caption, Ctrl+K, dialog semantics,
switch/aria roles, compact widget cleanup, Whisper small recommended for French
- docs/ROADMAP.md: audit results, e2e results and the plan towards a real JARVIS
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
- electron/voice: model catalog (Whisper base/small/turbo, SenseVoice, Kokoro v1.0,
Piper fr), streaming tar.bz2 downloader with progress, sherpa-onnx worker running in
an Electron utilityProcess (transcribe / synthesize / status / unload), IPC + bridge.
- Renderer: 'local' providers for STT and TTS, Settings → Modèles locaux (download,
progress, delete, activate), speaker selection, hands-free wake word with a tolerant
matcher (accents, punctuation, edit distance) and attention window after a bare wake word.
- Packaging: native addon and worker unpacked from the asar; release 2.1.0.
Validated on Linux: Kokoro (fr) → Whisper base round trip, Piper (fr) → Whisper round trip.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
electron-builder refuses a non-semver package.json version, so the published
number lives in package.json releaseVersion and scripts/dist.mjs passes it as
extraMetadata.version (plus the hotfix digit as Windows build number). The
workflow tags and names the release from the same field.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
- Version numbers gain a fourth hotfix component (2.0.0.1); the workflow maps it to the
Windows file build number.
- The system panel shows dashes instead of 0% until real measurements arrive from the
main process.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
The preload script imported ../shared/ipc at runtime; sandboxed preloads cannot
require local modules, so window.eveflow was never exposed in the installed app
and every request fell back to browser fetch (Failed to fetch, LOCALHOST metrics).
Main and preload are now bundled with esbuild into single files in dist-electron/.
Also: explicit error banner when the bridge is missing inside Electron, and a more
precise connection test message (transport path, capabilities status).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
Complete rewrite of the application:
- Toolchain: Electron 44, Vite 8, React 19, TypeScript 5.9, Zustand, Vitest;
main process rewritten in TypeScript (electron/), typed IPC contract (shared/),
sandboxed renderer with webSecurity on and an HTTP/SSE proxy in the main process.
- Voice: AudioWorklet microphone capture at 16 kHz with adaptive energy VAD and
auto-stop, WAV encoder, OpenAI-compatible STT, TTS queue with sentence-level
streaming, prefetch and Web Audio playback feeding an analyser; hands-free mode,
global push-to-talk hotkey, system-voice fallback.
- Hermes: full API client (capabilities, health, models, skills, toolsets,
sessions, jobs) with automatic transport selection: runs API (SSE lifecycle,
approvals, steer, stop) > sessions stream > chat completions with local tools;
tolerant event normalisation; webhook receiver hardened (secret, size limit).
- UI: JARVIS arc-reactor core on Canvas 2D reacting to the real audio signal,
holographic HUD layout (transcript, core, Hermes ops panel with tools/crons/
skills/sessions/link tabs, telemetry), settings drawer, approval modals,
compact floating widget, four themes, tray icon and global shortcuts.
- Removed the 3D Eve robot, three.js and legacy assets; migrated 1.x settings.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y