Compare commits

...
15 Commits
Author SHA1 Message Date
LogiFlow b8073c0be6 feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0) (#21)
feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0)
2026-09-04 17:40:11 +02:00
Claude 32551ccc0d feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0)
- Nouveau moteur « Microsoft Edge » (voix neuronales gratuites, sans clé) : Henri / Denise
  par défaut selon le genre, liste des voix fr-FR / fr-CA / fr-CH / fr-BE, WebSocket signé
  (Sec-MS-GEC) dans le processus principal, MP3 24 kHz. Moteur par défaut des nouvelles
  installations.
- Supertonic 3 ajouté au catalogue local (31 langues, 5 voix masculines + 5 féminines,
  44 kHz, 129 Mo) : langue transmise au worker, genres des voix vérifiés par mesure de F0.
- Kokoro déclassé pour le français (une seule voix féminine, accent) ; le choix
  masculin/féminin bascule sur le meilleur modèle installé et conserve le genre quand on
  change de modèle ; Google Translate signalé comme voix féminine uniquement.
- Deux préréglages JARVIS : en ligne (Edge Henri) et hors ligne (Supertonic 3).
- Traitement du micro par Chromium (écho, bruit, gain) débrayable pour de meilleures
  transcriptions au casque.
- Tests : protocole Edge (jeton, SSML, trames), sélection de voix ; README et feuille de
  route (mesures Supertonic, Parakeet vs Qwen3-ASR).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MMpgFriwxiBgurUVb21oCE
2026-09-04 15:37:00 +00:00
LogiFlow a95f91185f Merge pull request #20 from R0m1k3/claude/refonte-app-vocale-v91uz7
Claude/refonte app vocale v91uz7
2026-09-04 13:22:03 +02:00
LogiFlowandClaude Fable 5.1 5ef440e8cf feat: voix JARVIS, Parakeet v3 pour le français, worklets audio sous CSP stricte (v2.4.1) (#19)
* feat: voix JARVIS, Parakeet v3 pour le français, worklets audio sous CSP stricte (v2.4.1)

Voix
- Timbre « JARVIS » (Web Audio) : hauteur légèrement abaissée, chaleur dans
  les basses, présence, compression douce, courte réverbération d'intercom.
  Activé par défaut, réglable dans Paramètres → Voix.
- Préréglage « Voix JARVIS » en un clic : moteur local, voix masculine
  française (Piper Tom téléchargé automatiquement, Kokoro n'ayant pas de voix
  française masculine), débit calme. Affichage de la voix active.
- Boutons Masculine / Féminine désormais lisibles (style du sélecteur ajouté).

Reconnaissance
- Parakeet TDT 0.6B v3 (NVIDIA NeMo, int8, 25 langues européennes dont le
  français) ajouté au catalogue et recommandé : plus précis et plus rapide que
  Whisper sur processeur, ponctuation incluse.
- 0,4 s de silence ajoutées avant et après chaque énoncé avant la
  reconnaissance (syllabes coupées, hallucinations de Whisper sur les clips
  courts).

Écoute permanente
- Les modules AudioWorklet sont livrés en fichiers statiques (public/worklets)
  chargés depuis l'application : en version installée, la CSP stricte
  (script-src 'self') refusait les URL blob et l'écoute permanente échouait
  avec « Unable to load a worklet's module ». Repli blob conservé pour le dev.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

* docs: mesures Parakeet vs Whisper et worklets dans la feuille de route

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-04 13:02:32 +02:00
Claude e55d70bd22 docs: mesures Parakeet vs Whisper et worklets dans la feuille de route
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-04 11:02:08 +00:00
Claude 813425b622 feat: voix JARVIS, Parakeet v3 pour le français, worklets audio sous CSP stricte (v2.4.1)
Voix
- Timbre « JARVIS » (Web Audio) : hauteur légèrement abaissée, chaleur dans
  les basses, présence, compression douce, courte réverbération d'intercom.
  Activé par défaut, réglable dans Paramètres → Voix.
- Préréglage « Voix JARVIS » en un clic : moteur local, voix masculine
  française (Piper Tom téléchargé automatiquement, Kokoro n'ayant pas de voix
  française masculine), débit calme. Affichage de la voix active.
- Boutons Masculine / Féminine désormais lisibles (style du sélecteur ajouté).

Reconnaissance
- Parakeet TDT 0.6B v3 (NVIDIA NeMo, int8, 25 langues européennes dont le
  français) ajouté au catalogue et recommandé : plus précis et plus rapide que
  Whisper sur processeur, ponctuation incluse.
- 0,4 s de silence ajoutées avant et après chaque énoncé avant la
  reconnaissance (syllabes coupées, hallucinations de Whisper sur les clips
  courts).

Écoute permanente
- Les modules AudioWorklet sont livrés en fichiers statiques (public/worklets)
  chargés depuis l'application : en version installée, la CSP stricte
  (script-src 'self') refusait les URL blob et l'écoute permanente échouait
  avec « Unable to load a worklet's module ». Repli blob conservé pour le dev.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-04 10:57:55 +00:00
LogiFlowandClaude d1b12c16c5 fix: portail laissant passer /health mais bloquant /v1 traité comme liaison en échec (v2.4.0.3) (#18)
- Une page web reçue sur /v1/capabilities fait échouer la connexion et le test
  de liaison même si /health répond, ce qui déclenche la recherche de l'API.
- La découverte sonde /v1/capabilities puis /v1/models (JSON exigé, 401/403
  accepté) avant /health, pour ne pas prendre un portail pour l'API.


Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-04 08:22:19 +02:00
LogiFlowandClaude df25dd591c fix: page web reçue à la place de l'API Hermes détectée, URL découverte automatiquement (v2.4.0.2) (#17)
- Toute réponse HTML (portail de connexion, tableau de bord, erreur de proxy)
  est reconnue et expliquée avec le titre de la page, au lieu d'être prise
  pour un état « ok » ou de produire une réponse vide.
- Découverte automatique : quand la liaison échoue, EveFlow essaie l'API sur le
  même hôte (port 8642, /api, /v1, sous-domaines api. et hermes.) et corrige
  l'URL si un serveur Hermes répond. Le bouton « Tester la liaison » fait de
  même et affiche les essais.

Tests : 47 tests unitaires ; e2e Electron vert.


Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-04 08:06:05 +02:00
LogiFlow c75c1d674c Merge pull request #16 from R0m1k3/claude/refonte-app-vocale-v91uz7
fix: liaison « dégradé », réponses vides, transcriptions parasites, v…
2026-09-04 07:53:40 +02:00
LogiFlowandClaude 6d3073b083 fix: liaison « dégradé », réponses vides, transcriptions parasites, voix masculine (v2.4.0.1) (#15)
- Liaison Hermes : l'échec de l'API des crons (/api/jobs absente ou refusée)
  ne fait plus passer la liaison en « dégradé » ; l'onglet Crons affiche la
  raison. Les états « healthy / ready / ok » sont reconnus et un état
  « degraded » détaille les contrôles en échec.
- Chat completions : une réponse sans fragment n'affiche plus « … » ; le corps
  est relu (JSON non streamé, erreur renvoyée en HTTP 200) et sinon l'erreur
  explicite « Réponse vide de Hermes » est affichée avec le début du corps.
- Voix : transcriptions parasites de Whisper (« (cliquant) », « *Claire* »,
  « [Musique] », génériques de sous-titres) ignorées au lieu d'être envoyées.
- Préférence de voix masculine / féminine (Paramètres → Voix), appliquée à
  tous les moteurs : Piper Tom / Pierre en local (téléchargement automatique),
  onyx / nova en API OpenAI, Paul / Hortense en voix Windows. Genre des voix
  du catalogue déduit des libellés.

Tests : 42 tests unitaires (santé, récupération de réponse, filtre de bruit,
préférence de voix) ; e2e Electron inchangé et vert.


Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-04 07:44:07 +02:00
Claude 31659500ab fix: liaison « dégradé », réponses vides, transcriptions parasites, voix masculine (v2.4.0.1)
- Liaison Hermes : l'échec de l'API des crons (/api/jobs absente ou refusée)
  ne fait plus passer la liaison en « dégradé » ; l'onglet Crons affiche la
  raison. Les états « healthy / ready / ok » sont reconnus et un état
  « degraded » détaille les contrôles en échec.
- Chat completions : une réponse sans fragment n'affiche plus « … » ; le corps
  est relu (JSON non streamé, erreur renvoyée en HTTP 200) et sinon l'erreur
  explicite « Réponse vide de Hermes » est affichée avec le début du corps.
- Voix : transcriptions parasites de Whisper (« (cliquant) », « *Claire* »,
  « [Musique] », génériques de sous-titres) ignorées au lieu d'être envoyées.
- Préférence de voix masculine / féminine (Paramètres → Voix), appliquée à
  tous les moteurs : Piper Tom / Pierre en local (téléchargement automatique),
  onyx / nova en API OpenAI, Paul / Hortense en voix Windows. Genre des voix
  du catalogue déduit des libellés.

Tests : 42 tests unitaires (santé, récupération de réponse, filtre de bruit,
préférence de voix) ; e2e Electron inchangé et vert.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-04 05:43:41 +00:00
LogiFlowandClaude 6fc38d335b feat: serveur MCP, heures calmes, mode mission, widget glanceable, barge-in (v2.4.0) (#14)
Serveur MCP local
- Endpoint /mcp (JSON-RPC 2.0, Streamable HTTP, réponses JSON) sur le serveur
  webhook : initialize, tools/list, tools/call, ping. Même secret que le
  webhook. 14 outils : capture_screen (image MCP), lock_session, open_app,
  open_url, media_key, clipboard_get/set, find_files, speak_text,
  notify_user, set_hud_state, get_app_status, get_conversation_history,
  show_message. Les outils UI transitent par IPC vers le renderer.
- Les mêmes actions système sont proposées au modèle en chat completions.

Notifications
- Heures calmes (plage horaire, franchissement de minuit), mots prioritaires,
  échecs de crons toujours lus, résumé vocal des messages entrants (n premières
  phrases), thème nuit automatique, compteur de non-lus.

Conversation
- Mode mission : bouton dans la barre de commande, second modèle Hermes
  (hermes.missionModel) pour les tâches longues.
- Barge-in : la parole de l'utilisateur coupe la voix ; seuil d'énergie et
  ratio relevés pendant la synthèse pour ignorer l'écho.

Widget compact
- Bandeau glanceable : état, dernière phrase, non-lus, indicateurs heures
  calmes et mission.

Tests : 28 tests unitaires (heures calmes, priorité, résumé), e2e Electron
sous Xvfb : MCP initialize/tools/list/tools/call (presse-papiers, capture en
image, état et message via le renderer), bandeau compact, aucune erreur.


Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-04 00:47:34 +02:00
LogiFlowandClaude af8bdc5e9d feat: fin de phrase Silero, vision d'écran et actions locales (v2.3.0) (#13)
Écoute permanente
- Silero VAD dans le worker sherpa-onnx (vad.start/audio/stop) ; le renderer
  envoie des trames de 128 ms pendant la commande et reçoit le segment WAV
  complet (speech-start / segment). Repli sur le VAD énergétique si le modèle
  n'est pas installé ou si l'option est désactivée.
- Modèle « silero-vad » (0,6 Mo) dans le catalogue, section « Fin de phrase »
  dans Modèles locaux, réglage dans Paramètres → Micro.

Vision d'écran
- IPC system:screen-capture (desktopCapturer, JPEG 1600 px) ; bouton dans la
  barre de commande ; « regarde mon écran… » joint la capture à la question
  envoyée à Hermes (transport chat completions pour les images).

Actions locales
- IPC system:action à liste blanche : verrouiller la session, touches média
  (volume, mute, lecture, piste), ouvrir une application connue ou une URL
  http(s), presse-papiers, recherche de fichiers dans les dossiers utilisateur.
- Routeur d'intentions FR/EN exécuté avant Hermes (« coupe le son », « ouvre
  Spotify », « verrouille la session »…), résultat affiché et lu ; désactivable.

Correctif
- Chat completions : les premiers fragments SSE arrivés avant l'événement de
  démarrage étaient perdus (premier mot manquant) ; ils sont maintenant rejoués.

Tests : 26 tests unitaires (intentions locales ajoutées), e2e Electron sous
Xvfb avec les vrais modèles (Silero : un segment par phrase, 6,7 s d'audio
traités en 190 ms ; capture 109 ko reçue par Hermes ; intention locale
traitée sans Hermes).


Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-03 20:46:26 +02:00
LogiFlow 4f6074183f feat: écoute permanente par détection de mot-clé (v2.2.0) (#12)
feat: écoute permanente par détection de mot-clé (v2.2.0)
2026-09-03 20:22:00 +02:00
Claude 0d284fea9e feat: always-on wake word with sherpa-onnx keyword spotting (v2.2.0)
- Catalog: 3.3 MB zipformer keyword-spotting model (kws-en).
- shared/keywords: BPE encoding of wake phrases (SentencePiece table for common words,
  greedy longest-match fallback over the model vocabulary) and keywords file builder.
- Worker: KeywordSpotter stream fed with 16-bit PCM, detections pushed as unsolicited
  messages; engine derives the keywords file, maps sensitivity to threshold/score,
  forwards detections to the renderer and re-arms after a worker restart.
- Renderer: WakeListener keeps one microphone stream, batches 256 ms frames to the
  spotter and captures the command on the same stream after detection (pre-roll, VAD),
  then resumes spotting; the wake word also interrupts speech. Settings: wake mode
  (off / always-on / transcript filter), keyword, sensitivity, status and one-click
  model download; HUD caption shows the active keyword.
- Validated: detection in the worker (fork) and through the real Electron IPC path on
  Kokoro audio, no false positive on an English recording; 23 unit tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-03 18:21:35 +00:00
56 changed files with 3484 additions and 134 deletions

No files matched your search

+35 -4
View File
@@ -1,7 +1,7 @@
# EveFlow 2 — Interface vocale JARVIS pour Hermes Agent
[![Build](https://img.shields.io/github/actions/workflow/status/R0m1k3/EveFlow/windows-release.yml?style=flat-square)](https://github.com/R0m1k3/EveFlow/actions)
[![Version](https://img.shields.io/badge/version-2.1.0-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases)
[![Version](https://img.shields.io/badge/version-2.5.0-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases)
[![License](https://img.shields.io/badge/license-MIT-lightgrey.svg?style=flat-square)](LICENSE)
**EveFlow** est un compagnon de bureau Windows qui transforme [Hermes Agent](https://hermes-agent.nousresearch.com/) en assistant vocal à la JARVIS : un noyau holographique réactif au son, une conversation en streaming, les outils, sous-agents, approbations, crons, skills et sessions d'Hermes pilotés depuis un seul HUD.
@@ -22,10 +22,23 @@ La version 2 est une réécriture complète : plus de robot 3D, un pipeline voca
* **Capture micro** via AudioWorklet à 16 kHz, sans monitoring du micro dans les haut-parleurs, avec annulation d'écho et réduction de bruit.
* **Détection d'activité vocale** (seuil adaptatif, sensibilité et silence de fin réglables) : l'enregistrement s'arrête tout seul quand vous avez fini de parler.
* **Mains libres** : le micro se réactive après chaque réponse.
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse Kokoro v1.0 (voix française Siwis et voix anglaises) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
* **Mot d'activation** en mains libres : seules les phrases commençant par « Jarvis » (configurable) partent vers Hermes, le reste est ignoré ; un « Jarvis » seul ouvre une fenêtre d'écoute.
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Parakeet v3, Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse **Supertonic 3** (31 langues dont le français, cinq voix masculines et cinq féminines, 44 kHz, 129 Mo, environ 7× plus rapide que le temps réel sur 4 cœurs), Kokoro v1.0 (excellent en anglais ; en français une seule voix féminine avec accent) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
* **Fin de phrase neuronale** : en écoute permanente, Silero VAD (0,6 Mo, sherpa-onnx) décide du début et de la fin de la commande à la place du seuil d'énergie ; moins de faux départs sur le bruit, coupure plus nette. Repli automatique sur le VAD énergétique si le modèle n'est pas installé.
* **Vision d'écran** : « Jarvis, regarde mon écran » (ou le bouton de la barre de commande) joint une capture de l'écran principal à la question envoyée à Hermes.
* **Actions locales instantanées** : « verrouille la session », « monte le son », « coupe le son », « piste suivante », « ouvre Spotify », « ouvre github.com »… exécutées sur le PC sans passer par Hermes, résultat lu à voix haute. Liste blanche d'actions dans le processus principal, désactivable dans les paramètres.
* **Serveur MCP intégré** : Hermes se connecte à `http://<pc>:7842/mcp` et obtient les outils du PC (capture d'écran renvoyée en image, verrouillage, applications, URL, touches média, presse-papiers, recherche de fichiers, voix, notifications, état du HUD, affichage dans le fil). Même port et même secret que le webhook ; en mode chat completions, les mêmes outils sont proposés directement au modèle.
* **Heures calmes et priorités** : plage horaire pendant laquelle les messages poussés s'affichent sans être lus ni faire clignoter le noyau (badge « non lus » à la place), thème nuit automatique, mots prioritaires lus quand même, résumé vocal des rapports longs (les premières phrases seulement).
* **Mode mission** : un bouton dans la barre de commande bascule sur un second modèle Hermes (plus puissant) pour les tâches longues ; le modèle rapide reste utilisé pour la conversation courante.
* **Widget compact « glanceable »** : état (veille, écoute, réflexion, parle), dernière phrase de l'assistant, badge de non-lus, indicateurs heures calmes et mission.
* **Voix Microsoft Edge** (moteur par défaut) : les voix neuronales de la lecture à voix haute d'Edge, gratuites, sans clé ni installation : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR) et les voix fr-CA, fr-CH, fr-BE, plus de 300 voix dans 74 langues. Le rendu le plus naturel disponible ; nécessite une connexion.
* **Voix JARVIS** : deux préréglages en un clic (Paramètres → Voix) : en ligne (Edge Henri) ou hors ligne (Supertonic 3, voix masculine grave, téléchargé automatiquement), timbre « JARVIS » (légèrement plus grave et posé, chaleur, présence, courte réverbération d'intercom), débit calme.
* **Reconnaissance française de référence** : Parakeet TDT 0.6B v3 (NVIDIA NeMo, 25 langues européennes) dans le catalogue, plus précis et bien plus rapide que Whisper sur processeur, avec ponctuation. Whisper base/small/turbo restent disponibles.
* **Voix masculine ou féminine** : un réglage unique (Paramètres → Voix) appliqué à tous les moteurs. Henri / Denise sur Edge ; en local Supertonic 3 (ou Piper Tom / Pierre) pour le masculin, téléchargé automatiquement si aucune voix masculine n'est installée, et le genre est conservé quand on change de modèle ; onyx / nova pour les API compatibles OpenAI ; Paul / Hortense pour les voix Windows. Google Translate n'a qu'une voix féminine par langue : le réglage l'indique.
* **Barge-in** : en mains libres, parler par-dessus l'assistant coupe sa voix ; le seuil est relevé pendant qu'il parle pour ignorer l'écho du haut-parleur.
* **Écoute permanente** : un détecteur de mot-clé de 3 Mo (sherpa-onnx, keyword spotting) tourne en continu sur le micro, quasi gratuit en CPU. « Jarvis » (ou n'importe quel mot-clé) ouvre l'écoute, « Jarvis, allume… » envoie directement la commande, et le mot coupe la voix en cours. Alternative : filtre du mot après transcription en mains libres.
* **STT externe** : n'importe quelle API `/v1/audio/transcriptions` compatible OpenAI (Qwen3-ASR, Whisper, Speaches, faster-whisper-server, LocalAI, OpenAI). Repli sur la reconnaissance Chromium.
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI, voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI (Qwen3-TTS via un serveur compatible, Kokoro-FastAPI, OpenAI, LocalAI…), voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
* **Traitement du micro débrayable** : l'annulation d'écho, la réduction de bruit et le gain automatique de Chromium peuvent être coupés (Paramètres → Reconnaissance) ; avec un casque, le signal brut est souvent mieux transcrit.
* Raccourcis globaux : `Ctrl+Shift+Espace` (micro), `Ctrl+Shift+J` (afficher/masquer), `Ctrl+Shift+Échap` (couper la voix).
### Hermes, toute la puissance
@@ -89,6 +102,24 @@ Pour recevoir les résultats de crons ou le miroir d'autres canaux dans EveFlow,
---
### Donner à Hermes les outils du PC (MCP)
Dans `~/.hermes/config.yaml` côté Hermes :
```yaml
mcp_servers:
eveflow:
url: "http://<ip-du-pc>:7842/mcp"
headers:
Authorization: "Bearer <secret du webhook EveFlow>"
```
Sans secret, EveFlow n'écoute qu'en local (`127.0.0.1`) ; définissez un secret dans Paramètres → Webhook pour un Hermes distant. Outils exposés : `capture_screen`, `lock_session`, `open_app`, `open_url`, `media_key`, `clipboard_get`, `clipboard_set`, `find_files`, `speak_text`, `notify_user`, `set_hud_state`, `get_app_status`, `get_conversation_history`, `show_message`.
### L'URL répond par une page web ?
Si « Tester la liaison » signale « Le serveur renvoie une page web … au lieu de l'API Hermes », l'URL saisie pointe vers un portail (page de connexion, tableau de bord) et non vers le serveur API. EveFlow cherche alors automatiquement l'API sur le même hôte (port 8642, chemins `/api` et `/v1`, sous-domaines `api.` ou `hermes.`) et corrige l'URL s'il la trouve. Sinon, ouvrez le port 8642 du serveur Hermes ou exposez-le sur un chemin dédié de votre proxy, sans authentification web devant lui (la clé `API_SERVER_KEY` suffit).
## Développement
```bash
+50 -23
View File
@@ -8,9 +8,18 @@
|---|---|---|
| HUD arc-reactor réactif au son | Fait | Canvas 2D optimisé (pas d'ombres, couleurs en cache, 30 fps en veille, arrêt fenêtre masquée) |
| Reconnaissance vocale locale | Fait | Whisper base/small/turbo via sherpa-onnx dans un processus utilitaire |
| Synthèse vocale locale | Fait | Kokoro v1.0 (voix française Siwis) et Piper fr |
| Mot d'activation | Fait, mode « après transcription » | Filtre « Jarvis … » en mains libres, tolérant aux erreurs de transcription |
| Détection de fin de phrase | Fait | VAD énergétique adaptatif, pré-roll 400 ms |
| Synthèse vocale locale | Fait (2.5.0) | Supertonic 3 (31 langues, 5 voix masculines + 5 féminines, 44 kHz), Kokoro v1.0 (anglais ; français féminin avec accent) et Piper fr |
| Synthèse vocale en ligne | Fait (2.5.0) | Voix neuronales Microsoft Edge (Henri, Denise, Rémy, Vivienne…), gratuites, sans clé, via WebSocket signé dans le processus principal |
| Mot d'activation permanent | Fait (2.2.0) | Keyword spotting sherpa-onnx en continu ; mot-clé libre encodé en BPE ; validé sur audio réel (détection, zéro faux positif sur le test anglais) |
| Mot d'activation après transcription | Fait | Filtre « Jarvis … » en mains libres, tolérant aux erreurs de transcription |
| Détection de fin de phrase | Fait (2.3.0) | Silero VAD neuronal dans le worker (segment renvoyé au renderer), repli sur le VAD énergétique si le modèle manque |
| Vision d'écran | Fait (2.3.0) | Capture `desktopCapturer` jointe à la requête Hermes (bouton, ou « regarde mon écran ») |
| Serveur MCP local (outils du PC pour Hermes) | Fait (2.4.0) | `/mcp` sur le serveur webhook, JSON-RPC Streamable HTTP, 14 outils dont la capture d'écran renvoyée en image |
| Heures calmes, priorités, résumé vocal | Fait (2.4.0) | Paramètres → Notifications ; thème nuit automatique ; badge non-lus |
| Mode mission (second modèle) | Fait (2.4.0) | Bouton dans la barre de commande ; modèle dédié aux tâches longues |
| Widget compact glanceable | Fait (2.4.0) | État, dernière phrase, non-lus, indicateurs |
| Barge-in | Fait (2.4.0) | Coupe la voix dès que l'utilisateur parle ; seuil relevé pendant la synthèse (à valider avec l'annulation d'écho Windows) |
| Actions système locales | Fait (2.3.0) | Verrouillage, volume et touches média, ouvrir une application ou une URL, presse-papiers, recherche de fichiers ; intentions courtes exécutées sans passer par Hermes |
| Hermes : runs, sessions, chat completions | Fait | Transport choisi selon `/v1/capabilities` |
| Approbations, steer, stop | Fait | Modales, injection de consigne en cours de run |
| Crons, skills, toolsets, sessions | Fait | Panneau Hermes Ops |
@@ -21,35 +30,41 @@
Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai), [JarvisAi](https://github.com/PanPenek/JarvisAi), [bertrandmbanwi/Jarvis](https://github.com/bertrandmbanwi/Jarvis), [InterGenJLU/jarvis](https://github.com/InterGenJLU/jarvis), [livekit-wakeword](https://livekit.com/blog/livekit-wakeword), [sherpa-onnx keyword spotting](https://k2-fsa.github.io/sherpa/onnx/kws/index.html), [Hermes Agent features](https://hermes-agent.nousresearch.com/docs/user-guide/features/overview).
1. **Mot d'activation permanent, quasi gratuit en CPU.** Les projets de référence utilisent openWakeWord (« hey jarvis ») ou un modèle de keyword spotting qui écoute en continu, au lieu de transcrire chaque phrase. sherpa-onnx fournit un modèle KWS anglais de 3,3 Mo qui accepte n'importe quel mot-clé sans réentraînement ; « JARVIS » s'encode `▁JA R VI S` avec son modèle BPE (vérifié). C'est le prochain chantier prioritaire : streaming du micro vers le worker, détection en continu, puis capture de la commande.
2. **VAD neuronal (Silero) au lieu du seuil d'énergie.** Fin de phrase plus nette (environ 500 ms gagnés) et beaucoup moins de faux départs sur le bruit ambiant. Silero est déjà livré dans sherpa-onnx (`silero_vad.onnx`, 0,6 Mo).
1. **Mot d'activation permanent, quasi gratuit en CPU.** Les projets de référence utilisent openWakeWord (« hey jarvis ») ou un modèle de keyword spotting qui écoute en continu, au lieu de transcrire chaque phrase. sherpa-onnx fournit un modèle KWS anglais de 3,3 Mo qui accepte n'importe quel mot-clé sans réentraînement ; « JARVIS » s'encode `▁JA R VI S` avec son modèle BPE (vérifié). Livré en 2.2.0 (voir l'étape 1 ci-dessous).
2. **VAD neuronal (Silero) au lieu du seuil d'énergie.** Fin de phrase plus nette (environ 500 ms gagnés) et beaucoup moins de faux départs sur le bruit ambiant. Silero est déjà livré dans sherpa-onnx (`silero_vad.onnx`, 0,6 Mo). Livré en 2.3.0.
3. **Latence perçue sous la seconde.** Les références visent 1 s entre la fin de parole et le premier mot prononcé : STT rapide, premier token en streaming, TTS phrase par phrase (déjà en place), et un modèle Hermes rapide pour la conversation courante.
4. **Vision d'écran.** Capture d'écran à la demande (« Jarvis, qu'est-ce que je regarde ? ») envoyée à Hermes comme image, ou lecture d'une fenêtre. Hermes accepte déjà les images inline.
5. **Actions système locales.** Ouvrir une application, régler le volume, verrouiller la session, chercher un fichier ; ce sont des outils EveFlow côté client à exposer à Hermes (mode chat completions) ou un petit serveur MCP local que Hermes appelle.
4. **Vision d'écran.** Capture d'écran à la demande (« Jarvis, qu'est-ce que je regarde ? ») envoyée à Hermes comme image, ou lecture d'une fenêtre. Hermes accepte déjà les images inline. Livré en 2.3.0 (transport chat completions pour les images).
5. **Actions système locales.** Ouvrir une application, régler le volume, verrouiller la session, chercher un fichier. Livré en 2.3.0 sous forme d'intentions courtes exécutées localement avant Hermes ; reste à exposer les mêmes actions à Hermes via un serveur MCP local.
6. **Proactivité.** Notifications parlées à l'arrivée d'un cron, rappel, événement webhook, avec un résumé plutôt que la lecture intégrale ; c'est en partie fait via le webhook, à enrichir avec des règles (heures calmes, priorité).
7. **Mémoire et personnalisation.** Hermes gère la mémoire longue durée (`X-Hermes-Session-Key`) ; côté EveFlow, un profil (nom, préférences de voix, style de réponse) déjà transmis dans les instructions.
## Plan proposé
### Étape 1 (courte) : écoute permanente
- Streamer l'audio du micro (16 kHz, blocs de 128 ms) du renderer vers le worker via IPC.
- Dans le worker : `KeywordSpotter` sherpa-onnx (modèle gigaspeech 3,3 Mo, mot-clé configurable encodé automatiquement avec `bpe.model` via un petit encodeur BPE côté Node ou un dictionnaire pré-encodé pour « jarvis », « eve », « hey jarvis », « ok jarvis »).
- À la détection : chime, capture de la commande avec Silero VAD, transcription, envoi.
- Consommation attendue : quelques pourcents d'un cœur, pas de transcription en continu.
### Étape 1 : écoute permanente — livrée en 2.2.0
- Le renderer garde un seul flux micro (AudioWorklet 16 kHz) et envoie des blocs de 256 ms au processus principal, qui alimente le `KeywordSpotter` sherpa-onnx dans le worker.
- Mots-clés encodés en BPE (table SentencePiece pour les mots courants, repli glouton sur le vocabulaire du modèle), sensibilité réglable (seuil 0,45 → 0,12).
- À la détection : chime, capture de la commande sur le même flux (VAD énergétique, pré-roll 400 ms), transcription locale ou API, envoi à Hermes ; retour automatique à l'écoute.
- Fin de phrase Silero livrée en 2.3.0 : le renderer envoie des trames de 128 ms au worker pendant la commande, le worker renvoie le segment WAV complet ; VAD énergétique en repli.
### Étape 2 : vision et actions locales
- Outil `capture_screen` (Electron `desktopCapturer`) qui joint une capture à la requête Hermes.
- Outils système : `open_app`, `set_volume`, `lock_session`, `find_file`, `clipboard` ; exposés en chat completions et via un serveur MCP local pour les transports runs/sessions.
### Étape 2 : vision et actions locales — livrée en 2.3.0
- Capture d'écran (`desktopCapturer`, JPEG 1600 px) jointe à la requête Hermes : bouton dans la barre de commande, ou phrase « regarde mon écran… » à l'oral comme à l'écrit.
- Actions locales (processus principal, liste blanche) : verrouiller, volume/mute/lecture/piste, ouvrir une application connue ou une URL http(s), presse-papiers, recherche de fichiers dans Documents/Bureau/Téléchargements/Images.
- Routeur d'intentions FR/EN (`src/services/localCommands.ts`) exécuté avant l'envoi à Hermes ; désactivable dans Paramètres → Micro. Les phrases composées (« ouvre X puis… ») partent à Hermes.
- Exposition à Hermes via le serveur MCP local livrée en 2.4.0.
### Étape 3 : conversation plus naturelle
- Barge-in réel : couper la voix dès que l'utilisateur parle (déjà préparé, à valider avec l'annulation d'écho Windows).
- Réponses courtes à l'oral, détails à l'écran : instruction Hermes dédiée déjà en place, à affiner avec des consignes de format (« deux phrases à l'oral, détails en Markdown »).
- Modèle rapide pour le bavardage, modèle puissant pour les missions (choix par transport ou par mot-clé).
### Étape 3 : conversation plus naturelle — livrée en 2.4.0
- Barge-in : l'utilisateur qui parle par-dessus coupe la voix ; seuil d'énergie relevé pendant la synthèse pour ignorer l'écho (à valider sur Windows avec l'annulation d'écho du micro).
- Mode mission : second modèle Hermes choisi d'un clic pour les tâches longues.
- Serveur MCP local : Hermes enchaîne lui-même les actions du PC dans ses runs.
### Étape 4 : présence
- Widget compact « glanceable » : dernière phrase, état, badge d'alertes.
- Heures calmes, priorité des notifications, résumé vocal des crons.
- Thèmes et voix par contexte (nuit, travail).
### Étape 4 : présence — livrée en 2.4.0
- Widget compact glanceable : état, dernière phrase, non-lus, indicateurs heures calmes / mission.
- Heures calmes, mots prioritaires, résumé vocal des messages entrants, thème nuit.
### Pistes suivantes
- Voix différente par contexte (nuit, travail) et profils de réponse.
- Mémoire locale des préférences transmise à Hermes (`X-Hermes-Session-Key` déjà en place).
- Validation du barge-in et des touches média sur Windows réel (retours utilisateurs).
## Résultats du test de bout en bout (Linux, Xvfb, 4 cœurs lents)
@@ -60,5 +75,17 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai)
| Kokoro (fr) → Whisper base, phrase courte de 2 s | synthèse 2,9 s, transcription 2,5 s, texte approximatif |
| Kokoro (fr) → Whisper base, phrase de 7 s | transcription correcte à un mot près |
| Piper (fr) → Whisper base | transcription exacte |
| Silero VAD sur phrase Kokoro de 2,3 s (2.3.0) | un seul segment, début et fin détectés, 6,7 s d'audio traités en 190 ms |
| Capture d'écran → Hermes (2.3.0) | JPEG de 107 ko reçu côté Hermes (mock chat completions) |
| Intention locale « coupe le son » (2.3.0) | traitée sans Hermes, résultat affiché dans le fil |
| Voix JARVIS et Parakeet (2.4.1) | timbre JARVIS (Web Audio : pitch, EQ, compression, réverbération courte), préréglage voix masculine française, Parakeet TDT v3 pour le français, worklets audio livrés en fichiers statiques (l'écoute permanente échouait sous la CSP stricte en version installée), rembourrage de silence avant la reconnaissance |
| Correctif 2.4.0.3 | un portail qui laisse passer /health mais renvoie une page web sur /v1 est traité comme une liaison en échec et déclenche la recherche de l'API (sondes /v1/capabilities, /v1/models) |
| Correctif 2.4.0.2 | une page web (portail de connexion) reçue à la place de l'API Hermes est reconnue et expliquée ; l'API est recherchée automatiquement sur le même hôte (port 8642, /api, hermes.…) et l'URL corrigée |
| Correctifs 2.4.0.1 | réponse vide en chat completions désormais expliquée (JSON non streamé ou erreur HTTP 200), la liaison ne passe plus en « dégradé » quand seule l'API des crons échoue, transcriptions parasites (« (cliquant) », « *Claire* ») ignorées, préférence de voix masculine/féminine |
| Parakeet v3 vs Whisper base sur trois phrases Piper (fr) (2.4.1) | Parakeet : 3/3 exactes avec ponctuation, 0,4 à 0,5 s à chaud (6,7 s au premier appel) ; Whisper base : erreurs sur « Jarvis », « Peux-tu », 0,7 à 0,9 s |
| Worklets audio sous CSP stricte (2.4.1) | chargement des modules statiques OK dans l'application empaquetée |
| Voix françaises (2.5.0) | Edge Henri : MP3 reçu de bout en bout (32 ko pour 4 s). Supertonic 3 en français : 10 voix, RTF 0,14 sur 4 cœurs (8,7 s d'audio en 1,3 s), genres déterminés par mesure de la fréquence fondamentale (voix 0-4 : 170-210 Hz, voix 5-9 : 92-137 Hz) ; le worker compilé accepte `language` et bascule sur l'anglais pour une langue inconnue |
| Reconnaissance française : Parakeet v3 vs Qwen3-ASR 0.6B int8 (2.5.0) | Six phrases Supertonic (3,8 s) : Parakeet WER 8,5 % (erreurs surtout de forme : « 14h30 »), 384 ms par phrase ; Qwen3-ASR WER 15,3 % (« mémoires vivres »), 1 477 ms, 940 Mo. Qwen3-ASR n'est pas ajouté au catalogue |
| Serveur MCP (2.4.0) | initialize, tools/list (14 outils), tools/call côté principal (presse-papiers, capture image) et côté renderer (état, message dans le fil) |
Sur un PC à 28 cœurs les temps sont nettement plus courts. Whisper small est maintenant recommandé pour le français.
+184
View File
@@ -0,0 +1,184 @@
/**
* Local "JARVIS" actions: screenshot for vision, and an allow-list of system actions
* (lock, open an application or URL, media keys, clipboard, file search).
*/
import { app, clipboard, desktopCapturer, ipcMain, screen, shell } from 'electron';
import { execFile } from 'node:child_process';
import fs from 'node:fs';
import os from 'node:os';
import path from 'node:path';
import { IPC } from '../../shared/ipc';
import type { SystemAction, SystemActionResult } from '../../shared/bridge';
import { log } from '../logger';
const MEDIA_KEYS: Record<string, number> = {
'volume-up': 175,
'volume-down': 174,
mute: 173,
'play-pause': 179,
next: 176,
previous: 177
};
/** Applications the assistant may launch by name (Windows aliases + common Linux/macOS names). */
const APP_ALIASES: Record<string, string[]> = {
'bloc-notes': ['notepad.exe', 'gedit', 'TextEdit'],
notepad: ['notepad.exe', 'gedit', 'TextEdit'],
calculatrice: ['calc.exe', 'gnome-calculator', 'Calculator'],
calc: ['calc.exe', 'gnome-calculator', 'Calculator'],
explorateur: ['explorer.exe', 'nautilus', 'Finder'],
explorer: ['explorer.exe', 'nautilus', 'Finder'],
terminal: ['wt.exe', 'cmd.exe', 'gnome-terminal', 'Terminal'],
cmd: ['cmd.exe'],
powershell: ['powershell.exe'],
paint: ['mspaint.exe'],
chrome: ['chrome', 'google-chrome', 'Google Chrome'],
edge: ['msedge', 'microsoft-edge', 'Microsoft Edge'],
firefox: ['firefox', 'Firefox'],
vscode: ['code', 'Visual Studio Code'],
code: ['code', 'Visual Studio Code'],
spotify: ['spotify', 'Spotify'],
discord: ['discord', 'Discord'],
steam: ['steam', 'Steam'],
word: ['winword', 'Microsoft Word'],
excel: ['excel', 'Microsoft Excel'],
outlook: ['outlook', 'Microsoft Outlook'],
teams: ['ms-teams', 'Microsoft Teams'],
'task manager': ['taskmgr.exe'],
'gestionnaire des tâches': ['taskmgr.exe'],
paramètres: ['ms-settings:'],
settings: ['ms-settings:']
};
function run(cmd: string, args: string[], timeoutMs = 8000): Promise<{ code: number; out: string }> {
return new Promise((resolve) => {
execFile(cmd, args, { timeout: timeoutMs, windowsHide: true }, (err, stdout, stderr) => {
resolve({ code: err ? 1 : 0, out: `${stdout}${stderr}`.trim() });
});
});
}
async function openApp(name: string): Promise<SystemActionResult> {
const key = name.trim().toLowerCase();
if (!key || key.length > 40) return { ok: false, message: 'Nom d’application invalide' };
const candidates = APP_ALIASES[key] ?? [key.replace(/[^a-z0-9 ._-]/gi, '')];
if (process.platform === 'win32') {
for (const candidate of candidates) {
if (candidate.endsWith(':')) {
await shell.openExternal(candidate);
return { ok: true, message: `${name} ouvert` };
}
// `start` resolves App Paths, PATH and Start Menu names.
const result = await run('cmd.exe', ['/c', 'start', '', candidate]);
if (result.code === 0) return { ok: true, message: `${name} lancé` };
}
return { ok: false, message: `Impossible de lancer ${name}` };
}
for (const candidate of candidates) {
const result = process.platform === 'darwin' ? await run('open', ['-a', candidate]) : await run('sh', ['-c', `command -v ${JSON.stringify(candidate)} >/dev/null && (nohup ${JSON.stringify(candidate)} >/dev/null 2>&1 &)`]);
if (result.code === 0) return { ok: true, message: `${name} lancé` };
}
return { ok: false, message: `Application introuvable : ${name}` };
}
async function lockSession(): Promise<SystemActionResult> {
if (process.platform === 'win32') {
const r = await run('rundll32.exe', ['user32.dll,LockWorkStation']);
return { ok: r.code === 0, message: r.code === 0 ? 'Session verrouillée' : r.out };
}
if (process.platform === 'darwin') {
const r = await run('osascript', ['-e', 'tell application "System Events" to keystroke "q" using {command down, control down}']);
return { ok: r.code === 0, message: r.out };
}
const r = await run('sh', ['-c', 'loginctl lock-session || xdg-screensaver lock || gnome-screensaver-command -l']);
return { ok: r.code === 0, message: r.code === 0 ? 'Session verrouillée' : r.out };
}
async function mediaKey(key: string): Promise<SystemActionResult> {
const code = MEDIA_KEYS[key];
if (!code) return { ok: false, message: 'Touche inconnue' };
if (process.platform === 'win32') {
const script = `$s=Add-Type -MemberDefinition '[DllImport("user32.dll")] public static extern void keybd_event(byte b,byte s,uint f,UIntPtr e);' -Name K -Namespace W -PassThru; $s::keybd_event(${code},0,0,[UIntPtr]::Zero); $s::keybd_event(${code},0,2,[UIntPtr]::Zero)`;
const r = await run('powershell.exe', ['-NoProfile', '-NonInteractive', '-Command', script]);
return { ok: r.code === 0, message: r.code === 0 ? key : r.out };
}
const xdo: Record<string, string> = { 'volume-up': 'XF86AudioRaiseVolume', 'volume-down': 'XF86AudioLowerVolume', mute: 'XF86AudioMute', 'play-pause': 'XF86AudioPlay', next: 'XF86AudioNext', previous: 'XF86AudioPrev' };
const r = await run('sh', ['-c', `xdotool key ${xdo[key]}`]);
return { ok: r.code === 0, message: r.code === 0 ? key : 'xdotool indisponible' };
}
async function findFiles(query: string): Promise<SystemActionResult> {
const needle = query.trim().toLowerCase();
if (needle.length < 2) return { ok: false, message: 'Recherche trop courte' };
const roots = ['Documents', 'Desktop', 'Downloads', 'Pictures'].map((d) => path.join(os.homedir(), d));
const hits: string[] = [];
const deadline = Date.now() + 4000;
const walk = async (dir: string, depth: number) => {
if (depth > 4 || hits.length >= 25 || Date.now() > deadline) return;
let entries: fs.Dirent[] = [];
try {
entries = await fs.promises.readdir(dir, { withFileTypes: true });
} catch {
return;
}
for (const entry of entries) {
if (entry.name.startsWith('.') || entry.name === 'node_modules') continue;
const full = path.join(dir, entry.name);
if (entry.name.toLowerCase().includes(needle)) hits.push(full);
if (entry.isDirectory()) await walk(full, depth + 1);
if (hits.length >= 25) return;
}
};
for (const root of roots) await walk(root, 0);
return { ok: true, message: `${hits.length} résultat(s)`, data: hits };
}
export async function captureScreen(maxWidth = 1600): Promise<string> {
const display = screen.getPrimaryDisplay();
const scale = Math.min(1, maxWidth / display.size.width);
const sources = await desktopCapturer.getSources({
types: ['screen'],
thumbnailSize: { width: Math.round(display.size.width * scale), height: Math.round(display.size.height * scale) }
});
const primary = sources.find((s) => s.display_id === String(display.id)) ?? sources[0];
if (!primary) throw new Error('Aucun écran capturable');
return primary.thumbnail.toJPEG(82).toString('base64').replace(/^/, 'data:image/jpeg;base64,');
}
export async function runSystemAction(action: SystemAction): Promise<SystemActionResult> {
switch (action.type) {
case 'lock':
return lockSession();
case 'open-app':
return openApp(String(action.name ?? ''));
case 'open-url': {
const url = String(action.url ?? '');
if (!/^https?:\/\//i.test(url) || url.length > 2048) return { ok: false, message: 'URL refusée' };
await shell.openExternal(url);
return { ok: true, message: 'Ouvert dans le navigateur' };
}
case 'media':
return mediaKey(String(action.key));
case 'clipboard-read':
return { ok: true, data: (await clipboard.readText()).slice(0, 20_000) };
case 'clipboard-write':
await clipboard.writeText(String(action.text ?? '').slice(0, 100_000));
return { ok: true, message: 'Copié dans le presse-papiers' };
case 'find-files':
return findFiles(String(action.query ?? ''));
default:
return { ok: false, message: 'Action inconnue' };
}
}
export function registerSystemIpc(): void {
ipcMain.handle(IPC.screenCapture, (_e, maxWidth?: number) => captureScreen(Number(maxWidth) || 1600));
ipcMain.handle(IPC.systemAction, async (_e, action: SystemAction) => {
if (!action || typeof action.type !== 'string') throw new Error('Action invalide');
log('INFO', 'system', `action ${action.type}`);
const result = await runSystemAction(action);
if (!result.ok) log('WARN', 'system', `action ${action.type} failed: ${result.message}`);
return result;
});
void app;
}
+2
View File
@@ -9,6 +9,7 @@ import { getSharedDirectory, registerFilesIpc } from './ipc/files';
import { createMainWindow, getMainWindow, getWindowMode, setWindowMode, toggleWindowVisibility, windowEvents } from './window';
import { getWebhookStatus, startWebhookServer, stopWebhookServer } from './webhook';
import { registerVoiceIpc } from './voice/ipc';
import { registerSystemIpc } from './ipc/system';
import { stopEngine } from './voice/engine';
// Audio playback must never be blocked behind a user gesture (TTS starts on incoming events).
@@ -159,6 +160,7 @@ async function boot(): Promise<void> {
registerFilesIpc();
registerCoreIpc();
registerVoiceIpc();
registerSystemIpc();
createMainWindow();
createTray();
registerShortcuts();
+191
View File
@@ -0,0 +1,191 @@
/**
* Minimal MCP server (Streamable HTTP, JSON-RPC 2.0) mounted on the webhook HTTP server at /mcp.
* Hermes Agent connects to it as a remote MCP server and gains the PC-side tools: screen capture,
* system actions, voice, notifications and HUD state. No SDK: initialize / tools/list / tools/call
* are the only methods a client needs, and every reply is a plain JSON body.
*/
import { ipcMain, type BrowserWindow } from 'electron';
import type { IncomingMessage, ServerResponse } from 'node:http';
import { IPC, type McpToolRequest, type McpToolResponse } from '../shared/ipc';
import type { SystemAction } from '../shared/bridge';
import { log } from './logger';
import { captureScreen, runSystemAction } from './ipc/system';
export const MCP_PATH = '/mcp';
const PROTOCOL = '2025-03-26';
const RENDERER_TIMEOUT_MS = 20_000;
interface ToolSpec {
name: string;
description: string;
inputSchema: Record<string, unknown>;
/** 'main' = executed here; 'renderer' = forwarded to the UI process. */
where: 'main' | 'renderer';
}
const obj = (properties: Record<string, unknown>, required: string[] = []) => ({ type: 'object', properties, required, additionalProperties: false });
export const MCP_TOOLS: ToolSpec[] = [
{ name: 'capture_screen', description: "Capture l'écran principal de l'utilisateur et renvoie l'image (JPEG). À utiliser quand l'utilisateur parle de ce qu'il voit ou demande de l'aide sur son écran.", inputSchema: obj({ max_width: { type: 'integer', description: 'Largeur maximale en pixels (défaut 1600)' } }), where: 'main' },
{ name: 'lock_session', description: "Verrouille la session Windows/Linux/macOS de l'utilisateur.", inputSchema: obj({}), where: 'main' },
{ name: 'open_app', description: "Lance une application sur le PC de l'utilisateur (bloc-notes, calculatrice, chrome, spotify, vscode, terminal, explorateur…).", inputSchema: obj({ name: { type: 'string' } }, ['name']), where: 'main' },
{ name: 'open_url', description: "Ouvre une URL http(s) dans le navigateur par défaut de l'utilisateur.", inputSchema: obj({ url: { type: 'string' } }, ['url']), where: 'main' },
{ name: 'media_key', description: 'Envoie une touche média : volume-up, volume-down, mute, play-pause, next, previous.', inputSchema: obj({ key: { type: 'string', enum: ['volume-up', 'volume-down', 'mute', 'play-pause', 'next', 'previous'] } }, ['key']), where: 'main' },
{ name: 'clipboard_get', description: 'Lit le texte du presse-papiers.', inputSchema: obj({}), where: 'main' },
{ name: 'clipboard_set', description: 'Place un texte dans le presse-papiers.', inputSchema: obj({ text: { type: 'string' } }, ['text']), where: 'main' },
{ name: 'find_files', description: "Cherche des fichiers par nom dans Documents, Bureau, Téléchargements et Images de l'utilisateur (25 résultats max).", inputSchema: obj({ query: { type: 'string' } }, ['query']), where: 'main' },
{ name: 'speak_text', description: "Fait prononcer un texte par la voix d'EveFlow, immédiatement.", inputSchema: obj({ text: { type: 'string' } }, ['text']), where: 'renderer' },
{ name: 'notify_user', description: 'Affiche une notification système (titre + corps).', inputSchema: obj({ title: { type: 'string' }, body: { type: 'string' } }, ['body']), where: 'renderer' },
{ name: 'set_hud_state', description: 'Change momentanément l’état visuel du HUD : neutral, happy, thinking, alert, error.', inputSchema: obj({ state: { type: 'string', enum: ['neutral', 'happy', 'thinking', 'alert', 'error'] } }, ['state']), where: 'renderer' },
{ name: 'get_app_status', description: "État d'EveFlow : transport, liaison, nom de l'assistant, mains libres, voix en cours, heures calmes.", inputSchema: obj({}), where: 'renderer' },
{ name: 'get_conversation_history', description: 'Derniers messages de la conversation EveFlow (n ≤ 30).', inputSchema: obj({ n: { type: 'integer' } }), where: 'renderer' },
{ name: 'show_message', description: "Affiche un message dans le fil EveFlow (sans le prononcer) — pour les rapports longs ou les résultats de crons.", inputSchema: obj({ text: { type: 'string' }, title: { type: 'string' } }, ['text']), where: 'renderer' }
];
type Rec = Record<string, unknown>;
interface RpcRequest {
jsonrpc?: string;
id?: number | string | null;
method?: string;
params?: Rec;
}
const pending = new Map<string, { resolve: (r: McpToolResponse) => void; timer: ReturnType<typeof setTimeout> }>();
let seq = 0;
let ipcRegistered = false;
function ensureIpc(): void {
if (ipcRegistered) return;
ipcRegistered = true;
ipcMain.on(IPC.mcpResponse, (_e, res: McpToolResponse) => {
if (!res || typeof res.id !== 'string') return;
const entry = pending.get(res.id);
if (!entry) return;
clearTimeout(entry.timer);
pending.delete(res.id);
entry.resolve(res);
});
}
function askRenderer(win: BrowserWindow | null, name: string, args: Rec): Promise<McpToolResponse> {
ensureIpc();
if (!win || win.isDestroyed()) return Promise.resolve({ id: '', ok: false, error: 'EveFlow window unavailable' });
const id = `mcp-${++seq}-${Date.now()}`;
const req: McpToolRequest = { id, name, args };
return new Promise((resolve) => {
const timer = setTimeout(() => {
pending.delete(id);
resolve({ id, ok: false, error: 'renderer timeout' });
}, RENDERER_TIMEOUT_MS);
pending.set(id, { resolve, timer });
win.webContents.send(IPC.mcpRequest, req);
});
}
function text(value: unknown): { content: Array<Rec>; isError?: boolean } {
return { content: [{ type: 'text', text: typeof value === 'string' ? value : JSON.stringify(value) }] };
}
function failure(message: string): { content: Array<Rec>; isError: boolean } {
return { content: [{ type: 'text', text: message }], isError: true };
}
async function callTool(win: BrowserWindow | null, name: string, args: Rec): Promise<{ content: Array<Rec>; isError?: boolean }> {
const spec = MCP_TOOLS.find((t) => t.name === name);
if (!spec) return failure(`Outil inconnu : ${name}`);
if (spec.where === 'renderer') {
const res = await askRenderer(win, name, args);
return res.ok ? text(res.result ?? { ok: true }) : failure(res.error ?? 'échec');
}
const sys = async (action: SystemAction) => {
const r = await runSystemAction(action);
return r.ok ? text(r.data !== undefined ? { ok: true, message: r.message, data: r.data } : { ok: true, message: r.message }) : failure(r.message ?? 'échec');
};
switch (name) {
case 'capture_screen': {
const dataUrl = await captureScreen(Number(args.max_width) || 1600);
const base64 = dataUrl.slice(dataUrl.indexOf(',') + 1);
return { content: [{ type: 'image', data: base64, mimeType: 'image/jpeg' }, { type: 'text', text: `Capture de l'écran principal (${Math.round(base64.length * 0.75 / 1024)} ko).` }] };
}
case 'lock_session':
return sys({ type: 'lock' });
case 'open_app':
return sys({ type: 'open-app', name: String(args.name ?? '') });
case 'open_url':
return sys({ type: 'open-url', url: String(args.url ?? '') });
case 'media_key':
return sys({ type: 'media', key: String(args.key ?? '') as 'mute' });
case 'clipboard_get':
return sys({ type: 'clipboard-read' });
case 'clipboard_set':
return sys({ type: 'clipboard-write', text: String(args.text ?? '') });
case 'find_files':
return sys({ type: 'find-files', query: String(args.query ?? '') });
default:
return failure(`Outil non implémenté : ${name}`);
}
}
/** Handle one JSON-RPC message; returns the response object (or null for notifications). */
export async function handleMcpMessage(win: BrowserWindow | null, msg: RpcRequest): Promise<Rec | null> {
const id = msg.id ?? null;
const reply = (result: unknown) => ({ jsonrpc: '2.0', id, result });
const error = (code: number, message: string) => ({ jsonrpc: '2.0', id, error: { code, message } });
if (!msg.method) return error(-32600, 'Invalid Request');
if (msg.method.startsWith('notifications/')) return null;
switch (msg.method) {
case 'initialize':
return reply({ protocolVersion: PROTOCOL, capabilities: { tools: { listChanged: false } }, serverInfo: { name: 'eveflow', version: process.env.npm_package_version ?? '2.4.0' }, instructions: "Outils du PC de l'utilisateur via EveFlow : écran, applications, volume, presse-papiers, voix et notifications." });
case 'ping':
return reply({});
case 'tools/list':
return reply({ tools: MCP_TOOLS.map(({ name, description, inputSchema }) => ({ name, description, inputSchema })) });
case 'tools/call': {
const params = (msg.params ?? {}) as Rec;
const name = String(params.name ?? '');
const args = (params.arguments && typeof params.arguments === 'object' ? params.arguments : {}) as Rec;
log('INFO', 'mcp', `tools/call ${name}`);
try {
return reply(await callTool(win, name, args));
} catch (err) {
return reply(failure((err as Error).message));
}
}
case 'resources/list':
return reply({ resources: [] });
case 'prompts/list':
return reply({ prompts: [] });
default:
return error(-32601, `Method not found: ${msg.method}`);
}
}
/** HTTP entry point: POST /mcp with one JSON-RPC message or a batch. */
export async function handleMcpHttp(win: BrowserWindow | null, req: IncomingMessage, res: ServerResponse, body: string): Promise<void> {
const send = (code: number, payload: unknown) => {
if (payload === undefined) {
res.writeHead(code, { 'Cache-Control': 'no-store' });
res.end();
return;
}
res.writeHead(code, { 'Content-Type': 'application/json', 'Cache-Control': 'no-store' });
res.end(JSON.stringify(payload));
};
if (req.method === 'GET') return send(405, { error: 'SSE stream not supported; use POST' });
if (req.method === 'DELETE') return send(204, undefined);
if (req.method !== 'POST') return send(405, { error: 'Method Not Allowed' });
let parsed: unknown;
try {
parsed = body.trim() ? JSON.parse(body) : {};
} catch {
return send(400, { jsonrpc: '2.0', id: null, error: { code: -32700, message: 'Parse error' } });
}
const messages = Array.isArray(parsed) ? (parsed as RpcRequest[]) : [parsed as RpcRequest];
const replies: Rec[] = [];
for (const msg of messages) {
const out = await handleMcpMessage(win, msg && typeof msg === 'object' ? msg : {});
if (out) replies.push(out);
}
if (replies.length === 0) return send(202, undefined);
return send(200, Array.isArray(parsed) ? replies : replies[0]);
}
+19 -4
View File
@@ -13,8 +13,9 @@ import {
type WebhookStatus,
type WindowMode
} from '../shared/ipc';
import type { EveFlowBridge, Unsubscribe } from '../shared/bridge';
import { VOICE_IPC, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
import type { EveFlowBridge, SystemAction, SystemActionResult, Unsubscribe } from '../shared/bridge';
import type { McpToolRequest, McpToolResponse } from '../shared/ipc';
import { VOICE_IPC, type EdgeSynthesizeRequest, type EdgeSynthesizeResult, type EdgeVoice, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
function subscribe<T>(channel: string, callback: (payload: T) => void): Unsubscribe {
const listener = (_event: Electron.IpcRendererEvent, payload: T) => callback(payload);
@@ -43,7 +44,9 @@ const api: EveFlowBridge = {
},
system: {
metrics: () => ipcRenderer.invoke(IPC.metrics) as Promise<SystemMetrics>,
appInfo: () => ipcRenderer.invoke(IPC.appInfo) as Promise<AppInfo>
appInfo: () => ipcRenderer.invoke(IPC.appInfo) as Promise<AppInfo>,
captureScreen: (maxWidth?: number) => ipcRenderer.invoke(IPC.screenCapture, maxWidth) as Promise<string>,
action: (action: SystemAction) => ipcRenderer.invoke(IPC.systemAction, action) as Promise<SystemActionResult>
},
files: {
readLocal: (filePath: string) => ipcRenderer.invoke(IPC.readLocalFile, filePath) as Promise<string>,
@@ -55,6 +58,8 @@ const api: EveFlowBridge = {
hermes: {
onPush: (cb: (event: HermesPushEvent) => void) => subscribe<HermesPushEvent>(IPC.hermesPush, cb),
webhookStatus: () => ipcRenderer.invoke(IPC.webhookStatus) as Promise<WebhookStatus>,
onToolRequest: (cb: (req: McpToolRequest) => void) => subscribe<McpToolRequest>(IPC.mcpRequest, cb),
toolResponse: (res: McpToolResponse) => ipcRenderer.send(IPC.mcpResponse, res),
webhookRestart: () => ipcRenderer.invoke(IPC.webhookRestart) as Promise<WebhookStatus>
},
hotkeys: {
@@ -69,7 +74,17 @@ const api: EveFlowBridge = {
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => subscribe<VoiceDownloadProgress>(VOICE_IPC.modelsProgress, cb),
transcribe: (req: TranscribeRequest) => ipcRenderer.invoke(VOICE_IPC.transcribe, req) as Promise<TranscribeResult>,
synthesize: (req: SynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.synthesize, req) as Promise<SynthesizeResult>,
unload: (id?: string) => ipcRenderer.invoke(VOICE_IPC.unload, id) as Promise<unknown>
edgeSynthesize: (req: EdgeSynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.edgeSynthesize, req) as Promise<EdgeSynthesizeResult>,
edgeVoices: () => ipcRenderer.invoke(VOICE_IPC.edgeVoices) as Promise<EdgeVoice[]>,
unload: (id?: string) => ipcRenderer.invoke(VOICE_IPC.unload, id) as Promise<unknown>,
kwsStart: (req: KwsStartRequest) => ipcRenderer.invoke(VOICE_IPC.kwsStart, req) as Promise<{ accepted: string[]; rejected: string[] }>,
kwsStop: () => ipcRenderer.invoke(VOICE_IPC.kwsStop) as Promise<void>,
kwsAudio: (pcm: Uint8Array, sampleRate: number) => ipcRenderer.send(VOICE_IPC.kwsAudio, pcm, sampleRate),
onKwsDetected: (cb: (detection: KwsDetection) => void) => subscribe<KwsDetection>(VOICE_IPC.kwsDetected, cb),
vadStart: (req: VadStartRequest) => ipcRenderer.invoke(VOICE_IPC.vadStart, req) as Promise<void>,
vadStop: () => ipcRenderer.invoke(VOICE_IPC.vadStop) as Promise<void>,
vadAudio: (pcm: Uint8Array, sampleRate: number) => ipcRenderer.send(VOICE_IPC.vadAudio, pcm, sampleRate),
onVadEvent: (cb: (event: VadEvent) => void) => subscribe<VadEvent>(VOICE_IPC.vadEvent, cb)
}
};
+94 -5
View File
@@ -1,6 +1,7 @@
import type { VoiceModelSpec, VoiceSpeaker } from '../../shared/voice';
const ASR = 'https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models';
const KWS = 'https://github.com/k2-fsa/sherpa-onnx/releases/download/kws-models';
const TTS = 'https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models';
const KOKORO_SPEAKERS: VoiceSpeaker[] = [
@@ -16,6 +17,22 @@ const KOKORO_SPEAKERS: VoiceSpeaker[] = [
{ id: 26, name: 'George (homme, anglais UK)', lang: 'en' }
];
// Supertonic 3 ships ten voice styles in voice.bin (five feminine, five masculine). The ordering
// was checked by measuring the fundamental frequency of French synthesis: sid 0-4 around
// 170-210 Hz, sid 5-9 around 90-140 Hz.
const SUPERTONIC_SPEAKERS: VoiceSpeaker[] = [
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
{ id: 9, name: 'Homme 5 (grave)', lang: 'multi', gender: 'm' },
{ id: 7, name: 'Homme 3', lang: 'multi', gender: 'm' },
{ id: 8, name: 'Homme 4', lang: 'multi', gender: 'm' },
{ id: 5, name: 'Homme 1 (clair)', lang: 'multi', gender: 'm' },
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' },
{ id: 1, name: 'Femme 2', lang: 'multi', gender: 'f' },
{ id: 2, name: 'Femme 3', lang: 'multi', gender: 'f' },
{ id: 3, name: 'Femme 4', lang: 'multi', gender: 'f' },
{ id: 4, name: 'Femme 5', lang: 'multi', gender: 'f' }
];
export const VOICE_CATALOG: VoiceModelSpec[] = [
{
id: 'whisper-base',
@@ -43,12 +60,25 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [
files: ['small-encoder.int8.onnx', 'small-decoder.int8.onnx', 'small-tokens.txt'],
recommended: true
},
{
id: 'parakeet-v3',
kind: 'stt',
engine: 'nemo-transducer',
name: 'Parakeet TDT 0.6B v3 (25 langues européennes, français)',
description: 'Le plus précis en français et nettement plus rapide que Whisper sur processeur (NVIDIA NeMo, int8). Ponctuation et majuscules incluses.',
languages: ['fr', 'en', 'de', 'es', 'it', 'multi'],
sizeMb: 640,
url: `${ASR}/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8.tar.bz2`,
dir: 'sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8',
files: ['encoder.int8.onnx', 'decoder.int8.onnx', 'joiner.int8.onnx', 'tokens.txt'],
recommended: true
},
{
id: 'whisper-turbo',
kind: 'stt',
engine: 'whisper',
name: 'Whisper large-v3 turbo (multilingue)',
description: 'La meilleure précision ; demande un processeur puissant.',
description: 'Très bonne précision multilingue ; 2 à 4 s par phrase sur un processeur récent. Parakeet est plus rapide en français.',
languages: ['fr', 'en', 'multi'],
sizeMb: 564,
url: `${ASR}/sherpa-onnx-whisper-turbo.tar.bz2`,
@@ -67,20 +97,67 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [
dir: 'sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17',
files: ['model.int8.onnx', 'tokens.txt']
},
{
id: 'kws-en',
kind: 'kws',
engine: 'kws-transducer',
name: 'Détecteur de mot-clé (zipformer, 3 Mo)',
description: 'Écoute permanente du mot d’activation, quasi gratuite en CPU. Mot-clé libre (« jarvis », « hey jarvis »…).',
languages: ['en', 'fr'],
sizeMb: 4,
url: `${KWS}/sherpa-onnx-kws-zipformer-gigaspeech-3.3M-2024-01-01.tar.bz2`,
dir: 'sherpa-onnx-kws-zipformer-gigaspeech-3.3M-2024-01-01',
files: [
'encoder-epoch-12-avg-2-chunk-16-left-64.int8.onnx',
'decoder-epoch-12-avg-2-chunk-16-left-64.int8.onnx',
'joiner-epoch-12-avg-2-chunk-16-left-64.int8.onnx',
'tokens.txt'
],
recommended: true
},
{
id: 'silero-vad',
kind: 'vad',
engine: 'silero',
name: 'Silero VAD (fin de phrase neuronale, 0,6 Mo)',
description: 'Détecte précisément le début et la fin de la parole pendant l’écoute permanente ; moins de faux départs sur le bruit.',
languages: ['multi'],
sizeMb: 1,
url: `${ASR}/silero_vad.onnx`,
dir: 'silero-vad',
files: ['silero_vad.onnx'],
recommended: true
},
{
id: 'supertonic-3',
kind: 'tts',
engine: 'supertonic',
name: 'Supertonic 3 (31 langues, 5 voix masculines et 5 féminines)',
description:
'La meilleure voix française locale : naturelle, sans accent, dix voix au choix, 44 kHz. Environ 7 fois plus rapide que le temps réel sur 4 cœurs, 100 M de paramètres.',
languages: ['fr', 'en', 'de', 'es', 'it', 'pt', 'multi'],
sizeMb: 129,
url: `${TTS}/sherpa-onnx-supertonic-3-tts-int8-2026-05-11.tar.bz2`,
dir: 'sherpa-onnx-supertonic-3-tts-int8-2026-05-11',
files: ['duration_predictor.int8.onnx', 'text_encoder.int8.onnx', 'vector_estimator.int8.onnx', 'vocoder.int8.onnx', 'tts.json', 'unicode_indexer.bin', 'voice.bin'],
speakers: SUPERTONIC_SPEAKERS,
sampleRate: 44100,
recommended: true
},
{
id: 'kokoro-v1',
kind: 'tts',
engine: 'kokoro',
name: 'Kokoro v1.0 multilingue',
description: 'Voix très naturelle, une voix française (Siwis) et de nombreuses voix anglaises. 24 kHz.',
languages: ['fr', 'en', 'multi'],
description:
'Excellent en anglais (nombreuses voix). En français : une seule voix, féminine (Siwis), avec un accent marqué (phonémisation espeak). Préférez Supertonic 3 pour le français. 24 kHz.',
languages: ['en', 'fr', 'multi'],
sizeMb: 349,
url: `${TTS}/kokoro-multi-lang-v1_0.tar.bz2`,
dir: 'kokoro-multi-lang-v1_0',
files: ['model.onnx', 'voices.bin', 'tokens.txt', 'lexicon-us-en.txt', 'lexicon-zh.txt', 'espeak-ng-data/phontab'],
speakers: KOKORO_SPEAKERS,
sampleRate: 24000,
recommended: true
sampleRate: 24000
},
{
id: 'piper-fr-siwis',
@@ -132,3 +209,15 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [
export function findModel(id: string): VoiceModelSpec | undefined {
return VOICE_CATALOG.find((m) => m.id === id);
}
// Speaker gender from the label ("(homme, …)", "(femme, …)", known first names) so the UI and the
// voice preference can pick a masculine or feminine voice without a lookup table per model.
const MALE = /\b(homme|tom|pierre|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|santa)\b/i;
/** Languages accepted by the Supertonic 3 text front-end (2-letter codes). */
export const SUPERTONIC_LANGS = new Set(['ar', 'bg', 'hr', 'cs', 'da', 'nl', 'en', 'et', 'fi', 'fr', 'de', 'el', 'hi', 'hu', 'id', 'it', 'ja', 'ko', 'lv', 'lt', 'pl', 'pt', 'ro', 'ru', 'sk', 'sl', 'es', 'sv', 'tr', 'uk', 'vi']);
for (const spec of VOICE_CATALOG) {
for (const sp of spec.speakers ?? []) {
if (!sp.gender) sp.gender = /\b(femme|female)\b/i.test(sp.name) ? 'f' : MALE.test(sp.name) ? 'm' : /\bfemme\b/i.test(spec.name) ? 'f' : /\bhomme\b/i.test(spec.name) ? 'm' : undefined;
}
}
+146
View File
@@ -0,0 +1,146 @@
/**
* Microsoft Edge "Read aloud" neural voices (the service behind the Edge browser's read-aloud
* feature): free, no key, very natural French voices with a real masculine/feminine choice
* (Henri, Denise, Rémy, Vivienne…). Runs in the main process: one WebSocket per sentence,
* MP3 back to the renderer.
*/
import { createHash, randomBytes, randomUUID } from 'node:crypto';
import type { EdgeSynthesizeRequest, EdgeSynthesizeResult, EdgeVoice } from '../../shared/voice';
import {
EDGE_CHROMIUM_VERSION,
EDGE_VOICES_URL,
EDGE_WSS_URL,
edgeConfigMessage,
edgeConnectionId,
edgeHeaders,
edgeSsml,
edgeSsmlMessage,
edgeTextFramePath,
edgeTokenInput,
parseEdgeBinaryFrame
} from '../../shared/edgeTts';
import { log } from '../logger';
/** Seconds to add to the local clock so the signed token matches the server's time window. */
let clockSkewSec = 0;
let voicesCache: { at: number; voices: EdgeVoice[] } | null = null;
const VOICES_TTL_MS = 6 * 60 * 60 * 1000;
const SYNTH_TIMEOUT_MS = 20_000;
function token(): string {
return createHash('sha256').update(edgeTokenInput(Date.now(), clockSkewSec), 'ascii').digest('hex').toUpperCase();
}
function signedUrl(): string {
return `${EDGE_WSS_URL}&ConnectionId=${edgeConnectionId(randomUUID())}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
}
function headers(): Record<string, string> {
return { ...edgeHeaders(), Cookie: `muid=${randomBytes(16).toString('hex')};` };
}
/** Learn the server clock from a plain HTTPS response (the WebSocket handshake hides its headers). */
async function syncClock(): Promise<void> {
try {
const res = await fetch(EDGE_VOICES_URL, { method: 'HEAD', headers: headers() });
const date = res.headers.get('date');
if (!date) return;
const server = Date.parse(date);
if (Number.isFinite(server)) {
clockSkewSec = (server - Date.now()) / 1000;
log('INFO', 'edge-tts', `clock skew ${clockSkewSec.toFixed(0)} s`);
}
} catch (err) {
log('WARN', 'edge-tts', `clock sync failed: ${(err as Error).message}`);
}
}
function synthesizeOnce(req: EdgeSynthesizeRequest): Promise<Uint8Array> {
return new Promise((resolve, reject) => {
const chunks: Uint8Array[] = [];
let settled = false;
let ws: WebSocket;
try {
// Node's global WebSocket accepts extra handshake headers (undici), which the service checks.
ws = new (WebSocket as unknown as new (url: string, options: { headers: Record<string, string> }) => WebSocket)(signedUrl(), { headers: headers() });
} catch (err) {
reject(err as Error);
return;
}
ws.binaryType = 'arraybuffer';
const finish = (err?: Error) => {
if (settled) return;
settled = true;
clearTimeout(timer);
try {
ws.close();
} catch {
/* already closed */
}
if (err) reject(err);
else {
const total = chunks.reduce((n, c) => n + c.byteLength, 0);
const out = new Uint8Array(total);
let o = 0;
for (const c of chunks) {
out.set(c, o);
o += c.byteLength;
}
resolve(out);
}
};
const timer = setTimeout(() => finish(new Error('Edge TTS : délai dépassé')), SYNTH_TIMEOUT_MS);
ws.onopen = () => {
ws.send(edgeConfigMessage());
ws.send(edgeSsmlMessage(edgeConnectionId(randomUUID()), edgeSsml(req.text, req.voice, req.speed)));
};
ws.onmessage = (event: MessageEvent) => {
if (typeof event.data === 'string') {
if (edgeTextFramePath(event.data) === 'turn.end') finish();
return;
}
const frame = new Uint8Array(event.data as ArrayBuffer);
const { path, payload } = parseEdgeBinaryFrame(frame);
if (path === 'audio' && payload.byteLength) chunks.push(payload);
};
ws.onerror = (event: Event) => finish(new Error(`Edge TTS : connexion refusée (${(event as { message?: string }).message ?? 'erreur réseau'})`));
ws.onclose = (event: CloseEvent) => {
if (!settled) finish(chunks.length ? undefined : new Error(`Edge TTS : connexion fermée (${event.code}${event.reason ? ' ' + event.reason : ''})`));
};
});
}
export async function edgeSynthesize(req: EdgeSynthesizeRequest): Promise<EdgeSynthesizeResult> {
const started = Date.now();
let mp3: Uint8Array;
try {
mp3 = await synthesizeOnce(req);
} catch (err) {
// A refused handshake is almost always a stale signature: resync the clock and retry once.
log('WARN', 'edge-tts', `first attempt failed (${(err as Error).message}), resyncing clock`);
await syncClock();
mp3 = await synthesizeOnce(req);
}
if (!mp3.byteLength) throw new Error('Edge TTS : aucun audio reçu');
return { mp3, durationMs: Date.now() - started };
}
/** Voice list from the service (cached six hours); falls back to an empty list offline. */
export async function edgeVoices(): Promise<EdgeVoice[]> {
if (voicesCache && Date.now() - voicesCache.at < VOICES_TTL_MS) return voicesCache.voices;
const url = `${EDGE_VOICES_URL}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
const res = await fetch(url, { headers: headers() });
if (!res.ok) throw new Error(`Edge TTS : liste des voix HTTP ${res.status}`);
const raw = (await res.json()) as Array<{ ShortName?: string; FriendlyName?: string; Locale?: string; Gender?: string }>;
const voices: EdgeVoice[] = raw
.filter((v) => typeof v.ShortName === 'string' && typeof v.Locale === 'string')
.map((v) => ({
shortName: v.ShortName!,
name: v.ShortName!.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') || v.FriendlyName || v.ShortName!,
locale: v.Locale!,
gender: v.Gender === 'Male' ? ('m' as const) : ('f' as const)
}))
.sort((a, b) => a.locale.localeCompare(b.locale) || a.name.localeCompare(b.name));
voicesCache = { at: Date.now(), voices };
return voices;
}
+120 -4
View File
@@ -2,13 +2,23 @@
* Host side of the voice worker: spawns the utility process on demand, correlates
* requests and responses, restarts the worker if it crashes.
*/
import { utilityProcess, type UtilityProcess } from 'electron';
import { utilityProcess, type UtilityProcess, type WebContents } from 'electron';
import { createHash } from 'node:crypto';
import fs from 'node:fs';
import path from 'node:path';
import type { SynthesizeRequest, SynthesizeResult, TranscribeRequest, TranscribeResult, VoiceEngineStatus } from '../../shared/voice';
import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VadEvent, type VadStartRequest, type VoiceEngineStatus } from '../../shared/voice';
import { buildKeywordsFile, parseTokens } from '../../shared/keywords';
import { findModel } from './catalog';
import { isInstalled, modelDir, modelsDir } from './models';
import { log } from '../logger';
/** Renderer that receives keyword detections while spotting is active. */
let kwsSubscriber: WebContents | null = null;
let kwsActive = false;
let kwsRequest: KwsStartRequest | null = null;
let vadSubscriber: WebContents | null = null;
let vadActive = false;
interface Pending {
resolve: (value: unknown) => void;
reject: (err: Error) => void;
@@ -34,7 +44,28 @@ function spawn(): UtilityProcess {
child = proc;
proc.stdout?.on('data', (d: Buffer) => log('DEBUG', 'voice-worker', d.toString().trim()));
proc.stderr?.on('data', (d: Buffer) => log('WARN', 'voice-worker', d.toString().trim()));
proc.on('message', (msg: { id: number; ok: boolean; result?: unknown; error?: string }) => {
proc.on('message', (msg: { id?: number; type?: string; ok?: boolean; result?: unknown; error?: string; keyword?: string; at?: number; event?: { type: string; wav?: string; durationSec?: number } }) => {
if (msg.type === 'vad.event' && msg.event) {
if (vadSubscriber && !vadSubscriber.isDestroyed()) {
const ev = msg.event;
const payload: VadEvent =
ev.type === 'segment' && ev.wav
? { type: 'segment', wav: new Uint8Array(Buffer.from(ev.wav, 'base64')), durationSec: ev.durationSec ?? 0 }
: ev.type === 'speech-start'
? { type: 'speech-start' }
: { type: 'error', message: 'événement VAD inconnu' };
vadSubscriber.send(VOICE_IPC.vadEvent, payload);
}
return;
}
if (msg.type === 'kws.detected') {
if (kwsSubscriber && !kwsSubscriber.isDestroyed()) {
kwsSubscriber.send(VOICE_IPC.kwsDetected, { keyword: msg.keyword ?? '', at: msg.at ?? Date.now() } satisfies KwsDetection);
}
log('INFO', 'voice', `wake word detected: ${msg.keyword}`);
return;
}
if (typeof msg.id !== 'number') return;
const p = pending.get(msg.id);
if (!p) return;
pending.delete(msg.id);
@@ -48,6 +79,8 @@ function spawn(): UtilityProcess {
if (child !== proc) return;
child = null;
rejectAll('Le moteur vocal local s’est arrêté de façon inattendue.');
// Keyword spotting survives a worker restart: re-arm on the next audio frame.
if (kwsActive) kwsArmed = false;
});
log('INFO', 'voice', 'voice worker started');
return proc;
@@ -101,13 +134,96 @@ export function transcribe(req: TranscribeRequest): Promise<TranscribeResult> {
export async function synthesize(req: SynthesizeRequest): Promise<SynthesizeResult> {
const result = await request<Omit<SynthesizeResult, 'wav'> & { wav: string }>(
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed },
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed, language: req.language },
180_000
);
const buffer = Buffer.from(result.wav, 'base64');
return { ...result, wav: new Uint8Array(buffer.buffer, buffer.byteOffset, buffer.byteLength) };
}
let kwsArmed = false;
const SENSITIVITY_THRESHOLD: Record<number, number> = { 1: 0.45, 2: 0.35, 3: 0.25, 4: 0.18, 5: 0.12 };
const SENSITIVITY_SCORE: Record<number, number> = { 1: 1.0, 2: 1.0, 3: 1.2, 4: 1.5, 5: 2.0 };
/** Start keyword spotting; the keywords file is derived from the phrases and the model vocabulary. */
export async function kwsStart(req: KwsStartRequest, sender: WebContents): Promise<{ accepted: string[]; rejected: string[] }> {
const spec = findModel(req.modelId);
if (!spec || spec.kind !== 'kws') throw new Error('Modèle de détection introuvable');
if (!isInstalled(spec)) throw new Error('Détecteur de mot-clé non installé (Paramètres → Modèles locaux).');
const dir = modelDir(spec);
const vocab = parseTokens(fs.readFileSync(path.join(dir, 'tokens.txt'), 'utf8'));
const phrases = req.keywords.map((k) => String(k).slice(0, 40)).filter(Boolean).slice(0, 8);
const file = buildKeywordsFile(phrases, vocab);
if (file.accepted.length === 0) throw new Error(`Aucun mot d’activation encodable : ${file.rejected.join(', ')}`);
const hash = createHash('sha1').update(file.content).digest('hex').slice(0, 10);
const keywordsFile = path.join(modelsDir(), `keywords-${hash}.txt`);
fs.writeFileSync(keywordsFile, file.content, 'utf8');
const sensitivity = Math.min(5, Math.max(1, Math.round(req.sensitivity))) as 1 | 2 | 3 | 4 | 5;
kwsSubscriber = sender;
kwsRequest = req;
await request({ type: 'kws.start', model: { id: spec.id, engine: spec.engine, dir, files: spec.files }, keywordsFile, threshold: SENSITIVITY_THRESHOLD[sensitivity], score: SENSITIVITY_SCORE[sensitivity] }, 60_000);
kwsActive = true;
kwsArmed = true;
log('INFO', 'voice', `keyword spotting on: ${file.accepted.join(', ')} (threshold ${SENSITIVITY_THRESHOLD[sensitivity]})`);
return { accepted: file.accepted, rejected: file.rejected };
}
export async function kwsStop(): Promise<void> {
kwsActive = false;
kwsArmed = false;
kwsSubscriber = null;
if (child) await request({ type: 'kws.stop' }, 10_000).catch(() => undefined);
}
/** Feed 16-bit PCM from the renderer (fire-and-forget). */
export function kwsFeed(pcm: Uint8Array, sampleRate: number): void {
if (!kwsActive) return;
const proc = spawn();
if (!kwsArmed) {
// Worker restarted: re-create the spotter before feeding audio.
if (kwsRequest && kwsSubscriber) {
kwsArmed = true;
kwsStart(kwsRequest, kwsSubscriber).catch((err) => log('WARN', 'voice', `kws re-arm failed: ${(err as Error).message}`));
}
return;
}
const b64 = Buffer.from(pcm.buffer, pcm.byteOffset, pcm.byteLength).toString('base64');
try {
proc.postMessage({ id: -1, type: 'kws.audio', pcm: b64, sampleRate });
} catch (err) {
log('WARN', 'voice', `kws feed failed: ${(err as Error).message}`);
}
}
/** Neural end-of-speech detection (Silero) on frames streamed by the renderer. */
export async function vadStart(req: VadStartRequest, sender: WebContents): Promise<void> {
const spec = findModel(req.modelId);
if (!spec || spec.kind !== 'vad') throw new Error('Modèle VAD introuvable');
if (!isInstalled(spec)) throw new Error('Silero VAD non installé (Paramètres → Modèles locaux).');
vadSubscriber = sender;
await request(
{ type: 'vad.start', model: { id: spec.id, engine: spec.engine, dir: modelDir(spec), files: spec.files }, silenceMs: req.silenceMs, threshold: req.threshold, maxUtteranceSec: req.maxUtteranceSec },
30_000
);
vadActive = true;
}
export async function vadStop(): Promise<void> {
vadActive = false;
vadSubscriber = null;
if (child) await request({ type: 'vad.stop' }, 10_000).catch(() => undefined);
}
export function vadFeed(pcm: Uint8Array, sampleRate: number): void {
if (!vadActive || !child) return;
const b64 = Buffer.from(pcm.buffer, pcm.byteOffset, pcm.byteLength).toString('base64');
try {
child.postMessage({ id: -1, type: 'vad.audio', pcm: b64, sampleRate });
} catch (err) {
log('WARN', 'voice', `vad feed failed: ${(err as Error).message}`);
}
}
/** Native model memory is only released deterministically by restarting the worker. */
export function unload(_modelId?: string): Promise<unknown> {
stopEngine();
+33 -3
View File
@@ -1,6 +1,7 @@
import { ipcMain } from 'electron';
import { VOICE_IPC, type SynthesizeRequest, type TranscribeRequest } from '../../shared/voice';
import { engineStatus, synthesize, transcribe, unload } from './engine';
import { VOICE_IPC, type EdgeSynthesizeRequest, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice';
import { edgeSynthesize, edgeVoices } from './edgeTts';
import { engineStatus, kwsFeed, kwsStart, kwsStop, synthesize, transcribe, unload, vadFeed, vadStart, vadStop } from './engine';
import { cancelDownload, downloadModel, listModels, removeModel } from './models';
export function registerVoiceIpc(): void {
@@ -24,7 +25,36 @@ export function registerVoiceIpc(): void {
ipcMain.handle(VOICE_IPC.synthesize, (_e, req: SynthesizeRequest) => {
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
if (typeof req.modelId !== 'string') throw new Error('Modèle invalide');
return synthesize({ ...req, speaker: Number.isFinite(req.speaker) ? req.speaker : 0, speed: Number.isFinite(req.speed) ? req.speed : 1 });
return synthesize({
...req,
speaker: Number.isFinite(req.speaker) ? req.speaker : 0,
speed: Number.isFinite(req.speed) ? req.speed : 1,
language: typeof req.language === 'string' ? req.language.slice(0, 8) : undefined
});
});
ipcMain.handle(VOICE_IPC.edgeSynthesize, (_e, req: EdgeSynthesizeRequest) => {
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
if (typeof req.voice !== 'string' || !/^[a-z]{2,3}-[A-Za-z]{2,4}-[A-Za-z0-9]+$/.test(req.voice)) throw new Error('Voix Edge invalide');
return edgeSynthesize({ text: req.text, voice: req.voice, speed: Number.isFinite(req.speed) ? req.speed : 1 });
});
ipcMain.handle(VOICE_IPC.edgeVoices, () => edgeVoices());
ipcMain.handle(VOICE_IPC.unload, (_e, id?: string) => unload(id));
ipcMain.handle(VOICE_IPC.kwsStart, (event, req: KwsStartRequest) => {
if (!req || !Array.isArray(req.keywords) || typeof req.modelId !== 'string') throw new Error('Requête invalide');
return kwsStart({ ...req, keywords: req.keywords.filter((k) => typeof k === 'string'), sensitivity: Number(req.sensitivity) || 3 }, event.sender);
});
ipcMain.handle(VOICE_IPC.kwsStop, () => kwsStop());
ipcMain.handle(VOICE_IPC.vadStart, (event, req: VadStartRequest) => {
if (!req || typeof req.modelId !== 'string') throw new Error('Requête invalide');
return vadStart({ modelId: req.modelId, silenceMs: Number(req.silenceMs) || 700, threshold: Number(req.threshold) || 0.5, maxUtteranceSec: Number(req.maxUtteranceSec) || 30 }, event.sender);
});
ipcMain.handle(VOICE_IPC.vadStop, () => vadStop());
ipcMain.on(VOICE_IPC.vadAudio, (_e, pcm: unknown, sampleRate: unknown) => {
if (!(pcm instanceof Uint8Array) || pcm.byteLength === 0 || pcm.byteLength > 1024 * 1024) return;
vadFeed(pcm, typeof sampleRate === 'number' && sampleRate > 0 ? sampleRate : 16000);
});
ipcMain.on(VOICE_IPC.kwsAudio, (_e, pcm: unknown, sampleRate: unknown) => {
if (!(pcm instanceof Uint8Array) || pcm.byteLength === 0 || pcm.byteLength > 1024 * 1024) return;
kwsFeed(pcm, typeof sampleRate === 'number' && sampleRate > 0 ? sampleRate : 16000);
});
}
+187 -7
View File
@@ -6,6 +6,7 @@
import os from 'node:os';
import path from 'node:path';
import type { VoiceEngineKind } from '../../shared/voice';
import { SUPERTONIC_LANGS } from './catalog';
interface ModelRef {
id: string;
@@ -17,8 +18,14 @@ interface ModelRef {
type Request =
| { id: number; type: 'status' }
| { id: number; type: 'transcribe'; model: ModelRef; wav: Uint8Array | string; language: string }
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number }
| { id: number; type: 'unload'; modelId?: string };
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number; language?: string }
| { id: number; type: 'unload'; modelId?: string }
| { id: number; type: 'kws.start'; model: ModelRef; keywordsFile: string; threshold: number; score: number }
| { id: number; type: 'kws.audio'; pcm: string; sampleRate: number }
| { id: number; type: 'kws.stop' }
| { id: number; type: 'vad.start'; model: ModelRef; silenceMs: number; threshold: number; maxUtteranceSec: number }
| { id: number; type: 'vad.audio'; pcm: string; sampleRate: number }
| { id: number; type: 'vad.stop' };
type Response = { id: number; ok: true; result: unknown } | { id: number; ok: false; error: string };
@@ -29,15 +36,38 @@ type Sherpa = {
decode: (s: unknown) => void;
getResult: (s: unknown) => { text: string; lang?: string };
};
Vad: new (config: unknown, bufferSizeInSeconds: number) => {
acceptWaveform: (samples: Float32Array) => void;
isEmpty: () => boolean;
isDetected: () => boolean;
pop: () => void;
clear: () => void;
front: (enableExternalBuffer?: boolean) => { start: number; samples: Float32Array };
reset: () => void;
flush: () => void;
};
KeywordSpotter: new (config: unknown) => {
createStream: () => KwsStream;
isReady: (s: KwsStream) => boolean;
decode: (s: KwsStream) => void;
reset: (s: KwsStream) => void;
getResult: (s: KwsStream) => { keyword?: string };
};
OfflineTts: new (config: unknown) => {
numSpeakers: number;
sampleRate: number;
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean }) => { samples: Float32Array; sampleRate: number };
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean; generationConfig?: unknown }) => { samples: Float32Array; sampleRate: number };
};
/** Per-request options for the newer engines (Supertonic reads `extra.lang`). */
GenerationConfig: new (opts: { sid: number; speed: number; numSteps?: number; extra?: Record<string, string | number> }) => unknown;
version: string;
};
type KwsStream = { acceptWaveform: (w: { sampleRate: number; samples: Float32Array }) => void };
let sherpa: Sherpa | null = null;
let kws: { spotter: InstanceType<Sherpa['KeywordSpotter']>; stream: KwsStream } | null = null;
let vad: { detector: InstanceType<Sherpa['Vad']>; speaking: boolean; windowSize: number; carry: Float32Array } | null = null;
let notify: ((message: unknown) => void) | null = null;
let loadError: string | null = null;
function loadSherpa(): Sherpa {
@@ -125,6 +155,19 @@ function getSynthesizer(model: ModelRef) {
ttsModel = { vits: { model: p(onnx), tokens: p('tokens.txt'), dataDir: p('espeak-ng-data') } };
break;
}
case 'supertonic':
ttsModel = {
supertonic: {
durationPredictor: p('duration_predictor.int8.onnx'),
textEncoder: p('text_encoder.int8.onnx'),
vectorEstimator: p('vector_estimator.int8.onnx'),
vocoder: p('vocoder.int8.onnx'),
ttsJson: p('tts.json'),
unicodeIndexer: p('unicode_indexer.bin'),
voiceStyle: p('voice.bin')
}
};
break;
default:
throw new Error(`Moteur TTS non supporté : ${model.engine}`);
}
@@ -133,6 +176,12 @@ function getSynthesizer(model: ModelRef) {
return tts;
}
/** 2-letter code accepted by Supertonic 3 ("fr-FR" → "fr"); English when unknown, as upstream does. */
function supertonicLang(language: string | undefined): string {
const code = (language ?? '').toLowerCase().split(/[-_]/)[0];
return SUPERTONIC_LANGS.has(code) ? code : 'en';
}
// ── audio helpers ──────────────────────────────────────────────────────────
function decodeWav(bytes: Uint8Array): { samples: Float32Array; sampleRate: number } {
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
@@ -225,13 +274,124 @@ function encodeWav(samples: Float32Array, sampleRate: number): Uint8Array {
return new Uint8Array(buffer);
}
function startKws(model: ModelRef, keywordsFile: string, threshold: number, score: number): void {
const s = loadSherpa();
kws = null;
const p = (f: string) => path.join(model.dir, f);
const enc = model.files.find((f) => f.startsWith('encoder')) ?? 'encoder.int8.onnx';
const dec = model.files.find((f) => f.startsWith('decoder')) ?? 'decoder.int8.onnx';
const join = model.files.find((f) => f.startsWith('joiner')) ?? 'joiner.int8.onnx';
const spotter = new s.KeywordSpotter({
featConfig: { sampleRate: 16000, featureDim: 80 },
modelConfig: { transducer: { encoder: p(enc), decoder: p(dec), joiner: p(join) }, tokens: p('tokens.txt'), numThreads: 1, provider: 'cpu', debug: 0 },
maxActivePaths: 4,
numTrailingBlanks: 1,
keywordsScore: score,
keywordsThreshold: threshold,
keywordsFile
});
kws = { spotter, stream: spotter.createStream() };
}
function feedKws(pcmBase64: string, sampleRate: number): void {
if (!kws) return;
const bytes = Buffer.from(pcmBase64, 'base64');
const int16 = new Int16Array(bytes.buffer, bytes.byteOffset, Math.floor(bytes.byteLength / 2));
let samples: Float32Array = new Float32Array(int16.length);
for (let i = 0; i < int16.length; i++) samples[i] = int16[i] / 32768;
if (sampleRate !== 16000) samples = resampleTo16k(samples, sampleRate);
kws.stream.acceptWaveform({ sampleRate: 16000, samples });
while (kws.spotter.isReady(kws.stream)) {
kws.spotter.decode(kws.stream);
const result = kws.spotter.getResult(kws.stream);
if (result.keyword) {
kws.spotter.reset(kws.stream);
notify?.({ type: 'kws.detected', keyword: result.keyword, at: Date.now() });
}
}
}
function startVad(model: ModelRef, silenceMs: number, threshold: number, maxUtteranceSec: number): void {
const s = loadSherpa();
const windowSize = 512;
const detector = new s.Vad(
{
sileroVad: {
model: path.join(model.dir, 'silero_vad.onnx'),
threshold: Math.max(0.1, Math.min(0.95, threshold)),
minSilenceDuration: Math.max(0.15, silenceMs / 1000),
minSpeechDuration: 0.2,
windowSize,
maxSpeechDuration: Math.max(3, maxUtteranceSec)
},
sampleRate: 16000,
numThreads: 1,
provider: 'cpu',
debug: 0
},
Math.max(10, maxUtteranceSec + 5)
);
vad = { detector, speaking: false, windowSize, carry: new Float32Array(0) };
}
function feedVad(pcmBase64: string, sampleRate: number): void {
if (!vad) return;
const bytes = Buffer.from(pcmBase64, 'base64');
const int16 = new Int16Array(bytes.buffer, bytes.byteOffset, Math.floor(bytes.byteLength / 2));
let samples: Float32Array = new Float32Array(int16.length);
for (let i = 0; i < int16.length; i++) samples[i] = int16[i] / 32768;
if (sampleRate !== 16000) samples = resampleTo16k(samples, sampleRate);
// Silero expects fixed windows: keep the remainder for the next frame.
const merged = new Float32Array(vad.carry.length + samples.length);
merged.set(vad.carry, 0);
merged.set(samples, vad.carry.length);
const usable = merged.length - (merged.length % vad.windowSize);
for (let i = 0; i < usable; i += vad.windowSize) {
vad.detector.acceptWaveform(merged.subarray(i, i + vad.windowSize));
const detected = vad.detector.isDetected();
if (detected && !vad.speaking) {
vad.speaking = true;
notify?.({ type: 'vad.event', event: { type: 'speech-start' } });
}
while (!vad.detector.isEmpty()) {
const segment = vad.detector.front(false);
vad.detector.pop();
vad.speaking = false;
const wav = encodeWav(segment.samples, 16000);
notify?.({
type: 'vad.event',
event: { type: 'segment', wav: Buffer.from(wav.buffer, wav.byteOffset, wav.byteLength).toString('base64'), durationSec: segment.samples.length / 16000 }
});
}
}
vad.carry = merged.slice(usable);
}
// ── request handling ───────────────────────────────────────────────────────
function handle(req: Request): unknown {
switch (req.type) {
case 'vad.start':
startVad(req.model, req.silenceMs, req.threshold, req.maxUtteranceSec);
return { ok: true };
case 'vad.audio':
feedVad(req.pcm, req.sampleRate);
return { ok: true };
case 'vad.stop':
vad = null;
return { ok: true };
case 'kws.start':
startKws(req.model, req.keywordsFile, req.threshold, req.score);
return { ok: true };
case 'kws.audio':
feedKws(req.pcm, req.sampleRate);
return { ok: true };
case 'kws.stop':
kws = null;
return { ok: true };
case 'status': {
try {
const s = loadSherpa();
return { available: true, version: s.version, loaded: [...recognizers.keys(), ...synthesizers.keys()] };
return { available: true, version: s.version, loaded: [...recognizers.keys(), ...synthesizers.keys(), ...(kws ? ['kws'] : [])] };
} catch (err) {
return { available: false, error: (err as Error).message, loaded: [] };
}
@@ -241,8 +401,13 @@ function handle(req: Request): unknown {
// Audio crosses the process boundary as base64: V8 refuses to serialize external buffers.
const bytes = typeof req.wav === 'string' ? new Uint8Array(Buffer.from(req.wav, 'base64')) : req.wav;
const { samples, sampleRate } = decodeWav(bytes);
const pcm = resampleTo16k(samples, sampleRate);
if (pcm.length < 1600) throw new Error('Audio trop court');
const raw = resampleTo16k(samples, sampleRate);
if (raw.length < 1600) throw new Error('Audio trop court');
// 0.4 s of silence on both sides: utterances cut close to the words lose the first/last syllable
// and short clips make Whisper hallucinate.
const pad = 6400;
const pcm = new Float32Array(raw.length + 2 * pad);
pcm.set(raw, pad);
const recognizer = getRecognizer(req.model, req.language);
const stream = recognizer.createStream();
stream.acceptWaveform({ sampleRate: 16000, samples: pcm });
@@ -254,8 +419,19 @@ function handle(req: Request): unknown {
const started = Date.now();
const tts = getSynthesizer(req.model);
const sid = Math.max(0, Math.min(tts.numSpeakers - 1, Math.floor(req.speaker)));
const speed = Math.max(0.5, Math.min(2, req.speed || 1));
// Electron forbids N-API external buffers: ask sherpa-onnx to copy the samples into a V8 buffer.
const audio = tts.generate({ text: req.text, sid, speed: Math.max(0.5, Math.min(2, req.speed || 1)), enableExternalBuffer: false });
const audio =
req.model.engine === 'supertonic'
? tts.generate({
text: req.text,
sid,
speed,
enableExternalBuffer: false,
// Supertonic needs the language of the text; 5 denoising steps is the quality/speed sweet spot.
generationConfig: new (loadSherpa().GenerationConfig)({ sid, speed, numSteps: 5, extra: { lang: supertonicLang(req.language) } })
})
: tts.generate({ text: req.text, sid, speed, enableExternalBuffer: false });
const wav = encodeWav(audio.samples, audio.sampleRate);
return {
wav: Buffer.from(wav.buffer, wav.byteOffset, wav.byteLength).toString('base64'),
@@ -271,6 +447,8 @@ function handle(req: Request): unknown {
} else {
recognizers.clear();
synthesizers.clear();
kws = null;
vad = null;
}
return { ok: true };
}
@@ -290,7 +468,9 @@ function respond(req: Request): Response {
// Electron utility process transport, with a child_process fallback for tests.
const parentPort = (process as unknown as { parentPort?: { on: (ev: 'message', cb: (e: { data: Request }) => void) => void; postMessage: (m: unknown) => void } }).parentPort;
if (parentPort) {
notify = (m) => parentPort.postMessage(m);
parentPort.on('message', (event) => parentPort.postMessage(respond(event.data)));
} else if (process.send) {
notify = (m) => process.send!(m);
process.on('message', (msg: Request) => process.send!(respond(msg)));
}
+11 -2
View File
@@ -8,6 +8,7 @@ import type { BrowserWindow } from 'electron';
import { IPC, type WebhookStatus } from '../shared/ipc';
import { normalizeHermesPush } from '../shared/hermesPush';
import { log } from './logger';
import { handleMcpHttp, MCP_PATH } from './mcp';
export const WEBHOOK_PATH = '/eveflow/hook';
const MAX_BODY_BYTES = 2 * 1024 * 1024;
@@ -57,9 +58,10 @@ export async function startWebhookServer(getWindow: () => BrowserWindow | null,
};
if (req.method === 'GET' && (req.url === '/health' || req.url === '/')) {
return json(200, { ok: true, app: 'eveflow', path: WEBHOOK_PATH });
return json(200, { ok: true, app: 'eveflow', path: WEBHOOK_PATH, mcp: MCP_PATH });
}
if (req.method !== 'POST' || !req.url?.startsWith(WEBHOOK_PATH)) {
const isMcp = !!req.url && (req.url === MCP_PATH || req.url.startsWith(`${MCP_PATH}?`));
if (!isMcp && (req.method !== 'POST' || !req.url?.startsWith(WEBHOOK_PATH))) {
return json(404, { error: 'Not Found' });
}
@@ -91,6 +93,13 @@ export async function startWebhookServer(getWindow: () => BrowserWindow | null,
if (tooLarge) return;
// Concatenate before decoding so multibyte UTF-8 (accents, emoji) split across chunks stays intact.
const body = Buffer.concat(chunks).toString('utf8');
if (isMcp) {
handleMcpHttp(getWindow(), req, res, body).catch((err: Error) => {
log('ERROR', 'mcp', err.message);
if (!res.headersSent) json(500, { error: err.message });
});
return;
}
try {
const raw: unknown = body.trim() ? JSON.parse(body) : {};
const events = normalizeHermesPush(raw);
+2 -2
View File
@@ -1,7 +1,7 @@
{
"name": "eveflow",
"version": "2.1.0",
"releaseVersion": "2.1.0.1",
"version": "2.5.0",
"releaseVersion": "2.5.0",
"description": "JARVIS-style desktop HUD for Hermes Agent: voice, streaming runs, scheduled jobs, skills and telemetry",
"main": "dist-electron/main.js",
"private": true,
+27
View File
@@ -0,0 +1,27 @@
class EveFlowCaptureProcessor extends AudioWorkletProcessor {
constructor() {
super();
this.buffer = new Float32Array(2048);
this.offset = 0;
}
process(inputs) {
const input = inputs[0];
if (!input || input.length === 0) return true;
const channel = input[0];
if (!channel) return true;
let i = 0;
while (i < channel.length) {
const n = Math.min(channel.length - i, this.buffer.length - this.offset);
this.buffer.set(channel.subarray(i, i + n), this.offset);
this.offset += n;
i += n;
if (this.offset === this.buffer.length) {
this.port.postMessage(this.buffer, [this.buffer.buffer]);
this.buffer = new Float32Array(2048);
this.offset = 0;
}
}
return true;
}
}
registerProcessor('eveflow-capture', EveFlowCaptureProcessor);
+19
View File
@@ -0,0 +1,19 @@
class EveFlowWakeProcessor extends AudioWorkletProcessor {
constructor() { super(); this.buffer = new Float32Array(2048); this.offset = 0; }
process(inputs) {
const channel = inputs[0] && inputs[0][0];
if (!channel) return true;
let i = 0;
while (i < channel.length) {
const n = Math.min(channel.length - i, this.buffer.length - this.offset);
this.buffer.set(channel.subarray(i, i + n), this.offset);
this.offset += n; i += n;
if (this.offset === this.buffer.length) {
this.port.postMessage(this.buffer, [this.buffer.buffer]);
this.buffer = new Float32Array(2048); this.offset = 0;
}
}
return true;
}
}
registerProcessor('eveflow-wake', EveFlowWakeProcessor);
+46 -4
View File
@@ -11,7 +11,15 @@ import type {
WebhookStatus,
WindowMode
} from './ipc';
import type { McpToolRequest, McpToolResponse } from './ipc';
import type {
EdgeSynthesizeRequest,
EdgeSynthesizeResult,
EdgeVoice,
KwsDetection,
KwsStartRequest,
VadEvent,
VadStartRequest,
SynthesizeRequest,
SynthesizeResult,
TranscribeRequest,
@@ -44,10 +52,6 @@ export interface EveFlowBridge {
streamAbort: (id: string) => void;
onStreamEvent: (cb: (event: HttpStreamEvent) => void) => Unsubscribe;
};
system: {
metrics: () => Promise<SystemMetrics>;
appInfo: () => Promise<AppInfo>;
};
files: {
readLocal: (filePath: string) => Promise<string>;
writeShared: (filename: string, content: string, isBase64?: boolean) => Promise<{ path: string; url: string }>;
@@ -57,6 +61,9 @@ export interface EveFlowBridge {
hermes: {
onPush: (cb: (event: HermesPushEvent) => void) => Unsubscribe;
webhookStatus: () => Promise<WebhookStatus>;
/** Tool calls from Hermes via the local MCP endpoint that need the renderer. */
onToolRequest: (cb: (req: McpToolRequest) => void) => Unsubscribe;
toolResponse: (res: McpToolResponse) => void;
webhookRestart: () => Promise<WebhookStatus>;
};
hotkeys: {
@@ -71,6 +78,41 @@ export interface EveFlowBridge {
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => Unsubscribe;
transcribe: (req: TranscribeRequest) => Promise<TranscribeResult>;
synthesize: (req: SynthesizeRequest) => Promise<SynthesizeResult>;
/** Microsoft Edge neural voices (online, free): MP3 for one sentence. */
edgeSynthesize: (req: EdgeSynthesizeRequest) => Promise<EdgeSynthesizeResult>;
edgeVoices: () => Promise<EdgeVoice[]>;
unload: (id?: string) => Promise<unknown>;
kwsStart: (req: KwsStartRequest) => Promise<{ accepted: string[]; rejected: string[] }>;
kwsStop: () => Promise<void>;
/** Fire-and-forget 16-bit PCM frames for the keyword spotter. */
kwsAudio: (pcm: Uint8Array, sampleRate: number) => void;
onKwsDetected: (cb: (detection: KwsDetection) => void) => Unsubscribe;
vadStart: (req: VadStartRequest) => Promise<void>;
vadStop: () => Promise<void>;
vadAudio: (pcm: Uint8Array, sampleRate: number) => void;
onVadEvent: (cb: (event: VadEvent) => void) => Unsubscribe;
};
system: {
metrics: () => Promise<SystemMetrics>;
appInfo: () => Promise<AppInfo>;
/** Screenshot of the primary display as a JPEG data URL. */
captureScreen: (maxWidth?: number) => Promise<string>;
/** Allow-listed local actions (lock session, open app/url, media keys, clipboard). */
action: (action: SystemAction) => Promise<SystemActionResult>;
};
}
export type SystemAction =
| { type: 'lock' }
| { type: 'open-app'; name: string }
| { type: 'open-url'; url: string }
| { type: 'media'; key: 'volume-up' | 'volume-down' | 'mute' | 'play-pause' | 'next' | 'previous' }
| { type: 'clipboard-read' }
| { type: 'clipboard-write'; text: string }
| { type: 'find-files'; query: string };
export interface SystemActionResult {
ok: boolean;
message?: string;
data?: unknown;
}
+119
View File
@@ -0,0 +1,119 @@
/**
* Pure helpers for the Microsoft Edge "Read aloud" speech service (the endpoint used by the
* Edge browser, no API key): request signing input, SSML building, frame parsing and the
* default French voices. No Node or DOM dependency so it is shared by main and tests.
*/
export const EDGE_TRUSTED_CLIENT_TOKEN = '6A5AA1D4EAFF4E9FB37E23D68491D6F4';
export const EDGE_CHROMIUM_VERSION = '143.0.3650.75';
export const EDGE_WSS_URL = `wss://speech.platform.bing.com/consumer/speech/synthesize/readaloud/edge/v1?TrustedClientToken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
export const EDGE_VOICES_URL = `https://speech.platform.bing.com/consumer/speech/synthesize/readaloud/voices/list?trustedclienttoken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
export const EDGE_OUTPUT_FORMAT = 'audio-24khz-48kbitrate-mono-mp3';
const WIN_EPOCH_SEC = 11644473600;
/**
* String whose SHA-256 (upper-case hex) is the `Sec-MS-GEC` value: Windows file time (100 ns
* ticks since 1601) rounded down to 5 minutes, followed by the trusted client token.
* `nowMs` is the client clock, `skewSec` the correction learned from the server's Date header.
*/
export function edgeTokenInput(nowMs: number, skewSec = 0): string {
let seconds = Math.floor(nowMs / 1000 + skewSec) + WIN_EPOCH_SEC;
seconds -= seconds % 300;
// 10 million ticks per second; BigInt keeps the 18-digit value exact.
const ticks = BigInt(seconds) * 10_000_000n;
return `${ticks}${EDGE_TRUSTED_CLIENT_TOKEN}`;
}
export function edgeHeaders(): Record<string, string> {
const major = EDGE_CHROMIUM_VERSION.split('.')[0];
return {
'User-Agent': `Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36 Edg/${major}.0.0.0`,
'Accept-Language': 'en-US,en;q=0.9',
Pragma: 'no-cache',
'Cache-Control': 'no-cache',
Origin: 'chrome-extension://jdiccldimpdaibmpdkjnbmckianbfold'
};
}
export function escapeXml(text: string): string {
return text.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;').replace(/'/g, '&apos;');
}
/** Prosody rate attribute for a playback speed multiplier (1 = "+0%", 1.25 = "+25%"). */
export function edgeRate(speed: number): string {
const pct = Math.round((Math.max(0.5, Math.min(2, speed || 1)) - 1) * 100);
return `${pct >= 0 ? '+' : ''}${pct}%`;
}
export function edgeSsml(text: string, voice: string, speed: number): string {
const lang = voice.split('-').slice(0, 2).join('-') || 'fr-FR';
return (
`<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xml:lang='${lang}'>` +
`<voice name='${escapeXml(voice)}'><prosody pitch='+0Hz' rate='${edgeRate(speed)}' volume='+0%'>${escapeXml(text)}</prosody></voice></speak>`
);
}
/** Timestamp header the service expects ("JavaScript date string"). */
export function edgeTimestamp(date = new Date()): string {
return date.toUTCString().replace('GMT', 'GMT+0000 (Coordinated Universal Time)');
}
export function edgeConfigMessage(date = new Date()): string {
return (
`X-Timestamp:${edgeTimestamp(date)}\r\nContent-Type:application/json; charset=utf-8\r\nPath:speech.config\r\n\r\n` +
`{"context":{"synthesis":{"audio":{"metadataoptions":{"sentenceBoundaryEnabled":"false","wordBoundaryEnabled":"false"},"outputFormat":"${EDGE_OUTPUT_FORMAT}"}}}}`
);
}
export function edgeSsmlMessage(requestId: string, ssml: string, date = new Date()): string {
return `X-RequestId:${requestId}\r\nContent-Type:application/ssml+xml\r\nX-Timestamp:${edgeTimestamp(date)}Z\r\nPath:ssml\r\n\r\n${ssml}`;
}
/** Split a binary frame into its text headers and payload (2-byte big-endian header length). */
export function parseEdgeBinaryFrame(frame: Uint8Array): { path: string; payload: Uint8Array } {
if (frame.byteLength < 2) return { path: '', payload: new Uint8Array(0) };
const headerLength = (frame[0] << 8) | frame[1];
const end = Math.min(frame.byteLength, 2 + headerLength);
let header = '';
for (let i = 2; i < end; i++) header += String.fromCharCode(frame[i]);
const path = /Path:\s*([^\r\n]+)/i.exec(header)?.[1]?.trim() ?? '';
return { path, payload: frame.subarray(end) };
}
/** Path of a text frame ("turn.start", "response", "audio.metadata", "turn.end"). */
export function edgeTextFramePath(message: string): string {
return /Path:\s*([^\r\n]+)/i.exec(message)?.[1]?.trim() ?? '';
}
export function edgeConnectionId(hex32: string): string {
return hex32.replace(/-/g, '').toLowerCase();
}
/** Well-known voices per language so the choice works before the voice list is fetched. */
export const EDGE_DEFAULT_VOICES: Record<string, { male: string; female: string }> = {
fr: { male: 'fr-FR-HenriNeural', female: 'fr-FR-DeniseNeural' },
en: { male: 'en-US-AndrewMultilingualNeural', female: 'en-US-AvaMultilingualNeural' },
de: { male: 'de-DE-ConradNeural', female: 'de-DE-KatjaNeural' },
es: { male: 'es-ES-AlvaroNeural', female: 'es-ES-ElviraNeural' },
it: { male: 'it-IT-DiegoNeural', female: 'it-IT-ElsaNeural' },
pt: { male: 'pt-BR-AntonioNeural', female: 'pt-BR-FranciscaNeural' }
};
/** Default Edge voice for a language ("fr-FR", "fr", "en-GB") and gender; French when unknown. */
export function defaultEdgeVoice(language: string, gender: 'male' | 'female'): string {
const lang = (language || 'fr').toLowerCase().split(/[-_]/)[0];
return (EDGE_DEFAULT_VOICES[lang] ?? EDGE_DEFAULT_VOICES.fr)[gender];
}
/** Gender of an Edge voice from its short name, using the built-in table then common first names. */
export function edgeVoiceGender(shortName: string): 'male' | 'female' | undefined {
for (const pair of Object.values(EDGE_DEFAULT_VOICES)) {
if (pair.male === shortName) return 'male';
if (pair.female === shortName) return 'female';
}
const name = shortName.split('-')[2]?.replace(/(Multilingual)?Neural$/i, '') ?? '';
if (/^(Henri|Remy|Rémy|Gerard|Antoine|Jean|Thierry|Fabrice|Claude|Andrew|Brian|Guy|Christopher|Eric|Roger|Steffan|Ryan|Thomas|Conrad|Alvaro|Diego|Antonio)$/i.test(name)) return 'male';
if (/^(Denise|Eloise|Vivienne|Charline|Sylvie|Ariane|Ava|Emma|Jenny|Aria|Michelle|Ana|Sonia|Libby|Katja|Elvira|Elsa|Francisca)$/i.test(name)) return 'female';
return undefined;
}
+18
View File
@@ -111,6 +111,10 @@ export const IPC = {
httpStreamEvent: 'http:stream:event',
metrics: 'system:metrics',
appInfo: 'app:info',
screenCapture: 'system:screen-capture',
systemAction: 'system:action',
mcpRequest: 'mcp:request',
mcpResponse: 'mcp:response',
readLocalFile: 'files:read-local',
writeSharedFile: 'files:write-shared',
openPath: 'files:open-path',
@@ -122,3 +126,17 @@ export const IPC = {
} as const;
export type HotkeyEvent = 'ptt-toggle' | 'toggle-window' | 'stop-speaking';
/** A tool call arriving from Hermes through the local MCP endpoint that needs the renderer (UI, voice, chat). */
export interface McpToolRequest {
id: string;
name: string;
args: Record<string, unknown>;
}
export interface McpToolResponse {
id: string;
ok: boolean;
result?: unknown;
error?: string;
}
+108
View File
@@ -0,0 +1,108 @@
/**
* Encodes wake phrases into the BPE token sequences expected by the sherpa-onnx keyword
* spotter (gigaspeech BPE-500 model). Common phrases use sequences produced by the real
* SentencePiece model; anything else falls back to a greedy longest-match over the vocabulary,
* which is a close approximation for short words.
*/
const KNOWN: Record<string, string> = {
'JARVIS': '▁JA R VI S',
'HEY JARVIS': '▁HE Y ▁JA R VI S',
'OK JARVIS': '▁O K ▁JA R VI S',
'EVE': '▁E VE',
'HEY EVE': '▁HE Y ▁E VE',
'COMPUTER': '▁COMP U TER',
'HEY COMPUTER': '▁HE Y ▁COMP U TER',
'FRIDAY': '▁F RI DAY',
'HEY FRIDAY': '▁HE Y ▁F RI DAY',
'ALFRED': '▁A L F RE D',
'HERMES': '▁HER ME S',
'HEY HERMES': '▁HE Y ▁HER ME S',
'ASSISTANT': '▁AS S IST ANT',
'OK GOOGLE': '▁O K ▁GO O G LE',
'ALEXA': '▁A LE X A',
'SIRI': '▁S I RI',
'HEY SIRI': '▁HE Y ▁S I RI',
'NOVA': '▁NO V A',
'ATLAS': '▁AT LA S',
'HAL': '▁HA L'
};
export function normalizeKeyword(phrase: string): string {
return phrase
.normalize('NFD')
.replace(/[̀-ͯ]/g, '')
.toUpperCase()
.replace(/[^A-Z ]+/g, ' ')
.replace(/\s+/g, ' ')
.trim();
}
/** Parse a sherpa tokens.txt ("piece id" per line) into the set of pieces. */
export function parseTokens(tokensFile: string): Set<string> {
const pieces = new Set<string>();
for (const line of tokensFile.split(/\r?\n/)) {
const piece = line.trim().split(/\s+/)[0];
if (piece && !piece.startsWith('<')) pieces.add(piece);
}
return pieces;
}
function greedyWord(word: string, vocab: Set<string>): string[] | null {
const out: string[] = [];
let i = 0;
let first = true;
while (i < word.length) {
let matched = '';
for (let len = word.length - i; len >= 1; len--) {
const candidate = (first ? '▁' : '') + word.slice(i, i + len);
if (vocab.has(candidate)) {
matched = candidate;
break;
}
}
if (!matched && first) {
// No word-initial piece: use a bare "▁" if available, then continue without the prefix.
if (vocab.has('▁')) out.push('▁');
first = false;
continue;
}
if (!matched) return null;
out.push(matched);
i += matched.length - (first ? 1 : 0);
first = false;
}
return out;
}
/** Encode a phrase into space-separated BPE pieces, or null when it cannot be represented. */
export function encodeKeyword(phrase: string, vocab: Set<string>): string | null {
const normalized = normalizeKeyword(phrase);
if (!normalized) return null;
if (KNOWN[normalized]) return KNOWN[normalized];
const pieces: string[] = [];
for (const word of normalized.split(' ')) {
const encoded = greedyWord(word, vocab);
if (!encoded) return null;
pieces.push(...encoded);
}
return pieces.join(' ');
}
/** Build the keywords file content: one line per phrase, with the readable label. */
export function buildKeywordsFile(phrases: string[], vocab: Set<string>): { content: string; accepted: string[]; rejected: string[] } {
const lines: string[] = [];
const accepted: string[] = [];
const rejected: string[] = [];
for (const phrase of phrases) {
const encoded = encodeKeyword(phrase, vocab);
const label = normalizeKeyword(phrase).toLowerCase().replace(/ /g, '_');
if (!encoded || !label) {
rejected.push(phrase);
continue;
}
lines.push(`${encoded} @${label}`);
accepted.push(label);
}
return { content: lines.join('\n') + '\n', accepted, rejected };
}
+67 -3
View File
@@ -1,12 +1,14 @@
/** Local voice engine contract (sherpa-onnx in a utility process). Shared by main and renderer. */
export type VoiceModelKind = 'stt' | 'tts';
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper';
export type VoiceModelKind = 'stt' | 'tts' | 'kws' | 'vad';
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'supertonic' | 'kws-transducer' | 'silero';
export interface VoiceSpeaker {
id: number;
name: string;
lang: string;
/** m = masculine, f = feminine (from the catalog label). */
gender?: 'm' | 'f';
}
export interface VoiceModelSpec {
@@ -60,6 +62,31 @@ export interface SynthesizeRequest {
text: string;
speaker: number;
speed: number;
/** BCP-47 tag or 2-letter code of the text (multilingual engines such as Supertonic need it). */
language?: string;
}
/** Microsoft Edge "Read aloud" neural voices (online, no key). */
export interface EdgeVoice {
/** e.g. fr-FR-HenriNeural */
shortName: string;
/** Display name, e.g. Henri */
name: string;
locale: string;
gender: 'm' | 'f';
}
export interface EdgeSynthesizeRequest {
text: string;
voice: string;
/** Playback speed multiplier (0.5 .. 2). */
speed: number;
}
export interface EdgeSynthesizeResult {
/** MP3, 24 kHz mono 48 kbit/s. */
mp3: Uint8Array;
durationMs: number;
}
export interface SynthesizeResult {
@@ -69,6 +96,33 @@ export interface SynthesizeResult {
audioSec: number;
}
export interface KwsStartRequest {
modelId: string;
/** Wake phrases in plain text (e.g. "jarvis", "hey jarvis"). */
keywords: string[];
/** 1 (strict) .. 5 (eager) */
sensitivity: number;
}
export interface KwsDetection {
keyword: string;
at: number;
}
export interface VadStartRequest {
modelId: string;
/** Silence that ends an utterance, in ms. */
silenceMs: number;
/** Detection threshold 0..1 (0.5 default). */
threshold: number;
maxUtteranceSec: number;
}
export type VadEvent =
| { type: 'speech-start' }
| { type: 'segment'; wav: Uint8Array; durationSec: number }
| { type: 'error'; message: string };
export interface VoiceEngineStatus {
available: boolean;
error?: string;
@@ -86,5 +140,15 @@ export const VOICE_IPC = {
modelsProgress: 'voice:models:progress',
transcribe: 'voice:transcribe',
synthesize: 'voice:synthesize',
unload: 'voice:unload'
edgeSynthesize: 'voice:edge:synthesize',
edgeVoices: 'voice:edge:voices',
unload: 'voice:unload',
kwsStart: 'voice:kws:start',
kwsStop: 'voice:kws:stop',
kwsAudio: 'voice:kws:audio',
kwsDetected: 'voice:kws:detected',
vadStart: 'voice:vad:start',
vadStop: 'voice:vad:stop',
vadAudio: 'voice:vad:audio',
vadEvent: 'voice:vad:event'
} as const;
+36 -2
View File
@@ -3,7 +3,10 @@ import { AlertTriangle, X } from 'lucide-react';
import type { WindowMode } from '../shared/ipc';
import { bridge } from './lib/bridge';
import { Log } from './lib/log';
import { handlePush, stopGeneration } from './services/conversation';
import { handlePush, localToolContext, stopGeneration } from './services/conversation';
import { initMcpBridge } from './services/mcpBridge';
import { ensurePreferredVoice } from './services/voice/voicePreference';
import { isQuietTime } from './lib/quietHours';
import { speech } from './services/voice/speech';
import { voiceController } from './services/voice/voiceController';
import { useChat } from './state/chat';
@@ -43,8 +46,21 @@ function useBoot(): boolean {
}
if (api) {
disposers.push(useVoiceModels.getState().subscribe());
void useVoiceModels.getState().refresh();
void useVoiceModels.getState().refresh().then(() => ensurePreferredVoice());
if (settings.voice.wakeMode === 'kws') void voiceController.startWakeMode();
let lastWake = JSON.stringify([settings.voice.wakeMode, settings.voice.wakeWord, settings.voice.kwsSensitivity, settings.voice.micDeviceId]);
disposers.push(
useSettings.subscribe((s) => {
const v = s.settings.voice;
const key = JSON.stringify([v.wakeMode, v.wakeWord, v.kwsSensitivity, v.micDeviceId]);
if (key === lastWake) return;
lastWake = key;
if (v.wakeMode === 'kws') void voiceController.restartWakeMode();
else void voiceController.stopWakeMode();
})
);
disposers.push(api.hermes.onPush(handlePush));
initMcpBridge(localToolContext);
disposers.push(
api.hotkeys.on((event) => {
if (event === 'ptt-toggle') voiceController.toggle();
@@ -83,6 +99,24 @@ function useBoot(): boolean {
function useThemeSync(): void {
const theme = useSettings((s) => s.settings.theme);
const reduce = useSettings((s) => s.settings.ui.reduceMotion);
// Quiet hours: computed every 30 s; dims the HUD (night theme) and silences pushes.
const notif = useSettings((s) => s.settings.notifications);
useEffect(() => {
const tick = () => {
const quiet = notif.quietEnabled && isQuietTime(notif.quietStart, notif.quietEnd);
useChat.getState().setQuiet(quiet);
document.documentElement.toggleAttribute('data-night', quiet && notif.nightTheme);
};
tick();
const timer = setInterval(tick, 30_000);
return () => clearInterval(timer);
}, [notif]);
useEffect(() => {
const onVisible = () => { if (!document.hidden) useChat.getState().markRead(); };
document.addEventListener('visibilitychange', onVisible);
return () => document.removeEventListener('visibilitychange', onVisible);
}, []);
useEffect(() => {
document.documentElement.dataset.theme = theme;
document.documentElement.dataset.motion = reduce ? 'reduce' : 'full';
+21 -5
View File
@@ -1,6 +1,6 @@
import { useEffect, useRef, useState } from 'react';
import { Mic, Paperclip, Send, Square, X, Radio, Loader2, Volume2, VolumeX, Navigation } from 'lucide-react';
import { sendMessage, steer, stopGeneration } from '../../services/conversation';
import { Mic, Paperclip, Send, Square, X, Radio, Loader2, Volume2, VolumeX, Navigation, ScanEye, Crosshair } from 'lucide-react';
import { sendMessage, steer, stopGeneration , captureScreen } from '../../services/conversation';
import { speech } from '../../services/voice/speech';
import { voiceController } from '../../services/voice/voiceController';
import { useChat } from '../../state/chat';
@@ -50,6 +50,8 @@ export function CommandBar({ compact }: Props) {
const autoSpeak = useSettings((s) => s.settings.speech.autoSpeak);
const update = useSettings((s) => s.update);
const [images, setImages] = useState<string[]>([]);
const missionMode = useChat((s) => s.missionMode);
const missionModel = useSettings((s) => s.settings.hermes.missionModel);
const [steerMode, setSteerMode] = useState(false);
const fileRef = useRef<HTMLInputElement>(null);
const inputRef = useRef<HTMLTextAreaElement>(null);
@@ -167,9 +169,20 @@ export function CommandBar({ compact }: Props) {
spellCheck={false}
/>
{!compact && (
<button className="icon-btn" title="Joindre une image" onClick={() => fileRef.current?.click()} style={{ height: 46, width: 40 }}>
<Paperclip size={18} />
</button>
<>
<button
className="icon-btn"
title="Joindre une capture de l’écran (Hermes la voit)"
aria-label="Capturer l’écran"
onClick={() => void captureScreen().then((shot) => { if (shot) setImages((prev) => [...prev, shot].slice(-4)); })}
style={{ height: 46, width: 40 }}
>
<ScanEye size={18} />
</button>
<button className="icon-btn" title="Joindre une image" onClick={() => fileRef.current?.click()} style={{ height: 46, width: 40 }}>
<Paperclip size={18} />
</button>
</>
)}
<input ref={fileRef} type="file" accept="image/*" multiple hidden onChange={(e) => void onFiles(e.target.files)} />
{isSending && !steerMode ? (
@@ -187,6 +200,9 @@ export function CommandBar({ compact }: Props) {
<button className={`toggle-btn${handsFree ? ' on' : ''}`} onClick={() => voiceController.setHandsFree(!handsFree)} title="Écoute continue : le micro se réactive après chaque réponse">
<Radio size={12} /> mains libres
</button>
<button className={`toggle-btn${missionMode ? ' on' : ''}`} onClick={() => useChat.getState().setMissionMode(!missionMode)} title={missionModel ? `Mode mission : modèle ${missionModel} pour les tâches longues` : 'Mode mission (définissez un modèle « mission » dans Paramètres → Hermes)'}>
<Crosshair size={14} /> Mission
</button>
<button className={`toggle-btn${autoSpeak && speechProvider !== 'off' ? ' on' : ''}`} onClick={() => update({ speech: { autoSpeak: !autoSpeak } })} title="Lire les réponses à voix haute">
{autoSpeak && speechProvider !== 'off' ? <Volume2 size={12} /> : <VolumeX size={12} />} voix
</button>
+34 -1
View File
@@ -1,13 +1,37 @@
import { Maximize2, X, Minus } from 'lucide-react';
import { Maximize2, X, Minus, Moon, Crosshair } from 'lucide-react';
import { useShallow } from 'zustand/react/shallow';
import { bridge } from '../../lib/bridge';
import { useSettings } from '../../state/settings';
import { useChat } from '../../state/chat';
import { useVoice } from '../../state/voice';
import { CoreStage } from '../hud/CoreStage';
import { ChatPanel } from '../chat/ChatPanel';
const STATE_LABEL: Record<string, string> = {
idle: 'veille',
listening: 'écoute',
thinking: 'réflexion',
speaking: 'parle',
alert: 'alerte',
error: 'erreur',
success: 'terminé'
};
/** Compact mode: a glanceable strip (state, unread, last sentence) above the mini chat. */
export function CompactWidget() {
const assistantName = useSettings((s) => s.settings.assistantName);
const compactOpacity = useSettings((s) => s.settings.ui.compactOpacity);
const { hud, hudOverride, unread, missionMode, quiet, last } = useChat(
useShallow((s) => {
const lastAssistant = [...s.messages].reverse().find((m) => m.role === 'assistant' && m.content.trim());
return { hud: s.hud, hudOverride: s.hudOverride, unread: s.unread, missionMode: s.missionMode, quiet: s.quiet, last: lastAssistant?.content ?? '' };
})
);
const wake = useVoice((s) => s.wake);
const api = bridge();
const state = hudOverride ?? hud;
const label = state === 'idle' && wake === 'spotting' ? 'à l’écoute du mot-clé' : STATE_LABEL[state] ?? state;
const sentence = last.replace(/```[\s\S]*?```/g, ' ').replace(/[#*_>`|]/g, '').replace(/\s+/g, ' ').trim().slice(0, 140);
return (
<div className="compact-root" style={{ ['--compact-opacity' as string]: compactOpacity }}>
<div className="compact-head">
@@ -21,6 +45,15 @@ export function CompactWidget() {
)}
</div>
<CoreStage compact />
<div className="compact-glance" onClick={() => useChat.getState().markRead()} title={sentence}>
<span className={`state${state === 'alert' || state === 'error' ? ' alert' : ''}`}>{label}</span>
<span className="last">{sentence || 'Aucun message pour l’instant.'}</span>
<span className="badges">
{missionMode && <Crosshair size={12} className="mission" aria-label="mode mission" />}
{quiet && <Moon size={12} aria-label="heures calmes" />}
{unread > 0 && <span className="unread" aria-label={`${unread} non lus`}>{unread}</span>}
</span>
</div>
<div className="compact-body">
<ChatPanel compact />
</div>
+3
View File
@@ -35,6 +35,8 @@ export function CoreStage({ compact }: Props) {
const transcript = useVoice((s) => s.lastTranscript);
const interim = useVoice((s) => s.interim);
const voiceError = useVoice((s) => s.error);
const wake = useVoice((s) => s.wake);
const wakeKeywords = useVoice((s) => s.wakeKeywords);
const assistantName = useSettings((s) => s.settings.assistantName);
const reduceMotion = useSettings((s) => s.settings.ui.reduceMotion);
const transport = useHermes((s) => s.transport);
@@ -57,6 +59,7 @@ export function CoreStage({ compact }: Props) {
else if (hud === 'error') sub = previewText(error ?? voiceError ?? 'erreur', 90);
else if (hud === 'speaking') sub = 'synthèse vocale';
else if (transcript) sub = `« ${previewText(transcript, 80)} »`;
else if (wake === 'spotting') sub = `dites « ${wakeKeywords[0] ?? 'jarvis'} »${link === 'online' ? ` · ${transport}` : link === 'offline' ? ' · Hermes hors ligne' : ''}`;
else sub = link === 'online' ? `liaison Hermes · ${transport}` : link === 'offline' ? 'Hermes hors ligne' : '';
return (
+4 -1
View File
@@ -1,6 +1,7 @@
import { Fragment, useEffect, useState } from 'react';
import { useShallow } from 'zustand/react/shallow';
import {
AlertTriangle,
Wrench, CalendarClock, Sparkles, Layers, History, Play, Pause, Trash2, Pencil, Plus, Save, X, RefreshCw,
CheckCircle2, XCircle, Loader2, Bot, Cpu, GitBranch, Webhook, FolderOpen
} from 'lucide-react';
@@ -41,6 +42,7 @@ function ActivityTab() {
function JobsTab() {
const jobs = useHermes((s) => s.jobs);
const jobsError = useHermes((s) => s.jobsError);
const busy = useHermes((s) => s.busy);
const lastSyncAt = useHermes((s) => s.lastSyncAt);
const refreshJobs = useHermes((s) => s.refreshJobs);
@@ -95,8 +97,9 @@ function JobsTab() {
</div>
</div>
)}
{jobsError && <div className="test-result fail" style={{ marginBottom: 8 }}><AlertTriangle size={13} /> Crons indisponibles : {jobsError}</div>}
{jobs.length === 0 ? (
<div className="empty">Aucun cron Hermes. Créez une mission planifiée en langage naturel.</div>
<div className="empty">{jobsError ? 'La liaison Hermes reste utilisable ; seule la planification est indisponible.' : 'Aucun cron Hermes. Créez une mission planifiée en langage naturel.'}</div>
) : (
<div className="list">
{jobs.map((job) => {
+25 -5
View File
@@ -1,8 +1,9 @@
import { useEffect } from 'react';
import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, RefreshCw } from 'lucide-react';
import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, RefreshCw, Ear, Activity } from 'lucide-react';
import type { VoiceModelStatus } from '../../../shared/voice';
import { bridge } from '../../lib/bridge';
import { useShallow } from 'zustand/react/shallow';
import { pickSpeaker } from '../../lib/voicePreference';
import { useVoiceModels } from '../../state/voiceModels';
import { useSettings } from '../../state/settings';
@@ -19,20 +20,28 @@ function ModelRow({ model }: { model: VoiceModelStatus }) {
const isActive = model.kind === 'stt' ? activeStt === model.id : activeTts === model.id;
const busy = !!progress && (progress.phase === 'download' || progress.phase === 'extract');
const wakeMode = useSettings((s) => s.settings.voice.wakeMode);
const activate = () => {
if (model.kind === 'stt') update({ voice: { localModel: model.id, provider: 'local' } });
else update({ speech: { localModel: model.id, provider: 'local', localSpeaker: model.speakers?.[0]?.id ?? 0 } });
else if (model.kind === 'kws') update({ voice: { wakeMode: 'kws' } });
else if (model.kind === 'vad') update({ voice: { neuralVad: true } });
else {
const { speech } = useSettings.getState().settings;
update({ speech: { localModel: model.id, provider: 'local', localSpeaker: pickSpeaker(model, speech.language || 'fr-FR', speech.voiceGender ?? 'male') } });
}
};
const neuralVad = useSettings((s) => s.settings.voice.neuralVad);
const activeNow = model.kind === 'kws' ? wakeMode === 'kws' : model.kind === 'vad' ? neuralVad : isActive;
return (
<div className="row-item" style={isActive && model.installed ? { borderColor: 'var(--accent)' } : undefined}>
<div className="row-item" style={activeNow && model.installed ? { borderColor: 'var(--accent)' } : undefined}>
<div className="main">
<div className="row wrap" style={{ gap: 6 }}>
<span className="name">{model.name}</span>
{model.recommended && <span className="status-pill ok">recommandé</span>}
{model.installed && <span className="status-pill online">installé · {formatMb(model.installedBytes)}</span>}
{!model.installed && !busy && <span className="status-pill">{model.sizeMb} Mo</span>}
{isActive && model.installed && <span className="status-pill active">actif</span>}
{activeNow && model.installed && <span className="status-pill active">actif</span>}
</div>
<span className="desc">{model.description}</span>
<span className="meta">langues : {model.languages.join(', ')}{model.speakers ? ` · ${model.speakers.length} voix` : ''}</span>
@@ -49,7 +58,7 @@ function ModelRow({ model }: { model: VoiceModelStatus }) {
<button className="icon-btn danger" title="Annuler" onClick={() => void cancel(model.id)}><X size={13} /></button>
) : model.installed ? (
<>
{!isActive && <button className="btn small" onClick={activate}>Utiliser</button>}
{!activeNow && <button className="btn small" onClick={activate}>Utiliser</button>}
<button className="icon-btn danger" title="Supprimer du disque" onClick={() => { if (confirm(`Supprimer « ${model.name} » ?`)) void remove(model.id); }}><Trash2 size={13} /></button>
</>
) : (
@@ -72,6 +81,8 @@ export function ModelsSection() {
const stt = models.filter((m) => m.kind === 'stt');
const tts = models.filter((m) => m.kind === 'tts');
const kws = models.filter((m) => m.kind === 'kws');
const vad = models.filter((m) => m.kind === 'vad');
return (
<>
<div className="card">
@@ -91,6 +102,15 @@ export function ModelsSection() {
<button className="btn small ghost" onClick={() => { void refresh(); void checkEngine(); }}><RefreshCw size={13} /> Actualiser</button>
</div>
</div>
<div className="card">
<div className="section-title"><Ear size={12} /> Mot d’activation</div>
<div className="list">{kws.map((m) => <ModelRow key={m.id} model={m} />)}</div>
</div>
<div className="card">
<div className="section-title"><Activity size={12} /> Fin de phrase</div>
<div className="list">{vad.map((m) => <ModelRow key={m.id} model={m} />)}</div>
<span className="hint" style={{ display: 'block', marginTop: 6 }}>Utilisé automatiquement en écoute permanente dès qu’il est installé ; sinon le VAD énergétique prend le relais.</span>
</div>
<div className="card">
<div className="section-title"><Mic size={12} /> Reconnaissance vocale</div>
<div className="list">{stt.map((m) => <ModelRow key={m.id} model={m} />)}</div>
+215 -15
View File
@@ -1,9 +1,12 @@
import { useEffect, useState } from 'react';
import { X, Settings, CheckCircle2, XCircle, Loader2, Mic, Volume2, RotateCcw, Play } from 'lucide-react';
import { bridge } from '../../lib/bridge';
import { HermesClient } from '../../services/hermes/client';
import { HermesClient, discoverHermesUrl, hermesUrlCandidates } from '../../services/hermes/client';
import { listSystemVoices } from '../../services/voice/tts';
import { speech } from '../../services/voice/speech';
import { ensurePreferredVoice } from '../../services/voice/voicePreference';
import { pickSpeaker, resolveEdgeVoice } from '../../lib/voicePreference';
import type { EdgeVoice } from '../../../shared/voice';
import { listMicrophones } from '../../services/voice/capture';
import { transcribeWav } from '../../services/voice/stt';
import { encodeWav } from '../../services/voice/wav';
@@ -13,7 +16,7 @@ import { useVoice } from '../../state/voice';
import { installedModels, useVoiceModels } from '../../state/voiceModels';
import { ModelsSection } from './ModelsSection';
type Section = 'general' | 'hermes' | 'voice' | 'speech' | 'models' | 'webhook' | 'ui';
type Section = 'general' | 'hermes' | 'voice' | 'speech' | 'models' | 'webhook' | 'notifications' | 'ui';
interface Props {
onClose: () => void;
@@ -33,6 +36,33 @@ function Toggle({ on, onChange, label, hint }: { on: boolean; onChange: (v: bool
type TestState = { status: 'idle' | 'running' | 'ok' | 'fail'; message: string };
function WakeStatus() {
const wake = useVoice((s) => s.wake);
const keywords = useVoice((s) => s.wakeKeywords);
const error = useVoice((s) => s.error);
const models = useVoiceModels((s) => s.models);
const download = useVoiceModels((s) => s.download);
const progress = useVoiceModels((s) => s.progress['kws-en']);
const installed = models.find((m) => m.id === 'kws-en')?.installed;
if (installed === false) {
return (
<div className="test-result fail" style={{ marginTop: 8 }}>
<XCircle size={13} />
<span>Détecteur non installé.</span>
<button className="btn small primary" style={{ marginLeft: 'auto' }} disabled={!!progress} onClick={() => void download('kws-en')}>
{progress ? `${progress.percent}%` : 'Télécharger (4 Mo)'}
</button>
</div>
);
}
return (
<div className={`test-result ${wake === 'spotting' ? 'ok' : wake === 'error' ? 'fail' : ''}`} style={{ marginTop: 8 }}>
{wake === 'spotting' ? <CheckCircle2 size={13} /> : wake === 'error' ? <XCircle size={13} /> : <Loader2 size={13} className="spin" />}
<span>{wake === 'spotting' ? `à l’écoute de « ${keywords.join(' », « ')} »` : wake === 'error' ? error ?? 'erreur' : 'démarrage…'}</span>
</div>
);
}
function TestResult({ t }: { t: TestState }) {
if (t.status === 'idle') return null;
return (
@@ -58,8 +88,18 @@ export function SettingsDrawer({ onClose }: Props) {
const [sttTest, setSttTest] = useState<TestState>({ status: 'idle', message: '' });
const [ttsTest, setTtsTest] = useState<TestState>({ status: 'idle', message: '' });
const [voices, setVoices] = useState(listSystemVoices());
const [edgeVoices, setEdgeVoices] = useState<EdgeVoice[]>([]);
const [webhookSecretVisible, setWebhookSecretVisible] = useState(false);
const speechProvider = settings.speech.provider;
useEffect(() => {
if (speechProvider !== 'edge' || edgeVoices.length) return;
bridge()
?.voice.edgeVoices()
.then(setEdgeVoices)
.catch((err: Error) => setTtsTest({ status: 'fail', message: `Liste des voix Edge indisponible : ${err.message}` }));
}, [speechProvider, edgeVoices.length]);
useEffect(() => {
const refresh = () => setVoices(listSystemVoices());
refresh();
@@ -80,6 +120,7 @@ export function SettingsDrawer({ onClose }: Props) {
caps = ` · modèle ${String(c.model ?? '?')} · ${Object.entries(c.features ?? {}).filter(([, v]) => v === true).length} fonctions`;
} catch (err) {
const message = (err as Error).message;
if (/page web/.test(message)) throw err;
caps = /404/.test(message)
? ' · /v1/capabilities absent (Hermes ancien : transport chat completions)'
: ` · capabilities : ${message}`;
@@ -88,7 +129,16 @@ export function SettingsDrawer({ onClose }: Props) {
setHermesTest({ status: 'ok', message: `${health.status} (${via})${caps}` });
void hermesConnect();
} catch (err) {
setHermesTest({ status: 'fail', message: (err as Error).message });
const message = (err as Error).message;
setHermesTest({ status: 'running', message: `${message} — recherche de l’API sur le même hôte…` });
const found = await discoverHermesUrl(settings.hermes, (u) => setHermesTest({ status: 'running', message: `essai ${u}…` })).catch(() => null);
if (found) {
update({ hermes: { url: found } });
setHermesTest({ status: 'ok', message: `API Hermes trouvée : ${found} (URL corrigée automatiquement). Relancez le test.` });
void hermesConnect();
} else {
setHermesTest({ status: 'fail', message: `${message} Aucune API trouvée sur ${hermesUrlCandidates(settings.hermes.url).slice(0, 4).join(', ')}…` });
}
}
};
@@ -124,6 +174,7 @@ export function SettingsDrawer({ onClose }: Props) {
['speech', 'Voix / TTS'],
['models', 'Modèles locaux'],
['webhook', 'Webhook'],
['notifications', 'Notifications'],
['general', 'Général'],
['ui', 'Interface']
];
@@ -188,6 +239,11 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
</div>
<div className="field">
<label>Modèle « mission » (tâches longues)</label>
<input className="input" list="hermes-models" value={settings.hermes.missionModel} placeholder="(même modèle)" onChange={(e) => update({ hermes: { missionModel: e.target.value } })} />
<span className="hint">Utilisé quand le mode Mission est activé dans la barre de commande. Laissez vide pour garder le modèle principal.</span>
</div>
<div className="field">
<label>Instructions EveFlow (superposées au prompt Hermes)</label>
<textarea className="textarea" value={settings.hermes.instructions} onChange={(e) => update({ hermes: { instructions: e.target.value } })} />
</div>
@@ -278,13 +334,49 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
<Toggle on={settings.voice.handsFree} onChange={(v) => { update({ voice: { handsFree: v } }); useVoice.getState().setHandsFree(v); }} label="Mains libres au démarrage" hint="Le micro se réactive automatiquement après chaque réponse." />
<Toggle on={settings.voice.wakeChime} onChange={(v) => update({ voice: { wakeChime: v } })} label="Signal sonore d’écoute" />
<Toggle on={settings.voice.wakeWordEnabled} onChange={(v) => update({ voice: { wakeWordEnabled: v } })} label="Mot d’activation en mains libres" hint="Seules les phrases commençant par ce mot sont envoyées à Hermes ; le reste est ignoré. Recommandé avec la reconnaissance locale." />
{settings.voice.wakeWordEnabled && (
<div className="field" style={{ marginTop: 8 }}>
<label>Mot d’activation</label>
<input className="input" value={settings.voice.wakeWord} placeholder="jarvis" onChange={(e) => update({ voice: { wakeWord: e.target.value.toLowerCase() } })} />
<div className="field" style={{ marginTop: 10 }}>
<label>Mot d’activation</label>
<select className="select" value={settings.voice.wakeMode} onChange={(e) => update({ voice: { wakeMode: e.target.value as typeof settings.voice.wakeMode, wakeWordEnabled: e.target.value === 'transcript' } })}>
<option value="off">Désactivé (bouton, raccourci ou mains libres)</option>
<option value="kws">Écoute permanente : le micro reste ouvert et réagit au mot-clé (recommandé)</option>
<option value="transcript">Filtre après transcription (mains libres) : les phrases sans le mot sont ignorées</option>
</select>
<span className="hint">
{settings.voice.wakeMode === 'kws'
? 'Détection locale par un modèle de 3 Mo, quasi gratuite en CPU. Dire le mot seul ouvre l’écoute ; dire le mot puis la commande envoie directement. Le mot coupe aussi la voix en cours.'
: settings.voice.wakeMode === 'transcript'
? 'Chaque phrase est transcrite puis filtrée : plus coûteux, à réserver à la reconnaissance locale.'
: 'Le micro s’active avec le bouton, Ctrl+Shift+Espace ou la boucle mains libres.'}
</span>
</div>
{settings.voice.wakeMode !== 'off' && (
<div className="grid-2">
<div className="field">
<label>Mot-clé</label>
<input className="input" value={settings.voice.wakeWord} placeholder="jarvis" onChange={(e) => update({ voice: { wakeWord: e.target.value.toLowerCase().slice(0, 40) } })} />
<span className="hint">Prononciation anglaise conseillée (« jarvis », « hey jarvis », « computer », « friday »…).</span>
</div>
{settings.voice.wakeMode === 'kws' && (
<div className="field">
<label>Sensibilité du mot-clé : {settings.voice.kwsSensitivity}/5</label>
<input className="range" type="range" min={1} max={5} step={1} value={settings.voice.kwsSensitivity} onChange={(e) => update({ voice: { kwsSensitivity: Number(e.target.value) } })} />
</div>
)}
</div>
)}
{settings.voice.wakeMode === 'kws' && (
<>
<WakeStatus />
<Toggle on={settings.voice.neuralVad} onChange={(v) => update({ voice: { neuralVad: v } })} label="Fin de phrase neuronale (Silero)" hint="Coupe l’écoute au bon moment, même avec du bruit de fond. Nécessite le modèle Silero VAD (0,6 Mo) dans Modèles locaux." />
</>
)}
<Toggle on={settings.voice.localCommands} onChange={(v) => update({ voice: { localCommands: v } })} label="Commandes locales instantanées" hint="« Verrouille la session », « monte le son », « ouvre Spotify », « regarde mon écran »… exécutées sur ce PC sans passer par Hermes." />
<Toggle
on={settings.voice.micProcessing ?? true}
onChange={(v) => update({ voice: { micProcessing: v } })}
label="Traitement du micro par Chromium (écho, bruit, gain automatique)"
hint="Désactivez-le si les transcriptions sont approximatives avec un casque ou un bon micro : ces filtres déforment la voix avant la reconnaissance. Gardez-le activé avec des haut-parleurs (sinon la voix de l’assistant est réentendue par le micro)."
/>
<div className="row" style={{ marginTop: 10 }}>
<button className="btn small" onClick={() => void testStt()} disabled={settings.voice.provider === 'browser'}><Mic size={13} /> Tester la reconnaissance</button>
</div>
@@ -294,16 +386,86 @@ export function SettingsDrawer({ onClose }: Props) {
{section === 'speech' && (
<div className="card">
<div className="field">
<label>Voix</label>
<div className="segmented">
<button className={(settings.speech.voiceGender ?? 'male') === 'male' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'male' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Masculine</button>
<button className={settings.speech.voiceGender === 'female' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'female' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Féminine</button>
</div>
<span className="hint">
S’applique à tous les moteurs : Edge (Henri / Denise), voix locale (Supertonic 3, cinq voix de chaque genre, téléchargé automatiquement si besoin), API OpenAI (onyx / nova), voix système Windows (Paul / Hortense).
{settings.speech.provider === 'google-free' && ' Google Translate n’a qu’une voix féminine : ce choix n’a pas d’effet avec ce moteur.'}
{settings.speech.provider === 'local' && settings.speech.localModel === 'kokoro-v1' && ' Kokoro n’a qu’une voix française, féminine et avec accent : la voix masculine bascule sur Supertonic 3.'}
</span>
</div>
<div className="field">
<label>Timbre</label>
<div className="segmented">
<button className={settings.speech.timbre === 'jarvis' ? 'active' : ''} onClick={() => update({ speech: { timbre: 'jarvis' } })}>JARVIS</button>
<button className={settings.speech.timbre !== 'jarvis' ? 'active' : ''} onClick={() => update({ speech: { timbre: 'natural' } })}>Naturel</button>
</div>
<span className="hint">JARVIS : voix légèrement plus grave et posée, chaleur dans les basses, présence, courte réverbération d’intercom, comme dans le film.</span>
</div>
<div className="row" style={{ marginBottom: 10 }}>
<button
className="btn small primary"
onClick={() => {
update({ speech: { provider: 'edge', edgeVoice: '', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
void ensurePreferredVoice().then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
}}
>
<Volume2 size={13} /> Voix JARVIS en ligne (Edge Henri, masculine)
</button>
<button
className="btn small"
onClick={() => {
update({ speech: { provider: 'local', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
setTtsTest({ status: 'running', message: 'préparation de la voix JARVIS locale (téléchargement de Supertonic 3, 129 Mo, si nécessaire)…' });
void ensurePreferredVoice({ upgrade: true }).then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS locale prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
}}
>
<Volume2 size={13} /> Voix JARVIS hors ligne (Supertonic 3, masculine)
</button>
<span className="status-pill">
{settings.speech.provider === 'local'
? `voix active : ${voiceModels.find((m) => m.id === settings.speech.localModel)?.speakers?.find((s) => s.id === settings.speech.localSpeaker)?.name ?? settings.speech.localModel}`
: settings.speech.provider === 'edge'
? `voix active : ${resolveEdgeVoice(settings.speech.edgeVoice ?? '', settings.speech.language, settings.speech.voiceGender ?? 'male')}`
: `moteur : ${settings.speech.provider}`}
</span>
</div>
<div className="field">
<label>Moteur de synthèse</label>
<select className="select" value={settings.speech.provider} onChange={(e) => update({ speech: { provider: e.target.value as typeof settings.speech.provider } })}>
<option value="local">Local dans l’application (Kokoro / Piper via sherpa-onnx, hors ligne)</option>
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Kokoro, Piper, OpenAI, LocalAI…)</option>
<select
className="select"
value={settings.speech.provider}
onChange={(e) => {
update({ speech: { provider: e.target.value as typeof settings.speech.provider } });
void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m }));
}}
>
<option value="edge">Microsoft Edge (voix neuronales, gratuit, en ligne, sans clé) — recommandé</option>
<option value="local">Local dans l’application (Supertonic 3 / Kokoro / Piper via sherpa-onnx, hors ligne)</option>
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Qwen3-TTS, Kokoro, OpenAI, LocalAI…)</option>
<option value="system">Voix système Windows</option>
<option value="google-free">Google Translate (gratuit, en ligne)</option>
<option value="google-free">Google Translate (gratuit, en ligne, voix féminine uniquement)</option>
<option value="off">Désactivée</option>
</select>
</div>
{settings.speech.provider === 'edge' && (
<div className="field">
<label>Voix Edge</label>
<select className="select" value={settings.speech.edgeVoice ?? ''} onChange={(e) => update({ speech: { edgeVoice: e.target.value } })}>
<option value="">Automatique ({resolveEdgeVoice('', settings.speech.language, settings.speech.voiceGender ?? 'male')})</option>
{edgeVoices
.filter((v) => v.locale.toLowerCase().startsWith(settings.speech.language.toLowerCase().split('-')[0]))
.map((v) => (
<option key={v.shortName} value={v.shortName}>{v.name} · {v.locale} · {v.gender === 'm' ? 'homme' : 'femme'}</option>
))}
</select>
<span className="hint">Mêmes voix que la lecture à voix haute d’Edge : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR), plus les voix canadiennes, suisses et belges. Aucune donnée locale ; chaque phrase est synthétisée en ligne.</span>
</div>
)}
{settings.speech.provider === 'openai-compatible' && (
<>
<div className="field">
@@ -317,7 +479,7 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
<div className="field">
<label>Voix</label>
<input className="input" value={settings.speech.voice} placeholder="alloy, onyx, af_heart…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
<input className="input" value={settings.speech.voice} placeholder="onyx, nova, af_heart, Ryan (Qwen3-TTS)…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
</div>
<div className="field">
<label>Clé API</label>
@@ -337,7 +499,7 @@ export function SettingsDrawer({ onClose }: Props) {
{settings.speech.provider === 'local' && (
installedModels(voiceModels, 'tts').length === 0 ? (
<div className="field">
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Kokoro dans « Modèles locaux »</a>.</span>
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Supertonic 3 dans « Modèles locaux »</a>.</span>
</div>
) : (
<div className="grid-2">
@@ -345,7 +507,8 @@ export function SettingsDrawer({ onClose }: Props) {
<label>Modèle local</label>
<select className="select" value={settings.speech.localModel} onChange={(e) => {
const m = voiceModels.find((x) => x.id === e.target.value);
update({ speech: { localModel: e.target.value, localSpeaker: m?.speakers?.[0]?.id ?? 0 } });
// Keep the preferred gender when switching models (Kokoro has no masculine French voice: it falls back to Siwis).
update({ speech: { localModel: e.target.value, localSpeaker: m ? pickSpeaker(m, settings.speech.language, settings.speech.voiceGender ?? 'male') : 0 } });
}}>
{installedModels(voiceModels, 'tts').map((m) => <option key={m.id} value={m.id}>{m.name}</option>)}
</select>
@@ -414,6 +577,14 @@ export function SettingsDrawer({ onClose }: Props) {
-d '{"role":"assistant","text":"Rapport terminé","source":"telegram"}'`}</pre>
<span className="hint">Formats acceptés : {'{role,text}'}, {'{event:"run.completed",input,output}'}, {'{event:"job.completed",job:{name},output,status}'}, {'{type:"message",payload:{text}}'}.</span>
</div>
<div className="field">
<label>Serveur MCP pour Hermes (outils du PC : écran, applications, volume, presse-papiers, voix)</label>
<pre className="input" style={{ whiteSpace: 'pre-wrap', fontSize: 11.5, userSelect: 'text' }}>{`# ~/.hermes/config.yaml côté Hermes
mcp_servers:
eveflow:
url: "http://<ip-de-ce-pc>:${settings.webhook.port}/mcp"${settings.webhook.secret ? '\n headers:\n Authorization: "Bearer <secret>"' : ''}`}</pre>
<span className="hint">Même port et même secret que le webhook. Sans secret, EveFlow n’écoute qu’en local (127.0.0.1) : définissez un secret pour un Hermes distant.</span>
</div>
<div className="row">
<button className="btn primary small" onClick={() => void applyWebhook()} disabled={!bridge()}>Appliquer et redémarrer</button>
{hermesWebhook && <span className={`status-pill ${hermesWebhook.listening ? 'online' : 'offline'}`}>{hermesWebhook.listening ? `port ${hermesWebhook.port}` : hermesWebhook.error ?? 'inactif'}</span>}
@@ -421,6 +592,35 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
)}
{section === 'notifications' && (
<div className="card">
<Toggle on={settings.notifications.quietEnabled} onChange={(v) => update({ notifications: { quietEnabled: v } })} label="Heures calmes" hint="Les messages poussés (crons, Telegram) s’affichent sans être lus à voix haute ni faire clignoter le noyau ; le HUD passe en mode nuit." />
<div className="grid-2">
<div className="field">
<label>Début</label>
<input className="input" type="time" value={settings.notifications.quietStart} onChange={(e) => update({ notifications: { quietStart: e.target.value } })} />
</div>
<div className="field">
<label>Fin</label>
<input className="input" type="time" value={settings.notifications.quietEnd} onChange={(e) => update({ notifications: { quietEnd: e.target.value } })} />
</div>
</div>
<Toggle on={settings.notifications.nightTheme} onChange={(v) => update({ notifications: { nightTheme: v } })} label="Thème nuit pendant les heures calmes" />
<div className="field" style={{ marginTop: 10 }}>
<label>Mots prioritaires (lus même en heures calmes)</label>
<input className="input" value={settings.notifications.priorityKeywords} onChange={(e) => update({ notifications: { priorityKeywords: e.target.value } })} placeholder="urgent, alerte, panne" />
<span className="hint">Séparés par des virgules ; comparés au texte et au nom du cron. Les échecs de crons sont toujours prioritaires.</span>
</div>
<Toggle on={settings.notifications.summarizeIncoming} onChange={(v) => update({ notifications: { summarizeIncoming: v } })} label="Résumé vocal des messages entrants" hint="Seules les premières phrases sont lues ; le message complet reste dans le fil." />
{settings.notifications.summarizeIncoming && (
<div className="field">
<label>Phrases lues : {settings.notifications.summarySentences}</label>
<input className="range" type="range" min={1} max={5} step={1} value={settings.notifications.summarySentences} onChange={(e) => update({ notifications: { summarySentences: Number(e.target.value) } })} />
</div>
)}
</div>
)}
{section === 'general' && (
<div className="card">
<div className="grid-2">
+43
View File
@@ -0,0 +1,43 @@
/** Quiet-hours and spoken-summary helpers (pure, unit tested). */
function minutesOf(hhmm: string): number | null {
const m = /^(\d{1,2}):(\d{2})$/.exec(hhmm.trim());
if (!m) return null;
const h = Number(m[1]);
const min = Number(m[2]);
if (h > 23 || min > 59) return null;
return h * 60 + min;
}
/** True when `now` falls inside [start, end), with ranges crossing midnight ("22:30" → "07:30"). */
export function isQuietTime(start: string, end: string, now: Date = new Date()): boolean {
const s = minutesOf(start);
const e = minutesOf(end);
if (s === null || e === null || s === e) return false;
const cur = now.getHours() * 60 + now.getMinutes();
return s < e ? cur >= s && cur < e : cur >= s || cur < e;
}
/** True when the text or job name contains one of the comma-separated priority words. */
export function isPriority(text: string, keywords: string, jobName = ''): boolean {
const words = keywords
.split(/[,;\n]/)
.map((w) => w.trim().toLowerCase())
.filter(Boolean);
if (!words.length) return false;
const hay = `${jobName} ${text}`.toLowerCase();
return words.some((w) => hay.includes(w));
}
/** First `count` sentences of a text, for a short spoken summary of a long push. */
export function summarize(text: string, count: number): string {
const clean = text
.replace(/```[\s\S]*?```/g, ' ')
.replace(/[#*_>`|]/g, '')
.replace(/\s+/g, ' ')
.trim();
if (!clean) return '';
const sentences = clean.match(/[^.!?…]+[.!?…]+["»)]?\s*|[^.!?…]+$/g) ?? [clean];
const picked = sentences.slice(0, Math.max(1, count)).join('').trim();
return picked.length < clean.length ? picked : clean;
}
+24
View File
@@ -154,3 +154,27 @@ export function formatDuration(seconds: number): string {
if (h > 0) return `${h}h ${m.toString().padStart(2, '0')}m`;
return `${m}m ${(s % 60).toString().padStart(2, '0')}s`;
}
const NOISE_PHRASES = [
'sous-titres réalisés par la communauté d\'amara.org',
'sous-titrage société radio-canada',
'merci d\'avoir regardé',
'abonnez-vous',
'thank you for watching',
'thanks for watching',
'...'
];
/**
* Whisper-style hallucinations on silence or clicks: "(cliquant)", "*Claire*", "[Musique]",
* "Sous-titres réalisés par…". Such transcripts must not be sent to Hermes.
*/
export function isTranscriptNoise(text: string): boolean {
const t = text.trim();
if (!t) return true;
if (/^[\s\p{P}\p{S}]*$/u.test(t)) return true;
// whole transcript wrapped in brackets/asterisks: a sound description
if (/^[(\[*«"'\s]+[^()\[\]*]{0,60}[)\]*»"'\s]+$/u.test(t) && !/[a-zà-ÿ]{3,}\s+[a-zà-ÿ]{3,}\s+[a-zà-ÿ]{3,}/i.test(t)) return true;
const lower = t.toLowerCase().replace(/[.!?…\s]+$/u, '');
return NOISE_PHRASES.some((p) => lower === p.replace(/[.!?…\s]+$/u, ''));
}
+110
View File
@@ -0,0 +1,110 @@
/** Voice gender preference applied to every TTS provider (pure helpers, unit tested). */
import type { VoiceModelStatus, VoiceSpeaker } from '../../shared/voice';
import { defaultEdgeVoice, edgeVoiceGender } from '../../shared/edgeTts';
export type VoiceGender = 'male' | 'female';
const MALE_HINTS =
/\b(homme|male|masculin|paul|thomas|claude|henri|remy|rémy|gerard|guillaume|mathieu|antoine|nicolas|denis|pierre|tom|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|david|mark|richard|james|ryan|guy|dylan|aiden|uncle fu|andrew|brian|fabrice|jean|thierry)\b/i;
const FEMALE_HINTS =
/\b(femme|female|f[ée]minin|hortense|julie|denise|eloise|vivienne|charline|sylvie|ariane|am[ée]lie|audrey|siwis|jessica|heart|bella|sarah|nicole|sky|alloy|nova|shimmer|zira|aria|jenny|emma|ava|isabella|sophie|charlotte|coral|sage|vivian|serena|sohee|ono anna)\b/i;
/** Best-effort gender from a voice or speaker label ("Piper Tom (homme, français)", "Microsoft Paul", "am_adam", "fr-FR-HenriNeural"). */
export function inferGender(name: string): VoiceGender | undefined {
const edge = edgeVoiceGender(name);
if (edge) return edge;
const n = name.replace(/_/g, ' ');
if (/^(am|bm|em|hm|im|jm|pm|zm)\b/i.test(n) || MALE_HINTS.test(n)) return 'male';
if (/^(af|bf|ef|ff|hf|if|jf|pf|zf)\b/i.test(n) || FEMALE_HINTS.test(n)) return 'female';
return undefined;
}
/** OpenAI-compatible default voice for the preferred gender (used when the user left the field empty). */
export function defaultOpenAiVoice(gender: VoiceGender): string {
return gender === 'male' ? 'onyx' : 'nova';
}
/** Edge voice to use: the explicit choice when it matches the gender, otherwise the language default. */
export function resolveEdgeVoice(explicit: string, language: string, gender: VoiceGender): string {
const chosen = explicit.trim();
if (chosen && (edgeVoiceGender(chosen) ?? gender) === gender) return chosen;
return defaultEdgeVoice(language, gender);
}
/** Rank system voices: language first, then gender, then quality hints. */
export function rankSystemVoice(v: { name: string; lang: string; localService?: boolean }, lang: string, gender: VoiceGender): number {
const name = v.name.toLowerCase();
let score = v.lang.toLowerCase().startsWith(lang) ? 100 : 0;
const g = inferGender(v.name);
if (g === gender) score += 50;
else if (g && g !== gender) score -= 30;
if (name.includes('natural') || name.includes('neural') || name.includes('online')) score += 20;
if (name.includes('google')) score += 10;
if (name.includes('microsoft')) score += 5;
if (v.localService) score += 2;
return score;
}
export interface LocalVoiceChoice {
modelId: string;
speaker: number;
}
/** Local models from best to worst French rendering; unknown models come last. */
const LOCAL_QUALITY: Record<string, number> = { 'supertonic-3': 0, 'kokoro-v1': 1, 'piper-fr-upmc': 2, 'piper-fr-tom': 3, 'piper-fr-siwis': 4 };
export function localQualityRank(modelId: string): number {
return LOCAL_QUALITY[modelId] ?? 9;
}
function speakerGender(sp: VoiceSpeaker): VoiceGender | undefined {
return sp.gender === 'm' ? 'male' : sp.gender === 'f' ? 'female' : inferGender(sp.name);
}
function speakerSpeaks(sp: VoiceSpeaker, lang: string): boolean {
const spLang = (sp.lang || '').toLowerCase();
return !spLang || spLang === lang || spLang === 'multi';
}
/** First speaker of a model matching language + gender; falls back to any speaker of the language, then the first one. */
export function pickSpeaker(model: Pick<VoiceModelStatus, 'speakers'>, lang: string, gender: VoiceGender): number {
const l = lang.toLowerCase().split('-')[0];
const speakers = model.speakers ?? [];
const exact = speakers.find((sp) => speakerSpeaks(sp, l) && speakerGender(sp) === gender);
const sameLang = speakers.find((sp) => speakerSpeaks(sp, l));
return (exact ?? sameLang ?? speakers[0])?.id ?? 0;
}
/** Installed local voice matching language + gender (best model first), or null. */
export function findLocalVoice(models: VoiceModelStatus[], lang: string, gender: VoiceGender): LocalVoiceChoice | null {
const l = lang.toLowerCase().split('-')[0];
const installed = models.filter((m) => m.kind === 'tts' && m.installed).sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id));
for (const model of installed) {
for (const sp of model.speakers ?? []) {
if (!speakerSpeaks(sp, l)) continue;
if (speakerGender(sp) === gender) return { modelId: model.id, speaker: sp.id };
}
}
return null;
}
/** Catalog model to download when nothing installed offers the preferred gender (French). */
export function suggestedDownload(models: VoiceModelStatus[], lang: string, gender: VoiceGender): string | null {
const l = lang.toLowerCase().split('-')[0];
if (l !== 'fr') return null;
const wanted = gender === 'male' ? ['supertonic-3', 'piper-fr-tom', 'piper-fr-upmc'] : ['supertonic-3', 'kokoro-v1', 'piper-fr-siwis'];
for (const id of wanted) {
const m = models.find((x) => x.id === id);
if (m && !m.installed) return id;
}
return null;
}
/** Best local model in the catalog for the language that is not installed yet (null when the best is already there). */
export function bestLocalUpgrade(models: VoiceModelStatus[], lang: string): string | null {
const l = lang.toLowerCase().split('-')[0];
const best = models
.filter((m) => m.kind === 'tts' && (m.languages.includes(l) || m.languages.includes('multi')))
.sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id))[0];
return best && !best.installed ? best.id : null;
}
+66 -8
View File
@@ -9,9 +9,12 @@ import { previewText } from '../lib/text';
import { useChat, type PendingRequest } from '../state/chat';
import { useHermes } from '../state/hermes';
import { useSettings } from '../state/settings';
import { executeLocalTool, LOCAL_TOOL_DEFINITIONS } from './hermes/localTools';
import { executeLocalTool, LOCAL_TOOL_DEFINITIONS, type LocalToolContext } from './hermes/localTools';
import { isPriority, summarize } from '../lib/quietHours';
import type { HermesStreamEvent, SendHandle } from './hermes/types';
import { speech } from './voice/speech';
import { bridge } from '../lib/bridge';
import { parseLocalIntent, runLocalIntent } from './localCommands';
let active: SendHandle | null = null;
let activeMessageId: string | null = null;
@@ -46,7 +49,7 @@ export function pingCore(): void {
useChat.getState().ping();
}
function localToolContext() {
export function localToolContext(): LocalToolContext {
const chat = useChat.getState();
const settings = useSettings.getState().settings;
return {
@@ -58,13 +61,20 @@ function localToolContext() {
setTimeout(() => useChat.getState().setHudOverride(null), 6000);
},
speak: (text: string) => speech.say(text),
showMessage: (text: string, title?: string) => {
if (!text.trim()) return;
useChat.getState().addMessage({ role: 'assistant', content: text, source: 'hermes', jobName: title, status: 'done' });
},
getStatus: () => ({
transport: useHermes.getState().transport,
link: useHermes.getState().link,
assistant: settings.assistantName,
messages: chat.messages.length,
speaking: speech.isSpeaking(),
handsFree: settings.voice.handsFree
handsFree: settings.voice.handsFree,
quietHours: chat.quiet,
missionMode: chat.missionMode,
wakeMode: settings.voice.wakeMode
}),
getHistory: (n: number) => chat.messages.slice(-n).map((m) => ({ role: m.role, content: m.content })),
notify: (title: string, body: string) => {
@@ -165,9 +175,49 @@ function addPending(request: Omit<PendingRequest, 'id' | 'createdAt'>): void {
speech.say(request.kind === 'approval' ? 'Autorisation requise.' : request.kind === 'clarify' ? request.description : 'Saisie requise.', { interrupt: false });
}
/** Screenshot of the primary display as a data URL (Electron only). */
export async function captureScreen(): Promise<string | null> {
const api = bridge();
if (!api) return null;
try {
return await api.system.captureScreen(1600);
} catch (err) {
useChat.getState().setError(`Capture d’écran : ${(err as Error).message}`);
return null;
}
}
/** Short system intents handled on the machine without a Hermes round trip. Returns true when consumed. */
async function tryLocalIntent(text: string, source: string): Promise<{ handled: boolean; images?: string[]; text?: string }> {
const settings = useSettings.getState().settings;
if (!settings.voice.localCommands || !bridge()) return { handled: false };
const intent = parseLocalIntent(text);
if (!intent) return { handled: false };
if (intent.kind === 'screenshot') {
const shot = await captureScreen();
if (!shot) return { handled: false };
return { handled: false, images: [shot], text: intent.question || text };
}
const chat = useChat.getState();
chat.addMessage({ role: 'user', content: text, source, status: 'done' });
const result = await runLocalIntent(intent);
chat.addMessage({ role: 'assistant', content: result.message, source: 'local', status: result.ok ? 'done' : 'error' });
chat.setHud(result.ok ? 'idle' : 'error');
if (settings.speech.autoSpeak) speech.say(result.message, { interrupt: true });
return { handled: true };
}
export async function sendMessage(text: string, images: string[] = [], source = 'eveflow'): Promise<void> {
const trimmed = text.trim();
let trimmed = text.trim();
if ((!trimmed && images.length === 0) || active) return;
if (trimmed && images.length === 0) {
const local = await tryLocalIntent(trimmed, source);
if (local.handled) return;
if (local.images) {
images = local.images;
trimmed = (local.text ?? trimmed).trim();
}
}
const chat = useChat.getState();
const hermes = useHermes.getState();
const settings = useSettings.getState().settings;
@@ -189,7 +239,9 @@ export async function sendMessage(text: string, images: string[] = [], source =
const timing = { firstToken: null as number | null, startedAt };
const runIdRef = { id: null as string | null };
const client = hermes.client();
const mission = chat.missionMode && settings.hermes.missionModel.trim();
const client = hermes.client(mission ? settings.hermes.missionModel : undefined);
if (mission) Log.info('hermes', `mission mode → model ${settings.hermes.missionModel}`);
const handle = client.send(
{
text: trimmed || 'Analyse cette image.',
@@ -321,11 +373,17 @@ export function handlePush(event: HermesPushEvent): void {
chat.pushActivity({ kind: 'job', name: event.jobName || 'cron', status: event.status?.includes('fail') ? 'error' : 'done', detail: previewText(event.text, 160) });
void useHermes.getState().refreshJobs();
}
if (event.role !== 'user' && settings.speech.speakIncoming) {
speech.say(event.jobName ? `Résultat de ${event.jobName}. ${event.text}` : event.text);
const n = settings.notifications;
const priority = isPriority(event.text, n.priorityKeywords, event.jobName) || !!event.status?.includes('fail');
const silenced = chat.quiet && !priority;
if (document.hidden || silenced) chat.incUnread();
if (event.role !== 'user' && settings.speech.speakIncoming && !silenced) {
const body = n.summarizeIncoming ? summarize(event.text, n.summarySentences) : event.text;
speech.say(event.jobName ? `Résultat de ${event.jobName}. ${body}` : body);
}
if (silenced) Log.info('webhook', 'quiet hours: push shown silently');
chat.setHud(event.status?.includes('fail') ? 'alert' : 'success');
pingCore();
if (!silenced) pingCore();
setTimeout(() => {
const s = useChat.getState();
if (s.hud === 'success' || s.hud === 'alert') s.setHud(speech.isSpeaking() ? 'speaking' : 'idle');
+121 -4
View File
@@ -31,6 +31,75 @@ const isRec = (v: unknown): v is Rec => !!v && typeof v === 'object' && !Array.i
export type ResolvedTransport = Exclude<HermesTransport, 'auto'>;
/**
* A web page (login portal, dashboard, reverse-proxy error) instead of JSON means the URL does not
* point at the Hermes API. Returns a human explanation, or null when the body is not HTML.
*/
export function describeHtml(body: string): string | null {
const head = body.slice(0, 600).trimStart().toLowerCase();
if (!head.startsWith('<!doctype html') && !head.startsWith('<html') && !/^<\?xml[^>]*>\s*<html/.test(head)) return null;
const title = /<title[^>]*>([^<]{1,120})<\/title>/i.exec(body)?.[1]?.trim();
const login = /connecter|login|sign in|authentif|mot de passe|password/i.test(body.slice(0, 20_000));
return `Le serveur renvoie une page web${title ? ` « ${title} »` : ''} au lieu de l'API Hermes${login ? ' (page de connexion : l’URL passe par un portail web)' : ''}. Utilisez l'URL directe du serveur API Hermes (port 8642 par défaut, ou le chemin /v1 exposé par votre proxy).`;
}
/** Candidate API URLs derived from what the user typed (same host, other port or path). */
export function hermesUrlCandidates(url: string): string[] {
const base = hermesBaseUrl(url);
if (!base) return [];
const out = new Set<string>();
const add = (u: string) => out.add(u.replace(/\/+$/, ''));
try {
const u = new URL(base.includes('://') ? base : `http://${base}`);
const host = u.hostname;
const scheme = u.protocol.replace(':', '');
const path = u.pathname.replace(/\/+$/, '');
if (path) add(`${scheme}://${u.host}`);
for (const suffix of ['/api', '/hermes', '/hermes/api', '/v1', '/api/v1']) add(`${scheme}://${u.host}${path}${suffix}`);
if (!u.port) {
add(`${scheme}://${host}:8642`);
add(`http://${host}:8642`);
add(`https://${host}:8642`);
}
for (const sub of ['api', 'hermes-api']) {
if (!host.startsWith(`${sub}.`) && host.includes('.')) add(`${scheme}://${sub}.${host}`);
}
if (host.startsWith('jarvis.')) add(`${scheme}://hermes.${host.slice('jarvis.'.length)}`);
} catch {
return [];
}
out.delete(base);
return [...out];
}
/** True when this base URL answers like the Hermes API (JSON on /v1/*, or 401/403 = key required). */
export async function probeHermesUrl(config: HermesConfig, url: string): Promise<boolean> {
const client = new HermesClient({ ...config, url });
for (const path of ['/v1/capabilities', '/v1/models', '/health']) {
try {
const payload = await client.request<unknown>(path, { timeoutMs: 4000 });
if (payload && typeof payload === 'object') return true;
} catch (err) {
if (err instanceof HttpError) {
if (err.status === 401 || err.status === 403) return true; // the API is there, only the key is missing
if (/page web/.test(err.message)) return false; // a portal answers on this host: not the API
if (err.status === 404) continue; // older Hermes without this endpoint
}
return false;
}
}
return false;
}
/** Probe candidate URLs until one answers like the Hermes API. */
export async function discoverHermesUrl(config: HermesConfig, onProgress?: (url: string) => void): Promise<string | null> {
for (const candidate of hermesUrlCandidates(config.url)) {
onProgress?.(candidate);
if (await probeHermesUrl(config, candidate)) return candidate;
}
return null;
}
export function hermesBaseUrl(url: string): string {
let base = url.trim().replace(/\/+$/, '');
base = base.replace(/\/chat\/completions$/, '');
@@ -93,7 +162,7 @@ export class HermesClient {
return h;
}
private async request<T>(path: string, init: { method?: 'GET' | 'POST' | 'PATCH' | 'DELETE'; body?: unknown; timeoutMs?: number } = {}): Promise<T> {
async request<T>(path: string, init: { method?: 'GET' | 'POST' | 'PATCH' | 'DELETE'; body?: unknown; timeoutMs?: number } = {}): Promise<T> {
if (!this.config.url.trim()) throw new Error("URL Hermes non configurée.");
const url = `${this.base}${path}`;
const res = await httpFetch({
@@ -103,11 +172,13 @@ export class HermesClient {
body: init.body !== undefined ? JSON.stringify(init.body) : undefined,
timeoutMs: init.timeoutMs ?? 20_000
});
if (!res.ok) throw new HttpError(res.status, errorMessage(res.status, res.text ?? ''), res.text);
const text = res.text ?? '';
const html = describeHtml(text);
if (html) throw new HttpError(res.status, html, text);
if (!res.ok) throw new HttpError(res.status, errorMessage(res.status, res.text ?? ''), res.text);
if (!text.trim()) return null as T;
const parsed = tryParseJson<T>(text);
if (parsed === null) throw new Error(`Réponse Hermes illisible depuis ${path}`);
if (parsed === null) throw new Error(`Réponse Hermes illisible depuis ${path} : ${text.slice(0, 120).replace(/\s+/g, ' ')}`);
return parsed;
}
@@ -430,6 +501,7 @@ export class HermesClient {
let iterationText = '';
const errorChunks: string[] = [];
let ready = false;
let rawBody = '';
const parser = new SseParser((message) => {
if (message.data === '[DONE]') return;
const data = tryParseJson<unknown>(message.data);
@@ -451,7 +523,13 @@ export class HermesClient {
if (sessionId) headers['X-Hermes-Session-Id'] = sessionId;
const handle = await httpStream(
{ url: `${this.base}/v1/chat/completions`, method: 'POST', headers, body: JSON.stringify(payload), timeoutMs: 10 * 60_000 },
{ onChunk: (text) => (ready ? parser.feed(text) : errorChunks.push(text)) }
{
onChunk: (text) => {
if (rawBody.length < 512_000) rawBody += text;
if (ready) parser.feed(text);
else errorChunks.push(text);
}
}
);
currentHandle = handle;
if (!handle.start.ok) {
@@ -467,11 +545,24 @@ export class HermesClient {
throw new HttpError(handle.start.status, detail);
}
ready = true;
// Chunks that raced ahead of the start event were buffered as potential error bodies: replay them.
for (const chunk of errorChunks.splice(0)) parser.feed(chunk);
const rotated = handle.start.headers['x-hermes-session-id'];
if (rotated && rotated !== options.sessionId) onEvent({ kind: 'session', sessionId: rotated });
await handle.done;
parser.end();
if (!iterationText && toolCalls.size === 0 && !aborted()) {
// Nothing streamed: the server may have answered with a plain JSON completion (stream ignored)
// or with a 200 carrying an error object. Surface it instead of an empty bubble.
const recovered = recoverCompletion(rawBody);
if (recovered.text) {
iterationText = recovered.text;
onEvent({ kind: 'delta', text: recovered.text });
} else {
throw new Error(recovered.error ?? `Réponse vide de Hermes (${rawBody.length} octets reçus${rawBody ? ` : ${rawBody.slice(0, 160).replace(/\s+/g, ' ')}` : ''})`);
}
}
fullText += iterationText;
if (finishReason === 'tool_calls' && toolsAllowed && toolCalls.size > 0 && !aborted()) {
@@ -496,6 +587,32 @@ export class HermesClient {
}
}
/** Extract text or an error message from a non-streamed chat completion body. */
export function recoverCompletion(raw: string): { text?: string; error?: string } {
const body = raw.trim();
if (!body) return {};
const html = describeHtml(body);
if (html) return { error: html };
const candidates = body.startsWith('data:') || body.startsWith('event:')
? body.split(/\n+/).filter((l) => l.startsWith('data:')).map((l) => l.replace(/^data:\s*/, '')).filter((l) => l && l !== '[DONE]')
: [body];
let text = '';
for (const c of candidates) {
const data = tryParseJson<Rec>(c);
if (!data || !isRec(data)) continue;
if (data.error) {
const e = data.error as Rec | string;
return { error: `Hermes : ${typeof e === 'string' ? e : String((e as Rec).message ?? JSON.stringify(e))}` };
}
const choice = Array.isArray(data.choices) && isRec(data.choices[0]) ? (data.choices[0] as Rec) : null;
const msg = choice && isRec(choice.message) ? (choice.message as Rec) : choice && isRec(choice.delta) ? (choice.delta as Rec) : null;
if (msg && typeof msg.content === 'string') text += msg.content;
else if (typeof data.output === 'string') text += data.output;
else if (typeof data.text === 'string') text += data.text;
}
return text ? { text } : {};
}
/** Session ids created by the sessions transport carry an `hs:` prefix; other endpoints get the bare id. */
export function plainSession(id: string): string {
return id.startsWith('hs:') ? id.slice(3) : id;
+42
View File
@@ -3,16 +3,35 @@
* With the runs/sessions transports Hermes uses its own server-side toolsets instead.
*/
import { bridge } from '../../lib/bridge';
import type { SystemAction } from '../../../shared/bridge';
export interface LocalToolContext {
setEmotion: (emotion: string) => void;
showMessage?: (text: string, title?: string) => void;
speak: (text: string) => void;
getStatus: () => Record<string, unknown>;
getHistory: (n: number) => Array<{ role: string; content: string }>;
notify: (title: string, body: string) => void;
}
const fn = (name: string, description: string, properties: Record<string, unknown> = {}, required: string[] = []) => ({
type: 'function',
function: { name, description, parameters: { type: 'object', properties, required } }
});
/** Tools executed on the user's PC through the Electron main process (allow-listed). */
export const SYSTEM_TOOL_DEFINITIONS = [
fn('lock_session', "Verrouille la session de l'utilisateur."),
fn('open_app', "Lance une application du PC de l'utilisateur (bloc-notes, calculatrice, chrome, spotify, vscode, terminal, explorateur…).", { name: { type: 'string' } }, ['name']),
fn('open_url', "Ouvre une URL http(s) dans le navigateur de l'utilisateur.", { url: { type: 'string' } }, ['url']),
fn('media_key', 'Touche média : volume-up, volume-down, mute, play-pause, next, previous.', { key: { type: 'string', enum: ['volume-up', 'volume-down', 'mute', 'play-pause', 'next', 'previous'] } }, ['key']),
fn('clipboard_get', 'Lit le texte du presse-papiers.'),
fn('clipboard_set', 'Place un texte dans le presse-papiers.', { text: { type: 'string' } }, ['text']),
fn('find_files', "Cherche des fichiers par nom dans Documents, Bureau, Téléchargements et Images (25 max).", { query: { type: 'string' } }, ['query'])
];
export const LOCAL_TOOL_DEFINITIONS = [
...SYSTEM_TOOL_DEFINITIONS,
{
type: 'function',
function: {
@@ -95,6 +114,29 @@ export async function executeLocalTool(name: string, rawArgs: string, ctx: Local
case 'notify_user':
ctx.notify(String(args.title ?? 'EveFlow'), String(args.body ?? ''));
return JSON.stringify({ ok: true });
case 'show_message':
ctx.showMessage?.(String(args.text ?? ''), typeof args.title === 'string' ? args.title : undefined);
return JSON.stringify({ ok: true });
case 'lock_session':
case 'open_app':
case 'open_url':
case 'media_key':
case 'clipboard_get':
case 'clipboard_set':
case 'find_files': {
const api = bridge();
if (!api) return JSON.stringify({ error: 'system actions unavailable outside Electron' });
const action: SystemAction =
name === 'lock_session' ? { type: 'lock' }
: name === 'open_app' ? { type: 'open-app', name: String(args.name ?? '') }
: name === 'open_url' ? { type: 'open-url', url: String(args.url ?? '') }
: name === 'media_key' ? { type: 'media', key: String(args.key ?? '') as 'mute' }
: name === 'clipboard_get' ? { type: 'clipboard-read' }
: name === 'clipboard_set' ? { type: 'clipboard-write', text: String(args.text ?? '') }
: { type: 'find-files', query: String(args.query ?? '') };
const result = await api.system.action(action);
return JSON.stringify(result.ok ? { ok: true, message: result.message, data: result.data } : { error: result.message ?? 'failed' });
}
case 'write_shared_file': {
const api = bridge();
if (!api) return JSON.stringify({ error: 'file system unavailable outside Electron' });
+2
View File
@@ -10,6 +10,8 @@ export interface HermesConfig {
reasoningEffort: '' | 'low' | 'medium' | 'high';
/** Extra instructions layered on top of the Hermes system prompt. */
instructions: string;
/** Model used in "mission" mode (long tasks); empty = same as `model`. */
missionModel: string;
/** Expose EveFlow client tools in chat-completions mode. */
localTools: boolean;
}
+75
View File
@@ -0,0 +1,75 @@
/**
* Instant local commands: short French/English intents handled on the machine without a
* round trip to Hermes (lock, open app/url, volume, media, screenshot to Hermes).
* Anything unmatched goes to Hermes as usual.
*/
import type { SystemAction } from '../../shared/bridge';
import { bridge } from '../lib/bridge';
import { Log } from '../lib/log';
export interface LocalIntent {
kind: 'action' | 'screenshot';
action?: SystemAction;
/** Spoken confirmation. */
reply: string;
/** For screenshot: the question to send to Hermes with the image. */
question?: string;
}
const norm = (t: string) =>
t
.normalize('NFD')
.replace(/[̀-ͯ]/g, '')
.toLowerCase()
.replace(/[’']/g, ' ')
.replace(/[^a-z0-9 :/._-]+/g, ' ')
.replace(/\s+/g, ' ')
.trim();
const LOCK = /^(verrouille|verrouiller|bloque|lock)( (la |ma )?(session|le pc|l ordinateur|the (pc|computer|screen)))?$/;
const VOL_UP = /^(monte|augmente|hausse|plus fort|volume plus|turn up|raise)( (le|the)? ?(son|volume))?( de \d+)?$/;
const VOL_DOWN = /^(baisse|diminue|moins fort|volume moins|turn down|lower)( (le|the)? ?(son|volume))?( de \d+)?$/;
const MUTE = /^(coupe|couper|mute|silence|desactive)( (le|the)? ?(son|volume|audio))?$/;
const PLAY = /^(pause|play|lecture|reprends|reprendre|mets en pause|met en pause|stop la musique|arrete la musique)$/;
const NEXT = /^(suivant|suivante|piste suivante|musique suivante|next|skip)$/;
const PREV = /^(precedent|precedente|piste precedente|previous)$/;
const OPEN_APP = /^(ouvre|ouvrir|lance|lancer|demarre|open|launch|start) ((l application|l appli|le logiciel|le programme|the app|le|la|les|l|un|une|the|moi) )*(.+)$/;
const OPEN_URL = /^(ouvre|ouvrir|va sur|open|go to) (https?:\/\/\S+|(www\.)?[a-z0-9-]+\.[a-z]{2,}(\/\S*)?)$/;
const SCREEN = /(regarde|regardes|analyse|decris|decrit|lis|explique|qu est ce qu il y a sur|que vois tu sur|what is on|look at|read) (mon |l |the )?(ecran|screen)|capture (d )?ecran|screenshot/;
export function parseLocalIntent(text: string): LocalIntent | null {
const t = norm(text).replace(/^(jarvis|hey jarvis|ok jarvis)[ ,]*/, '');
if (!t) return null;
if (LOCK.test(t)) return { kind: 'action', action: { type: 'lock' }, reply: 'Session verrouillée.' };
if (MUTE.test(t)) return { kind: 'action', action: { type: 'media', key: 'mute' }, reply: 'Son coupé.' };
if (VOL_UP.test(t)) return { kind: 'action', action: { type: 'media', key: 'volume-up' }, reply: 'Volume augmenté.' };
if (VOL_DOWN.test(t)) return { kind: 'action', action: { type: 'media', key: 'volume-down' }, reply: 'Volume baissé.' };
if (PLAY.test(t)) return { kind: 'action', action: { type: 'media', key: 'play-pause' }, reply: 'Lecture.' };
if (NEXT.test(t)) return { kind: 'action', action: { type: 'media', key: 'next' }, reply: 'Piste suivante.' };
if (PREV.test(t)) return { kind: 'action', action: { type: 'media', key: 'previous' }, reply: 'Piste précédente.' };
if (SCREEN.test(t)) return { kind: 'screenshot', reply: 'Je regarde votre écran.', question: text.trim() };
const url = OPEN_URL.exec(t);
if (url) {
const target = url[2].startsWith('http') ? url[2] : `https://${url[2]}`;
return { kind: 'action', action: { type: 'open-url', url: target }, reply: `J’ouvre ${url[2]}.` };
}
const app = OPEN_APP.exec(t);
if (app) {
const name = app[app.length - 1].trim();
if (name.length <= 40 && !/\s(et|puis|and)\s/.test(name)) return { kind: 'action', action: { type: 'open-app', name }, reply: `J’ouvre ${name}.` };
}
return null;
}
/** Execute an action intent through the bridge. Returns the spoken outcome. */
export async function runLocalIntent(intent: LocalIntent): Promise<{ ok: boolean; message: string }> {
const api = bridge();
if (!api || !intent.action) return { ok: false, message: 'Actions locales indisponibles hors Electron.' };
try {
const result = await api.system.action(intent.action);
Log.info('local', `${intent.action.type}: ${result.ok ? 'ok' : result.message}`);
return { ok: result.ok, message: result.ok ? intent.reply : result.message || 'Échec de la commande locale.' };
} catch (err) {
return { ok: false, message: (err as Error).message };
}
}
+42
View File
@@ -0,0 +1,42 @@
/**
* Answers tool calls that Hermes sends through the local MCP endpoint and that need the renderer
* (voice, notifications, HUD, chat). System tools are executed in the main process directly.
*/
import type { McpToolRequest } from '../../shared/ipc';
import { bridge } from '../lib/bridge';
import { Log } from '../lib/log';
import { useChat } from '../state/chat';
import { executeLocalTool, type LocalToolContext } from './hermes/localTools';
let unsubscribe: (() => void) | null = null;
export function initMcpBridge(context: () => LocalToolContext): void {
const api = bridge();
if (!api || unsubscribe) return;
unsubscribe = api.hermes.onToolRequest((req: McpToolRequest) => {
void handle(req, context)
.then((result) => api.hermes.toolResponse({ id: req.id, ok: true, result }))
.catch((err: Error) => api.hermes.toolResponse({ id: req.id, ok: false, error: err.message }));
});
}
async function handle(req: McpToolRequest, context: () => LocalToolContext): Promise<unknown> {
Log.info('mcp', `tool ${req.name}`);
const chat = useChat.getState();
if (req.name === 'show_message') {
const text = String(req.args.text ?? '');
if (!text.trim()) throw new Error('texte vide');
chat.addMessage({ role: 'assistant', content: text, source: 'mcp', jobName: typeof req.args.title === 'string' ? req.args.title : undefined, status: 'done' });
chat.pushActivity({ kind: 'system', name: 'hermes → eveflow', status: 'done', detail: text.slice(0, 160) });
return { ok: true };
}
const raw = await executeLocalTool(req.name, JSON.stringify(req.args ?? {}), context());
try {
const parsed = JSON.parse(raw) as { error?: string };
if (parsed && typeof parsed === 'object' && parsed.error) throw new Error(parsed.error);
return parsed;
} catch (err) {
if ((err as Error).message && !(err instanceof SyntaxError)) throw err;
return raw;
}
}
+21 -4
View File
@@ -49,6 +49,8 @@ export interface CaptureCallbacks {
export interface CaptureOptions {
deviceId?: string;
/** Chromium's echo cancellation / noise suppression / auto gain (default on). Off keeps the raw signal for the recogniser. */
micProcessing?: boolean;
mode: CaptureMode;
vad?: Partial<VadOptions>;
callbacks: CaptureCallbacks;
@@ -66,6 +68,21 @@ function getWorkletUrl(): string {
return workletUrl;
}
/**
* Load an AudioWorklet module: first the static file shipped with the app (allowed by the strict
* CSP, script-src 'self'), then a blob URL for dev servers that do not serve /worklets.
*/
export async function loadWorklet(ctx: AudioContext, name: string, blobUrl: () => string): Promise<void> {
const staticUrl = new URL(`worklets/${name}.js`, document.baseURI).href;
try {
await ctx.audioWorklet.addModule(staticUrl);
return;
} catch (err) {
Log.debug('audio', `static worklet ${name} unavailable (${(err as Error).message}), trying blob`);
}
await ctx.audioWorklet.addModule(blobUrl());
}
export async function listMicrophones(): Promise<MicDevice[]> {
try {
const devices = await navigator.mediaDevices.enumerateDevices();
@@ -109,9 +126,9 @@ export class MicCapture {
const constraints: MediaStreamConstraints = {
audio: {
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
echoCancellation: options.micProcessing !== false,
noiseSuppression: options.micProcessing !== false,
autoGainControl: options.micProcessing !== false,
channelCount: 1
}
};
@@ -147,7 +164,7 @@ export class MicCapture {
audioBus.setInputAnalyser(analyser);
try {
await this.ctx.audioWorklet.addModule(getWorkletUrl());
await loadWorklet(this.ctx, 'eveflow-capture', getWorkletUrl);
this.worklet = new AudioWorkletNode(this.ctx, 'eveflow-capture', { numberOfInputs: 1, numberOfOutputs: 0, channelCount: 1 });
this.worklet.port.onmessage = (event: MessageEvent<Float32Array>) => this.onSamples(event.data);
source.connect(this.worklet);
+92 -9
View File
@@ -1,15 +1,17 @@
/**
* Text-to-speech engine with a sentence queue, bounded prefetching and Web Audio playback so
* the HUD reacts to the actual waveform. Providers: in-app sherpa-onnx (Kokoro / Piper),
* OpenAI-compatible /v1/audio/speech, system voices, and the legacy Google Translate endpoint.
* the HUD reacts to the actual waveform. Providers: Microsoft Edge neural voices (online, free),
* in-app sherpa-onnx (Supertonic / Kokoro / Piper), OpenAI-compatible /v1/audio/speech, system
* voices, and the legacy Google Translate endpoint (one feminine voice per language).
*/
import { Log } from '../../lib/log';
import { defaultOpenAiVoice, rankSystemVoice, resolveEdgeVoice } from '../../lib/voicePreference';
import { bridge } from '../../lib/bridge';
import { chunkForSpeech, cleanForSpeech, extractSentences } from '../../lib/text';
import { httpFetch } from '../../lib/transport';
import { audioBus } from './audioBus';
export type TtsProvider = 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
export type TtsProvider = 'edge' | 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
export interface TtsConfig {
provider: TtsProvider;
@@ -26,6 +28,12 @@ export interface TtsConfig {
localModel: string;
/** Speaker id inside the local model. */
localSpeaker: number;
/** Microsoft Edge voice short name (provider 'edge'); empty = automatic from language + gender. */
edgeVoice?: string;
/** Preferred voice gender, applied to every provider's default voice. */
voiceGender?: 'male' | 'female';
/** 'jarvis' adds a subtle AI timbre: slightly lower pitch, warm/crisp EQ, short room reverb. */
timbre?: 'natural' | 'jarvis';
}
export type TtsState = 'idle' | 'loading' | 'speaking';
@@ -174,8 +182,9 @@ export class TtsEngine {
}
private prefetch(text: string, gen: number): Promise<AudioBuffer | null> {
const provider = this.config.provider;
const task =
this.config.provider === 'google-free' ? this.fetchGoogle(text) : this.config.provider === 'local' ? this.fetchLocal(text) : this.fetchOpenAi(text);
provider === 'google-free' ? this.fetchGoogle(text) : provider === 'local' ? this.fetchLocal(text) : provider === 'edge' ? this.fetchEdge(text) : this.fetchOpenAi(text);
return task
.then(async (bytes) => {
if (gen !== this.generation || !bytes) return null;
@@ -201,7 +210,7 @@ export class TtsEngine {
body: JSON.stringify({
model: model || 'tts-1',
input: text,
voice: voice || 'alloy',
voice: voice || defaultOpenAiVoice(this.config.voiceGender ?? 'male'),
speed: Math.max(0.5, Math.min(2, speed || 1)),
response_format: format || 'mp3'
}),
@@ -223,12 +232,23 @@ export class TtsEngine {
modelId: this.config.localModel,
text,
speaker: this.config.localSpeaker,
speed: Math.max(0.5, Math.min(2, this.config.speed || 1))
speed: Math.max(0.5, Math.min(2, this.config.speed || 1)),
language: this.config.language || 'fr-FR'
});
Log.debug('tts', `local synthesis ${result.durationMs} ms for ${result.audioSec.toFixed(1)}s of audio`);
return result.wav;
}
/** Microsoft Edge neural voices through the main process (no key, online). */
private async fetchEdge(text: string): Promise<Uint8Array> {
const api = bridge();
if (!api) throw new Error('Les voix Edge nécessitent l’application Electron.');
const voice = resolveEdgeVoice(this.config.edgeVoice ?? '', this.config.language || 'fr-FR', this.config.voiceGender ?? 'male');
const result = await api.voice.edgeSynthesize({ text, voice, speed: Math.max(0.5, Math.min(2, this.config.speed || 1)) });
Log.debug('tts', `edge synthesis (${voice}) ${result.durationMs} ms, ${result.mp3.byteLength} bytes`);
return result.mp3;
}
private async fetchGoogle(text: string): Promise<Uint8Array> {
const url = `https://translate.google.com/translate_tts?ie=UTF-8&tl=${encodeURIComponent(this.config.language.split('-')[0] || 'fr')}&client=tw-ob&q=${encodeURIComponent(text.slice(0, 200))}`;
const res = await httpFetch({
@@ -252,7 +272,13 @@ export class TtsEngine {
const source = ctx.createBufferSource();
source.buffer = buffer;
if (this.config.provider === 'google-free') source.playbackRate.value = Math.max(0.5, Math.min(2, this.config.speed || 1));
source.connect(audioBus.output);
if (this.config.timbre === 'jarvis') {
// Slightly lower and calmer, warm low end, crisp presence, a touch of room: the film's intercom feel.
source.playbackRate.value *= 0.94;
source.connect(jarvisChain(ctx));
} else {
source.connect(audioBus.output);
}
const finish = () => {
if (this.cancelCurrent === cancel) this.cancelCurrent = null;
resolve();
@@ -312,8 +338,8 @@ export class TtsEngine {
if (exact) return exact;
}
const lang = (this.config.language || 'fr').toLowerCase().split('-')[0];
const candidates = voices.filter((v) => v.lang.toLowerCase().startsWith(lang));
return candidates.sort((a, b) => scoreVoice(b) - scoreVoice(a))[0] ?? voices[0];
const gender = this.config.voiceGender ?? 'male';
return [...voices].sort((a, b) => rankSystemVoice(b, lang, gender) - rankSystemVoice(a, lang, gender))[0];
}
}
@@ -331,3 +357,60 @@ export function listSystemVoices(): SpeechSynthesisVoice[] {
if (typeof speechSynthesis === 'undefined') return [];
return [...speechSynthesis.getVoices()].sort((a, b) => a.lang.localeCompare(b.lang) || scoreVoice(b) - scoreVoice(a));
}
let jarvisInput: AudioNode | null = null;
let jarvisCtx: AudioContext | null = null;
/** Shared effect chain for the JARVIS timbre (built once per AudioContext). */
function jarvisChain(ctx: AudioContext): AudioNode {
if (jarvisInput && jarvisCtx === ctx) return jarvisInput;
const input = ctx.createGain();
const warmth = ctx.createBiquadFilter();
warmth.type = 'lowshelf';
warmth.frequency.value = 180;
warmth.gain.value = 3.5;
const presence = ctx.createBiquadFilter();
presence.type = 'peaking';
presence.frequency.value = 3200;
presence.Q.value = 0.9;
presence.gain.value = 2.5;
const air = ctx.createBiquadFilter();
air.type = 'highshelf';
air.frequency.value = 7000;
air.gain.value = -2;
const comp = ctx.createDynamicsCompressor();
comp.threshold.value = -20;
comp.ratio.value = 3;
comp.attack.value = 0.005;
comp.release.value = 0.12;
const dry = ctx.createGain();
dry.gain.value = 0.86;
const wet = ctx.createGain();
wet.gain.value = 0.14;
const reverb = ctx.createConvolver();
reverb.buffer = impulse(ctx, 0.22, 3.2);
input.connect(warmth);
warmth.connect(presence);
presence.connect(air);
air.connect(comp);
comp.connect(dry);
comp.connect(reverb);
reverb.connect(wet);
dry.connect(audioBus.output);
wet.connect(audioBus.output);
jarvisInput = input;
jarvisCtx = ctx;
return input;
}
/** Short synthetic room impulse (exponentially decaying noise). */
function impulse(ctx: AudioContext, seconds: number, decay: number): AudioBuffer {
const rate = ctx.sampleRate;
const length = Math.max(1, Math.floor(rate * seconds));
const buffer = ctx.createBuffer(2, length, rate);
for (let ch = 0; ch < 2; ch++) {
const data = buffer.getChannelData(ch);
for (let i = 0; i < length; i++) data[i] = (Math.random() * 2 - 1) * Math.pow(1 - i / length, decay);
}
return buffer;
}
+130 -4
View File
@@ -2,6 +2,7 @@
* Voice controller: microphone → VAD → STT → conversation, plus hands-free loop and barge-in.
*/
import { Log } from '../../lib/log';
import { isTranscriptNoise } from '../../lib/text';
import { useChat } from '../../state/chat';
import { useSettings } from '../../state/settings';
import { useVoice } from '../../state/voice';
@@ -10,7 +11,9 @@ import { audioBus } from './audioBus';
import { listMicrophones, MicCapture } from './capture';
import { speech } from './speech';
import { BrowserRecognizer, transcribeWav } from './stt';
import { WakeListener } from './wakeListener';
import type { WavResult } from './wav';
import { bridge } from '../../lib/bridge';
const SENSITIVITY_RATIO: Record<number, number> = { 1: 4.5, 2: 3.4, 3: 2.6, 4: 2.0, 5: 1.6 };
const SENSITIVITY_MIN_RMS: Record<number, number> = { 1: 0.03, 2: 0.02, 3: 0.012, 4: 0.008, 5: 0.005 };
@@ -46,6 +49,8 @@ export function matchWakeWord(text: string, wakeWord: string): { matched: boolea
class VoiceController {
private capture = new MicCapture();
private wake = new WakeListener();
private unsubscribeKws: (() => void) | null = null;
/** After a bare wake word, the next utterance is accepted without the wake word. */
private attentionUntil = 0;
private browser = new BrowserRecognizer();
@@ -74,17 +79,123 @@ class VoiceController {
dispose(): void {
this.unsubscribeVoice?.();
this.stop();
void this.stopWakeMode();
}
get isListening(): boolean {
return useVoice.getState().phase !== 'off';
}
get wakeActive(): boolean {
return this.wake.isActive;
}
toggle(): void {
if (this.wake.isActive) {
// Always-on mode: the button starts/cancels a command capture on the shared microphone.
if (this.wake.currentPhase === 'command' || this.wake.currentPhase === 'speech') this.wake.cancelCommand();
else this.onWakeDetected('manuel');
return;
}
if (this.isListening) this.stop();
else void this.start();
}
/** Always-on keyword spotting: the microphone stays open and the main process spots the wake word. */
async startWakeMode(): Promise<void> {
const api = bridge();
const voice = useVoice.getState();
if (!api || this.wake.isActive) return;
const settings = useSettings.getState().settings.voice;
voice.setWake('starting');
try {
const phrases = [settings.wakeWord || 'jarvis', `hey ${settings.wakeWord || 'jarvis'}`];
const result = await api.voice.kwsStart({ modelId: 'kws-en', keywords: phrases, sensitivity: settings.kwsSensitivity });
this.unsubscribeKws?.();
this.unsubscribeKws = api.voice.onKwsDetected((d) => this.onWakeDetected(d.keyword));
const sensitivity = Math.min(5, Math.max(1, Math.round(settings.sensitivity))) as 1 | 2 | 3 | 4 | 5;
let neuralVad = false;
if (settings.neuralVad) {
try {
const installed = (await api.voice.listModels()).some((m) => m.id === 'silero-vad' && m.installed);
if (installed) {
await api.voice.vadStart({ modelId: 'silero-vad', silenceMs: Math.max(300, Math.min(1500, settings.silenceMs - 200)), threshold: 0.5, maxUtteranceSec: 25 });
neuralVad = true;
}
} catch (err) {
Log.warn('voice', `silero unavailable, energy VAD fallback: ${(err as Error).message}`);
}
}
await this.wake.start({
deviceId: settings.micDeviceId || undefined,
micProcessing: settings.micProcessing ?? true,
neuralVad,
vad: { silenceMs: settings.silenceMs, speechRatio: SENSITIVITY_RATIO[sensitivity], minRms: SENSITIVITY_MIN_RMS[sensitivity] },
callbacks: {
onPhase: (phase) => {
const v = useVoice.getState();
if (phase === 'spotting') {
v.setPhase('off');
const chat = useChat.getState();
if (chat.hud === 'listening') chat.setHud(chat.isSending ? 'thinking' : 'idle');
} else if (phase === 'command') {
v.setPhase('listening');
useChat.getState().setHud('listening');
} else if (phase === 'speech') {
v.setPhase('speech');
useChat.getState().ping();
}
},
onLevel: (level) => {
const q = Math.round(level * 20) / 20;
if (q !== useVoice.getState().inputLevel) useVoice.getState().setInputLevel(q);
},
onUtterance: (wav) => void this.transcribe(wav, true),
onNoSpeech: () => useVoice.getState().setInputLevel(0),
onError: (message) => useVoice.getState().setError(message)
}
});
voice.setWake('spotting', result.accepted.map((k) => k.replace(/_/g, ' ')));
voice.setNeuralVad(neuralVad);
Log.info('voice', `wake mode on: ${result.accepted.join(', ')}${result.rejected.length ? ` (rejetés : ${result.rejected.join(', ')})` : ''}${neuralVad ? ' · Silero' : ''}`);
} catch (err) {
const message = (err as Error).message;
voice.setWake('error');
voice.setError(message);
useChat.getState().setError(`Mot d’activation : ${message}`);
Log.error('voice', `wake mode failed: ${message}`);
await api.voice.kwsStop().catch(() => undefined);
await api.voice.vadStop().catch(() => undefined);
}
}
async stopWakeMode(): Promise<void> {
this.unsubscribeKws?.();
this.unsubscribeKws = null;
this.wake.stop();
useVoice.getState().setWake('off');
useVoice.getState().setNeuralVad(false);
const api = bridge();
await api?.voice.kwsStop().catch(() => undefined);
await api?.voice.vadStop().catch(() => undefined);
}
/** Restart spotting with the current settings (wake word, sensitivity, microphone). */
async restartWakeMode(): Promise<void> {
await this.stopWakeMode();
if (useSettings.getState().settings.voice.wakeMode === 'kws') await this.startWakeMode();
}
private onWakeDetected(keyword: string): void {
if (!this.wake.isActive || this.wake.currentPhase !== 'spotting') return;
if (useChat.getState().isSending) return;
Log.info('voice', `wake: ${keyword}`);
speech.stop();
useChat.getState().ping();
this.playChime(true);
this.wake.beginCommand();
}
setHandsFree(on: boolean): void {
useVoice.getState().setHandsFree(on);
useSettings.getState().update({ voice: { handsFree: on } });
@@ -113,6 +224,11 @@ class VoiceController {
async start(auto = false): Promise<void> {
const voice = useVoice.getState();
if (voice.phase !== 'off') return;
if (this.wake.isActive) {
// The always-on listener owns the microphone: open a command capture on it instead.
if (!auto) this.onWakeDetected('manuel');
return;
}
const settings = useSettings.getState().settings.voice;
const seq = ++this.startSeq;
voice.setError(null);
@@ -128,15 +244,18 @@ class VoiceController {
}
const sensitivity = Math.min(5, Math.max(1, Math.round(settings.sensitivity))) as 1 | 2 | 3 | 4 | 5;
// Barge-in: while the assistant talks, demand a clearly louder signal so the speaker echo cannot trigger.
const barging = auto && settings.bargeIn && speech.isSpeaking();
try {
await audioBus.resume();
await this.capture.start({
deviceId: settings.micDeviceId || undefined,
micProcessing: settings.micProcessing ?? true,
mode: settings.captureMode,
vad: {
silenceMs: settings.silenceMs,
speechRatio: SENSITIVITY_RATIO[sensitivity],
minRms: SENSITIVITY_MIN_RMS[sensitivity],
speechRatio: barging ? Math.min(0.9, SENSITIVITY_RATIO[sensitivity] + 0.15) : SENSITIVITY_RATIO[sensitivity],
minRms: barging ? SENSITIVITY_MIN_RMS[sensitivity] * 2.5 : SENSITIVITY_MIN_RMS[sensitivity],
noSpeechTimeoutMs: useVoice.getState().handsFree ? Number.POSITIVE_INFINITY : 9_000
},
callbacks: {
@@ -145,6 +264,8 @@ class VoiceController {
if (q !== useVoice.getState().inputLevel) useVoice.getState().setInputLevel(q);
},
onSpeechStart: () => {
// The user started talking over the assistant: cut the voice (barge-in).
if (speech.isSpeaking()) speech.stop();
useVoice.getState().setPhase('speech');
useChat.getState().ping();
},
@@ -218,7 +339,7 @@ class VoiceController {
}
}
private async transcribe(wav: WavResult): Promise<void> {
private async transcribe(wav: WavResult, fromWake = false): Promise<void> {
const voice = useVoice.getState();
voice.setPhase('transcribing');
useChat.getState().setHud('thinking');
@@ -226,10 +347,15 @@ class VoiceController {
const settings = useSettings.getState().settings.voice;
try {
let text = await transcribeWav(wav.bytes, settings);
if (isTranscriptNoise(text)) {
Log.debug('voice', `transcript ignored as noise: ${text}`);
text = '';
}
voice.setTranscript(text);
voice.setPhase('off');
Log.info('voice', `transcript (${wav.durationSec.toFixed(1)}s): ${text}`);
if (voice.handsFree && settings.wakeWordEnabled && Date.now() > this.attentionUntil) {
const filterByTranscript = settings.wakeMode === 'transcript' || (settings.wakeMode === 'off' && settings.wakeWordEnabled);
if (!fromWake && voice.handsFree && filterByTranscript && Date.now() > this.attentionUntil) {
const { matched, rest } = matchWakeWord(text, settings.wakeWord || 'jarvis');
if (!matched) {
Log.debug('voice', 'utterance ignored: no wake word');
+68
View File
@@ -0,0 +1,68 @@
/**
* Applies the preferred voice gender to the active TTS provider: switches the local model/speaker,
* downloads a matching French voice when none is installed, resets an Edge voice of the other
* gender, and explains providers that cannot honour the choice (Google Translate).
*/
import { Log } from '../../lib/log';
import { bestLocalUpgrade, findLocalVoice, resolveEdgeVoice, suggestedDownload } from '../../lib/voicePreference';
import { useSettings } from '../../state/settings';
import { useVoiceModels } from '../../state/voiceModels';
let downloading: string | null = null;
const genderLabel = (g: 'male' | 'female') => (g === 'male' ? 'masculine' : 'féminine');
/**
* Make the active provider match `speech.voiceGender`. Returns a short status message.
* With `upgrade`, the best local model for the language is downloaded even if a lesser voice exists.
*/
export async function ensurePreferredVoice(options: { upgrade?: boolean } = {}): Promise<string> {
const { settings, update } = useSettings.getState();
const { speech } = settings;
const gender = speech.voiceGender ?? 'male';
const lang = speech.language || 'fr-FR';
if (speech.provider === 'edge') {
const voice = resolveEdgeVoice(speech.edgeVoice ?? '', lang, gender);
if (voice !== (speech.edgeVoice ?? '').trim() && speech.edgeVoice) update({ speech: { edgeVoice: '' } });
return `Voix ${genderLabel(gender)} : ${voice.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') ?? voice} (Edge)`;
}
if (speech.provider === 'google-free') {
return gender === 'male'
? 'Google Translate n’a qu’une voix féminine par langue : choisissez « Microsoft Edge » ou une voix locale pour une voix masculine.'
: '';
}
if (speech.provider !== 'local') return '';
const models = useVoiceModels.getState().models;
if (!models.length) return '';
const current = models.find((m) => m.id === speech.localModel);
const currentSpeaker = current?.speakers?.find((s) => s.id === speech.localSpeaker);
const currentGender = currentSpeaker ? (currentSpeaker.gender === 'm' ? 'male' : currentSpeaker.gender === 'f' ? 'female' : undefined) : undefined;
const upgrade = options.upgrade ? bestLocalUpgrade(models, lang) : null;
if (current?.installed && currentGender === gender && !upgrade) return '';
const choice = upgrade ? null : findLocalVoice(models, lang, gender);
if (choice) {
update({ speech: { localModel: choice.modelId, localSpeaker: choice.speaker } });
const name = models.find((m) => m.id === choice.modelId)?.speakers?.find((s) => s.id === choice.speaker)?.name ?? choice.modelId;
Log.info('tts', `voice preference ${gender}: ${name}`);
return `Voix ${genderLabel(gender)} : ${name}`;
}
const download = upgrade ?? suggestedDownload(models, lang, gender);
if (!download || downloading === download) return download ? 'Téléchargement de la voix en cours…' : '';
downloading = download;
Log.info('tts', `${upgrade ? 'upgrading local voice' : `no ${gender} voice installed`}, downloading ${download}`);
try {
await useVoiceModels.getState().download(download);
const after = findLocalVoice(useVoiceModels.getState().models, lang, gender);
if (after) update({ speech: { localModel: after.modelId, localSpeaker: after.speaker } });
const name = after ? useVoiceModels.getState().models.find((m) => m.id === after.modelId)?.name : undefined;
return after ? `Voix ${genderLabel(gender)} installée : ${name ?? after.modelId}.` : 'Voix téléchargée.';
} catch (err) {
Log.warn('tts', `voice download failed: ${(err as Error).message}`);
return `Téléchargement impossible : ${(err as Error).message}`;
} finally {
downloading = null;
}
}
+276
View File
@@ -0,0 +1,276 @@
/**
* Always-on microphone for the keyword spotter. Streams 16 kHz PCM to the main process
* (which feeds the sherpa-onnx spotter); once the wake word is detected it captures the
* following utterance with the energy VAD and hands back a WAV, then resumes spotting.
* One microphone stream, no re-opening between phrases.
*/
import { bridge } from '../../lib/bridge';
import { Log } from '../../lib/log';
import { audioBus } from './audioBus';
import { loadWorklet } from './capture';
import { DEFAULT_VAD, EnergyVad, type VadOptions } from './vad';
import { buildWav16k, rms, type WavResult } from './wav';
import type { VadEvent } from '../../../shared/voice';
const WORKLET_SOURCE = `
class EveFlowWakeProcessor extends AudioWorkletProcessor {
constructor() { super(); this.buffer = new Float32Array(2048); this.offset = 0; }
process(inputs) {
const channel = inputs[0] && inputs[0][0];
if (!channel) return true;
let i = 0;
while (i < channel.length) {
const n = Math.min(channel.length - i, this.buffer.length - this.offset);
this.buffer.set(channel.subarray(i, i + n), this.offset);
this.offset += n; i += n;
if (this.offset === this.buffer.length) {
this.port.postMessage(this.buffer, [this.buffer.buffer]);
this.buffer = new Float32Array(2048); this.offset = 0;
}
}
return true;
}
}
registerProcessor('eveflow-wake', EveFlowWakeProcessor);
`;
export type WakePhase = 'off' | 'spotting' | 'command' | 'speech';
export interface WakeCallbacks {
onPhase: (phase: WakePhase) => void;
onLevel?: (level: number) => void;
onUtterance: (wav: WavResult) => void;
onNoSpeech: () => void;
onError: (message: string) => void;
}
let workletUrl: string | null = null;
export class WakeListener {
private ctx: AudioContext | null = null;
private stream: MediaStream | null = null;
private node: AudioWorkletNode | null = null;
private phase: WakePhase = 'off';
private vad: EnergyVad | null = null;
private chunks: Float32Array[] = [];
private totalSamples = 0;
private sampleRate = 16_000;
private pending: Float32Array[] = [];
private pendingSamples = 0;
private callbacks: WakeCallbacks | null = null;
private vadOptions: Partial<VadOptions> = {};
private noSpeechTimer: ReturnType<typeof setTimeout> | null = null;
/** Neural end-of-speech (Silero in the main process) instead of the energy VAD. */
private neural = false;
private unsubscribeVad: (() => void) | null = null;
private vadPending: Float32Array[] = [];
private vadPendingSamples = 0;
get isActive(): boolean {
return this.phase !== 'off';
}
get currentPhase(): WakePhase {
return this.phase;
}
async start(options: { deviceId?: string; micProcessing?: boolean; vad?: Partial<VadOptions>; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise<void> {
if (this.phase !== 'off') return;
this.callbacks = options.callbacks;
this.vadOptions = options.vad ?? {};
this.neural = !!options.neuralVad;
if (this.neural) {
this.unsubscribeVad?.();
this.unsubscribeVad = bridge()?.voice.onVadEvent((event) => this.onVadEvent(event)) ?? null;
}
this.stream = await navigator.mediaDevices.getUserMedia({
audio: {
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
echoCancellation: options.micProcessing !== false,
noiseSuppression: options.micProcessing !== false,
autoGainControl: options.micProcessing !== false,
channelCount: 1
}
});
try {
this.ctx = new AudioContext({ sampleRate: 16_000, latencyHint: 'playback' });
} catch {
this.ctx = new AudioContext({ latencyHint: 'playback' });
}
if (this.ctx.state === 'suspended') await this.ctx.resume().catch(() => undefined);
this.sampleRate = this.ctx.sampleRate;
const source = this.ctx.createMediaStreamSource(this.stream);
const analyser = this.ctx.createAnalyser();
analyser.fftSize = 512;
source.connect(analyser);
audioBus.setInputAnalyser(analyser);
await loadWorklet(this.ctx, 'eveflow-wake', () => {
if (!workletUrl) workletUrl = URL.createObjectURL(new Blob([WORKLET_SOURCE], { type: 'application/javascript' }));
return workletUrl;
});
this.node = new AudioWorkletNode(this.ctx, 'eveflow-wake', { numberOfInputs: 1, numberOfOutputs: 0, channelCount: 1 });
this.node.port.onmessage = (event: MessageEvent<Float32Array>) => this.onSamples(event.data);
source.connect(this.node);
this.setPhase('spotting');
Log.info('wake', `listener started (${this.sampleRate} Hz, fin de phrase ${this.neural ? 'Silero' : 'énergie'})`);
}
private onVadEvent(event: VadEvent): void {
if (!this.neural || (this.phase !== 'command' && this.phase !== 'speech')) return;
if (event.type === 'speech-start') {
if (this.phase === 'command') {
this.setPhase('speech');
if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer);
this.noSpeechTimer = null;
}
} else if (event.type === 'segment') {
const wav: WavResult = { bytes: event.wav, sampleRate: 16_000, durationSec: event.durationSec };
this.resumeSpotting();
if (wav.durationSec >= 0.25) this.callbacks?.onUtterance(wav);
else this.callbacks?.onNoSpeech();
} else if (event.type === 'error') {
this.callbacks?.onError(event.message);
this.resumeSpotting();
}
}
/** Called by the controller when the main process reports the wake word (or on manual trigger). */
beginCommand(): void {
if (this.phase === 'off' || this.phase === 'command' || this.phase === 'speech') return;
this.vad = this.neural ? null : new EnergyVad({ ...DEFAULT_VAD, ...this.vadOptions, noSpeechTimeoutMs: Number.POSITIVE_INFINITY });
this.vadPending = [];
this.vadPendingSamples = 0;
this.chunks = [];
this.totalSamples = 0;
this.setPhase('command');
if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer);
this.noSpeechTimer = setTimeout(() => {
if (this.phase === 'command') {
this.resumeSpotting();
this.callbacks?.onNoSpeech();
}
}, 8000);
}
/** Abort a command capture and go back to spotting. */
cancelCommand(): void {
if (this.phase === 'command' || this.phase === 'speech') this.resumeSpotting();
}
stop(): void {
if (this.phase === 'off') return;
if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer);
this.noSpeechTimer = null;
this.unsubscribeVad?.();
this.unsubscribeVad = null;
audioBus.setInputAnalyser(null);
if (this.node) {
this.node.port.onmessage = null;
this.node.disconnect();
this.node = null;
}
this.stream?.getTracks().forEach((t) => t.stop());
this.stream = null;
void this.ctx?.close().catch(() => undefined);
this.ctx = null;
this.chunks = [];
this.pending = [];
this.pendingSamples = 0;
this.vadPending = [];
this.vadPendingSamples = 0;
this.setPhase('off');
Log.info('wake', 'listener stopped');
}
private setPhase(phase: WakePhase): void {
this.phase = phase;
this.callbacks?.onPhase(phase);
}
private resumeSpotting(): void {
if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer);
this.noSpeechTimer = null;
this.vad = null;
this.chunks = [];
this.totalSamples = 0;
this.setPhase('spotting');
}
private onSamples(samples: Float32Array): void {
if (this.phase === 'off') return;
const level = rms(samples);
this.callbacks?.onLevel?.(Math.min(1, level * 6));
if (this.phase === 'spotting') {
// Batch ~256 ms of audio per IPC message for the keyword spotter.
this.pending.push(samples);
this.pendingSamples += samples.length;
if (this.pendingSamples >= this.sampleRate * 0.25) this.flushToSpotter();
return;
}
if (this.neural) {
// command / speech with Silero: stream ~128 ms frames to the main process, which returns the segment.
this.vadPending.push(samples);
this.vadPendingSamples += samples.length;
if (this.vadPendingSamples >= this.sampleRate * 0.128) this.flushToVad();
return;
}
// command / speech: collect the utterance
this.chunks.push(samples);
this.totalSamples += samples.length;
if (this.vad && this.phase === 'command') {
// keep only a short pre-roll before speech starts
while (this.chunks.length > 1 && this.totalSamples - this.chunks[0].length > this.sampleRate * 0.4) {
this.totalSamples -= this.chunks.shift()!.length;
}
}
if (!this.vad) return;
const signal = this.vad.feed(level, (samples.length / this.sampleRate) * 1000);
if (signal === 'speech-start') {
this.setPhase('speech');
if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer);
this.noSpeechTimer = null;
} else if (signal === 'speech-end' || signal === 'max-length') {
const wav = buildWav16k(this.chunks, this.sampleRate);
this.resumeSpotting();
if (wav.durationSec >= 0.25) this.callbacks?.onUtterance(wav);
else this.callbacks?.onNoSpeech();
} else if (signal === 'too-short') {
this.resumeSpotting();
this.callbacks?.onNoSpeech();
}
}
private static toInt16(chunks: Float32Array[], total: number): Uint8Array {
const merged = new Int16Array(total);
let offset = 0;
for (const chunk of chunks) {
for (let i = 0; i < chunk.length; i++) {
const s = Math.max(-1, Math.min(1, chunk[i]));
merged[offset + i] = s < 0 ? s * 0x8000 : s * 0x7fff;
}
offset += chunk.length;
}
return new Uint8Array(merged.buffer);
}
private flushToSpotter(): void {
const api = bridge();
if (!api) return;
const bytes = WakeListener.toInt16(this.pending, this.pendingSamples);
this.pending = [];
this.pendingSamples = 0;
api.voice.kwsAudio(bytes, this.sampleRate);
}
private flushToVad(): void {
const api = bridge();
if (!api) return;
const bytes = WakeListener.toInt16(this.vadPending, this.vadPendingSamples);
this.vadPending = [];
this.vadPendingSamples = 0;
api.voice.vadAudio(bytes, this.sampleRate);
}
}
+17
View File
@@ -57,6 +57,16 @@ interface ChatStore {
pingCount: number;
/** Time to first token of the current reply, in ms. */
latencyMs: number | null;
/** Pushes received while the window was hidden / compact, cleared when the user looks. */
unread: number;
/** Mission mode routes the next messages to the "mission" model (long tasks). */
missionMode: boolean;
/** True while quiet hours apply (computed by App). */
quiet: boolean;
incUnread: () => void;
markRead: () => void;
setMissionMode: (on: boolean) => void;
setQuiet: (quiet: boolean) => void;
addMessage: (message: Omit<ChatMessage, 'id' | 'timestamp'> & Partial<Pick<ChatMessage, 'id' | 'timestamp'>>) => string;
updateMessage: (id: string, patch: Partial<ChatMessage> | ((m: ChatMessage) => Partial<ChatMessage>)) => void;
@@ -91,6 +101,13 @@ export const useChat = create<ChatStore>((set, get) => ({
draft: '',
pingCount: 0,
latencyMs: null,
unread: 0,
missionMode: false,
quiet: false,
incUnread: () => set((s) => ({ unread: Math.min(99, s.unread + 1) })),
markRead: () => set((s) => (s.unread ? { unread: 0 } : {})),
setMissionMode: (missionMode) => set({ missionMode }),
setQuiet: (quiet) => set((s) => (s.quiet === quiet ? {} : { quiet })),
addMessage: (message) => {
const id = message.id ?? uid('msg');
+61 -10
View File
@@ -2,7 +2,7 @@ import { create } from 'zustand';
import type { WebhookStatus } from '../../shared/ipc';
import { Log } from '../lib/log';
import { persistGet, persistSet } from '../lib/persist';
import { HermesClient, jobOutput, jobStatus, resolveTransport, type ResolvedTransport } from '../services/hermes/client';
import { HermesClient, jobOutput, jobStatus, resolveTransport, type ResolvedTransport, discoverHermesUrl } from '../services/hermes/client';
import type {
HermesCapabilities,
HermesHealth,
@@ -38,12 +38,16 @@ interface HermesStore {
sessions: HermesSession[];
jobs: HermesJob[];
jobRuns: JobRun[];
/** Last cron sync failure (the chat link stays independent of it). */
jobsError: string | null;
/** True while probing alternative API URLs after a failed connection. */
discovering: boolean;
transport: ResolvedTransport;
lastSyncAt: number | null;
webhook: WebhookStatus | null;
busy: boolean;
client: () => HermesClient;
client: (modelOverride?: string) => HermesClient;
connect: () => Promise<void>;
refreshJobs: () => Promise<void>;
refreshSessions: () => Promise<void>;
@@ -88,15 +92,17 @@ export const useHermes = create<HermesStore>((set, get) => ({
sessions: [],
jobs: [],
jobRuns: [],
jobsError: null,
discovering: false,
transport: 'completions',
lastSyncAt: null,
webhook: null,
busy: false,
client: () => {
client: (modelOverride) => {
const config = useSettings.getState().settings.hermes;
// Without an explicit model, use the alias advertised by /v1/models (Hermes rejects unknown names).
const model = config.model.trim() || get().models[0]?.id || '';
const model = (modelOverride ?? '').trim() || config.model.trim() || get().models[0]?.id || '';
return new HermesClient({ ...config, model });
},
@@ -116,17 +122,35 @@ export const useHermes = create<HermesStore>((set, get) => ({
try {
capabilities = await client.capabilities();
} catch (err) {
// A web page on /v1/* means the whole API is behind a portal even if /health passed through.
if (/page web/.test((err as Error).message)) throw err;
Log.warn('hermes', `capabilities unavailable: ${(err as Error).message}`);
}
const transport = resolveTransport(config, capabilities);
const degraded = String(health.status ?? 'ok').toLowerCase() !== 'ok';
set({ health, capabilities, transport, link: degraded ? 'degraded' : 'online', linkDetail: degraded ? `état ${health.status}` : '' });
const degraded = !isHealthyStatus(health.status);
set({ health, capabilities, transport, link: degraded ? 'degraded' : 'online', linkDetail: degraded ? describeHealth(health) : '' });
Log.info('hermes', `connected (${transport})`, { status: health.status, model: capabilities?.model });
void get().refreshCatalog();
void get().refreshJobs();
void get().refreshSessions();
} catch (err) {
const message = (err as Error).message;
// The URL answers with a web page (portal, dashboard) or nothing: look for the API on the same host.
if (!get().discovering && /page web|illisible|fetch failed|ECONNREFUSED|404/i.test(message)) {
set({ discovering: true, linkDetail: 'recherche de l’API Hermes…' });
try {
const found = await discoverHermesUrl(config);
if (found) {
Log.info('hermes', `API found at ${found} (was ${config.url})`);
useSettings.getState().update({ hermes: { url: found } });
set({ discovering: false, linkDetail: `URL corrigée automatiquement : ${found}` });
await get().connect();
return;
}
} finally {
set({ discovering: false });
}
}
set({ link: 'offline', linkDetail: message, transport: resolveTransport(config, null) });
Log.warn('hermes', `connection failed: ${message}`);
}
@@ -165,7 +189,7 @@ export const useHermes = create<HermesStore>((set, get) => ({
const merged = [...incoming, ...previous.filter((r) => !incoming.some((i) => i.id === r.id))]
.sort((a, b) => new Date(b.at).getTime() - new Date(a.at).getTime())
.slice(0, 100);
set({ jobs, jobRuns: merged, lastSyncAt: Date.now(), link: get().link === 'offline' || get().link === 'degraded' ? 'online' : get().link });
set({ jobs, jobRuns: merged, jobsError: null, lastSyncAt: Date.now(), link: get().link === 'offline' ? 'online' : get().link });
const snapshot = JSON.stringify({ jobs, jobRuns: merged });
if (snapshot !== lastCacheSnapshot) {
lastCacheSnapshot = snapshot;
@@ -174,9 +198,8 @@ export const useHermes = create<HermesStore>((set, get) => ({
} catch (err) {
const message = (err as Error).message;
Log.warn('hermes', `jobs sync failed: ${message}`);
if (/HTTP 404/.test(message)) return; // jobs API disabled on this server
// A failing jobs poll does not mean chat is down: degrade, and let the next health probe decide.
set({ linkDetail: message, link: get().link === 'online' ? 'degraded' : get().link });
// The cron API can be absent or restricted on a given Hermes: the chat link is unaffected.
set({ jobsError: /HTTP 404/.test(message) ? 'API des crons absente sur ce serveur Hermes' : message });
}
},
@@ -233,3 +256,31 @@ export const useHermes = create<HermesStore>((set, get) => ({
}));
export { jobStatus };
const HEALTHY = new Set(['ok', 'healthy', 'up', 'alive', 'pass', 'ready', 'running', 'online', 'true']);
/** Hermes /health reports "ok"; /health/detailed may report "healthy", "degraded" or "unhealthy". */
export function isHealthyStatus(status: unknown): boolean {
if (status === undefined || status === null || status === '') return true;
return HEALTHY.has(String(status).toLowerCase());
}
/** Human summary of a degraded health payload: failing checks by name. */
export function describeHealth(health: Record<string, unknown>): string {
const failing: string[] = [];
const visit = (node: unknown, prefix: string, depth: number) => {
if (!node || typeof node !== 'object' || depth > 3) return;
for (const [key, value] of Object.entries(node as Record<string, unknown>)) {
if (key === 'status') continue;
if (value && typeof value === 'object') {
const rec = value as Record<string, unknown>;
const st = rec.status ?? rec.ok ?? rec.healthy;
if (st !== undefined && !isHealthyStatus(st)) failing.push(prefix + key);
else visit(value, `${prefix}${key}.`, depth + 1);
} else if (typeof value === 'boolean' && !value && /ok|healthy|ready|connected|available/i.test(key)) failing.push(prefix + key);
}
};
visit(health, '', 0);
const base = `état ${String(health.status)}`;
return failing.length ? `${base} · ${failing.slice(0, 4).join(', ')}` : base;
}
+40 -4
View File
@@ -19,6 +19,15 @@ export interface VoiceSettings extends SttConfig {
/** Hands-free: only react to utterances starting with this word (local STT recommended). */
wakeWordEnabled: boolean;
wakeWord: string;
/** off = push-to-talk / hands-free; transcript = filter after transcription; kws = always-on keyword spotting. */
wakeMode: 'off' | 'transcript' | 'kws';
kwsSensitivity: number; // 1..5
/** Use Silero VAD for end-of-speech in always-on mode when the model is installed. */
neuralVad: boolean;
/** Execute short system intents locally (lock, volume, open app…) instead of asking Hermes. */
localCommands: boolean;
/** Chromium mic processing (echo cancellation, noise suppression, auto gain). Off often transcribes better on a headset. */
micProcessing: boolean;
}
export interface SpeechSettings extends TtsConfig {
@@ -32,6 +41,20 @@ export interface WebhookSettings {
secret: string;
}
export interface NotificationSettings {
/** Quiet hours: no spoken pushes, no chime, dimmed HUD (24 h "HH:MM"). */
quietEnabled: boolean;
quietStart: string;
quietEnd: string;
/** Pushes whose text or job name contains one of these words are spoken even during quiet hours. */
priorityKeywords: string;
/** Speak only the first sentences of incoming pushes (cron reports can be long). */
summarizeIncoming: boolean;
summarySentences: number;
/** Dim the HUD (night theme) during quiet hours. */
nightTheme: boolean;
}
export interface Settings {
version: 2;
assistantName: string;
@@ -42,6 +65,7 @@ export interface Settings {
voice: VoiceSettings;
speech: SpeechSettings;
webhook: WebhookSettings;
notifications: NotificationSettings;
ui: {
showTelemetry: boolean;
showReasoning: boolean;
@@ -66,7 +90,8 @@ export const DEFAULT_SETTINGS: Settings = {
reasoningEffort: '',
instructions:
"Tu es l'interface vocale EveFlow (style JARVIS). Réponds en français, de façon concise et orale quand la question est simple; utilise le Markdown uniquement pour le contenu structuré (code, listes, tableaux). Les images doivent être des URL http(s) ou des fichiers du dossier partagé.",
localTools: true
localTools: true,
missionModel: ''
},
voice: {
provider: 'openai-compatible',
@@ -83,10 +108,16 @@ export const DEFAULT_SETTINGS: Settings = {
wakeChime: true,
wakeWordEnabled: false,
wakeWord: 'jarvis',
wakeMode: 'off',
kwsSensitivity: 3,
neuralVad: true,
localCommands: true,
micProcessing: true,
localModel: 'whisper-base'
},
speech: {
provider: 'openai-compatible',
// Edge neural voices speak out of the box (no server, no key) with a real masculine/feminine choice.
provider: 'edge',
apiUrl: 'http://127.0.0.1:8000/v1',
apiKey: '',
model: 'tts-1',
@@ -96,12 +127,16 @@ export const DEFAULT_SETTINGS: Settings = {
systemVoice: '',
language: 'fr-FR',
volume: 1,
localModel: 'kokoro-v1',
localSpeaker: 30,
localModel: 'supertonic-3',
localSpeaker: 6,
edgeVoice: '',
voiceGender: 'male',
timbre: 'jarvis',
autoSpeak: true,
speakIncoming: true
},
webhook: { enabled: true, port: 7842, secret: '' },
notifications: { quietEnabled: false, quietStart: '22:30', quietEnd: '07:30', priorityKeywords: 'urgent, alerte, alarme, panne', summarizeIncoming: true, summarySentences: 2, nightTheme: true },
ui: { showTelemetry: true, showReasoning: false, reduceMotion: false, compactOpacity: 0.92 },
hermesSessionId: ''
};
@@ -181,6 +216,7 @@ export const useSettings = create<SettingsStore>((set, get) => ({
}
const merged = merge(DEFAULT_SETTINGS, saved);
if (!merged.hermesSessionId) merged.hermesSessionId = uid('eveflow');
if (saved && saved.voice && saved.voice.wakeMode === undefined && saved.voice.wakeWordEnabled) merged.voice.wakeMode = 'transcript';
set({ settings: merged, loaded: true });
persistSet(STORAGE_KEY, merged);
},
+11
View File
@@ -2,6 +2,7 @@ import { create } from 'zustand';
import type { TtsState } from '../services/voice/tts';
export type ListenPhase = 'off' | 'arming' | 'listening' | 'speech' | 'transcribing';
export type WakeState = 'off' | 'starting' | 'spotting' | 'error';
interface VoiceStore {
phase: ListenPhase;
@@ -12,6 +13,11 @@ interface VoiceStore {
interim: string;
error: string | null;
micDevices: Array<{ deviceId: string; label: string }>;
wake: WakeState;
wakeKeywords: string[];
neuralVad: boolean;
setNeuralVad: (on: boolean) => void;
setWake: (state: WakeState, keywords?: string[]) => void;
setPhase: (phase: ListenPhase) => void;
setInputLevel: (level: number) => void;
setTts: (state: TtsState) => void;
@@ -31,6 +37,11 @@ export const useVoice = create<VoiceStore>((set) => ({
interim: '',
error: null,
micDevices: [],
wake: 'off',
wakeKeywords: [],
neuralVad: false,
setNeuralVad: (neuralVad) => set({ neuralVad }),
setWake: (wake, wakeKeywords) => set(wakeKeywords ? { wake, wakeKeywords } : { wake }),
setPhase: (phase) => set({ phase }),
setInputLevel: (inputLevel) => set({ inputLevel }),
setTts: (tts) => set({ tts }),
+51
View File
@@ -67,3 +67,54 @@
font-size: 13px;
padding: 8px 11px;
}
/* Glanceable strip: state, unread badge, last sentence. */
.compact-glance {
display: grid;
grid-template-columns: auto 1fr auto;
align-items: center;
gap: 8px;
padding: 4px 12px 8px;
font-family: var(--font-mono);
font-size: 10.5px;
letter-spacing: 0.12em;
text-transform: uppercase;
color: var(--ink-2);
}
.compact-glance .state {
color: var(--accent);
}
.compact-glance .state.alert {
color: var(--danger);
}
.compact-glance .last {
font-family: var(--font-body);
font-size: 12.5px;
letter-spacing: 0;
text-transform: none;
color: var(--ink-1);
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
.compact-glance .badges {
display: flex;
gap: 6px;
align-items: center;
}
.compact-glance .unread {
min-width: 18px;
height: 18px;
padding: 0 5px;
border-radius: 9px;
background: var(--accent);
color: #03101c;
font-weight: 700;
font-size: 11px;
display: inline-flex;
align-items: center;
justify-content: center;
}
.compact-glance .mission {
color: var(--warn, #f5c451);
}
+32
View File
@@ -115,3 +115,35 @@
text-align: left;
font: inherit;
}
/* Segmented choice (voice gender, capture mode…) */
.segmented {
display: inline-flex;
border: 1px solid var(--line);
border-radius: 10px;
overflow: hidden;
background: rgba(255, 255, 255, 0.03);
}
.segmented button {
padding: 8px 16px;
background: transparent;
border: none;
color: var(--ink-1);
font: inherit;
font-size: 13px;
letter-spacing: 0.02em;
cursor: pointer;
transition: background 0.15s, color 0.15s;
}
.segmented button + button {
border-left: 1px solid var(--line);
}
.segmented button:hover {
color: var(--ink-0);
background: rgba(var(--accent-rgb), 0.08);
}
.segmented button.active {
color: #03101c;
background: var(--accent);
font-weight: 600;
}
+12
View File
@@ -159,3 +159,15 @@ textarea {
.kv dd {
user-select: text;
}
/* Quiet hours: dimmer HUD, softer glow. */
:root[data-night] {
--accent-glow: rgba(var(--accent-rgb), 0.18);
}
:root[data-night] .hud-root,
:root[data-night] .compact-root {
filter: brightness(0.72) saturate(0.85);
}
:root[data-night] .core-stage canvas {
opacity: 0.8;
}
+90
View File
@@ -0,0 +1,90 @@
import { describe, expect, it } from 'vitest';
import {
defaultEdgeVoice,
edgeConfigMessage,
edgeRate,
edgeSsml,
edgeSsmlMessage,
edgeTextFramePath,
edgeTokenInput,
edgeVoiceGender,
escapeXml,
parseEdgeBinaryFrame
} from '../shared/edgeTts';
describe('edge token input', () => {
it('uses Windows file time rounded down to 5 minutes followed by the client token', () => {
// 2026-09-04T15:23:47Z → 15:20:00Z = 1788535200 s since 1970 → +11644473600 = 13433008800 s → ×1e7 ticks.
const nowMs = Date.UTC(2026, 8, 4, 15, 23, 47);
expect(edgeTokenInput(nowMs)).toBe('1343300880000000006A5AA1D4EAFF4E9FB37E23D68491D6F4');
});
it('is stable inside a 5-minute window and applies the clock skew', () => {
const a = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1));
const b = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 24, 59));
expect(a).toBe(b);
expect(edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1), 300)).not.toBe(a);
});
});
describe('edge ssml', () => {
it('escapes text and derives the language from the voice', () => {
const ssml = edgeSsml('Tom & Jerry <3 "ok"', 'fr-FR-HenriNeural', 1.25);
expect(ssml).toContain("xml:lang='fr-FR'");
expect(ssml).toContain("<voice name='fr-FR-HenriNeural'>");
expect(ssml).toContain("rate='+25%'");
expect(ssml).toContain('Tom &amp; Jerry &lt;3 &quot;ok&quot;');
expect(escapeXml("l'été")).toBe('l&apos;été');
});
it('formats rates with a sign', () => {
expect(edgeRate(1)).toBe('+0%');
expect(edgeRate(0.8)).toBe('-20%');
expect(edgeRate(3)).toBe('+100%');
});
it('builds the config and ssml messages with the expected headers', () => {
const date = new Date(Date.UTC(2026, 8, 4, 15, 23, 47));
const config = edgeConfigMessage(date);
expect(config.startsWith('X-Timestamp:Fri, 04 Sep 2026 15:23:47 GMT+0000 (Coordinated Universal Time)\r\n')).toBe(true);
expect(config).toContain('Path:speech.config\r\n\r\n{');
expect(config).toContain('audio-24khz-48kbitrate-mono-mp3');
const ssml = edgeSsmlMessage('abc123', '<speak/>', date);
expect(ssml).toContain('X-RequestId:abc123\r\n');
expect(ssml).toContain('Content-Type:application/ssml+xml\r\n');
expect(ssml.endsWith('Path:ssml\r\n\r\n<speak/>')).toBe(true);
expect(edgeTextFramePath('X-RequestId:1\r\nContent-Type:application/json\r\nPath:turn.end\r\n\r\n{}')).toBe('turn.end');
});
});
describe('edge binary frames', () => {
it('splits the header (2-byte big-endian length) from the audio payload', () => {
const header = 'X-RequestId:1\r\nContent-Type:audio/mpeg\r\nX-StreamId:2\r\nPath:audio\r\n';
const payload = new Uint8Array([0xff, 0xfb, 0x90, 0x00]);
const frame = new Uint8Array(2 + header.length + payload.length);
frame[0] = header.length >> 8;
frame[1] = header.length & 0xff;
for (let i = 0; i < header.length; i++) frame[2 + i] = header.charCodeAt(i);
frame.set(payload, 2 + header.length);
const parsed = parseEdgeBinaryFrame(frame);
expect(parsed.path).toBe('audio');
expect([...parsed.payload]).toEqual([...payload]);
});
it('tolerates truncated frames', () => {
expect(parseEdgeBinaryFrame(new Uint8Array([0x00])).payload.byteLength).toBe(0);
expect(parseEdgeBinaryFrame(new Uint8Array([0x10, 0x00, 0x41])).path).toBe('');
});
});
describe('edge voices', () => {
it('picks Henri / Denise for French and falls back to French for unknown languages', () => {
expect(defaultEdgeVoice('fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(defaultEdgeVoice('fr', 'female')).toBe('fr-FR-DeniseNeural');
expect(defaultEdgeVoice('en-GB', 'male')).toBe('en-US-AndrewMultilingualNeural');
expect(defaultEdgeVoice('xx', 'female')).toBe('fr-FR-DeniseNeural');
});
it('knows the gender of the common French voices', () => {
expect(edgeVoiceGender('fr-FR-HenriNeural')).toBe('male');
expect(edgeVoiceGender('fr-FR-RemyMultilingualNeural')).toBe('male');
expect(edgeVoiceGender('fr-FR-VivienneMultilingualNeural')).toBe('female');
expect(edgeVoiceGender('fr-CA-SylvieNeural')).toBe('female');
expect(edgeVoiceGender('zz-ZZ-NobodyNeural')).toBeUndefined();
});
});
+33
View File
@@ -0,0 +1,33 @@
import { describe, expect, it } from 'vitest';
import { describeHtml, hermesUrlCandidates, recoverCompletion } from '../src/services/hermes/client';
describe('describeHtml', () => {
it('explains a login page instead of the API', () => {
const msg = describeHtml('<!doctype html><html lang="fr-FR"><head><title>Jarvis – Se connecter</title></head><body>Mot de passe</body></html>');
expect(msg).toContain('Jarvis – Se connecter');
expect(msg).toContain('page de connexion');
expect(msg).toContain('8642');
});
it('ignores JSON and SSE', () => {
expect(describeHtml('{"status":"ok"}')).toBeNull();
expect(describeHtml('data: {"choices":[]}')).toBeNull();
});
it('is used by recoverCompletion', () => {
expect(recoverCompletion('<html><head><title>Portal</title></head></html>').error).toContain('Portal');
});
});
describe('hermesUrlCandidates', () => {
it('tries the API port, common paths and sibling hosts', () => {
const c = hermesUrlCandidates('http://jarvis.vonrodbox.eu');
expect(c).toContain('http://jarvis.vonrodbox.eu:8642');
expect(c).toContain('http://jarvis.vonrodbox.eu/api');
expect(c).toContain('http://api.jarvis.vonrodbox.eu');
expect(c).toContain('http://hermes.vonrodbox.eu');
expect(c).not.toContain('http://jarvis.vonrodbox.eu');
});
it('keeps an explicit port and handles garbage', () => {
expect(hermesUrlCandidates('http://10.0.0.5:8642').some((u) => u.includes(':8642:'))).toBe(false);
expect(hermesUrlCandidates('')).toEqual([]);
});
});
+27
View File
@@ -0,0 +1,27 @@
import { describe, expect, it } from 'vitest';
import { recoverCompletion } from '../src/services/hermes/client';
import { describeHealth, isHealthyStatus } from '../src/state/hermes';
describe('health status', () => {
it('accepts the usual healthy words', () => {
for (const s of ['ok', 'OK', 'healthy', 'ready', undefined]) expect(isHealthyStatus(s)).toBe(true);
for (const s of ['degraded', 'unhealthy', 'error']) expect(isHealthyStatus(s)).toBe(false);
});
it('names the failing checks', () => {
expect(describeHealth({ status: 'degraded', checks: { memory: { status: 'ok' }, sessions_db: { status: 'error' } } })).toBe('état degraded · checks.sessions_db');
expect(describeHealth({ status: 'degraded' })).toBe('état degraded');
});
});
describe('recoverCompletion', () => {
it('reads a plain JSON completion when the server ignored streaming', () => {
expect(recoverCompletion(JSON.stringify({ choices: [{ message: { role: 'assistant', content: 'Bonjour.' } }] }))).toEqual({ text: 'Bonjour.' });
});
it('surfaces an error object returned with HTTP 200', () => {
expect(recoverCompletion(JSON.stringify({ error: { message: 'model not found' } })).error).toContain('model not found');
});
it('handles SSE bodies and empty input', () => {
expect(recoverCompletion('data: {"choices":[{"delta":{"content":"A"}}]}\n\ndata: {"choices":[{"delta":{"content":"B"}}]}\n\ndata: [DONE]\n\n')).toEqual({ text: 'AB' });
expect(recoverCompletion('')).toEqual({});
});
});
+25
View File
@@ -0,0 +1,25 @@
import { describe, expect, it } from 'vitest';
import { buildKeywordsFile, encodeKeyword, normalizeKeyword, parseTokens } from '../shared/keywords';
const vocab = parseTokens(['<blk> 0', '▁ 3', '▁JA 4', 'R 5', 'VI 6', 'S 7', '▁HE 8', 'Y 9', '▁NO 10', 'V 11', 'A 12', '▁MA 13', 'X 14', '▁T 15', 'ON 16', 'Y 17'].join('\n'));
describe('keywords', () => {
it('normalises accents, case and punctuation', () => {
expect(normalizeKeyword(' Hé, Jarvis ! ')).toBe('HE JARVIS');
});
it('uses the known SentencePiece encodings', () => {
expect(encodeKeyword('jarvis', vocab)).toBe('▁JA R VI S');
expect(encodeKeyword('Hey Jarvis', vocab)).toBe('▁HE Y ▁JA R VI S');
});
it('falls back to greedy longest match', () => {
expect(encodeKeyword('max', vocab)).toBe('▁MA X');
expect(encodeKeyword('tony', vocab)).toBe('▁T ON Y');
expect(encodeKeyword('zzz', vocab)).toBeNull();
});
it('builds a keywords file with labels', () => {
const file = buildKeywordsFile(['jarvis', 'hey jarvis', 'zzz'], vocab);
expect(file.content).toBe('▁JA R VI S @jarvis\n▁HE Y ▁JA R VI S @hey_jarvis\n');
expect(file.accepted).toEqual(['jarvis', 'hey_jarvis']);
expect(file.rejected).toEqual(['zzz']);
});
});
+24
View File
@@ -0,0 +1,24 @@
import { describe, expect, it } from 'vitest';
import { parseLocalIntent } from '../src/services/localCommands';
describe('parseLocalIntent', () => {
it('recognises system actions in French and English', () => {
expect(parseLocalIntent('Jarvis, verrouille la session')?.action).toEqual({ type: 'lock' });
expect(parseLocalIntent('monte le son')?.action).toEqual({ type: 'media', key: 'volume-up' });
expect(parseLocalIntent('Baisse le volume')?.action).toEqual({ type: 'media', key: 'volume-down' });
expect(parseLocalIntent('coupe le son')?.action).toEqual({ type: 'media', key: 'mute' });
expect(parseLocalIntent('piste suivante')?.action).toEqual({ type: 'media', key: 'next' });
expect(parseLocalIntent('ouvre le bloc-notes')?.action).toEqual({ type: 'open-app', name: 'bloc-notes' });
expect(parseLocalIntent('open spotify')?.action).toEqual({ type: 'open-app', name: 'spotify' });
expect(parseLocalIntent('ouvre github.com')?.action).toEqual({ type: 'open-url', url: 'https://github.com' });
});
it('detects screen questions', () => {
expect(parseLocalIntent('Jarvis, regarde mon écran et dis-moi ce que tu vois')?.kind).toBe('screenshot');
expect(parseLocalIntent('fais une capture d’écran')?.kind).toBe('screenshot');
});
it('leaves everything else to Hermes', () => {
expect(parseLocalIntent('Quelle est la météo à Paris demain ?')).toBeNull();
expect(parseLocalIntent('ouvre le fichier puis envoie-le à Marc')).toBeNull();
expect(parseLocalIntent('')).toBeNull();
});
});
+37
View File
@@ -0,0 +1,37 @@
import { describe, expect, it } from 'vitest';
import { isPriority, isQuietTime, summarize } from '../src/lib/quietHours';
const at = (h: number, m = 0) => new Date(2026, 0, 1, h, m);
describe('isQuietTime', () => {
it('handles ranges crossing midnight', () => {
expect(isQuietTime('22:30', '07:30', at(23))).toBe(true);
expect(isQuietTime('22:30', '07:30', at(3, 15))).toBe(true);
expect(isQuietTime('22:30', '07:30', at(7, 30))).toBe(false);
expect(isQuietTime('22:30', '07:30', at(12))).toBe(false);
});
it('handles same-day ranges and invalid input', () => {
expect(isQuietTime('13:00', '14:00', at(13, 30))).toBe(true);
expect(isQuietTime('13:00', '14:00', at(14))).toBe(false);
expect(isQuietTime('bad', '14:00', at(13))).toBe(false);
expect(isQuietTime('10:00', '10:00', at(10))).toBe(false);
});
});
describe('isPriority', () => {
it('matches keywords in text or job name', () => {
expect(isPriority('Serveur en panne depuis 5 min', 'urgent, panne')).toBe(true);
expect(isPriority('Rapport quotidien', 'urgent', 'Alerte disque')).toBe(false);
expect(isPriority('Rapport quotidien', 'alerte', 'Alerte disque')).toBe(true);
expect(isPriority('x', '')).toBe(false);
});
});
describe('summarize', () => {
it('keeps the first sentences and strips markdown', () => {
const text = '## Rapport\n\nTrois annonces **majeures** aujourd’hui. Le marché monte de 2 %. Détails ci-dessous :\n\n```\ncode\n```\n- point 1';
expect(summarize(text, 2)).toBe('Rapport Trois annonces majeures aujourd’hui. Le marché monte de 2 %.');
expect(summarize('Une seule phrase', 2)).toBe('Une seule phrase');
expect(summarize('', 2)).toBe('');
});
});
+16 -1
View File
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { chunkForSpeech, cleanForSpeech, extractSentences, preprocessMedia } from '../src/lib/text';
import { chunkForSpeech, cleanForSpeech, extractSentences, isTranscriptNoise, preprocessMedia } from '../src/lib/text';
describe('text', () => {
it('cleans markdown for speech', () => {
@@ -25,3 +25,18 @@ describe('text', () => {
expect(preprocessMedia('MEDIA:/tmp/a.png')).toContain('![image](/tmp/a.png)');
});
});
describe('isTranscriptNoise', () => {
it('drops Whisper hallucinations on silence', () => {
expect(isTranscriptNoise('(cliquant)')).toBe(true);
expect(isTranscriptNoise('*Claire*')).toBe(true);
expect(isTranscriptNoise('[Musique]')).toBe(true);
expect(isTranscriptNoise('...')).toBe(true);
expect(isTranscriptNoise("Sous-titres réalisés par la communauté d'Amara.org")).toBe(true);
});
it('keeps real sentences', () => {
expect(isTranscriptNoise('Jarvis, allume la lumière du salon.')).toBe(false);
expect(isTranscriptNoise('(Jarvis) quelle heure est-il maintenant ?')).toBe(false);
expect(isTranscriptNoise('Oui')).toBe(false);
});
});
+80
View File
@@ -0,0 +1,80 @@
import { describe, expect, it } from 'vitest';
import { bestLocalUpgrade, defaultOpenAiVoice, findLocalVoice, inferGender, pickSpeaker, rankSystemVoice, resolveEdgeVoice, suggestedDownload } from '../src/lib/voicePreference';
import type { VoiceModelStatus, VoiceSpeaker } from '../shared/voice';
const model = (id: string, installed: boolean, speakers: VoiceSpeaker[], languages = ['fr']): VoiceModelStatus =>
({ id, kind: 'tts', engine: 'piper', name: id, description: '', languages, sizeMb: 1, url: '', dir: id, files: [], speakers, installed, installedBytes: 0 }) as unknown as VoiceModelStatus;
describe('inferGender', () => {
it('reads catalog labels, system voices, kokoro ids and Edge short names', () => {
expect(inferGender('Piper Tom (homme, français)')).toBe('male');
expect(inferGender('Siwis (femme, français)')).toBe('female');
expect(inferGender('Microsoft Paul - French (France)')).toBe('male');
expect(inferGender('Microsoft Hortense - French (France)')).toBe('female');
expect(inferGender('am_adam')).toBe('male');
expect(inferGender('fr-FR-HenriNeural')).toBe('male');
expect(inferGender('fr-FR-DeniseNeural')).toBe('female');
expect(inferGender('Voix 3')).toBeUndefined();
});
});
describe('local voice selection', () => {
const supertonic: VoiceSpeaker[] = [
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' }
];
const models = [
model('kokoro-v1', true, [{ id: 30, name: 'Siwis (femme, français)', lang: 'fr' }, { id: 4, name: 'Adam (homme, anglais US)', lang: 'en' }], ['en', 'fr', 'multi']),
model('piper-fr-tom', false, [{ id: 0, name: 'Tom', lang: 'fr' }]),
model('piper-fr-upmc', true, [{ id: 0, name: 'Jessica (femme)', lang: 'fr' }, { id: 1, name: 'Pierre (homme)', lang: 'fr' }]),
model('supertonic-3', false, supertonic, ['fr', 'en', 'multi'])
];
it('prefers an installed speaker of the wanted gender in the right language', () => {
expect(findLocalVoice(models, 'fr-FR', 'male')).toEqual({ modelId: 'piper-fr-upmc', speaker: 1 });
expect(findLocalVoice(models, 'fr-FR', 'female')).toEqual({ modelId: 'kokoro-v1', speaker: 30 });
expect(findLocalVoice(models.slice(0, 2), 'fr', 'male')).toBeNull();
});
it('ranks Supertonic above Kokoro and Piper once installed', () => {
const installed = models.map((m) => (m.id === 'supertonic-3' ? { ...m, installed: true } : m));
expect(findLocalVoice(installed, 'fr-FR', 'male')).toEqual({ modelId: 'supertonic-3', speaker: 6 });
expect(findLocalVoice(installed, 'fr-FR', 'female')).toEqual({ modelId: 'supertonic-3', speaker: 0 });
});
it('suggests Supertonic first, then the Piper voices, for a French download', () => {
expect(suggestedDownload(models, 'fr', 'male')).toBe('supertonic-3');
expect(suggestedDownload(models.filter((m) => m.id !== 'supertonic-3'), 'fr', 'male')).toBe('piper-fr-tom');
expect(suggestedDownload(models, 'en', 'male')).toBeNull();
});
it('reports the best local model still to download', () => {
expect(bestLocalUpgrade(models, 'fr-FR')).toBe('supertonic-3');
expect(bestLocalUpgrade(models.map((m) => ({ ...m, installed: true })), 'fr-FR')).toBeNull();
});
it('keeps the gender when switching models and falls back to the language', () => {
expect(pickSpeaker(models[3], 'fr-FR', 'male')).toBe(6);
expect(pickSpeaker(models[3], 'fr-FR', 'female')).toBe(0);
expect(pickSpeaker(models[2], 'fr', 'male')).toBe(1);
// Kokoro: no masculine French voice → the French voice, not an English one.
expect(pickSpeaker(models[0], 'fr', 'male')).toBe(30);
expect(pickSpeaker({ speakers: [] }, 'fr', 'male')).toBe(0);
});
});
describe('provider defaults', () => {
it('picks onyx for male and nova for female on OpenAI-compatible APIs', () => {
expect(defaultOpenAiVoice('male')).toBe('onyx');
expect(defaultOpenAiVoice('female')).toBe('nova');
});
it('keeps an explicit Edge voice only when it matches the gender', () => {
expect(resolveEdgeVoice('', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(resolveEdgeVoice('fr-FR-RemyMultilingualNeural', 'fr-FR', 'male')).toBe('fr-FR-RemyMultilingualNeural');
expect(resolveEdgeVoice('fr-FR-DeniseNeural', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(resolveEdgeVoice('fr-CH-ArianeNeural', 'fr-FR', 'female')).toBe('fr-CH-ArianeNeural');
});
it('ranks system voices by language then gender', () => {
const paul = { name: 'Microsoft Paul - French (France)', lang: 'fr-FR', localService: true };
const hortense = { name: 'Microsoft Hortense - French (France)', lang: 'fr-FR', localService: true };
const david = { name: 'Microsoft David - English (US)', lang: 'en-US', localService: true };
expect(rankSystemVoice(paul, 'fr', 'male')).toBeGreaterThan(rankSystemVoice(hortense, 'fr', 'male'));
expect(rankSystemVoice(hortense, 'fr', 'female')).toBeGreaterThan(rankSystemVoice(paul, 'fr', 'female'));
expect(rankSystemVoice(paul, 'fr', 'male')).toBeGreaterThan(rankSystemVoice(david, 'fr', 'male'));
});
});