Compare commits

...
13 Commits
Author SHA1 Message Date
LogiFlow 1ea3df2ad1 Merge pull request #24 from R0m1k3/claude/hermes-context-loss-fayw61
fix: keep Hermes session across runs and speak long replies as a digest (2.5.4)
2026-09-05 11:08:12 +02:00
Claude dc491e59c3 Merge origin/main into claude/hermes-context-loss-fayw61 (2.5.4)
Keeps main's neutral default instructions (assistant name now comes from the
settings) together with the new guidance on long replies and continuity;
bumps to 2.5.4 since main already published 2.5.3.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ln4KL1feHtnV4nee4sZFhM
2026-09-05 08:24:29 +00:00
Claude c2aef4cbc7 fix: keep Hermes session across runs and speak long replies as a digest (2.5.3)
Over the runs transport Hermes never reports the session it attached to a
run in the SSE stream, only in GET /v1/runs/{id}. EveFlow only read it in the
polling fallback, so the locally generated id was sent again and again and
each message opened a fresh Hermes session, losing the conversation context.
The client now adopts the run's session id after every run.

Long answers are now spoken as a digest: the first sentences (configurable,
4 by default) plus the closing question, followed by a short notice that the
full text is on screen. The default instructions ask Hermes to open long
replies with the essentials and to rely on the ongoing conversation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ln4KL1feHtnV4nee4sZFhM
2026-09-05 08:22:48 +00:00
Michael 992693c9cf fix: use configurable assistant name in all Hermes conversations (2.5.3) 2026-09-05 09:39:06 +02:00
LogiFlow 5991410685 Merge pull request #23 from R0m1k3/main
Main
2026-09-05 09:18:23 +02:00
Michael 9f95f79943 ci: publish Windows releases from main 2026-09-05 09:18:06 +02:00
Michael c9b44433ac fix: discover actual Hermes provider models and route selections (2.5.2) 2026-09-05 09:06:36 +02:00
Michael 5201742f45 feat: add Hermes AI model selectors and discovery feedback (2.5.1) 2026-09-05 08:36:27 +02:00
LogiFlow b8073c0be6 feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0) (#21)
feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0)
2026-09-04 17:40:11 +02:00
Claude 32551ccc0d feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0)
- Nouveau moteur « Microsoft Edge » (voix neuronales gratuites, sans clé) : Henri / Denise
  par défaut selon le genre, liste des voix fr-FR / fr-CA / fr-CH / fr-BE, WebSocket signé
  (Sec-MS-GEC) dans le processus principal, MP3 24 kHz. Moteur par défaut des nouvelles
  installations.
- Supertonic 3 ajouté au catalogue local (31 langues, 5 voix masculines + 5 féminines,
  44 kHz, 129 Mo) : langue transmise au worker, genres des voix vérifiés par mesure de F0.
- Kokoro déclassé pour le français (une seule voix féminine, accent) ; le choix
  masculin/féminin bascule sur le meilleur modèle installé et conserve le genre quand on
  change de modèle ; Google Translate signalé comme voix féminine uniquement.
- Deux préréglages JARVIS : en ligne (Edge Henri) et hors ligne (Supertonic 3).
- Traitement du micro par Chromium (écho, bruit, gain) débrayable pour de meilleures
  transcriptions au casque.
- Tests : protocole Edge (jeton, SSML, trames), sélection de voix ; README et feuille de
  route (mesures Supertonic, Parakeet vs Qwen3-ASR).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MMpgFriwxiBgurUVb21oCE
2026-09-04 15:37:00 +00:00
LogiFlow a95f91185f Merge pull request #20 from R0m1k3/claude/refonte-app-vocale-v91uz7
Claude/refonte app vocale v91uz7
2026-09-04 13:22:03 +02:00
Claude e55d70bd22 docs: mesures Parakeet vs Whisper et worklets dans la feuille de route
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-04 11:02:08 +00:00
Claude 813425b622 feat: voix JARVIS, Parakeet v3 pour le français, worklets audio sous CSP stricte (v2.4.1)
Voix
- Timbre « JARVIS » (Web Audio) : hauteur légèrement abaissée, chaleur dans
  les basses, présence, compression douce, courte réverbération d'intercom.
  Activé par défaut, réglable dans Paramètres → Voix.
- Préréglage « Voix JARVIS » en un clic : moteur local, voix masculine
  française (Piper Tom téléchargé automatiquement, Kokoro n'ayant pas de voix
  française masculine), débit calme. Affichage de la voix active.
- Boutons Masculine / Féminine désormais lisibles (style du sélecteur ajouté).

Reconnaissance
- Parakeet TDT 0.6B v3 (NVIDIA NeMo, int8, 25 langues européennes dont le
  français) ajouté au catalogue et recommandé : plus précis et plus rapide que
  Whisper sur processeur, ponctuation incluse.
- 0,4 s de silence ajoutées avant et après chaque énoncé avant la
  reconnaissance (syllabes coupées, hallucinations de Whisper sur les clips
  courts).

Écoute permanente
- Les modules AudioWorklet sont livrés en fichiers statiques (public/worklets)
  chargés depuis l'application : en version installée, la CSP stricte
  (script-src 'self') refusait les URL blob et l'écoute permanente échouait
  avec « Unable to load a worklet's module ». Repli blob conservé pour le dev.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Wn5VX9HNbJ7N54hR24u9Y
2026-09-04 10:57:55 +00:00
36 changed files with 1251 additions and 121 deletions

No files matched your search

+1 -1
View File
@@ -67,7 +67,7 @@ jobs:
out/latest.yml
if-no-files-found: error
- name: Publish GitHub release
if: startsWith(github.ref, 'refs/tags/v') || github.event_name == 'workflow_dispatch' || (github.event_name == 'push' && github.ref == 'refs/heads/master' && steps.ver.outputs.released == 'false')
if: startsWith(github.ref, 'refs/tags/v') || github.event_name == 'workflow_dispatch' || (github.event_name == 'push' && (github.ref == 'refs/heads/master' || github.ref == 'refs/heads/main') && steps.ver.outputs.released == 'false')
uses: softprops/action-gh-release@v2
with:
tag_name: v${{ steps.ver.outputs.version }}
+8 -5
View File
@@ -1,7 +1,7 @@
# EveFlow 2 — Interface vocale JARVIS pour Hermes Agent
[![Build](https://img.shields.io/github/actions/workflow/status/R0m1k3/EveFlow/windows-release.yml?style=flat-square)](https://github.com/R0m1k3/EveFlow/actions)
[![Version](https://img.shields.io/badge/version-2.4.1-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases)
[![Version](https://img.shields.io/badge/version-2.5.0-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases)
[![License](https://img.shields.io/badge/license-MIT-lightgrey.svg?style=flat-square)](LICENSE)
**EveFlow** est un compagnon de bureau Windows qui transforme [Hermes Agent](https://hermes-agent.nousresearch.com/) en assistant vocal à la JARVIS : un noyau holographique réactif au son, une conversation en streaming, les outils, sous-agents, approbations, crons, skills et sessions d'Hermes pilotés depuis un seul HUD.
@@ -22,21 +22,24 @@ La version 2 est une réécriture complète : plus de robot 3D, un pipeline voca
* **Capture micro** via AudioWorklet à 16 kHz, sans monitoring du micro dans les haut-parleurs, avec annulation d'écho et réduction de bruit.
* **Détection d'activité vocale** (seuil adaptatif, sensibilité et silence de fin réglables) : l'enregistrement s'arrête tout seul quand vous avez fini de parler.
* **Mains libres** : le micro se réactive après chaque réponse.
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse Kokoro v1.0 (voix française Siwis et voix anglaises) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Parakeet v3, Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse **Supertonic 3** (31 langues dont le français, cinq voix masculines et cinq féminines, 44 kHz, 129 Mo, environ 7× plus rapide que le temps réel sur 4 cœurs), Kokoro v1.0 (excellent en anglais ; en français une seule voix féminine avec accent) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
* **Fin de phrase neuronale** : en écoute permanente, Silero VAD (0,6 Mo, sherpa-onnx) décide du début et de la fin de la commande à la place du seuil d'énergie ; moins de faux départs sur le bruit, coupure plus nette. Repli automatique sur le VAD énergétique si le modèle n'est pas installé.
* **Vision d'écran** : « Jarvis, regarde mon écran » (ou le bouton de la barre de commande) joint une capture de l'écran principal à la question envoyée à Hermes.
* **Actions locales instantanées** : « verrouille la session », « monte le son », « coupe le son », « piste suivante », « ouvre Spotify », « ouvre github.com »… exécutées sur le PC sans passer par Hermes, résultat lu à voix haute. Liste blanche d'actions dans le processus principal, désactivable dans les paramètres.
* **Serveur MCP intégré** : Hermes se connecte à `http://<pc>:7842/mcp` et obtient les outils du PC (capture d'écran renvoyée en image, verrouillage, applications, URL, touches média, presse-papiers, recherche de fichiers, voix, notifications, état du HUD, affichage dans le fil). Même port et même secret que le webhook ; en mode chat completions, les mêmes outils sont proposés directement au modèle.
* **Heures calmes et priorités** : plage horaire pendant laquelle les messages poussés s'affichent sans être lus ni faire clignoter le noyau (badge « non lus » à la place), thème nuit automatique, mots prioritaires lus quand même, résumé vocal des rapports longs (les premières phrases seulement).
* **Mode mission** : un bouton dans la barre de commande bascule sur un second modèle Hermes (plus puissant) pour les tâches longues ; le modèle rapide reste utilisé pour la conversation courante.
* **Résumé vocal des réponses longues** : seules les premières phrases (réglable) et la question finale sont lues, le reste s'affiche dans le fil ; la conversation Hermes est continue d'un message à l'autre (la session créée par Hermes est reprise à chaque run).
* **Widget compact « glanceable »** : état (veille, écoute, réflexion, parle), dernière phrase de l'assistant, badge de non-lus, indicateurs heures calmes et mission.
* **Voix JARVIS** : préréglage en un clic (Paramètres → Voix) : voix française masculine locale (Piper Tom, téléchargée automatiquement), timbre « JARVIS » (légèrement plus grave et posé, chaleur, présence, courte réverbération d'intercom), débit calme. Kokoro n'a pas de voix française masculine.
* **Voix Microsoft Edge** (moteur par défaut) : les voix neuronales de la lecture à voix haute d'Edge, gratuites, sans clé ni installation : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR) et les voix fr-CA, fr-CH, fr-BE, plus de 300 voix dans 74 langues. Le rendu le plus naturel disponible ; nécessite une connexion.
* **Voix JARVIS** : deux préréglages en un clic (Paramètres → Voix) : en ligne (Edge Henri) ou hors ligne (Supertonic 3, voix masculine grave, téléchargé automatiquement), timbre « JARVIS » (légèrement plus grave et posé, chaleur, présence, courte réverbération d'intercom), débit calme.
* **Reconnaissance française de référence** : Parakeet TDT 0.6B v3 (NVIDIA NeMo, 25 langues européennes) dans le catalogue, plus précis et bien plus rapide que Whisper sur processeur, avec ponctuation. Whisper base/small/turbo restent disponibles.
* **Voix masculine ou féminine** : un réglage unique (Paramètres → Voix) appliqué à tous les moteurs. En local, Piper Tom ou Pierre (UPMC) pour le masculin, téléchargé automatiquement si aucune voix masculine n'est installée ; onyx / nova pour les API compatibles OpenAI ; Paul / Hortense pour les voix Windows.
* **Voix masculine ou féminine** : un réglage unique (Paramètres → Voix) appliqué à tous les moteurs. Henri / Denise sur Edge ; en local Supertonic 3 (ou Piper Tom / Pierre) pour le masculin, téléchargé automatiquement si aucune voix masculine n'est installée, et le genre est conservé quand on change de modèle ; onyx / nova pour les API compatibles OpenAI ; Paul / Hortense pour les voix Windows. Google Translate n'a qu'une voix féminine par langue : le réglage l'indique.
* **Barge-in** : en mains libres, parler par-dessus l'assistant coupe sa voix ; le seuil est relevé pendant qu'il parle pour ignorer l'écho du haut-parleur.
* **Écoute permanente** : un détecteur de mot-clé de 3 Mo (sherpa-onnx, keyword spotting) tourne en continu sur le micro, quasi gratuit en CPU. « Jarvis » (ou n'importe quel mot-clé) ouvre l'écoute, « Jarvis, allume… » envoie directement la commande, et le mot coupe la voix en cours. Alternative : filtre du mot après transcription en mains libres.
* **STT externe** : n'importe quelle API `/v1/audio/transcriptions` compatible OpenAI (Qwen3-ASR, Whisper, Speaches, faster-whisper-server, LocalAI, OpenAI). Repli sur la reconnaissance Chromium.
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI, voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI (Qwen3-TTS via un serveur compatible, Kokoro-FastAPI, OpenAI, LocalAI…), voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
* **Traitement du micro débrayable** : l'annulation d'écho, la réduction de bruit et le gain automatique de Chromium peuvent être coupés (Paramètres → Reconnaissance) ; avec un casque, le signal brut est souvent mieux transcrit.
* Raccourcis globaux : `Ctrl+Shift+Espace` (micro), `Ctrl+Shift+J` (afficher/masquer), `Ctrl+Shift+Échap` (couper la voix).
### Hermes, toute la puissance
+4 -1
View File
@@ -8,7 +8,8 @@
|---|---|---|
| HUD arc-reactor réactif au son | Fait | Canvas 2D optimisé (pas d'ombres, couleurs en cache, 30 fps en veille, arrêt fenêtre masquée) |
| Reconnaissance vocale locale | Fait | Whisper base/small/turbo via sherpa-onnx dans un processus utilitaire |
| Synthèse vocale locale | Fait | Kokoro v1.0 (voix française Siwis) et Piper fr |
| Synthèse vocale locale | Fait (2.5.0) | Supertonic 3 (31 langues, 5 voix masculines + 5 féminines, 44 kHz), Kokoro v1.0 (anglais ; français féminin avec accent) et Piper fr |
| Synthèse vocale en ligne | Fait (2.5.0) | Voix neuronales Microsoft Edge (Henri, Denise, Rémy, Vivienne…), gratuites, sans clé, via WebSocket signé dans le processus principal |
| Mot d'activation permanent | Fait (2.2.0) | Keyword spotting sherpa-onnx en continu ; mot-clé libre encodé en BPE ; validé sur audio réel (détection, zéro faux positif sur le test anglais) |
| Mot d'activation après transcription | Fait | Filtre « Jarvis … » en mains libres, tolérant aux erreurs de transcription |
| Détection de fin de phrase | Fait (2.3.0) | Silero VAD neuronal dans le worker (segment renvoyé au renderer), repli sur le VAD énergétique si le modèle manque |
@@ -83,6 +84,8 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai)
| Correctifs 2.4.0.1 | réponse vide en chat completions désormais expliquée (JSON non streamé ou erreur HTTP 200), la liaison ne passe plus en « dégradé » quand seule l'API des crons échoue, transcriptions parasites (« (cliquant) », « *Claire* ») ignorées, préférence de voix masculine/féminine |
| Parakeet v3 vs Whisper base sur trois phrases Piper (fr) (2.4.1) | Parakeet : 3/3 exactes avec ponctuation, 0,4 à 0,5 s à chaud (6,7 s au premier appel) ; Whisper base : erreurs sur « Jarvis », « Peux-tu », 0,7 à 0,9 s |
| Worklets audio sous CSP stricte (2.4.1) | chargement des modules statiques OK dans l'application empaquetée |
| Voix françaises (2.5.0) | Edge Henri : MP3 reçu de bout en bout (32 ko pour 4 s). Supertonic 3 en français : 10 voix, RTF 0,14 sur 4 cœurs (8,7 s d'audio en 1,3 s), genres déterminés par mesure de la fréquence fondamentale (voix 0-4 : 170-210 Hz, voix 5-9 : 92-137 Hz) ; le worker compilé accepte `language` et bascule sur l'anglais pour une langue inconnue |
| Reconnaissance française : Parakeet v3 vs Qwen3-ASR 0.6B int8 (2.5.0) | Six phrases Supertonic (3,8 s) : Parakeet WER 8,5 % (erreurs surtout de forme : « 14h30 »), 384 ms par phrase ; Qwen3-ASR WER 15,3 % (« mémoires vivres »), 1 477 ms, 940 Mo. Qwen3-ASR n'est pas ajouté au catalogue |
| Serveur MCP (2.4.0) | initialize, tools/list (14 outils), tools/call côté principal (presse-papiers, capture image) et côté renderer (état, message dans le fil) |
Sur un PC à 28 cœurs les temps sont nettement plus courts. Whisper small est maintenant recommandé pour le français.
+3 -1
View File
@@ -15,7 +15,7 @@ import {
} from '../shared/ipc';
import type { EveFlowBridge, SystemAction, SystemActionResult, Unsubscribe } from '../shared/bridge';
import type { McpToolRequest, McpToolResponse } from '../shared/ipc';
import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
import { VOICE_IPC, type EdgeSynthesizeRequest, type EdgeSynthesizeResult, type EdgeVoice, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
function subscribe<T>(channel: string, callback: (payload: T) => void): Unsubscribe {
const listener = (_event: Electron.IpcRendererEvent, payload: T) => callback(payload);
@@ -74,6 +74,8 @@ const api: EveFlowBridge = {
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => subscribe<VoiceDownloadProgress>(VOICE_IPC.modelsProgress, cb),
transcribe: (req: TranscribeRequest) => ipcRenderer.invoke(VOICE_IPC.transcribe, req) as Promise<TranscribeResult>,
synthesize: (req: SynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.synthesize, req) as Promise<SynthesizeResult>,
edgeSynthesize: (req: EdgeSynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.edgeSynthesize, req) as Promise<EdgeSynthesizeResult>,
edgeVoices: () => ipcRenderer.invoke(VOICE_IPC.edgeVoices) as Promise<EdgeVoice[]>,
unload: (id?: string) => ipcRenderer.invoke(VOICE_IPC.unload, id) as Promise<unknown>,
kwsStart: (req: KwsStartRequest) => ipcRenderer.invoke(VOICE_IPC.kwsStart, req) as Promise<{ accepted: string[]; rejected: string[] }>,
kwsStop: () => ipcRenderer.invoke(VOICE_IPC.kwsStop) as Promise<void>,
+39 -4
View File
@@ -17,6 +17,22 @@ const KOKORO_SPEAKERS: VoiceSpeaker[] = [
{ id: 26, name: 'George (homme, anglais UK)', lang: 'en' }
];
// Supertonic 3 ships ten voice styles in voice.bin (five feminine, five masculine). The ordering
// was checked by measuring the fundamental frequency of French synthesis: sid 0-4 around
// 170-210 Hz, sid 5-9 around 90-140 Hz.
const SUPERTONIC_SPEAKERS: VoiceSpeaker[] = [
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
{ id: 9, name: 'Homme 5 (grave)', lang: 'multi', gender: 'm' },
{ id: 7, name: 'Homme 3', lang: 'multi', gender: 'm' },
{ id: 8, name: 'Homme 4', lang: 'multi', gender: 'm' },
{ id: 5, name: 'Homme 1 (clair)', lang: 'multi', gender: 'm' },
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' },
{ id: 1, name: 'Femme 2', lang: 'multi', gender: 'f' },
{ id: 2, name: 'Femme 3', lang: 'multi', gender: 'f' },
{ id: 3, name: 'Femme 4', lang: 'multi', gender: 'f' },
{ id: 4, name: 'Femme 5', lang: 'multi', gender: 'f' }
];
export const VOICE_CATALOG: VoiceModelSpec[] = [
{
id: 'whisper-base',
@@ -112,20 +128,36 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [
files: ['silero_vad.onnx'],
recommended: true
},
{
id: 'supertonic-3',
kind: 'tts',
engine: 'supertonic',
name: 'Supertonic 3 (31 langues, 5 voix masculines et 5 féminines)',
description:
'La meilleure voix française locale : naturelle, sans accent, dix voix au choix, 44 kHz. Environ 7 fois plus rapide que le temps réel sur 4 cœurs, 100 M de paramètres.',
languages: ['fr', 'en', 'de', 'es', 'it', 'pt', 'multi'],
sizeMb: 129,
url: `${TTS}/sherpa-onnx-supertonic-3-tts-int8-2026-05-11.tar.bz2`,
dir: 'sherpa-onnx-supertonic-3-tts-int8-2026-05-11',
files: ['duration_predictor.int8.onnx', 'text_encoder.int8.onnx', 'vector_estimator.int8.onnx', 'vocoder.int8.onnx', 'tts.json', 'unicode_indexer.bin', 'voice.bin'],
speakers: SUPERTONIC_SPEAKERS,
sampleRate: 44100,
recommended: true
},
{
id: 'kokoro-v1',
kind: 'tts',
engine: 'kokoro',
name: 'Kokoro v1.0 multilingue',
description: 'Voix très naturelle, une voix française (Siwis) et de nombreuses voix anglaises. 24 kHz.',
languages: ['fr', 'en', 'multi'],
description:
'Excellent en anglais (nombreuses voix). En français : une seule voix, féminine (Siwis), avec un accent marqué (phonémisation espeak). Préférez Supertonic 3 pour le français. 24 kHz.',
languages: ['en', 'fr', 'multi'],
sizeMb: 349,
url: `${TTS}/kokoro-multi-lang-v1_0.tar.bz2`,
dir: 'kokoro-multi-lang-v1_0',
files: ['model.onnx', 'voices.bin', 'tokens.txt', 'lexicon-us-en.txt', 'lexicon-zh.txt', 'espeak-ng-data/phontab'],
speakers: KOKORO_SPEAKERS,
sampleRate: 24000,
recommended: true
sampleRate: 24000
},
{
id: 'piper-fr-siwis',
@@ -181,6 +213,9 @@ export function findModel(id: string): VoiceModelSpec | undefined {
// Speaker gender from the label ("(homme, …)", "(femme, …)", known first names) so the UI and the
// voice preference can pick a masculine or feminine voice without a lookup table per model.
const MALE = /\b(homme|tom|pierre|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|santa)\b/i;
/** Languages accepted by the Supertonic 3 text front-end (2-letter codes). */
export const SUPERTONIC_LANGS = new Set(['ar', 'bg', 'hr', 'cs', 'da', 'nl', 'en', 'et', 'fi', 'fr', 'de', 'el', 'hi', 'hu', 'id', 'it', 'ja', 'ko', 'lv', 'lt', 'pl', 'pt', 'ro', 'ru', 'sk', 'sl', 'es', 'sv', 'tr', 'uk', 'vi']);
for (const spec of VOICE_CATALOG) {
for (const sp of spec.speakers ?? []) {
if (!sp.gender) sp.gender = /\b(femme|female)\b/i.test(sp.name) ? 'f' : MALE.test(sp.name) ? 'm' : /\bfemme\b/i.test(spec.name) ? 'f' : /\bhomme\b/i.test(spec.name) ? 'm' : undefined;
+146
View File
@@ -0,0 +1,146 @@
/**
* Microsoft Edge "Read aloud" neural voices (the service behind the Edge browser's read-aloud
* feature): free, no key, very natural French voices with a real masculine/feminine choice
* (Henri, Denise, Rémy, Vivienne…). Runs in the main process: one WebSocket per sentence,
* MP3 back to the renderer.
*/
import { createHash, randomBytes, randomUUID } from 'node:crypto';
import type { EdgeSynthesizeRequest, EdgeSynthesizeResult, EdgeVoice } from '../../shared/voice';
import {
EDGE_CHROMIUM_VERSION,
EDGE_VOICES_URL,
EDGE_WSS_URL,
edgeConfigMessage,
edgeConnectionId,
edgeHeaders,
edgeSsml,
edgeSsmlMessage,
edgeTextFramePath,
edgeTokenInput,
parseEdgeBinaryFrame
} from '../../shared/edgeTts';
import { log } from '../logger';
/** Seconds to add to the local clock so the signed token matches the server's time window. */
let clockSkewSec = 0;
let voicesCache: { at: number; voices: EdgeVoice[] } | null = null;
const VOICES_TTL_MS = 6 * 60 * 60 * 1000;
const SYNTH_TIMEOUT_MS = 20_000;
function token(): string {
return createHash('sha256').update(edgeTokenInput(Date.now(), clockSkewSec), 'ascii').digest('hex').toUpperCase();
}
function signedUrl(): string {
return `${EDGE_WSS_URL}&ConnectionId=${edgeConnectionId(randomUUID())}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
}
function headers(): Record<string, string> {
return { ...edgeHeaders(), Cookie: `muid=${randomBytes(16).toString('hex')};` };
}
/** Learn the server clock from a plain HTTPS response (the WebSocket handshake hides its headers). */
async function syncClock(): Promise<void> {
try {
const res = await fetch(EDGE_VOICES_URL, { method: 'HEAD', headers: headers() });
const date = res.headers.get('date');
if (!date) return;
const server = Date.parse(date);
if (Number.isFinite(server)) {
clockSkewSec = (server - Date.now()) / 1000;
log('INFO', 'edge-tts', `clock skew ${clockSkewSec.toFixed(0)} s`);
}
} catch (err) {
log('WARN', 'edge-tts', `clock sync failed: ${(err as Error).message}`);
}
}
function synthesizeOnce(req: EdgeSynthesizeRequest): Promise<Uint8Array> {
return new Promise((resolve, reject) => {
const chunks: Uint8Array[] = [];
let settled = false;
let ws: WebSocket;
try {
// Node's global WebSocket accepts extra handshake headers (undici), which the service checks.
ws = new (WebSocket as unknown as new (url: string, options: { headers: Record<string, string> }) => WebSocket)(signedUrl(), { headers: headers() });
} catch (err) {
reject(err as Error);
return;
}
ws.binaryType = 'arraybuffer';
const finish = (err?: Error) => {
if (settled) return;
settled = true;
clearTimeout(timer);
try {
ws.close();
} catch {
/* already closed */
}
if (err) reject(err);
else {
const total = chunks.reduce((n, c) => n + c.byteLength, 0);
const out = new Uint8Array(total);
let o = 0;
for (const c of chunks) {
out.set(c, o);
o += c.byteLength;
}
resolve(out);
}
};
const timer = setTimeout(() => finish(new Error('Edge TTS : délai dépassé')), SYNTH_TIMEOUT_MS);
ws.onopen = () => {
ws.send(edgeConfigMessage());
ws.send(edgeSsmlMessage(edgeConnectionId(randomUUID()), edgeSsml(req.text, req.voice, req.speed)));
};
ws.onmessage = (event: MessageEvent) => {
if (typeof event.data === 'string') {
if (edgeTextFramePath(event.data) === 'turn.end') finish();
return;
}
const frame = new Uint8Array(event.data as ArrayBuffer);
const { path, payload } = parseEdgeBinaryFrame(frame);
if (path === 'audio' && payload.byteLength) chunks.push(payload);
};
ws.onerror = (event: Event) => finish(new Error(`Edge TTS : connexion refusée (${(event as { message?: string }).message ?? 'erreur réseau'})`));
ws.onclose = (event: CloseEvent) => {
if (!settled) finish(chunks.length ? undefined : new Error(`Edge TTS : connexion fermée (${event.code}${event.reason ? ' ' + event.reason : ''})`));
};
});
}
export async function edgeSynthesize(req: EdgeSynthesizeRequest): Promise<EdgeSynthesizeResult> {
const started = Date.now();
let mp3: Uint8Array;
try {
mp3 = await synthesizeOnce(req);
} catch (err) {
// A refused handshake is almost always a stale signature: resync the clock and retry once.
log('WARN', 'edge-tts', `first attempt failed (${(err as Error).message}), resyncing clock`);
await syncClock();
mp3 = await synthesizeOnce(req);
}
if (!mp3.byteLength) throw new Error('Edge TTS : aucun audio reçu');
return { mp3, durationMs: Date.now() - started };
}
/** Voice list from the service (cached six hours); falls back to an empty list offline. */
export async function edgeVoices(): Promise<EdgeVoice[]> {
if (voicesCache && Date.now() - voicesCache.at < VOICES_TTL_MS) return voicesCache.voices;
const url = `${EDGE_VOICES_URL}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
const res = await fetch(url, { headers: headers() });
if (!res.ok) throw new Error(`Edge TTS : liste des voix HTTP ${res.status}`);
const raw = (await res.json()) as Array<{ ShortName?: string; FriendlyName?: string; Locale?: string; Gender?: string }>;
const voices: EdgeVoice[] = raw
.filter((v) => typeof v.ShortName === 'string' && typeof v.Locale === 'string')
.map((v) => ({
shortName: v.ShortName!,
name: v.ShortName!.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') || v.FriendlyName || v.ShortName!,
locale: v.Locale!,
gender: v.Gender === 'Male' ? ('m' as const) : ('f' as const)
}))
.sort((a, b) => a.locale.localeCompare(b.locale) || a.name.localeCompare(b.name));
voicesCache = { at: Date.now(), voices };
return voices;
}
+1 -1
View File
@@ -134,7 +134,7 @@ export function transcribe(req: TranscribeRequest): Promise<TranscribeResult> {
export async function synthesize(req: SynthesizeRequest): Promise<SynthesizeResult> {
const result = await request<Omit<SynthesizeResult, 'wav'> & { wav: string }>(
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed },
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed, language: req.language },
180_000
);
const buffer = Buffer.from(result.wav, 'base64');
+14 -2
View File
@@ -1,5 +1,6 @@
import { ipcMain } from 'electron';
import { VOICE_IPC, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice';
import { VOICE_IPC, type EdgeSynthesizeRequest, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice';
import { edgeSynthesize, edgeVoices } from './edgeTts';
import { engineStatus, kwsFeed, kwsStart, kwsStop, synthesize, transcribe, unload, vadFeed, vadStart, vadStop } from './engine';
import { cancelDownload, downloadModel, listModels, removeModel } from './models';
@@ -24,8 +25,19 @@ export function registerVoiceIpc(): void {
ipcMain.handle(VOICE_IPC.synthesize, (_e, req: SynthesizeRequest) => {
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
if (typeof req.modelId !== 'string') throw new Error('Modèle invalide');
return synthesize({ ...req, speaker: Number.isFinite(req.speaker) ? req.speaker : 0, speed: Number.isFinite(req.speed) ? req.speed : 1 });
return synthesize({
...req,
speaker: Number.isFinite(req.speaker) ? req.speaker : 0,
speed: Number.isFinite(req.speed) ? req.speed : 1,
language: typeof req.language === 'string' ? req.language.slice(0, 8) : undefined
});
});
ipcMain.handle(VOICE_IPC.edgeSynthesize, (_e, req: EdgeSynthesizeRequest) => {
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
if (typeof req.voice !== 'string' || !/^[a-z]{2,3}-[A-Za-z]{2,4}-[A-Za-z0-9]+$/.test(req.voice)) throw new Error('Voix Edge invalide');
return edgeSynthesize({ text: req.text, voice: req.voice, speed: Number.isFinite(req.speed) ? req.speed : 1 });
});
ipcMain.handle(VOICE_IPC.edgeVoices, () => edgeVoices());
ipcMain.handle(VOICE_IPC.unload, (_e, id?: string) => unload(id));
ipcMain.handle(VOICE_IPC.kwsStart, (event, req: KwsStartRequest) => {
if (!req || !Array.isArray(req.keywords) || typeof req.modelId !== 'string') throw new Error('Requête invalide');
+36 -3
View File
@@ -6,6 +6,7 @@
import os from 'node:os';
import path from 'node:path';
import type { VoiceEngineKind } from '../../shared/voice';
import { SUPERTONIC_LANGS } from './catalog';
interface ModelRef {
id: string;
@@ -17,7 +18,7 @@ interface ModelRef {
type Request =
| { id: number; type: 'status' }
| { id: number; type: 'transcribe'; model: ModelRef; wav: Uint8Array | string; language: string }
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number }
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number; language?: string }
| { id: number; type: 'unload'; modelId?: string }
| { id: number; type: 'kws.start'; model: ModelRef; keywordsFile: string; threshold: number; score: number }
| { id: number; type: 'kws.audio'; pcm: string; sampleRate: number }
@@ -55,8 +56,10 @@ type Sherpa = {
OfflineTts: new (config: unknown) => {
numSpeakers: number;
sampleRate: number;
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean }) => { samples: Float32Array; sampleRate: number };
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean; generationConfig?: unknown }) => { samples: Float32Array; sampleRate: number };
};
/** Per-request options for the newer engines (Supertonic reads `extra.lang`). */
GenerationConfig: new (opts: { sid: number; speed: number; numSteps?: number; extra?: Record<string, string | number> }) => unknown;
version: string;
};
@@ -152,6 +155,19 @@ function getSynthesizer(model: ModelRef) {
ttsModel = { vits: { model: p(onnx), tokens: p('tokens.txt'), dataDir: p('espeak-ng-data') } };
break;
}
case 'supertonic':
ttsModel = {
supertonic: {
durationPredictor: p('duration_predictor.int8.onnx'),
textEncoder: p('text_encoder.int8.onnx'),
vectorEstimator: p('vector_estimator.int8.onnx'),
vocoder: p('vocoder.int8.onnx'),
ttsJson: p('tts.json'),
unicodeIndexer: p('unicode_indexer.bin'),
voiceStyle: p('voice.bin')
}
};
break;
default:
throw new Error(`Moteur TTS non supporté : ${model.engine}`);
}
@@ -160,6 +176,12 @@ function getSynthesizer(model: ModelRef) {
return tts;
}
/** 2-letter code accepted by Supertonic 3 ("fr-FR" → "fr"); English when unknown, as upstream does. */
function supertonicLang(language: string | undefined): string {
const code = (language ?? '').toLowerCase().split(/[-_]/)[0];
return SUPERTONIC_LANGS.has(code) ? code : 'en';
}
// ── audio helpers ──────────────────────────────────────────────────────────
function decodeWav(bytes: Uint8Array): { samples: Float32Array; sampleRate: number } {
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
@@ -397,8 +419,19 @@ function handle(req: Request): unknown {
const started = Date.now();
const tts = getSynthesizer(req.model);
const sid = Math.max(0, Math.min(tts.numSpeakers - 1, Math.floor(req.speaker)));
const speed = Math.max(0.5, Math.min(2, req.speed || 1));
// Electron forbids N-API external buffers: ask sherpa-onnx to copy the samples into a V8 buffer.
const audio = tts.generate({ text: req.text, sid, speed: Math.max(0.5, Math.min(2, req.speed || 1)), enableExternalBuffer: false });
const audio =
req.model.engine === 'supertonic'
? tts.generate({
text: req.text,
sid,
speed,
enableExternalBuffer: false,
// Supertonic needs the language of the text; 5 denoising steps is the quality/speed sweet spot.
generationConfig: new (loadSherpa().GenerationConfig)({ sid, speed, numSteps: 5, extra: { lang: supertonicLang(req.language) } })
})
: tts.generate({ text: req.text, sid, speed, enableExternalBuffer: false });
const wav = encodeWav(audio.samples, audio.sampleRate);
return {
wav: Buffer.from(wav.buffer, wav.byteOffset, wav.byteLength).toString('base64'),
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "eveflow",
"version": "2.1.0",
"version": "2.5.4",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "eveflow",
"version": "2.1.0",
"version": "2.5.4",
"license": "MIT",
"dependencies": {
"@fontsource/orbitron": "^5.3.0",
+2 -2
View File
@@ -1,7 +1,7 @@
{
"name": "eveflow",
"version": "2.4.1",
"releaseVersion": "2.4.1",
"version": "2.5.4",
"releaseVersion": "2.5.4",
"description": "JARVIS-style desktop HUD for Hermes Agent: voice, streaming runs, scheduled jobs, skills and telemetry",
"main": "dist-electron/main.js",
"private": true,
+6
View File
@@ -13,6 +13,9 @@ import type {
} from './ipc';
import type { McpToolRequest, McpToolResponse } from './ipc';
import type {
EdgeSynthesizeRequest,
EdgeSynthesizeResult,
EdgeVoice,
KwsDetection,
KwsStartRequest,
VadEvent,
@@ -75,6 +78,9 @@ export interface EveFlowBridge {
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => Unsubscribe;
transcribe: (req: TranscribeRequest) => Promise<TranscribeResult>;
synthesize: (req: SynthesizeRequest) => Promise<SynthesizeResult>;
/** Microsoft Edge neural voices (online, free): MP3 for one sentence. */
edgeSynthesize: (req: EdgeSynthesizeRequest) => Promise<EdgeSynthesizeResult>;
edgeVoices: () => Promise<EdgeVoice[]>;
unload: (id?: string) => Promise<unknown>;
kwsStart: (req: KwsStartRequest) => Promise<{ accepted: string[]; rejected: string[] }>;
kwsStop: () => Promise<void>;
+119
View File
@@ -0,0 +1,119 @@
/**
* Pure helpers for the Microsoft Edge "Read aloud" speech service (the endpoint used by the
* Edge browser, no API key): request signing input, SSML building, frame parsing and the
* default French voices. No Node or DOM dependency so it is shared by main and tests.
*/
export const EDGE_TRUSTED_CLIENT_TOKEN = '6A5AA1D4EAFF4E9FB37E23D68491D6F4';
export const EDGE_CHROMIUM_VERSION = '143.0.3650.75';
export const EDGE_WSS_URL = `wss://speech.platform.bing.com/consumer/speech/synthesize/readaloud/edge/v1?TrustedClientToken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
export const EDGE_VOICES_URL = `https://speech.platform.bing.com/consumer/speech/synthesize/readaloud/voices/list?trustedclienttoken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
export const EDGE_OUTPUT_FORMAT = 'audio-24khz-48kbitrate-mono-mp3';
const WIN_EPOCH_SEC = 11644473600;
/**
* String whose SHA-256 (upper-case hex) is the `Sec-MS-GEC` value: Windows file time (100 ns
* ticks since 1601) rounded down to 5 minutes, followed by the trusted client token.
* `nowMs` is the client clock, `skewSec` the correction learned from the server's Date header.
*/
export function edgeTokenInput(nowMs: number, skewSec = 0): string {
let seconds = Math.floor(nowMs / 1000 + skewSec) + WIN_EPOCH_SEC;
seconds -= seconds % 300;
// 10 million ticks per second; BigInt keeps the 18-digit value exact.
const ticks = BigInt(seconds) * 10_000_000n;
return `${ticks}${EDGE_TRUSTED_CLIENT_TOKEN}`;
}
export function edgeHeaders(): Record<string, string> {
const major = EDGE_CHROMIUM_VERSION.split('.')[0];
return {
'User-Agent': `Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36 Edg/${major}.0.0.0`,
'Accept-Language': 'en-US,en;q=0.9',
Pragma: 'no-cache',
'Cache-Control': 'no-cache',
Origin: 'chrome-extension://jdiccldimpdaibmpdkjnbmckianbfold'
};
}
export function escapeXml(text: string): string {
return text.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;').replace(/'/g, '&apos;');
}
/** Prosody rate attribute for a playback speed multiplier (1 = "+0%", 1.25 = "+25%"). */
export function edgeRate(speed: number): string {
const pct = Math.round((Math.max(0.5, Math.min(2, speed || 1)) - 1) * 100);
return `${pct >= 0 ? '+' : ''}${pct}%`;
}
export function edgeSsml(text: string, voice: string, speed: number): string {
const lang = voice.split('-').slice(0, 2).join('-') || 'fr-FR';
return (
`<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xml:lang='${lang}'>` +
`<voice name='${escapeXml(voice)}'><prosody pitch='+0Hz' rate='${edgeRate(speed)}' volume='+0%'>${escapeXml(text)}</prosody></voice></speak>`
);
}
/** Timestamp header the service expects ("JavaScript date string"). */
export function edgeTimestamp(date = new Date()): string {
return date.toUTCString().replace('GMT', 'GMT+0000 (Coordinated Universal Time)');
}
export function edgeConfigMessage(date = new Date()): string {
return (
`X-Timestamp:${edgeTimestamp(date)}\r\nContent-Type:application/json; charset=utf-8\r\nPath:speech.config\r\n\r\n` +
`{"context":{"synthesis":{"audio":{"metadataoptions":{"sentenceBoundaryEnabled":"false","wordBoundaryEnabled":"false"},"outputFormat":"${EDGE_OUTPUT_FORMAT}"}}}}`
);
}
export function edgeSsmlMessage(requestId: string, ssml: string, date = new Date()): string {
return `X-RequestId:${requestId}\r\nContent-Type:application/ssml+xml\r\nX-Timestamp:${edgeTimestamp(date)}Z\r\nPath:ssml\r\n\r\n${ssml}`;
}
/** Split a binary frame into its text headers and payload (2-byte big-endian header length). */
export function parseEdgeBinaryFrame(frame: Uint8Array): { path: string; payload: Uint8Array } {
if (frame.byteLength < 2) return { path: '', payload: new Uint8Array(0) };
const headerLength = (frame[0] << 8) | frame[1];
const end = Math.min(frame.byteLength, 2 + headerLength);
let header = '';
for (let i = 2; i < end; i++) header += String.fromCharCode(frame[i]);
const path = /Path:\s*([^\r\n]+)/i.exec(header)?.[1]?.trim() ?? '';
return { path, payload: frame.subarray(end) };
}
/** Path of a text frame ("turn.start", "response", "audio.metadata", "turn.end"). */
export function edgeTextFramePath(message: string): string {
return /Path:\s*([^\r\n]+)/i.exec(message)?.[1]?.trim() ?? '';
}
export function edgeConnectionId(hex32: string): string {
return hex32.replace(/-/g, '').toLowerCase();
}
/** Well-known voices per language so the choice works before the voice list is fetched. */
export const EDGE_DEFAULT_VOICES: Record<string, { male: string; female: string }> = {
fr: { male: 'fr-FR-HenriNeural', female: 'fr-FR-DeniseNeural' },
en: { male: 'en-US-AndrewMultilingualNeural', female: 'en-US-AvaMultilingualNeural' },
de: { male: 'de-DE-ConradNeural', female: 'de-DE-KatjaNeural' },
es: { male: 'es-ES-AlvaroNeural', female: 'es-ES-ElviraNeural' },
it: { male: 'it-IT-DiegoNeural', female: 'it-IT-ElsaNeural' },
pt: { male: 'pt-BR-AntonioNeural', female: 'pt-BR-FranciscaNeural' }
};
/** Default Edge voice for a language ("fr-FR", "fr", "en-GB") and gender; French when unknown. */
export function defaultEdgeVoice(language: string, gender: 'male' | 'female'): string {
const lang = (language || 'fr').toLowerCase().split(/[-_]/)[0];
return (EDGE_DEFAULT_VOICES[lang] ?? EDGE_DEFAULT_VOICES.fr)[gender];
}
/** Gender of an Edge voice from its short name, using the built-in table then common first names. */
export function edgeVoiceGender(shortName: string): 'male' | 'female' | undefined {
for (const pair of Object.values(EDGE_DEFAULT_VOICES)) {
if (pair.male === shortName) return 'male';
if (pair.female === shortName) return 'female';
}
const name = shortName.split('-')[2]?.replace(/(Multilingual)?Neural$/i, '') ?? '';
if (/^(Henri|Remy|Rémy|Gerard|Antoine|Jean|Thierry|Fabrice|Claude|Andrew|Brian|Guy|Christopher|Eric|Roger|Steffan|Ryan|Thomas|Conrad|Alvaro|Diego|Antonio)$/i.test(name)) return 'male';
if (/^(Denise|Eloise|Vivienne|Charline|Sylvie|Ariane|Ava|Emma|Jenny|Aria|Michelle|Ana|Sonia|Libby|Katja|Elvira|Elsa|Francisca)$/i.test(name)) return 'female';
return undefined;
}
+28 -1
View File
@@ -1,7 +1,7 @@
/** Local voice engine contract (sherpa-onnx in a utility process). Shared by main and renderer. */
export type VoiceModelKind = 'stt' | 'tts' | 'kws' | 'vad';
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'kws-transducer' | 'silero';
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'supertonic' | 'kws-transducer' | 'silero';
export interface VoiceSpeaker {
id: number;
@@ -62,6 +62,31 @@ export interface SynthesizeRequest {
text: string;
speaker: number;
speed: number;
/** BCP-47 tag or 2-letter code of the text (multilingual engines such as Supertonic need it). */
language?: string;
}
/** Microsoft Edge "Read aloud" neural voices (online, no key). */
export interface EdgeVoice {
/** e.g. fr-FR-HenriNeural */
shortName: string;
/** Display name, e.g. Henri */
name: string;
locale: string;
gender: 'm' | 'f';
}
export interface EdgeSynthesizeRequest {
text: string;
voice: string;
/** Playback speed multiplier (0.5 .. 2). */
speed: number;
}
export interface EdgeSynthesizeResult {
/** MP3, 24 kHz mono 48 kbit/s. */
mp3: Uint8Array;
durationMs: number;
}
export interface SynthesizeResult {
@@ -115,6 +140,8 @@ export const VOICE_IPC = {
modelsProgress: 'voice:models:progress',
transcribe: 'voice:transcribe',
synthesize: 'voice:synthesize',
edgeSynthesize: 'voice:edge:synthesize',
edgeVoices: 'voice:edge:voices',
unload: 'voice:unload',
kwsStart: 'voice:kws:start',
kwsStop: 'voice:kws:stop',
+25
View File
@@ -0,0 +1,25 @@
import { useId, useState } from 'react';
import type { HermesModel } from '../../services/hermes/types';
export function ModelSelect({ label, value, models, defaultLabel, onChange }: {
label: string;
value: string;
models: HermesModel[];
defaultLabel: string;
onChange: (value: string) => void;
}) {
const id = useId();
const [manual, setManual] = useState(false);
return <div className="field">
<label htmlFor={id}>{label}</label>
<select id={id} className="select" value={value} onChange={(e) => onChange(e.target.value)}>
<option value="">{defaultLabel}</option>
{value && !models.some((m) => m.id === value) && <option value={value}>{value} (configuré)</option>}
{models.map((m) => <option key={m.id} value={m.id} disabled={m.available === false}>{m.name || m.id}{m.provider || m.owned_by ? ` · ${m.provider || m.owned_by}` : ''}{m.available === false ? ' (indisponible)' : ''}</option>)}
</select>
<button type="button" className="btn small" aria-expanded={manual} onClick={() => setManual(!manual)}>
{manual ? 'Masquer la saisie manuelle' : 'Saisir un autre identifiant'}
</button>
{manual && <input className="input" aria-label={`${label} : identifiant manuel`} value={value} placeholder={defaultLabel} onChange={(e) => onChange(e.target.value)} />}
</div>;
}
+5 -1
View File
@@ -3,6 +3,7 @@ import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, Re
import type { VoiceModelStatus } from '../../../shared/voice';
import { bridge } from '../../lib/bridge';
import { useShallow } from 'zustand/react/shallow';
import { pickSpeaker } from '../../lib/voicePreference';
import { useVoiceModels } from '../../state/voiceModels';
import { useSettings } from '../../state/settings';
@@ -24,7 +25,10 @@ function ModelRow({ model }: { model: VoiceModelStatus }) {
if (model.kind === 'stt') update({ voice: { localModel: model.id, provider: 'local' } });
else if (model.kind === 'kws') update({ voice: { wakeMode: 'kws' } });
else if (model.kind === 'vad') update({ voice: { neuralVad: true } });
else update({ speech: { localModel: model.id, provider: 'local', localSpeaker: model.speakers?.[0]?.id ?? 0 } });
else {
const { speech } = useSettings.getState().settings;
update({ speech: { localModel: model.id, provider: 'local', localSpeaker: pickSpeaker(model, speech.language || 'fr-FR', speech.voiceGender ?? 'male') } });
}
};
const neuralVad = useSettings((s) => s.settings.voice.neuralVad);
const activeNow = model.kind === 'kws' ? wakeMode === 'kws' : model.kind === 'vad' ? neuralVad : isActive;
+100 -23
View File
@@ -5,6 +5,8 @@ import { HermesClient, discoverHermesUrl, hermesUrlCandidates } from '../../serv
import { listSystemVoices } from '../../services/voice/tts';
import { speech } from '../../services/voice/speech';
import { ensurePreferredVoice } from '../../services/voice/voicePreference';
import { pickSpeaker, resolveEdgeVoice } from '../../lib/voicePreference';
import type { EdgeVoice } from '../../../shared/voice';
import { listMicrophones } from '../../services/voice/capture';
import { transcribeWav } from '../../services/voice/stt';
import { encodeWav } from '../../services/voice/wav';
@@ -13,6 +15,7 @@ import { DEFAULT_SETTINGS, useSettings, type HudTheme } from '../../state/settin
import { useVoice } from '../../state/voice';
import { installedModels, useVoiceModels } from '../../state/voiceModels';
import { ModelsSection } from './ModelsSection';
import { ModelSelect } from './ModelSelect';
type Section = 'general' | 'hermes' | 'voice' | 'speech' | 'models' | 'webhook' | 'notifications' | 'ui';
@@ -76,6 +79,10 @@ export function SettingsDrawer({ onClose }: Props) {
const update = useSettings((s) => s.update);
const reset = useSettings((s) => s.reset);
const hermesModels = useHermes((s) => s.models);
const modelsLoading = useHermes((s) => s.modelsLoading);
const modelsError = useHermes((s) => s.modelsError);
const modelsNotice = useHermes((s) => s.modelsNotice);
const refreshHermesModels = useHermes((s) => s.refreshModels);
const hermesWebhook = useHermes((s) => s.webhook);
const hermesConnect = useHermes((s) => s.connect);
const micDevices = useVoice((s) => s.micDevices);
@@ -86,8 +93,25 @@ export function SettingsDrawer({ onClose }: Props) {
const [sttTest, setSttTest] = useState<TestState>({ status: 'idle', message: '' });
const [ttsTest, setTtsTest] = useState<TestState>({ status: 'idle', message: '' });
const [voices, setVoices] = useState(listSystemVoices());
const [edgeVoices, setEdgeVoices] = useState<EdgeVoice[]>([]);
const [webhookSecretVisible, setWebhookSecretVisible] = useState(false);
useEffect(() => {
if (section !== 'hermes') return;
useHermes.setState({ models: [], modelsError: null });
const timer = setTimeout(() => void refreshHermesModels(), 500);
return () => clearTimeout(timer);
}, [section, settings.hermes.url, settings.hermes.apiKey, settings.hermes.sessionKey, refreshHermesModels]);
const speechProvider = settings.speech.provider;
useEffect(() => {
if (speechProvider !== 'edge' || edgeVoices.length) return;
bridge()
?.voice.edgeVoices()
.then(setEdgeVoices)
.catch((err: Error) => setTtsTest({ status: 'fail', message: `Liste des voix Edge indisponible : ${err.message}` }));
}, [speechProvider, edgeVoices.length]);
useEffect(() => {
const refresh = () => setVoices(listSystemVoices());
refresh();
@@ -194,18 +218,20 @@ export function SettingsDrawer({ onClose }: Props) {
<input className="input" type="password" value={settings.hermes.apiKey} onChange={(e) => update({ hermes: { apiKey: e.target.value } })} />
</div>
<div className="grid-2">
<div className="field">
<label>Modèle</label>
<input className="input" list="hermes-models" value={settings.hermes.model} placeholder="(défaut du serveur)" onChange={(e) => update({ hermes: { model: e.target.value } })} />
<datalist id="hermes-models">
{hermesModels.map((m) => <option key={m.id} value={m.id} />)}
</datalist>
</div>
<ModelSelect label="Modèle IA" models={hermesModels} value={settings.hermes.model} defaultLabel="Défaut du serveur" onChange={(model) => update({ hermes: { model } })} />
<div className="field">
<label>Clé de mémoire (X-Hermes-Session-Key)</label>
<input className="input" value={settings.hermes.sessionKey} onChange={(e) => update({ hermes: { sessionKey: e.target.value } })} />
</div>
</div>
<div className="field">
<button className="btn small" disabled={modelsLoading} onClick={() => void refreshHermesModels(true)}>
{modelsLoading ? <Loader2 size={13} className="spin" /> : <RotateCcw size={13} />}
{modelsLoading ? 'Chargement des modèles…' : 'Actualiser les modèles'}
</button>
<span className="hint" role="status">{modelsLoading ? 'Interrogation du serveur Hermes…' : modelsError || (hermesModels.length ? `${hermesModels.length} modèle(s) disponible(s).` : 'Aucun modèle annoncé par le serveur.')}</span>
<span className={modelsNotice ? 'test-result fail' : 'hint'}>{modelsNotice || 'Modèles des fournisseurs configurés dans Hermes. Les modèles sans accès sont indiqués comme indisponibles.'}</span>
</div>
<div className="grid-2">
<div className="field">
<label>Transport</label>
@@ -227,8 +253,7 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
</div>
<div className="field">
<label>Modèle « mission » (tâches longues)</label>
<input className="input" list="hermes-models" value={settings.hermes.missionModel} placeholder="(même modèle)" onChange={(e) => update({ hermes: { missionModel: e.target.value } })} />
<ModelSelect label="Modèle mission (tâches longues)" models={hermesModels} value={settings.hermes.missionModel} defaultLabel="Même modèle que le principal" onChange={(missionModel) => update({ hermes: { missionModel } })} />
<span className="hint">Utilisé quand le mode Mission est activé dans la barre de commande. Laissez vide pour garder le modèle principal.</span>
</div>
<div className="field">
@@ -359,6 +384,12 @@ export function SettingsDrawer({ onClose }: Props) {
</>
)}
<Toggle on={settings.voice.localCommands} onChange={(v) => update({ voice: { localCommands: v } })} label="Commandes locales instantanées" hint="« Verrouille la session », « monte le son », « ouvre Spotify », « regarde mon écran »… exécutées sur ce PC sans passer par Hermes." />
<Toggle
on={settings.voice.micProcessing ?? true}
onChange={(v) => update({ voice: { micProcessing: v } })}
label="Traitement du micro par Chromium (écho, bruit, gain automatique)"
hint="Désactivez-le si les transcriptions sont approximatives avec un casque ou un bon micro : ces filtres déforment la voix avant la reconnaissance. Gardez-le activé avec des haut-parleurs (sinon la voix de l’assistant est réentendue par le micro)."
/>
<div className="row" style={{ marginTop: 10 }}>
<button className="btn small" onClick={() => void testStt()} disabled={settings.voice.provider === 'browser'}><Mic size={13} /> Tester la reconnaissance</button>
</div>
@@ -374,7 +405,11 @@ export function SettingsDrawer({ onClose }: Props) {
<button className={(settings.speech.voiceGender ?? 'male') === 'male' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'male' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Masculine</button>
<button className={settings.speech.voiceGender === 'female' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'female' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Féminine</button>
</div>
<span className="hint">S’applique à tous les moteurs : voix locale (Piper Tom ou Pierre pour le masculin, téléchargée automatiquement si besoin), voix OpenAI (onyx / nova), voix système Windows (Paul / Hortense). Kokoro n’a pas de voix française masculine.</span>
<span className="hint">
S’applique à tous les moteurs : Edge (Henri / Denise), voix locale (Supertonic 3, cinq voix de chaque genre, téléchargé automatiquement si besoin), API OpenAI (onyx / nova), voix système Windows (Paul / Hortense).
{settings.speech.provider === 'google-free' && ' Google Translate n’a qu’une voix féminine : ce choix n’a pas d’effet avec ce moteur.'}
{settings.speech.provider === 'local' && settings.speech.localModel === 'kokoro-v1' && ' Kokoro n’a qu’une voix française, féminine et avec accent : la voix masculine bascule sur Supertonic 3.'}
</span>
</div>
<div className="field">
<label>Timbre</label>
@@ -388,29 +423,62 @@ export function SettingsDrawer({ onClose }: Props) {
<button
className="btn small primary"
onClick={() => {
update({ speech: { provider: 'local', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
setTtsTest({ status: 'running', message: 'préparation de la voix JARVIS (téléchargement de Piper Tom si nécessaire)…' });
update({ speech: { provider: 'edge', edgeVoice: '', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
void ensurePreferredVoice().then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
}}
>
<Volume2 size={13} /> Préréglage voix JARVIS (français, masculine, locale)
<Volume2 size={13} /> Voix JARVIS en ligne (Edge Henri, masculine)
</button>
<button
className="btn small"
onClick={() => {
update({ speech: { provider: 'local', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
setTtsTest({ status: 'running', message: 'préparation de la voix JARVIS locale (téléchargement de Supertonic 3, 129 Mo, si nécessaire)…' });
void ensurePreferredVoice({ upgrade: true }).then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS locale prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
}}
>
<Volume2 size={13} /> Voix JARVIS hors ligne (Supertonic 3, masculine)
</button>
<span className="status-pill">
{settings.speech.provider === 'local'
? `voix active : ${voiceModels.find((m) => m.id === settings.speech.localModel)?.speakers?.find((s) => s.id === settings.speech.localSpeaker)?.name ?? settings.speech.localModel}`
: `moteur : ${settings.speech.provider}`}
: settings.speech.provider === 'edge'
? `voix active : ${resolveEdgeVoice(settings.speech.edgeVoice ?? '', settings.speech.language, settings.speech.voiceGender ?? 'male')}`
: `moteur : ${settings.speech.provider}`}
</span>
</div>
<div className="field">
<label>Moteur de synthèse</label>
<select className="select" value={settings.speech.provider} onChange={(e) => update({ speech: { provider: e.target.value as typeof settings.speech.provider } })}>
<option value="local">Local dans l’application (Kokoro / Piper via sherpa-onnx, hors ligne)</option>
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Kokoro, Piper, OpenAI, LocalAI…)</option>
<select
className="select"
value={settings.speech.provider}
onChange={(e) => {
update({ speech: { provider: e.target.value as typeof settings.speech.provider } });
void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m }));
}}
>
<option value="edge">Microsoft Edge (voix neuronales, gratuit, en ligne, sans clé) — recommandé</option>
<option value="local">Local dans l’application (Supertonic 3 / Kokoro / Piper via sherpa-onnx, hors ligne)</option>
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Qwen3-TTS, Kokoro, OpenAI, LocalAI…)</option>
<option value="system">Voix système Windows</option>
<option value="google-free">Google Translate (gratuit, en ligne)</option>
<option value="google-free">Google Translate (gratuit, en ligne, voix féminine uniquement)</option>
<option value="off">Désactivée</option>
</select>
</div>
{settings.speech.provider === 'edge' && (
<div className="field">
<label>Voix Edge</label>
<select className="select" value={settings.speech.edgeVoice ?? ''} onChange={(e) => update({ speech: { edgeVoice: e.target.value } })}>
<option value="">Automatique ({resolveEdgeVoice('', settings.speech.language, settings.speech.voiceGender ?? 'male')})</option>
{edgeVoices
.filter((v) => v.locale.toLowerCase().startsWith(settings.speech.language.toLowerCase().split('-')[0]))
.map((v) => (
<option key={v.shortName} value={v.shortName}>{v.name} · {v.locale} · {v.gender === 'm' ? 'homme' : 'femme'}</option>
))}
</select>
<span className="hint">Mêmes voix que la lecture à voix haute d’Edge : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR), plus les voix canadiennes, suisses et belges. Aucune donnée locale ; chaque phrase est synthétisée en ligne.</span>
</div>
)}
{settings.speech.provider === 'openai-compatible' && (
<>
<div className="field">
@@ -424,7 +492,7 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
<div className="field">
<label>Voix</label>
<input className="input" value={settings.speech.voice} placeholder="alloy, onyx, af_heart…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
<input className="input" value={settings.speech.voice} placeholder="onyx, nova, af_heart, Ryan (Qwen3-TTS)…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
</div>
<div className="field">
<label>Clé API</label>
@@ -444,7 +512,7 @@ export function SettingsDrawer({ onClose }: Props) {
{settings.speech.provider === 'local' && (
installedModels(voiceModels, 'tts').length === 0 ? (
<div className="field">
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Kokoro dans « Modèles locaux »</a>.</span>
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Supertonic 3 dans « Modèles locaux »</a>.</span>
</div>
) : (
<div className="grid-2">
@@ -452,7 +520,8 @@ export function SettingsDrawer({ onClose }: Props) {
<label>Modèle local</label>
<select className="select" value={settings.speech.localModel} onChange={(e) => {
const m = voiceModels.find((x) => x.id === e.target.value);
update({ speech: { localModel: e.target.value, localSpeaker: m?.speakers?.[0]?.id ?? 0 } });
// Keep the preferred gender when switching models (Kokoro has no masculine French voice: it falls back to Siwis).
update({ speech: { localModel: e.target.value, localSpeaker: m ? pickSpeaker(m, settings.speech.language, settings.speech.voiceGender ?? 'male') : 0 } });
}}>
{installedModels(voiceModels, 'tts').map((m) => <option key={m.id} value={m.id}>{m.name}</option>)}
</select>
@@ -488,6 +557,13 @@ export function SettingsDrawer({ onClose }: Props) {
</div>
</div>
<Toggle on={settings.speech.autoSpeak} onChange={(v) => update({ speech: { autoSpeak: v } })} label="Lire les réponses en streaming" hint="Chaque phrase est prononcée dès qu’elle est complète." />
<Toggle on={settings.speech.summarizeReplies} onChange={(v) => update({ speech: { summarizeReplies: v } })} label="Résumé vocal des réponses longues" hint="Seules les premières phrases et la question finale sont lues ; la réponse complète reste dans le fil." />
{settings.speech.summarizeReplies && (
<div className="field">
<label>Phrases lues : {settings.speech.replySentences}</label>
<input className="range" type="range" min={1} max={10} step={1} value={settings.speech.replySentences} onChange={(e) => update({ speech: { replySentences: Number(e.target.value) } })} />
</div>
)}
<Toggle on={settings.speech.speakIncoming} onChange={(v) => update({ speech: { speakIncoming: v } })} label="Lire les messages entrants (webhook, crons)" />
<div className="row" style={{ marginTop: 10 }}>
<button className="btn small" onClick={testTts} disabled={settings.speech.provider === 'off'}><Volume2 size={13} /> Tester la voix</button>
@@ -569,8 +645,9 @@ mcp_servers:
<div className="card">
<div className="grid-2">
<div className="field">
<label>Nom de l’assistant</label>
<input className="input" value={settings.assistantName} onChange={(e) => update({ assistantName: e.target.value.toUpperCase().slice(0, 18) || 'JARVIS' })} />
<label htmlFor="assistant-name">Nom de l’assistant</label>
<input id="assistant-name" className="input" value={settings.assistantName} maxLength={60} placeholder="Jarvis, Nova, Alfred…" onChange={(e) => update({ assistantName: e.target.value })} onBlur={() => update({ assistantName: settings.assistantName.trim() || DEFAULT_SETTINGS.assistantName })} />
<span className="hint">Choisissez le nom que vous voulez. Il est sauvegardé automatiquement et utilisé par l’assistant dès votre prochain message, y compris dans la conversation en cours.</span>
</div>
<div className="field">
<label>Votre nom</label>
+27
View File
@@ -129,6 +129,33 @@ export function chunkForSpeech(text: string, maxLength = 220): string[] {
return out;
}
/** Last speakable chunk of a reply when it is a question (« Tu veux que je… ? »), else null. */
export function closingQuestion(text: string): string | null {
const parts = speakableChunks(text);
const last = parts[parts.length - 1];
return last && /\?\s*$/.test(last) ? last : null;
}
/** Chunks of the spoken form of a text (markdown and code already stripped), letters only. */
function speakableChunks(text: string): string[] {
return chunkForSpeech(cleanForSpeech(text)).filter((part) => /[\p{L}\p{N}]/u.test(part));
}
/**
* Spoken digest of a long reply: the first `maxSentences` speakable chunks plus the closing
* question when there is one, so a long answer stays short to listen to but the conversation
* can go on. `truncated` tells the caller that the screen holds more than what is spoken.
*/
export function spokenDigest(text: string, maxSentences: number): { text: string; truncated: boolean } {
const parts = speakableChunks(text);
const limit = Math.max(1, Math.floor(maxSentences));
if (parts.length <= limit) return { text: parts.join(' '), truncated: false };
const head = parts.slice(0, limit);
const question = closingQuestion(text);
if (question && !head.includes(question)) head.push(question);
return { text: head.join(' '), truncated: true };
}
export function previewText(value: string, max = 96): string {
const flat = value.replace(/\s+/g, ' ').trim();
return flat.length > max ? `${flat.slice(0, max - 1)}…` : flat;
+55 -11
View File
@@ -1,13 +1,18 @@
/** Voice gender preference applied to every TTS provider (pure helpers, unit tested). */
import type { VoiceModelStatus } from '../../shared/voice';
import type { VoiceModelStatus, VoiceSpeaker } from '../../shared/voice';
import { defaultEdgeVoice, edgeVoiceGender } from '../../shared/edgeTts';
export type VoiceGender = 'male' | 'female';
const MALE_HINTS = /\b(homme|male|masculin|paul|thomas|claude|henri|guillaume|mathieu|antoine|nicolas|denis|pierre|tom|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|david|mark|richard|james|ryan|guy)\b/i;
const FEMALE_HINTS = /\b(femme|female|f[ée]minin|hortense|julie|denise|eloise|am[ée]lie|audrey|siwis|jessica|heart|bella|sarah|nicole|sky|alloy|nova|shimmer|zira|aria|jenny|emma|isabella|sophie|charlotte|vivienne|coral|sage)\b/i;
const MALE_HINTS =
/\b(homme|male|masculin|paul|thomas|claude|henri|remy|rémy|gerard|guillaume|mathieu|antoine|nicolas|denis|pierre|tom|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|david|mark|richard|james|ryan|guy|dylan|aiden|uncle fu|andrew|brian|fabrice|jean|thierry)\b/i;
const FEMALE_HINTS =
/\b(femme|female|f[ée]minin|hortense|julie|denise|eloise|vivienne|charline|sylvie|ariane|am[ée]lie|audrey|siwis|jessica|heart|bella|sarah|nicole|sky|alloy|nova|shimmer|zira|aria|jenny|emma|ava|isabella|sophie|charlotte|coral|sage|vivian|serena|sohee|ono anna)\b/i;
/** Best-effort gender from a voice or speaker label ("Piper Tom (homme, français)", "Microsoft Paul", "am_adam"). */
/** Best-effort gender from a voice or speaker label ("Piper Tom (homme, français)", "Microsoft Paul", "am_adam", "fr-FR-HenriNeural"). */
export function inferGender(name: string): VoiceGender | undefined {
const edge = edgeVoiceGender(name);
if (edge) return edge;
const n = name.replace(/_/g, ' ');
if (/^(am|bm|em|hm|im|jm|pm|zm)\b/i.test(n) || MALE_HINTS.test(n)) return 'male';
if (/^(af|bf|ef|ff|hf|if|jf|pf|zf)\b/i.test(n) || FEMALE_HINTS.test(n)) return 'female';
@@ -19,6 +24,13 @@ export function defaultOpenAiVoice(gender: VoiceGender): string {
return gender === 'male' ? 'onyx' : 'nova';
}
/** Edge voice to use: the explicit choice when it matches the gender, otherwise the language default. */
export function resolveEdgeVoice(explicit: string, language: string, gender: VoiceGender): string {
const chosen = explicit.trim();
if (chosen && (edgeVoiceGender(chosen) ?? gender) === gender) return chosen;
return defaultEdgeVoice(language, gender);
}
/** Rank system voices: language first, then gender, then quality hints. */
export function rankSystemVoice(v: { name: string; lang: string; localService?: boolean }, lang: string, gender: VoiceGender): number {
const name = v.name.toLowerCase();
@@ -38,16 +50,39 @@ export interface LocalVoiceChoice {
speaker: number;
}
/** Installed local voice matching language + gender, or null. */
/** Local models from best to worst French rendering; unknown models come last. */
const LOCAL_QUALITY: Record<string, number> = { 'supertonic-3': 0, 'kokoro-v1': 1, 'piper-fr-upmc': 2, 'piper-fr-tom': 3, 'piper-fr-siwis': 4 };
export function localQualityRank(modelId: string): number {
return LOCAL_QUALITY[modelId] ?? 9;
}
function speakerGender(sp: VoiceSpeaker): VoiceGender | undefined {
return sp.gender === 'm' ? 'male' : sp.gender === 'f' ? 'female' : inferGender(sp.name);
}
function speakerSpeaks(sp: VoiceSpeaker, lang: string): boolean {
const spLang = (sp.lang || '').toLowerCase();
return !spLang || spLang === lang || spLang === 'multi';
}
/** First speaker of a model matching language + gender; falls back to any speaker of the language, then the first one. */
export function pickSpeaker(model: Pick<VoiceModelStatus, 'speakers'>, lang: string, gender: VoiceGender): number {
const l = lang.toLowerCase().split('-')[0];
const speakers = model.speakers ?? [];
const exact = speakers.find((sp) => speakerSpeaks(sp, l) && speakerGender(sp) === gender);
const sameLang = speakers.find((sp) => speakerSpeaks(sp, l));
return (exact ?? sameLang ?? speakers[0])?.id ?? 0;
}
/** Installed local voice matching language + gender (best model first), or null. */
export function findLocalVoice(models: VoiceModelStatus[], lang: string, gender: VoiceGender): LocalVoiceChoice | null {
const l = lang.toLowerCase().split('-')[0];
const installed = models.filter((m) => m.kind === 'tts' && m.installed);
const installed = models.filter((m) => m.kind === 'tts' && m.installed).sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id));
for (const model of installed) {
for (const sp of model.speakers ?? []) {
const spLang = (sp.lang || '').toLowerCase();
if (spLang && spLang !== l && spLang !== 'multi') continue;
const g = (sp as { gender?: string }).gender === 'm' ? 'male' : (sp as { gender?: string }).gender === 'f' ? 'female' : inferGender(sp.name);
if (g === gender) return { modelId: model.id, speaker: sp.id };
if (!speakerSpeaks(sp, l)) continue;
if (speakerGender(sp) === gender) return { modelId: model.id, speaker: sp.id };
}
}
return null;
@@ -57,10 +92,19 @@ export function findLocalVoice(models: VoiceModelStatus[], lang: string, gender:
export function suggestedDownload(models: VoiceModelStatus[], lang: string, gender: VoiceGender): string | null {
const l = lang.toLowerCase().split('-')[0];
if (l !== 'fr') return null;
const wanted = gender === 'male' ? ['piper-fr-tom', 'piper-fr-upmc'] : ['kokoro-v1', 'piper-fr-siwis'];
const wanted = gender === 'male' ? ['supertonic-3', 'piper-fr-tom', 'piper-fr-upmc'] : ['supertonic-3', 'kokoro-v1', 'piper-fr-siwis'];
for (const id of wanted) {
const m = models.find((x) => x.id === id);
if (m && !m.installed) return id;
}
return null;
}
/** Best local model in the catalog for the language that is not installed yet (null when the best is already there). */
export function bestLocalUpgrade(models: VoiceModelStatus[], lang: string): string | null {
const l = lang.toLowerCase().split('-')[0];
const best = models
.filter((m) => m.kind === 'tts' && (m.languages.includes(l) || m.languages.includes('multi')))
.sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id))[0];
return best && !best.installed ? best.id : null;
}
+2 -2
View File
@@ -269,8 +269,8 @@ export async function sendMessage(text: string, images: string[] = [], source =
if (handle.aborted) {
speech.discardStream();
} else if (settings.speech.autoSpeak) {
if (!streamedText && finalText) speech.say(finalText);
else speech.endStream();
if (!streamedText && finalText) speech.sayReply(finalText);
else speech.endStream(finalText);
} else {
speech.discardStream();
}
+65 -5
View File
@@ -31,6 +31,13 @@ const isRec = (v: unknown): v is Rec => !!v && typeof v === 'object' && !Array.i
export type ResolvedTransport = Exclude<HermesTransport, 'auto'>;
/** Stored picker IDs include the provider so identical model names remain distinct. */
export function modelSelection(value: string): { model?: string; provider?: string } {
const id = value.trim();
const separator = id.indexOf('::');
return separator > 0 ? { provider: id.slice(0, separator), model: id.slice(separator + 2) } : id ? { model: id } : {};
}
/**
* A web page (login portal, dashboard, reverse-proxy error) instead of JSON means the URL does not
* point at the Hermes API. Returns a human explanation, or null when the body is not HTML.
@@ -200,7 +207,44 @@ export class HermesClient {
async models(): Promise<HermesModel[]> {
const payload = await this.request<unknown>('/v1/models');
return extractArray<HermesModel>(payload, ['data', 'models']);
const readEntries = (value: unknown): unknown[] => {
if (Array.isArray(value)) return value;
if (isRec(value)) {
if (Array.isArray(value.models)) return value.models;
if (value.data !== undefined) return readEntries(value.data);
}
throw new Error('Le serveur a renvoyé une liste de modèles invalide.');
};
const entries = readEntries(payload);
const models = new Map<string, HermesModel>();
for (const entry of entries) {
const id = typeof entry === 'string' ? entry.trim() : isRec(entry) && typeof entry.id === 'string' ? entry.id.trim() : '';
if (id && !models.has(id)) models.set(id, { ...(isRec(entry) ? entry : {}), id });
}
return [...models.values()];
}
async modelCatalog(refresh = false): Promise<{ models: HermesModel[]; notice: string | null }> {
let payload: unknown;
try {
payload = await this.request<unknown>(`/api/model/options${refresh ? '?refresh=true' : ''}`, { timeoutMs: 30_000 });
} catch (err) {
if (!(err instanceof HttpError) || ![404, 405].includes(err.status)) throw err;
return { models: await this.models(), notice: 'Ce serveur ne propose pas le catalogue IA (/api/model/options). Mettez Hermes à jour pour choisir le fournisseur et son modèle. La liste ci-dessous contient uniquement les alias de connexion.' };
}
if (!isRec(payload) || !Array.isArray(payload.providers)) throw new Error('Catalogue IA Hermes invalide : liste des fournisseurs absente.');
const models = new Map<string, HermesModel>();
for (const row of payload.providers) {
if (!isRec(row) || typeof row.slug !== 'string' || !row.slug.trim() || !Array.isArray(row.models)) continue;
const unavailable = Array.isArray(row.unavailable_models) ? row.unavailable_models : [];
for (const model of row.models) {
if (typeof model !== 'string' || !model.trim()) continue;
const id = `${row.slug}::${model}`;
models.set(id, { id, name: model, provider: typeof row.name === 'string' ? row.name : row.slug,
available: row.authenticated !== false && !unavailable.includes(model) });
}
}
return { models: [...models.values()], notice: null };
}
async skills(): Promise<HermesSkill[]> {
@@ -290,11 +334,12 @@ export class HermesClient {
// ── Runs ──────────────────────────────────────────────────────────────────
async startRun(body: { input: string; session_id?: string; instructions?: string; model?: string }): Promise<{ run_id: string; status: string }> {
async startRun(body: { input: string; session_id?: string; instructions?: string; model?: string; provider?: string }): Promise<{ run_id: string; status: string }> {
const payload: Rec = { input: body.input };
if (body.session_id) payload.session_id = body.session_id;
if (body.instructions) payload.instructions = body.instructions;
if (body.model) payload.model = body.model;
if (body.provider) payload.provider = body.provider;
return this.request<{ run_id: string; status: string }>('/v1/runs', { method: 'POST', body: payload, timeoutMs: 30_000 });
}
@@ -388,7 +433,7 @@ export class HermesClient {
input: options.text,
session_id: plainSession(options.sessionId) || undefined,
instructions: this.config.instructions || undefined,
model: this.config.model || undefined
...modelSelection(this.config.model)
});
const runId = run.run_id;
if (isAborted()) {
@@ -402,6 +447,7 @@ export class HermesClient {
let finalText = '';
let streamedText = '';
let completed = false;
let sessionSeen = false;
let failure: string | null = null;
const handle = await this.streamRunEvents(runId, (event) => {
@@ -410,6 +456,7 @@ export class HermesClient {
completed = true;
finalText = event.text ?? '';
}
if (event.kind === 'session') sessionSeen = true;
if (event.kind === 'error') failure = event.message;
onEvent(event);
});
@@ -420,6 +467,9 @@ export class HermesClient {
});
await handle.done;
// The run events never carry the session id; only the run status does. Without adopting it,
// every message would start a fresh Hermes session and the conversation would lose its context.
if (!sessionSeen) await this.adoptRunSession(runId, onEvent);
if (stopped || isAborted()) return streamedText;
if (failure) throw new Error(failure);
@@ -440,6 +490,16 @@ export class HermesClient {
return finalText || streamedText;
}
/** Reads the session Hermes actually attached to a run and reports it, so the next run continues it. */
private async adoptRunSession(runId: string, onEvent: SendOptions['onEvent']): Promise<void> {
try {
const info = await this.getRun(runId);
if (info.session_id) onEvent({ kind: 'session', sessionId: String(info.session_id) });
} catch (err) {
Log.warn('hermes', `run ${runId}: session id unavailable (${(err as Error).message})`);
}
}
private async sendViaSessions(options: SendOptions, setAbort: (fn: () => void) => void, isAborted: () => boolean): Promise<string> {
const { onEvent } = options;
let sessionId = options.sessionId;
@@ -452,7 +512,7 @@ export class HermesClient {
const realId = sessionId.slice(3);
const body: Rec = { input: options.text };
if (this.config.instructions) body.instructions = this.config.instructions;
if (this.config.model) body.model = this.config.model;
Object.assign(body, modelSelection(this.config.model));
let streamed = '';
let finalText = '';
@@ -489,7 +549,7 @@ export class HermesClient {
let fullText = '';
let toolsAllowed = useTools;
for (let iteration = 0; iteration < 6 && !aborted(); iteration++) {
const payload: Rec = { model: this.config.model || 'hermes-agent', messages, stream: true };
const payload: Rec = { model: 'hermes-agent', ...modelSelection(this.config.model), messages, stream: true };
if (toolsAllowed) {
payload.tools = options.localToolDefinitions;
payload.tool_choice = 'auto';
+2
View File
@@ -35,6 +35,8 @@ export interface HermesHealth {
export interface HermesModel {
id: string;
name?: string;
available?: boolean;
owned_by?: string;
provider?: string;
[key: string]: unknown;
+5 -3
View File
@@ -49,6 +49,8 @@ export interface CaptureCallbacks {
export interface CaptureOptions {
deviceId?: string;
/** Chromium's echo cancellation / noise suppression / auto gain (default on). Off keeps the raw signal for the recogniser. */
micProcessing?: boolean;
mode: CaptureMode;
vad?: Partial<VadOptions>;
callbacks: CaptureCallbacks;
@@ -124,9 +126,9 @@ export class MicCapture {
const constraints: MediaStreamConstraints = {
audio: {
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
echoCancellation: options.micProcessing !== false,
noiseSuppression: options.micProcessing !== false,
autoGainControl: options.micProcessing !== false,
channelCount: 1
}
};
+56 -7
View File
@@ -1,16 +1,23 @@
/**
* Singleton facade over the TTS engine bound to the settings store, exposing speaking state
* to the voice and chat stores.
* to the voice and chat stores. Long replies are spoken as a digest (first sentences plus the
* closing question) so listening stays short while the full text remains on screen.
*/
import { useChat } from '../../state/chat';
import { useSettings } from '../../state/settings';
import { useVoice } from '../../state/voice';
import { chunkForSpeech, closingQuestion, extractSentences, spokenDigest } from '../../lib/text';
import { TtsEngine } from './tts';
export const DIGEST_NOTICE = 'Le détail complet est affiché à l’écran.';
class SpeechFacade {
private engine: TtsEngine | null = null;
private streaming = false;
private streamEnabled = true;
private streamBuffer = '';
private spokenCount = 0;
private truncated = false;
private get tts(): TtsEngine {
if (!this.engine) {
@@ -35,32 +42,74 @@ class SpeechFacade {
return this.engine?.isActive ?? false;
}
/** Number of spoken chunks allowed for one reply; unbounded when the digest is off. */
private replyLimit(): number {
const { summarizeReplies, replySentences } = useSettings.getState().settings.speech;
return summarizeReplies ? Math.max(1, Math.floor(replySentences || 1)) : Number.POSITIVE_INFINITY;
}
say(text: string, options: { interrupt?: boolean } = {}): void {
if (useSettings.getState().settings.speech.provider === 'off') return;
// While an answer streams, spoken notices are inserted without discarding the rest.
this.tts.speak(text, { interrupt: options.interrupt ?? !this.streaming });
}
/** Speak a finished (non-streamed) assistant reply, digested when it is long. */
sayReply(text: string): void {
const limit = this.replyLimit();
if (!Number.isFinite(limit)) return this.say(text);
const digest = spokenDigest(text, limit);
this.say(digest.truncated ? `${digest.text} ${DIGEST_NOTICE}` : digest.text);
}
pushStream(delta: string): void {
if (!useSettings.getState().settings.speech.autoSpeak) return;
this.streaming = true;
if (this.streamEnabled) this.tts.pushStream(delta);
if (!this.streamEnabled) return;
this.streamBuffer += delta;
const { sentences, rest } = extractSentences(this.streamBuffer);
this.streamBuffer = rest;
for (const sentence of sentences) this.enqueueWithinLimit(sentence);
}
endStream(): void {
if (this.streaming) this.tts.endStream();
this.streaming = false;
private enqueueWithinLimit(text: string): void {
if (this.truncated) return;
if (this.spokenCount >= this.replyLimit()) {
this.truncated = true;
return;
}
if (this.tts.enqueue(text)) this.spokenCount++;
}
/** Flush a streamed reply; `fullText` lets a digested reply end with its closing question. */
endStream(fullText = ''): void {
if (!this.streaming) return;
const rest = this.streamBuffer.trim();
if (rest) for (const chunk of chunkForSpeech(rest)) this.enqueueWithinLimit(chunk);
if (this.truncated) {
const question = closingQuestion(fullText);
if (question) this.tts.enqueue(question);
this.tts.enqueue(DIGEST_NOTICE);
}
this.resetStream();
}
discardStream(): void {
this.streaming = false;
this.resetStream();
this.tts.stop();
}
stop(): void {
this.streaming = false;
this.resetStream();
this.engine?.stop();
}
private resetStream(): void {
this.streaming = false;
this.streamBuffer = '';
this.spokenCount = 0;
this.truncated = false;
}
}
export const speech = new SpeechFacade();
+25 -8
View File
@@ -1,16 +1,17 @@
/**
* Text-to-speech engine with a sentence queue, bounded prefetching and Web Audio playback so
* the HUD reacts to the actual waveform. Providers: in-app sherpa-onnx (Kokoro / Piper),
* OpenAI-compatible /v1/audio/speech, system voices, and the legacy Google Translate endpoint.
* the HUD reacts to the actual waveform. Providers: Microsoft Edge neural voices (online, free),
* in-app sherpa-onnx (Supertonic / Kokoro / Piper), OpenAI-compatible /v1/audio/speech, system
* voices, and the legacy Google Translate endpoint (one feminine voice per language).
*/
import { Log } from '../../lib/log';
import { defaultOpenAiVoice, rankSystemVoice } from '../../lib/voicePreference';
import { defaultOpenAiVoice, rankSystemVoice, resolveEdgeVoice } from '../../lib/voicePreference';
import { bridge } from '../../lib/bridge';
import { chunkForSpeech, cleanForSpeech, extractSentences } from '../../lib/text';
import { httpFetch } from '../../lib/transport';
import { audioBus } from './audioBus';
export type TtsProvider = 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
export type TtsProvider = 'edge' | 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
export interface TtsConfig {
provider: TtsProvider;
@@ -27,6 +28,8 @@ export interface TtsConfig {
localModel: string;
/** Speaker id inside the local model. */
localSpeaker: number;
/** Microsoft Edge voice short name (provider 'edge'); empty = automatic from language + gender. */
edgeVoice?: string;
/** Preferred voice gender, applied to every provider's default voice. */
voiceGender?: 'male' | 'female';
/** 'jarvis' adds a subtle AI timbre: slightly lower pitch, warm/crisp EQ, short room reverb. */
@@ -112,12 +115,14 @@ export class TtsEngine {
if (rest) for (const chunk of chunkForSpeech(rest)) this.enqueue(chunk);
}
enqueue(text: string): void {
/** Queue one chunk; false when nothing speakable remains once markdown and code are stripped. */
enqueue(text: string): boolean {
const item = this.makeItem(text);
if (!item) return;
if (!item) return false;
this.queue.push(item);
this.fillPrefetch();
void this.drain();
return true;
}
private makeItem(text: string): QueueItem | null {
@@ -179,8 +184,9 @@ export class TtsEngine {
}
private prefetch(text: string, gen: number): Promise<AudioBuffer | null> {
const provider = this.config.provider;
const task =
this.config.provider === 'google-free' ? this.fetchGoogle(text) : this.config.provider === 'local' ? this.fetchLocal(text) : this.fetchOpenAi(text);
provider === 'google-free' ? this.fetchGoogle(text) : provider === 'local' ? this.fetchLocal(text) : provider === 'edge' ? this.fetchEdge(text) : this.fetchOpenAi(text);
return task
.then(async (bytes) => {
if (gen !== this.generation || !bytes) return null;
@@ -228,12 +234,23 @@ export class TtsEngine {
modelId: this.config.localModel,
text,
speaker: this.config.localSpeaker,
speed: Math.max(0.5, Math.min(2, this.config.speed || 1))
speed: Math.max(0.5, Math.min(2, this.config.speed || 1)),
language: this.config.language || 'fr-FR'
});
Log.debug('tts', `local synthesis ${result.durationMs} ms for ${result.audioSec.toFixed(1)}s of audio`);
return result.wav;
}
/** Microsoft Edge neural voices through the main process (no key, online). */
private async fetchEdge(text: string): Promise<Uint8Array> {
const api = bridge();
if (!api) throw new Error('Les voix Edge nécessitent l’application Electron.');
const voice = resolveEdgeVoice(this.config.edgeVoice ?? '', this.config.language || 'fr-FR', this.config.voiceGender ?? 'male');
const result = await api.voice.edgeSynthesize({ text, voice, speed: Math.max(0.5, Math.min(2, this.config.speed || 1)) });
Log.debug('tts', `edge synthesis (${voice}) ${result.durationMs} ms, ${result.mp3.byteLength} bytes`);
return result.mp3;
}
private async fetchGoogle(text: string): Promise<Uint8Array> {
const url = `https://translate.google.com/translate_tts?ie=UTF-8&tl=${encodeURIComponent(this.config.language.split('-')[0] || 'fr')}&client=tw-ob&q=${encodeURIComponent(text.slice(0, 200))}`;
const res = await httpFetch({
+2
View File
@@ -128,6 +128,7 @@ class VoiceController {
}
await this.wake.start({
deviceId: settings.micDeviceId || undefined,
micProcessing: settings.micProcessing ?? true,
neuralVad,
vad: { silenceMs: settings.silenceMs, speechRatio: SENSITIVITY_RATIO[sensitivity], minRms: SENSITIVITY_MIN_RMS[sensitivity] },
callbacks: {
@@ -249,6 +250,7 @@ class VoiceController {
await audioBus.resume();
await this.capture.start({
deviceId: settings.micDeviceId || undefined,
micProcessing: settings.micProcessing ?? true,
mode: settings.captureMode,
vad: {
silenceMs: settings.silenceMs,
+32 -12
View File
@@ -1,44 +1,64 @@
/**
* Applies the preferred voice gender to the active TTS provider: switches the local model/speaker,
* downloads a matching French voice when none is installed, and logs what it did.
* downloads a matching French voice when none is installed, resets an Edge voice of the other
* gender, and explains providers that cannot honour the choice (Google Translate).
*/
import { Log } from '../../lib/log';
import { findLocalVoice, inferGender, suggestedDownload } from '../../lib/voicePreference';
import { bestLocalUpgrade, findLocalVoice, resolveEdgeVoice, suggestedDownload } from '../../lib/voicePreference';
import { useSettings } from '../../state/settings';
import { useVoiceModels } from '../../state/voiceModels';
let downloading: string | null = null;
/** Make the local voice match `speech.voiceGender`. Returns a short status message. */
export async function ensurePreferredVoice(): Promise<string> {
const genderLabel = (g: 'male' | 'female') => (g === 'male' ? 'masculine' : 'féminine');
/**
* Make the active provider match `speech.voiceGender`. Returns a short status message.
* With `upgrade`, the best local model for the language is downloaded even if a lesser voice exists.
*/
export async function ensurePreferredVoice(options: { upgrade?: boolean } = {}): Promise<string> {
const { settings, update } = useSettings.getState();
const { speech } = settings;
const gender = speech.voiceGender ?? 'male';
const lang = speech.language || 'fr-FR';
if (speech.provider === 'edge') {
const voice = resolveEdgeVoice(speech.edgeVoice ?? '', lang, gender);
if (voice !== (speech.edgeVoice ?? '').trim() && speech.edgeVoice) update({ speech: { edgeVoice: '' } });
return `Voix ${genderLabel(gender)} : ${voice.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') ?? voice} (Edge)`;
}
if (speech.provider === 'google-free') {
return gender === 'male'
? 'Google Translate n’a qu’une voix féminine par langue : choisissez « Microsoft Edge » ou une voix locale pour une voix masculine.'
: '';
}
if (speech.provider !== 'local') return '';
const models = useVoiceModels.getState().models;
if (!models.length) return '';
const lang = speech.language || 'fr-FR';
const current = models.find((m) => m.id === speech.localModel);
const currentSpeaker = current?.speakers?.find((s) => s.id === speech.localSpeaker);
const currentGender = currentSpeaker ? (currentSpeaker as { gender?: string }).gender === 'm' ? 'male' : (currentSpeaker as { gender?: string }).gender === 'f' ? 'female' : inferGender(currentSpeaker.name) : undefined;
if (current?.installed && currentGender === gender) return '';
const currentGender = currentSpeaker ? (currentSpeaker.gender === 'm' ? 'male' : currentSpeaker.gender === 'f' ? 'female' : undefined) : undefined;
const upgrade = options.upgrade ? bestLocalUpgrade(models, lang) : null;
if (current?.installed && currentGender === gender && !upgrade) return '';
const choice = findLocalVoice(models, lang, gender);
const choice = upgrade ? null : findLocalVoice(models, lang, gender);
if (choice) {
update({ speech: { localModel: choice.modelId, localSpeaker: choice.speaker } });
const name = models.find((m) => m.id === choice.modelId)?.speakers?.find((s) => s.id === choice.speaker)?.name ?? choice.modelId;
Log.info('tts', `voice preference ${gender}: ${name}`);
return `Voix ${gender === 'male' ? 'masculine' : 'féminine'} : ${name}`;
return `Voix ${genderLabel(gender)} : ${name}`;
}
const download = suggestedDownload(models, lang, gender);
const download = upgrade ?? suggestedDownload(models, lang, gender);
if (!download || downloading === download) return download ? 'Téléchargement de la voix en cours…' : '';
downloading = download;
Log.info('tts', `no ${gender} voice installed, downloading ${download}`);
Log.info('tts', `${upgrade ? 'upgrading local voice' : `no ${gender} voice installed`}, downloading ${download}`);
try {
await useVoiceModels.getState().download(download);
const after = findLocalVoice(useVoiceModels.getState().models, lang, gender);
if (after) update({ speech: { localModel: after.modelId, localSpeaker: after.speaker } });
return after ? `Voix ${gender === 'male' ? 'masculine' : 'féminine'} installée.` : 'Voix téléchargée.';
const name = after ? useVoiceModels.getState().models.find((m) => m.id === after.modelId)?.name : undefined;
return after ? `Voix ${genderLabel(gender)} installée : ${name ?? after.modelId}.` : 'Voix téléchargée.';
} catch (err) {
Log.warn('tts', `voice download failed: ${(err as Error).message}`);
return `Téléchargement impossible : ${(err as Error).message}`;
+4 -4
View File
@@ -74,7 +74,7 @@ export class WakeListener {
return this.phase;
}
async start(options: { deviceId?: string; vad?: Partial<VadOptions>; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise<void> {
async start(options: { deviceId?: string; micProcessing?: boolean; vad?: Partial<VadOptions>; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise<void> {
if (this.phase !== 'off') return;
this.callbacks = options.callbacks;
this.vadOptions = options.vad ?? {};
@@ -86,9 +86,9 @@ export class WakeListener {
this.stream = await navigator.mediaDevices.getUserMedia({
audio: {
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
echoCancellation: true,
noiseSuppression: true,
autoGainControl: true,
echoCancellation: options.micProcessing !== false,
noiseSuppression: options.micProcessing !== false,
autoGainControl: options.micProcessing !== false,
channelCount: 1
}
});
+35 -7
View File
@@ -33,6 +33,10 @@ interface HermesStore {
capabilities: HermesCapabilities | null;
health: HermesHealth | null;
models: HermesModel[];
modelsLoading: boolean;
modelsError: string | null;
modelsNotice: string | null;
refreshModels: (refresh?: boolean) => Promise<void>;
skills: HermesSkill[];
toolsets: HermesToolset[];
sessions: HermesSession[];
@@ -63,6 +67,7 @@ interface HermesStore {
const CACHE_KEY = 'eveflow.hermes.cache.v2';
let connectInflight: Promise<void> | null = null;
let lastCacheSnapshot = '';
let modelsRequest = 0;
const isTerminal = (status: string) => ['ok', 'failed', 'delivery_failed', 'completed', 'error'].includes(status.toLowerCase());
function runFromJob(job: HermesJob): JobRun | null {
@@ -87,6 +92,9 @@ export const useHermes = create<HermesStore>((set, get) => ({
capabilities: null,
health: null,
models: [],
modelsLoading: false,
modelsError: null,
modelsNotice: null,
skills: [],
toolsets: [],
sessions: [],
@@ -100,10 +108,12 @@ export const useHermes = create<HermesStore>((set, get) => ({
busy: false,
client: (modelOverride) => {
const config = useSettings.getState().settings.hermes;
// Without an explicit model, use the alias advertised by /v1/models (Hermes rejects unknown names).
const model = (modelOverride ?? '').trim() || config.model.trim() || get().models[0]?.id || '';
return new HermesClient({ ...config, model });
const settings = useSettings.getState().settings;
const config = settings.hermes;
const model = (modelOverride ?? '').trim() || config.model.trim();
const name = settings.assistantName.trim() || 'JARVIS';
const identity = `Identité de l'assistant dans cette conversation : ton nom est ${JSON.stringify(name)}. Utilise ce nom pour te présenter et parler de toi. EveFlow est le nom de l'application, pas ton nom. Cette identité remplace les anciens noms ou personas présents dans l'historique, la mémoire ou les instructions précédentes. Ne rappelle pas ton nom dans chaque réponse.`;
return new HermesClient({ ...config, model, instructions: [config.instructions.trim(), identity].filter(Boolean).join('\n\n') });
},
connect: () => {
@@ -160,14 +170,32 @@ export const useHermes = create<HermesStore>((set, get) => ({
return connectInflight;
},
refreshModels: async (refresh = false) => {
const request = ++modelsRequest;
const config = useSettings.getState().settings.hermes;
const current = () => {
const now = useSettings.getState().settings.hermes;
return request === modelsRequest && now.url === config.url && now.apiKey === config.apiKey && now.sessionKey === config.sessionKey;
};
set({ models: [], modelsLoading: true, modelsError: null, modelsNotice: null });
try {
const { models, notice } = await new HermesClient(config).modelCatalog(refresh);
if (current()) set({ models, modelsNotice: notice });
} catch (err) {
if (current()) set({ modelsError: (err as Error).message });
} finally {
if (request === modelsRequest) set({ modelsLoading: false });
}
},
refreshCatalog: async () => {
const client = get().client();
const [models, skills, toolsets] = await Promise.all([
client.models().catch(() => [] as HermesModel[]),
const [, skills, toolsets] = await Promise.all([
get().refreshModels(),
client.skills().catch(() => [] as HermesSkill[]),
client.toolsets().catch(() => [] as HermesToolset[])
]);
set({ models, skills, toolsets });
set({ skills, toolsets });
},
refreshSessions: async () => {
+15 -5
View File
@@ -26,11 +26,16 @@ export interface VoiceSettings extends SttConfig {
neuralVad: boolean;
/** Execute short system intents locally (lock, volume, open app…) instead of asking Hermes. */
localCommands: boolean;
/** Chromium mic processing (echo cancellation, noise suppression, auto gain). Off often transcribes better on a headset. */
micProcessing: boolean;
}
export interface SpeechSettings extends TtsConfig {
autoSpeak: boolean;
speakIncoming: boolean;
/** Speak only the first sentences (and the closing question) of long replies; the full text stays on screen. */
summarizeReplies: boolean;
replySentences: number;
}
export interface WebhookSettings {
@@ -87,7 +92,7 @@ export const DEFAULT_SETTINGS: Settings = {
transport: 'auto',
reasoningEffort: '',
instructions:
"Tu es l'interface vocale EveFlow (style JARVIS). Réponds en français, de façon concise et orale quand la question est simple; utilise le Markdown uniquement pour le contenu structuré (code, listes, tableaux). Les images doivent être des URL http(s) ou des fichiers du dossier partagé.",
"Réponds en français, de façon concise et orale quand la question est simple; si la réponse est longue, commence par une ou deux phrases qui en donnent l'essentiel, puis le détail. Tu es dans une conversation continue : tiens compte des échanges précédents sans redemander ce qui a déjà été dit. Utilise le Markdown uniquement pour le contenu structuré (code, listes, tableaux). Les images doivent être des URL http(s) ou des fichiers du dossier partagé.",
localTools: true,
missionModel: ''
},
@@ -110,10 +115,12 @@ export const DEFAULT_SETTINGS: Settings = {
kwsSensitivity: 3,
neuralVad: true,
localCommands: true,
micProcessing: true,
localModel: 'whisper-base'
},
speech: {
provider: 'openai-compatible',
// Edge neural voices speak out of the box (no server, no key) with a real masculine/feminine choice.
provider: 'edge',
apiUrl: 'http://127.0.0.1:8000/v1',
apiKey: '',
model: 'tts-1',
@@ -123,12 +130,15 @@ export const DEFAULT_SETTINGS: Settings = {
systemVoice: '',
language: 'fr-FR',
volume: 1,
localModel: 'kokoro-v1',
localSpeaker: 30,
localModel: 'supertonic-3',
localSpeaker: 6,
edgeVoice: '',
voiceGender: 'male',
timbre: 'jarvis',
autoSpeak: true,
speakIncoming: true
speakIncoming: true,
summarizeReplies: true,
replySentences: 4
},
webhook: { enabled: true, port: 7842, secret: '' },
notifications: { quietEnabled: false, quietStart: '22:30', quietEnd: '07:30', priorityKeywords: 'urgent, alerte, alarme, panne', summarizeIncoming: true, summarySentences: 2, nightTheme: true },
+90
View File
@@ -0,0 +1,90 @@
import { describe, expect, it } from 'vitest';
import {
defaultEdgeVoice,
edgeConfigMessage,
edgeRate,
edgeSsml,
edgeSsmlMessage,
edgeTextFramePath,
edgeTokenInput,
edgeVoiceGender,
escapeXml,
parseEdgeBinaryFrame
} from '../shared/edgeTts';
describe('edge token input', () => {
it('uses Windows file time rounded down to 5 minutes followed by the client token', () => {
// 2026-09-04T15:23:47Z → 15:20:00Z = 1788535200 s since 1970 → +11644473600 = 13433008800 s → ×1e7 ticks.
const nowMs = Date.UTC(2026, 8, 4, 15, 23, 47);
expect(edgeTokenInput(nowMs)).toBe('1343300880000000006A5AA1D4EAFF4E9FB37E23D68491D6F4');
});
it('is stable inside a 5-minute window and applies the clock skew', () => {
const a = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1));
const b = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 24, 59));
expect(a).toBe(b);
expect(edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1), 300)).not.toBe(a);
});
});
describe('edge ssml', () => {
it('escapes text and derives the language from the voice', () => {
const ssml = edgeSsml('Tom & Jerry <3 "ok"', 'fr-FR-HenriNeural', 1.25);
expect(ssml).toContain("xml:lang='fr-FR'");
expect(ssml).toContain("<voice name='fr-FR-HenriNeural'>");
expect(ssml).toContain("rate='+25%'");
expect(ssml).toContain('Tom &amp; Jerry &lt;3 &quot;ok&quot;');
expect(escapeXml("l'été")).toBe('l&apos;été');
});
it('formats rates with a sign', () => {
expect(edgeRate(1)).toBe('+0%');
expect(edgeRate(0.8)).toBe('-20%');
expect(edgeRate(3)).toBe('+100%');
});
it('builds the config and ssml messages with the expected headers', () => {
const date = new Date(Date.UTC(2026, 8, 4, 15, 23, 47));
const config = edgeConfigMessage(date);
expect(config.startsWith('X-Timestamp:Fri, 04 Sep 2026 15:23:47 GMT+0000 (Coordinated Universal Time)\r\n')).toBe(true);
expect(config).toContain('Path:speech.config\r\n\r\n{');
expect(config).toContain('audio-24khz-48kbitrate-mono-mp3');
const ssml = edgeSsmlMessage('abc123', '<speak/>', date);
expect(ssml).toContain('X-RequestId:abc123\r\n');
expect(ssml).toContain('Content-Type:application/ssml+xml\r\n');
expect(ssml.endsWith('Path:ssml\r\n\r\n<speak/>')).toBe(true);
expect(edgeTextFramePath('X-RequestId:1\r\nContent-Type:application/json\r\nPath:turn.end\r\n\r\n{}')).toBe('turn.end');
});
});
describe('edge binary frames', () => {
it('splits the header (2-byte big-endian length) from the audio payload', () => {
const header = 'X-RequestId:1\r\nContent-Type:audio/mpeg\r\nX-StreamId:2\r\nPath:audio\r\n';
const payload = new Uint8Array([0xff, 0xfb, 0x90, 0x00]);
const frame = new Uint8Array(2 + header.length + payload.length);
frame[0] = header.length >> 8;
frame[1] = header.length & 0xff;
for (let i = 0; i < header.length; i++) frame[2 + i] = header.charCodeAt(i);
frame.set(payload, 2 + header.length);
const parsed = parseEdgeBinaryFrame(frame);
expect(parsed.path).toBe('audio');
expect([...parsed.payload]).toEqual([...payload]);
});
it('tolerates truncated frames', () => {
expect(parseEdgeBinaryFrame(new Uint8Array([0x00])).payload.byteLength).toBe(0);
expect(parseEdgeBinaryFrame(new Uint8Array([0x10, 0x00, 0x41])).path).toBe('');
});
});
describe('edge voices', () => {
it('picks Henri / Denise for French and falls back to French for unknown languages', () => {
expect(defaultEdgeVoice('fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(defaultEdgeVoice('fr', 'female')).toBe('fr-FR-DeniseNeural');
expect(defaultEdgeVoice('en-GB', 'male')).toBe('en-US-AndrewMultilingualNeural');
expect(defaultEdgeVoice('xx', 'female')).toBe('fr-FR-DeniseNeural');
});
it('knows the gender of the common French voices', () => {
expect(edgeVoiceGender('fr-FR-HenriNeural')).toBe('male');
expect(edgeVoiceGender('fr-FR-RemyMultilingualNeural')).toBe('male');
expect(edgeVoiceGender('fr-FR-VivienneMultilingualNeural')).toBe('female');
expect(edgeVoiceGender('fr-CA-SylvieNeural')).toBe('female');
expect(edgeVoiceGender('zz-ZZ-NobodyNeural')).toBeUndefined();
});
});
+122
View File
@@ -0,0 +1,122 @@
import { afterEach, describe, expect, it, vi } from 'vitest';
import { HermesClient } from '../src/services/hermes/client';
import { httpFetch, httpStream } from '../src/lib/transport';
import { DEFAULT_SETTINGS, useSettings } from '../src/state/settings';
import { useHermes } from '../src/state/hermes';
import { createElement, act } from 'react';
import { createRoot } from 'react-dom/client';
import { ModelSelect } from '../src/components/settings/ModelSelect';
vi.mock('../src/lib/transport', async (original) => ({ ...await original<typeof import('../src/lib/transport')>(), httpFetch: vi.fn(), httpStream: vi.fn() }));
const config = { ...DEFAULT_SETTINGS.hermes, url: 'https://example.test/v1/', apiKey: ' test-key ' };
function respond(payload: unknown, status = 200) {
vi.mocked(httpFetch).mockResolvedValue({ ok: status === 200, status, statusText: '', headers: {}, text: JSON.stringify(payload) });
}
afterEach(() => { vi.restoreAllMocks(); useSettings.setState({ settings: DEFAULT_SETTINGS }); });
describe('Hermes models', () => {
it.each(['runs', 'sessions', 'completions'] as const)('sends the configured assistant identity over %s, including after a rename', async (transport) => {
const requests: Record<string, unknown>[] = [];
const capture = async (req: { body?: string }) => { requests.push(JSON.parse(req.body!)); throw new Error('stop'); };
vi.mocked(httpFetch).mockImplementation(capture);
vi.mocked(httpStream).mockImplementation(capture);
for (const name of ['Jarvis', 'Nova']) {
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, assistantName: name, hermes: { ...config, instructions: 'Tu es Eve. Réponds en français.' } } });
const send = useHermes.getState().client().send({ text: 'Qui es-tu ?', sessionId: 'hs:test', history: [], onEvent: vi.fn() }, transport);
await expect(send.result).rejects.toThrow('stop');
const body = requests.at(-1)!;
const instructions = transport === 'completions' ? (body.messages as { content: string }[])[0].content : body.instructions as string;
expect(instructions).toContain(`ton nom est "${name}"`);
expect(instructions).toContain('Réponds en français.');
expect(instructions).toContain("EveFlow est le nom de l'application, pas ton nom");
}
expect(useSettings.getState().settings.hermes.instructions).toBe('Tu es Eve. Réponds en français.');
});
it('reads the real provider inventory, keeps providers distinct and marks inaccessible models', async () => {
respond({ providers: [
{ slug: 'first', name: 'First', authenticated: true, models: ['same', 'locked'], unavailable_models: ['locked'] },
{ slug: 'second', name: 'Second', authenticated: false, models: ['same'] }
] });
const result = await new HermesClient(config).modelCatalog(true);
expect(httpFetch).toHaveBeenCalledWith(expect.objectContaining({ url: 'https://example.test/api/model/options?refresh=true' }));
expect(result.models.map(m => [m.id, m.available])).toEqual([['first::same', true], ['first::locked', false], ['second::same', false]]);
expect(result.notice).toBeNull();
});
it('explains old servers instead of presenting the connection alias as an LLM catalog', async () => {
vi.mocked(httpFetch).mockResolvedValueOnce({ ok: false, status: 404, statusText: '', headers: {}, text: '' })
.mockResolvedValueOnce({ ok: true, status: 200, statusText: '', headers: {}, text: JSON.stringify({ data: [{ id: 'hermes-agent' }] }) });
const result = await new HermesClient(config).modelCatalog();
expect(result.models).toEqual([{ id: 'hermes-agent' }]);
expect(result.notice).toContain('Mettez Hermes à jour');
});
it.each(['runs', 'sessions', 'completions'] as const)('sends provider and actual model separately over %s', async (transport) => {
const requests: unknown[] = [];
vi.mocked(httpFetch).mockImplementation(async req => {
requests.push(JSON.parse(req.body as string));
throw new Error('stop');
});
vi.mocked(httpStream).mockImplementation(async req => {
requests.push(JSON.parse(req.body as string));
throw new Error('stop');
});
const send = new HermesClient({ ...config, model: 'custom:local::real-model' }).send({ text: 'Bonjour', sessionId: 'hs:test', history: [], onEvent: vi.fn() }, transport);
await expect(send.result).rejects.toThrow('stop');
expect(requests[0]).toMatchObject({ model: 'real-model', provider: 'custom:local' });
});
it('uses authenticated discovery and preserves provider labels, order and unique valid IDs', async () => {
respond({ data: [{ id: 'b', provider: 'provider-b' }, { id: 'a' }, { id: 'b' }, {}, ' c '] });
expect(await new HermesClient(config).models()).toEqual([{ id: 'b', provider: 'provider-b' }, { id: 'a' }, { id: 'c' }]);
expect(httpFetch).toHaveBeenCalledWith(expect.objectContaining({ url: 'https://example.test/v1/models', headers: expect.objectContaining({ Authorization: 'Bearer test-key' }) }));
});
it('accepts nested catalogs and empty lists, rejects malformed responses', async () => {
respond({ data: { models: ['a'] } });
expect(await new HermesClient(config).models()).toEqual([{ id: 'a' }]);
respond({ data: [] });
expect(await new HermesClient(config).models()).toEqual([]);
respond({ data: { error: 'unavailable' } });
await expect(new HermesClient(config).models()).rejects.toThrow('invalide');
});
it('surfaces authentication failures without clearing the selected model', async () => {
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, hermes: { ...config, model: 'chosen' } } });
respond({}, 401);
await useHermes.getState().refreshModels();
expect(useHermes.getState().modelsError).toContain('401');
expect(useHermes.getState().modelsLoading).toBe(false);
expect(useSettings.getState().settings.hermes.model).toBe('chosen');
});
it('discards results from a previous server or an older refresh', async () => {
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, hermes: config } });
let finish!: (value: { models: { id: string }[]; notice: null }) => void;
vi.spyOn(HermesClient.prototype, 'modelCatalog').mockImplementationOnce(() => new Promise(resolve => { finish = resolve; })).mockResolvedValue({ models: [{ id: 'new' }], notice: null });
const old = useHermes.getState().refreshModels();
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, hermes: { ...config, url: 'https://new.test' } } });
await useHermes.getState().refreshModels();
finish({ models: [{ id: 'old' }], notice: null });
await old;
expect(useHermes.getState().models).toEqual([{ id: 'new' }]);
});
it('sends the selected principal model and the mission override on new runs', async () => {
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, hermes: { ...config, model: 'principal', missionModel: 'mission' } } });
const start = vi.spyOn(HermesClient.prototype, 'startRun').mockRejectedValue(new Error('stop before network'));
for (const [override, expected] of [[undefined, 'principal'], ['mission', 'mission']] as const) {
const send = useHermes.getState().client(override).send({ text: 'Bonjour', sessionId: 'test-session', history: [], onEvent: vi.fn() }, 'runs');
await expect(send.result).rejects.toThrow('stop before network');
expect(start).toHaveBeenLastCalledWith(expect.objectContaining({ model: expected }));
}
});
it('offers a visible selection, preserves a custom value and supports the default', async () => {
Object.assign(globalThis, { IS_REACT_ACT_ENVIRONMENT: true });
const container = document.createElement('div');
const root = createRoot(container);
const onChange = vi.fn();
try {
await act(async () => root.render(createElement(ModelSelect, { label: 'Modèle IA', value: 'custom', models: [{ id: 'available' }], defaultLabel: 'Défaut', onChange })));
const select = container.querySelector('select')!;
expect(select.value).toBe('custom');
await act(async () => { select.value = 'available'; select.dispatchEvent(new Event('change', { bubbles: true })); });
expect(onChange).toHaveBeenLastCalledWith('available');
await act(async () => { select.value = ''; select.dispatchEvent(new Event('change', { bubbles: true })); });
expect(onChange).toHaveBeenLastCalledWith('');
} finally { await act(async () => root.unmount()); }
});
});
+57
View File
@@ -0,0 +1,57 @@
import { afterEach, describe, expect, it, vi } from 'vitest';
import { HermesClient } from '../src/services/hermes/client';
import { httpFetch, httpStream } from '../src/lib/transport';
import { DEFAULT_SETTINGS } from '../src/state/settings';
import type { HermesStreamEvent } from '../src/services/hermes/types';
vi.mock('../src/lib/transport', async (original) => ({ ...await original<typeof import('../src/lib/transport')>(), httpFetch: vi.fn(), httpStream: vi.fn() }));
const config = { ...DEFAULT_SETTINGS.hermes, url: 'https://example.test' };
afterEach(() => vi.restoreAllMocks());
/** Run started, streamed to completion over SSE (which never carries the session id), status read afterwards. */
function mockRun(sessionId: string) {
const bodies: Array<{ url: string; body: unknown }> = [];
vi.mocked(httpFetch).mockImplementation(async (req) => {
bodies.push({ url: req.url, body: req.body ? JSON.parse(req.body as string) : undefined });
if (req.url.endsWith('/v1/runs')) return { ok: true, status: 200, statusText: '', headers: {}, text: JSON.stringify({ run_id: 'run_1', status: 'started' }) };
return { ok: true, status: 200, statusText: '', headers: {}, text: JSON.stringify({ run_id: 'run_1', status: 'completed', session_id: sessionId, output: 'Salut.' }) };
});
vi.mocked(httpStream).mockImplementation(async (_req, handlers) => {
// Chunks arrive on a later tick, after the client has accepted the stream (as over IPC).
const done = new Promise<void>((resolve) => setTimeout(() => {
handlers.onChunk('event: message.delta\ndata: {"run_id":"run_1","delta":"Salut."}\n\n');
handlers.onChunk('event: run.completed\ndata: {"run_id":"run_1","status":"completed","output":"Salut."}\n\n');
resolve();
}, 0));
return { start: { id: 'stream-1', ok: true, status: 200, statusText: '', headers: {} }, done, abort: () => undefined };
});
return bodies;
}
describe('Hermes runs session continuity', () => {
it('adopts the session Hermes attached to the run so the next message continues it', async () => {
const bodies = mockRun('api_abc');
const events: HermesStreamEvent[] = [];
const client = new HermesClient(config);
const reply = await client.send({ text: 'Bonjour', sessionId: 'eveflow-local', history: [], onEvent: (e) => events.push(e) }, 'runs').result;
expect(reply).toBe('Salut.');
expect(events.map((e) => e.kind)).toEqual(['run.started', 'delta', 'completed', 'session']);
expect(events[3]).toEqual({ kind: 'session', sessionId: 'api_abc' });
expect(bodies.map((b) => b.url)).toEqual(['https://example.test/v1/runs', 'https://example.test/v1/runs/run_1']);
expect(bodies[0].body).toEqual(expect.objectContaining({ input: 'Bonjour', session_id: 'eveflow-local' }));
const again = await client.send({ text: 'Et ensuite ?', sessionId: 'api_abc', history: [], onEvent: () => undefined }, 'runs').result;
expect(again).toBe('Salut.');
expect(bodies[2].body).toEqual(expect.objectContaining({ session_id: 'api_abc' }));
});
it('does not fail the reply when the run status is unavailable', async () => {
mockRun('api_abc');
vi.mocked(httpFetch).mockImplementation(async (req) =>
req.url.endsWith('/v1/runs')
? { ok: true, status: 200, statusText: '', headers: {}, text: JSON.stringify({ run_id: 'run_1', status: 'started' }) }
: { ok: false, status: 500, statusText: '', headers: {}, text: 'boom' });
const events: HermesStreamEvent[] = [];
const reply = await new HermesClient(config).send({ text: 'Bonjour', sessionId: 'x', history: [], onEvent: (e) => events.push(e) }, 'runs').result;
expect(reply).toBe('Salut.');
expect(events.some((e) => e.kind === 'session')).toBe(false);
});
});
+59
View File
@@ -0,0 +1,59 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import { DEFAULT_SETTINGS, useSettings } from '../src/state/settings';
const enqueue = vi.fn((text: string) => /[\p{L}\p{N}]/u.test(text.replace(/```[\s\S]*?```/g, '')));
const speak = vi.fn();
const stop = vi.fn();
vi.mock('../src/services/voice/tts', () => ({
TtsEngine: class {
enqueue = enqueue;
speak = speak;
stop = stop;
onState() { return () => undefined; }
updateConfig() { /* noop */ }
get isActive() { return false; }
}
}));
const { speech, DIGEST_NOTICE } = await import('../src/services/voice/speech');
function stream(text: string, size = 7): void {
for (let i = 0; i < text.length; i += size) speech.pushStream(text.slice(i, i + size));
}
const spoken = () => enqueue.mock.calls.map((c) => c[0]);
describe('spoken digest of streamed replies', () => {
beforeEach(() => {
enqueue.mockClear();
speak.mockClear();
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, speech: { ...DEFAULT_SETTINGS.speech, summarizeReplies: true, replySentences: 2 } } });
});
afterEach(() => speech.stop());
it('stops after the configured sentences, then adds the closing question and a notice', () => {
const text = 'Première phrase du rapport. Deuxième phrase utile. Troisième phrase de détail. Quatrième phrase encore. Tu veux la suite ?';
stream(text);
speech.endStream(text);
expect(spoken()).toEqual(['Première phrase du rapport.', 'Deuxième phrase utile.', 'Tu veux la suite ?', DIGEST_NOTICE]);
});
it('speaks short replies in full without any notice', () => {
const text = 'Une phrase courte. Une seconde phrase';
stream(text);
speech.endStream(text);
expect(spoken()).toEqual(['Une phrase courte.', 'Une seconde phrase']);
});
it('speaks everything when the digest is disabled', () => {
useSettings.setState({ settings: { ...DEFAULT_SETTINGS, speech: { ...DEFAULT_SETTINGS.speech, summarizeReplies: false } } });
const text = 'Première phrase du rapport. Deuxième phrase utile. Troisième phrase de détail. Quatrième phrase encore.';
stream(text);
speech.endStream(text);
expect(spoken()).toHaveLength(4);
});
it('digests a reply delivered in one piece', () => {
speech.sayReply('Première phrase du rapport. Deuxième phrase utile. Troisième phrase de détail. On continue ?');
expect(speak).toHaveBeenCalledWith(`Première phrase du rapport. Deuxième phrase utile. On continue ? ${DIGEST_NOTICE}`, expect.anything());
});
});
+19 -1
View File
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { chunkForSpeech, cleanForSpeech, extractSentences, isTranscriptNoise, preprocessMedia } from '../src/lib/text';
import { chunkForSpeech, cleanForSpeech, closingQuestion, extractSentences, isTranscriptNoise, preprocessMedia, spokenDigest } from '../src/lib/text';
describe('text', () => {
it('cleans markdown for speech', () => {
@@ -40,3 +40,21 @@ describe('isTranscriptNoise', () => {
expect(isTranscriptNoise('Oui')).toBe(false);
});
});
describe('spokenDigest', () => {
const reply = 'Voici la réponse. Elle contient plusieurs phrases. Une troisième pour la forme. Et une quatrième.\n\n```js\nconsole.log(1)\n```\n\nTu veux que je continue ?';
it('speaks the first sentences and keeps the closing question', () => {
const digest = spokenDigest(reply, 2);
expect(digest.truncated).toBe(true);
expect(digest.text).toBe('Voici la réponse. Elle contient plusieurs phrases. Tu veux que je continue ?');
});
it('leaves short replies untouched', () => {
const digest = spokenDigest('Bonjour Michael. Tout va bien.', 4);
expect(digest).toEqual({ text: 'Bonjour Michael. Tout va bien.', truncated: false });
});
it('ignores code blocks when counting and only reports a real question', () => {
expect(spokenDigest('Une phrase.\n\n```\ncode\n```\n', 1).truncated).toBe(false);
expect(closingQuestion(reply)).toBe('Tu veux que je continue ?');
expect(closingQuestion('Aucune question ici.')).toBeNull();
});
});
+40 -9
View File
@@ -1,36 +1,61 @@
import { describe, expect, it } from 'vitest';
import { defaultOpenAiVoice, findLocalVoice, inferGender, rankSystemVoice, suggestedDownload } from '../src/lib/voicePreference';
import type { VoiceModelStatus } from '../shared/voice';
import { bestLocalUpgrade, defaultOpenAiVoice, findLocalVoice, inferGender, pickSpeaker, rankSystemVoice, resolveEdgeVoice, suggestedDownload } from '../src/lib/voicePreference';
import type { VoiceModelStatus, VoiceSpeaker } from '../shared/voice';
const model = (id: string, installed: boolean, speakers: Array<{ id: number; name: string; lang: string }>): VoiceModelStatus =>
({ id, kind: 'tts', engine: 'piper', name: id, description: '', languages: ['fr'], sizeMb: 1, url: '', dir: id, files: [], speakers, installed, installedBytes: 0 }) as unknown as VoiceModelStatus;
const model = (id: string, installed: boolean, speakers: VoiceSpeaker[], languages = ['fr']): VoiceModelStatus =>
({ id, kind: 'tts', engine: 'piper', name: id, description: '', languages, sizeMb: 1, url: '', dir: id, files: [], speakers, installed, installedBytes: 0 }) as unknown as VoiceModelStatus;
describe('inferGender', () => {
it('reads catalog labels, system voices and kokoro ids', () => {
it('reads catalog labels, system voices, kokoro ids and Edge short names', () => {
expect(inferGender('Piper Tom (homme, français)')).toBe('male');
expect(inferGender('Siwis (femme, français)')).toBe('female');
expect(inferGender('Microsoft Paul - French (France)')).toBe('male');
expect(inferGender('Microsoft Hortense - French (France)')).toBe('female');
expect(inferGender('am_adam')).toBe('male');
expect(inferGender('fr-FR-HenriNeural')).toBe('male');
expect(inferGender('fr-FR-DeniseNeural')).toBe('female');
expect(inferGender('Voix 3')).toBeUndefined();
});
});
describe('local voice selection', () => {
const supertonic: VoiceSpeaker[] = [
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' }
];
const models = [
model('kokoro-v1', true, [{ id: 30, name: 'Siwis (femme, français)', lang: 'fr' }, { id: 4, name: 'Adam (homme, anglais US)', lang: 'en' }]),
model('kokoro-v1', true, [{ id: 30, name: 'Siwis (femme, français)', lang: 'fr' }, { id: 4, name: 'Adam (homme, anglais US)', lang: 'en' }], ['en', 'fr', 'multi']),
model('piper-fr-tom', false, [{ id: 0, name: 'Tom', lang: 'fr' }]),
model('piper-fr-upmc', true, [{ id: 0, name: 'Jessica (femme)', lang: 'fr' }, { id: 1, name: 'Pierre (homme)', lang: 'fr' }])
model('piper-fr-upmc', true, [{ id: 0, name: 'Jessica (femme)', lang: 'fr' }, { id: 1, name: 'Pierre (homme)', lang: 'fr' }]),
model('supertonic-3', false, supertonic, ['fr', 'en', 'multi'])
];
it('prefers an installed speaker of the wanted gender in the right language', () => {
expect(findLocalVoice(models, 'fr-FR', 'male')).toEqual({ modelId: 'piper-fr-upmc', speaker: 1 });
expect(findLocalVoice(models, 'fr-FR', 'female')).toEqual({ modelId: 'kokoro-v1', speaker: 30 });
expect(findLocalVoice(models.slice(0, 2), 'fr', 'male')).toBeNull();
});
it('suggests a French download when nothing matches', () => {
expect(suggestedDownload(models, 'fr', 'male')).toBe('piper-fr-tom');
it('ranks Supertonic above Kokoro and Piper once installed', () => {
const installed = models.map((m) => (m.id === 'supertonic-3' ? { ...m, installed: true } : m));
expect(findLocalVoice(installed, 'fr-FR', 'male')).toEqual({ modelId: 'supertonic-3', speaker: 6 });
expect(findLocalVoice(installed, 'fr-FR', 'female')).toEqual({ modelId: 'supertonic-3', speaker: 0 });
});
it('suggests Supertonic first, then the Piper voices, for a French download', () => {
expect(suggestedDownload(models, 'fr', 'male')).toBe('supertonic-3');
expect(suggestedDownload(models.filter((m) => m.id !== 'supertonic-3'), 'fr', 'male')).toBe('piper-fr-tom');
expect(suggestedDownload(models, 'en', 'male')).toBeNull();
});
it('reports the best local model still to download', () => {
expect(bestLocalUpgrade(models, 'fr-FR')).toBe('supertonic-3');
expect(bestLocalUpgrade(models.map((m) => ({ ...m, installed: true })), 'fr-FR')).toBeNull();
});
it('keeps the gender when switching models and falls back to the language', () => {
expect(pickSpeaker(models[3], 'fr-FR', 'male')).toBe(6);
expect(pickSpeaker(models[3], 'fr-FR', 'female')).toBe(0);
expect(pickSpeaker(models[2], 'fr', 'male')).toBe(1);
// Kokoro: no masculine French voice → the French voice, not an English one.
expect(pickSpeaker(models[0], 'fr', 'male')).toBe(30);
expect(pickSpeaker({ speakers: [] }, 'fr', 'male')).toBe(0);
});
});
describe('provider defaults', () => {
@@ -38,6 +63,12 @@ describe('provider defaults', () => {
expect(defaultOpenAiVoice('male')).toBe('onyx');
expect(defaultOpenAiVoice('female')).toBe('nova');
});
it('keeps an explicit Edge voice only when it matches the gender', () => {
expect(resolveEdgeVoice('', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(resolveEdgeVoice('fr-FR-RemyMultilingualNeural', 'fr-FR', 'male')).toBe('fr-FR-RemyMultilingualNeural');
expect(resolveEdgeVoice('fr-FR-DeniseNeural', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
expect(resolveEdgeVoice('fr-CH-ArianeNeural', 'fr-FR', 'female')).toBe('fr-CH-ArianeNeural');
});
it('ranks system voices by language then gender', () => {
const paul = { name: 'Microsoft Paul - French (France)', lang: 'fr-FR', localService: true };
const hortense = { name: 'Microsoft Hortense - French (France)', lang: 'fr-FR', localService: true };