mirror of
https://github.com/R0m1k3/EveFlow.git
synced 2026-10-11 17:29:03 +02:00
feat: voix Edge neuronales, Supertonic 3 en local et genre respecté par tous les moteurs (v2.5.0)
- Nouveau moteur « Microsoft Edge » (voix neuronales gratuites, sans clé) : Henri / Denise par défaut selon le genre, liste des voix fr-FR / fr-CA / fr-CH / fr-BE, WebSocket signé (Sec-MS-GEC) dans le processus principal, MP3 24 kHz. Moteur par défaut des nouvelles installations. - Supertonic 3 ajouté au catalogue local (31 langues, 5 voix masculines + 5 féminines, 44 kHz, 129 Mo) : langue transmise au worker, genres des voix vérifiés par mesure de F0. - Kokoro déclassé pour le français (une seule voix féminine, accent) ; le choix masculin/féminin bascule sur le meilleur modèle installé et conserve le genre quand on change de modèle ; Google Translate signalé comme voix féminine uniquement. - Deux préréglages JARVIS : en ligne (Edge Henri) et hors ligne (Supertonic 3). - Traitement du micro par Chromium (écho, bruit, gain) débrayable pour de meilleures transcriptions au casque. - Tests : protocole Edge (jeton, SSML, trames), sélection de voix ; README et feuille de route (mesures Supertonic, Parakeet vs Qwen3-ASR). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MMpgFriwxiBgurUVb21oCE
This commit is contained in:
23 files changed
+735
-81
No files matched your search
@@ -1,7 +1,7 @@
|
||||
# EveFlow 2 — Interface vocale JARVIS pour Hermes Agent
|
||||
|
||||
[](https://github.com/R0m1k3/EveFlow/actions)
|
||||
[](https://github.com/R0m1k3/EveFlow/releases)
|
||||
[](https://github.com/R0m1k3/EveFlow/releases)
|
||||
[](LICENSE)
|
||||
|
||||
**EveFlow** est un compagnon de bureau Windows qui transforme [Hermes Agent](https://hermes-agent.nousresearch.com/) en assistant vocal à la JARVIS : un noyau holographique réactif au son, une conversation en streaming, les outils, sous-agents, approbations, crons, skills et sessions d'Hermes pilotés depuis un seul HUD.
|
||||
@@ -22,7 +22,7 @@ La version 2 est une réécriture complète : plus de robot 3D, un pipeline voca
|
||||
* **Capture micro** via AudioWorklet à 16 kHz, sans monitoring du micro dans les haut-parleurs, avec annulation d'écho et réduction de bruit.
|
||||
* **Détection d'activité vocale** (seuil adaptatif, sensibilité et silence de fin réglables) : l'enregistrement s'arrête tout seul quand vous avez fini de parler.
|
||||
* **Mains libres** : le micro se réactive après chaque réponse.
|
||||
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse Kokoro v1.0 (voix française Siwis et voix anglaises) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
|
||||
* **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Parakeet v3, Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse **Supertonic 3** (31 langues dont le français, cinq voix masculines et cinq féminines, 44 kHz, 129 Mo, environ 7× plus rapide que le temps réel sur 4 cœurs), Kokoro v1.0 (excellent en anglais ; en français une seule voix féminine avec accent) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur.
|
||||
* **Fin de phrase neuronale** : en écoute permanente, Silero VAD (0,6 Mo, sherpa-onnx) décide du début et de la fin de la commande à la place du seuil d'énergie ; moins de faux départs sur le bruit, coupure plus nette. Repli automatique sur le VAD énergétique si le modèle n'est pas installé.
|
||||
* **Vision d'écran** : « Jarvis, regarde mon écran » (ou le bouton de la barre de commande) joint une capture de l'écran principal à la question envoyée à Hermes.
|
||||
* **Actions locales instantanées** : « verrouille la session », « monte le son », « coupe le son », « piste suivante », « ouvre Spotify », « ouvre github.com »… exécutées sur le PC sans passer par Hermes, résultat lu à voix haute. Liste blanche d'actions dans le processus principal, désactivable dans les paramètres.
|
||||
@@ -30,13 +30,15 @@ La version 2 est une réécriture complète : plus de robot 3D, un pipeline voca
|
||||
* **Heures calmes et priorités** : plage horaire pendant laquelle les messages poussés s'affichent sans être lus ni faire clignoter le noyau (badge « non lus » à la place), thème nuit automatique, mots prioritaires lus quand même, résumé vocal des rapports longs (les premières phrases seulement).
|
||||
* **Mode mission** : un bouton dans la barre de commande bascule sur un second modèle Hermes (plus puissant) pour les tâches longues ; le modèle rapide reste utilisé pour la conversation courante.
|
||||
* **Widget compact « glanceable »** : état (veille, écoute, réflexion, parle), dernière phrase de l'assistant, badge de non-lus, indicateurs heures calmes et mission.
|
||||
* **Voix JARVIS** : préréglage en un clic (Paramètres → Voix) : voix française masculine locale (Piper Tom, téléchargée automatiquement), timbre « JARVIS » (légèrement plus grave et posé, chaleur, présence, courte réverbération d'intercom), débit calme. Kokoro n'a pas de voix française masculine.
|
||||
* **Voix Microsoft Edge** (moteur par défaut) : les voix neuronales de la lecture à voix haute d'Edge, gratuites, sans clé ni installation : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR) et les voix fr-CA, fr-CH, fr-BE, plus de 300 voix dans 74 langues. Le rendu le plus naturel disponible ; nécessite une connexion.
|
||||
* **Voix JARVIS** : deux préréglages en un clic (Paramètres → Voix) : en ligne (Edge Henri) ou hors ligne (Supertonic 3, voix masculine grave, téléchargé automatiquement), timbre « JARVIS » (légèrement plus grave et posé, chaleur, présence, courte réverbération d'intercom), débit calme.
|
||||
* **Reconnaissance française de référence** : Parakeet TDT 0.6B v3 (NVIDIA NeMo, 25 langues européennes) dans le catalogue, plus précis et bien plus rapide que Whisper sur processeur, avec ponctuation. Whisper base/small/turbo restent disponibles.
|
||||
* **Voix masculine ou féminine** : un réglage unique (Paramètres → Voix) appliqué à tous les moteurs. En local, Piper Tom ou Pierre (UPMC) pour le masculin, téléchargé automatiquement si aucune voix masculine n'est installée ; onyx / nova pour les API compatibles OpenAI ; Paul / Hortense pour les voix Windows.
|
||||
* **Voix masculine ou féminine** : un réglage unique (Paramètres → Voix) appliqué à tous les moteurs. Henri / Denise sur Edge ; en local Supertonic 3 (ou Piper Tom / Pierre) pour le masculin, téléchargé automatiquement si aucune voix masculine n'est installée, et le genre est conservé quand on change de modèle ; onyx / nova pour les API compatibles OpenAI ; Paul / Hortense pour les voix Windows. Google Translate n'a qu'une voix féminine par langue : le réglage l'indique.
|
||||
* **Barge-in** : en mains libres, parler par-dessus l'assistant coupe sa voix ; le seuil est relevé pendant qu'il parle pour ignorer l'écho du haut-parleur.
|
||||
* **Écoute permanente** : un détecteur de mot-clé de 3 Mo (sherpa-onnx, keyword spotting) tourne en continu sur le micro, quasi gratuit en CPU. « Jarvis » (ou n'importe quel mot-clé) ouvre l'écoute, « Jarvis, allume… » envoie directement la commande, et le mot coupe la voix en cours. Alternative : filtre du mot après transcription en mains libres.
|
||||
* **STT externe** : n'importe quelle API `/v1/audio/transcriptions` compatible OpenAI (Qwen3-ASR, Whisper, Speaches, faster-whisper-server, LocalAI, OpenAI). Repli sur la reconnaissance Chromium.
|
||||
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI, voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
|
||||
* **TTS externe** : API `/v1/audio/speech` compatible OpenAI (Qwen3-TTS via un serveur compatible, Kokoro-FastAPI, OpenAI, LocalAI…), voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée.
|
||||
* **Traitement du micro débrayable** : l'annulation d'écho, la réduction de bruit et le gain automatique de Chromium peuvent être coupés (Paramètres → Reconnaissance) ; avec un casque, le signal brut est souvent mieux transcrit.
|
||||
* Raccourcis globaux : `Ctrl+Shift+Espace` (micro), `Ctrl+Shift+J` (afficher/masquer), `Ctrl+Shift+Échap` (couper la voix).
|
||||
|
||||
### Hermes, toute la puissance
|
||||
|
||||
+4
-1
@@ -8,7 +8,8 @@
|
||||
|---|---|---|
|
||||
| HUD arc-reactor réactif au son | Fait | Canvas 2D optimisé (pas d'ombres, couleurs en cache, 30 fps en veille, arrêt fenêtre masquée) |
|
||||
| Reconnaissance vocale locale | Fait | Whisper base/small/turbo via sherpa-onnx dans un processus utilitaire |
|
||||
| Synthèse vocale locale | Fait | Kokoro v1.0 (voix française Siwis) et Piper fr |
|
||||
| Synthèse vocale locale | Fait (2.5.0) | Supertonic 3 (31 langues, 5 voix masculines + 5 féminines, 44 kHz), Kokoro v1.0 (anglais ; français féminin avec accent) et Piper fr |
|
||||
| Synthèse vocale en ligne | Fait (2.5.0) | Voix neuronales Microsoft Edge (Henri, Denise, Rémy, Vivienne…), gratuites, sans clé, via WebSocket signé dans le processus principal |
|
||||
| Mot d'activation permanent | Fait (2.2.0) | Keyword spotting sherpa-onnx en continu ; mot-clé libre encodé en BPE ; validé sur audio réel (détection, zéro faux positif sur le test anglais) |
|
||||
| Mot d'activation après transcription | Fait | Filtre « Jarvis … » en mains libres, tolérant aux erreurs de transcription |
|
||||
| Détection de fin de phrase | Fait (2.3.0) | Silero VAD neuronal dans le worker (segment renvoyé au renderer), repli sur le VAD énergétique si le modèle manque |
|
||||
@@ -83,6 +84,8 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai)
|
||||
| Correctifs 2.4.0.1 | réponse vide en chat completions désormais expliquée (JSON non streamé ou erreur HTTP 200), la liaison ne passe plus en « dégradé » quand seule l'API des crons échoue, transcriptions parasites (« (cliquant) », « *Claire* ») ignorées, préférence de voix masculine/féminine |
|
||||
| Parakeet v3 vs Whisper base sur trois phrases Piper (fr) (2.4.1) | Parakeet : 3/3 exactes avec ponctuation, 0,4 à 0,5 s à chaud (6,7 s au premier appel) ; Whisper base : erreurs sur « Jarvis », « Peux-tu », 0,7 à 0,9 s |
|
||||
| Worklets audio sous CSP stricte (2.4.1) | chargement des modules statiques OK dans l'application empaquetée |
|
||||
| Voix françaises (2.5.0) | Edge Henri : MP3 reçu de bout en bout (32 ko pour 4 s). Supertonic 3 en français : 10 voix, RTF 0,14 sur 4 cœurs (8,7 s d'audio en 1,3 s), genres déterminés par mesure de la fréquence fondamentale (voix 0-4 : 170-210 Hz, voix 5-9 : 92-137 Hz) ; le worker compilé accepte `language` et bascule sur l'anglais pour une langue inconnue |
|
||||
| Reconnaissance française : Parakeet v3 vs Qwen3-ASR 0.6B int8 (2.5.0) | Six phrases Supertonic (3,8 s) : Parakeet WER 8,5 % (erreurs surtout de forme : « 14h30 »), 384 ms par phrase ; Qwen3-ASR WER 15,3 % (« mémoires vivres »), 1 477 ms, 940 Mo. Qwen3-ASR n'est pas ajouté au catalogue |
|
||||
| Serveur MCP (2.4.0) | initialize, tools/list (14 outils), tools/call côté principal (presse-papiers, capture image) et côté renderer (état, message dans le fil) |
|
||||
|
||||
Sur un PC à 28 cœurs les temps sont nettement plus courts. Whisper small est maintenant recommandé pour le français.
|
||||
+3
-1
@@ -15,7 +15,7 @@ import {
|
||||
} from '../shared/ipc';
|
||||
import type { EveFlowBridge, SystemAction, SystemActionResult, Unsubscribe } from '../shared/bridge';
|
||||
import type { McpToolRequest, McpToolResponse } from '../shared/ipc';
|
||||
import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
|
||||
import { VOICE_IPC, type EdgeSynthesizeRequest, type EdgeSynthesizeResult, type EdgeVoice, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice';
|
||||
|
||||
function subscribe<T>(channel: string, callback: (payload: T) => void): Unsubscribe {
|
||||
const listener = (_event: Electron.IpcRendererEvent, payload: T) => callback(payload);
|
||||
@@ -74,6 +74,8 @@ const api: EveFlowBridge = {
|
||||
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => subscribe<VoiceDownloadProgress>(VOICE_IPC.modelsProgress, cb),
|
||||
transcribe: (req: TranscribeRequest) => ipcRenderer.invoke(VOICE_IPC.transcribe, req) as Promise<TranscribeResult>,
|
||||
synthesize: (req: SynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.synthesize, req) as Promise<SynthesizeResult>,
|
||||
edgeSynthesize: (req: EdgeSynthesizeRequest) => ipcRenderer.invoke(VOICE_IPC.edgeSynthesize, req) as Promise<EdgeSynthesizeResult>,
|
||||
edgeVoices: () => ipcRenderer.invoke(VOICE_IPC.edgeVoices) as Promise<EdgeVoice[]>,
|
||||
unload: (id?: string) => ipcRenderer.invoke(VOICE_IPC.unload, id) as Promise<unknown>,
|
||||
kwsStart: (req: KwsStartRequest) => ipcRenderer.invoke(VOICE_IPC.kwsStart, req) as Promise<{ accepted: string[]; rejected: string[] }>,
|
||||
kwsStop: () => ipcRenderer.invoke(VOICE_IPC.kwsStop) as Promise<void>,
|
||||
|
||||
@@ -17,6 +17,22 @@ const KOKORO_SPEAKERS: VoiceSpeaker[] = [
|
||||
{ id: 26, name: 'George (homme, anglais UK)', lang: 'en' }
|
||||
];
|
||||
|
||||
// Supertonic 3 ships ten voice styles in voice.bin (five feminine, five masculine). The ordering
|
||||
// was checked by measuring the fundamental frequency of French synthesis: sid 0-4 around
|
||||
// 170-210 Hz, sid 5-9 around 90-140 Hz.
|
||||
const SUPERTONIC_SPEAKERS: VoiceSpeaker[] = [
|
||||
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
|
||||
{ id: 9, name: 'Homme 5 (grave)', lang: 'multi', gender: 'm' },
|
||||
{ id: 7, name: 'Homme 3', lang: 'multi', gender: 'm' },
|
||||
{ id: 8, name: 'Homme 4', lang: 'multi', gender: 'm' },
|
||||
{ id: 5, name: 'Homme 1 (clair)', lang: 'multi', gender: 'm' },
|
||||
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' },
|
||||
{ id: 1, name: 'Femme 2', lang: 'multi', gender: 'f' },
|
||||
{ id: 2, name: 'Femme 3', lang: 'multi', gender: 'f' },
|
||||
{ id: 3, name: 'Femme 4', lang: 'multi', gender: 'f' },
|
||||
{ id: 4, name: 'Femme 5', lang: 'multi', gender: 'f' }
|
||||
];
|
||||
|
||||
export const VOICE_CATALOG: VoiceModelSpec[] = [
|
||||
{
|
||||
id: 'whisper-base',
|
||||
@@ -112,20 +128,36 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [
|
||||
files: ['silero_vad.onnx'],
|
||||
recommended: true
|
||||
},
|
||||
{
|
||||
id: 'supertonic-3',
|
||||
kind: 'tts',
|
||||
engine: 'supertonic',
|
||||
name: 'Supertonic 3 (31 langues, 5 voix masculines et 5 féminines)',
|
||||
description:
|
||||
'La meilleure voix française locale : naturelle, sans accent, dix voix au choix, 44 kHz. Environ 7 fois plus rapide que le temps réel sur 4 cœurs, 100 M de paramètres.',
|
||||
languages: ['fr', 'en', 'de', 'es', 'it', 'pt', 'multi'],
|
||||
sizeMb: 129,
|
||||
url: `${TTS}/sherpa-onnx-supertonic-3-tts-int8-2026-05-11.tar.bz2`,
|
||||
dir: 'sherpa-onnx-supertonic-3-tts-int8-2026-05-11',
|
||||
files: ['duration_predictor.int8.onnx', 'text_encoder.int8.onnx', 'vector_estimator.int8.onnx', 'vocoder.int8.onnx', 'tts.json', 'unicode_indexer.bin', 'voice.bin'],
|
||||
speakers: SUPERTONIC_SPEAKERS,
|
||||
sampleRate: 44100,
|
||||
recommended: true
|
||||
},
|
||||
{
|
||||
id: 'kokoro-v1',
|
||||
kind: 'tts',
|
||||
engine: 'kokoro',
|
||||
name: 'Kokoro v1.0 multilingue',
|
||||
description: 'Voix très naturelle, une voix française (Siwis) et de nombreuses voix anglaises. 24 kHz.',
|
||||
languages: ['fr', 'en', 'multi'],
|
||||
description:
|
||||
'Excellent en anglais (nombreuses voix). En français : une seule voix, féminine (Siwis), avec un accent marqué (phonémisation espeak). Préférez Supertonic 3 pour le français. 24 kHz.',
|
||||
languages: ['en', 'fr', 'multi'],
|
||||
sizeMb: 349,
|
||||
url: `${TTS}/kokoro-multi-lang-v1_0.tar.bz2`,
|
||||
dir: 'kokoro-multi-lang-v1_0',
|
||||
files: ['model.onnx', 'voices.bin', 'tokens.txt', 'lexicon-us-en.txt', 'lexicon-zh.txt', 'espeak-ng-data/phontab'],
|
||||
speakers: KOKORO_SPEAKERS,
|
||||
sampleRate: 24000,
|
||||
recommended: true
|
||||
sampleRate: 24000
|
||||
},
|
||||
{
|
||||
id: 'piper-fr-siwis',
|
||||
@@ -181,6 +213,9 @@ export function findModel(id: string): VoiceModelSpec | undefined {
|
||||
// Speaker gender from the label ("(homme, …)", "(femme, …)", known first names) so the UI and the
|
||||
// voice preference can pick a masculine or feminine voice without a lookup table per model.
|
||||
const MALE = /\b(homme|tom|pierre|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|santa)\b/i;
|
||||
|
||||
/** Languages accepted by the Supertonic 3 text front-end (2-letter codes). */
|
||||
export const SUPERTONIC_LANGS = new Set(['ar', 'bg', 'hr', 'cs', 'da', 'nl', 'en', 'et', 'fi', 'fr', 'de', 'el', 'hi', 'hu', 'id', 'it', 'ja', 'ko', 'lv', 'lt', 'pl', 'pt', 'ro', 'ru', 'sk', 'sl', 'es', 'sv', 'tr', 'uk', 'vi']);
|
||||
for (const spec of VOICE_CATALOG) {
|
||||
for (const sp of spec.speakers ?? []) {
|
||||
if (!sp.gender) sp.gender = /\b(femme|female)\b/i.test(sp.name) ? 'f' : MALE.test(sp.name) ? 'm' : /\bfemme\b/i.test(spec.name) ? 'f' : /\bhomme\b/i.test(spec.name) ? 'm' : undefined;
|
||||
|
||||
@@ -0,0 +1,146 @@
|
||||
/**
|
||||
* Microsoft Edge "Read aloud" neural voices (the service behind the Edge browser's read-aloud
|
||||
* feature): free, no key, very natural French voices with a real masculine/feminine choice
|
||||
* (Henri, Denise, Rémy, Vivienne…). Runs in the main process: one WebSocket per sentence,
|
||||
* MP3 back to the renderer.
|
||||
*/
|
||||
import { createHash, randomBytes, randomUUID } from 'node:crypto';
|
||||
import type { EdgeSynthesizeRequest, EdgeSynthesizeResult, EdgeVoice } from '../../shared/voice';
|
||||
import {
|
||||
EDGE_CHROMIUM_VERSION,
|
||||
EDGE_VOICES_URL,
|
||||
EDGE_WSS_URL,
|
||||
edgeConfigMessage,
|
||||
edgeConnectionId,
|
||||
edgeHeaders,
|
||||
edgeSsml,
|
||||
edgeSsmlMessage,
|
||||
edgeTextFramePath,
|
||||
edgeTokenInput,
|
||||
parseEdgeBinaryFrame
|
||||
} from '../../shared/edgeTts';
|
||||
import { log } from '../logger';
|
||||
|
||||
/** Seconds to add to the local clock so the signed token matches the server's time window. */
|
||||
let clockSkewSec = 0;
|
||||
let voicesCache: { at: number; voices: EdgeVoice[] } | null = null;
|
||||
const VOICES_TTL_MS = 6 * 60 * 60 * 1000;
|
||||
const SYNTH_TIMEOUT_MS = 20_000;
|
||||
|
||||
function token(): string {
|
||||
return createHash('sha256').update(edgeTokenInput(Date.now(), clockSkewSec), 'ascii').digest('hex').toUpperCase();
|
||||
}
|
||||
|
||||
function signedUrl(): string {
|
||||
return `${EDGE_WSS_URL}&ConnectionId=${edgeConnectionId(randomUUID())}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
|
||||
}
|
||||
|
||||
function headers(): Record<string, string> {
|
||||
return { ...edgeHeaders(), Cookie: `muid=${randomBytes(16).toString('hex')};` };
|
||||
}
|
||||
|
||||
/** Learn the server clock from a plain HTTPS response (the WebSocket handshake hides its headers). */
|
||||
async function syncClock(): Promise<void> {
|
||||
try {
|
||||
const res = await fetch(EDGE_VOICES_URL, { method: 'HEAD', headers: headers() });
|
||||
const date = res.headers.get('date');
|
||||
if (!date) return;
|
||||
const server = Date.parse(date);
|
||||
if (Number.isFinite(server)) {
|
||||
clockSkewSec = (server - Date.now()) / 1000;
|
||||
log('INFO', 'edge-tts', `clock skew ${clockSkewSec.toFixed(0)} s`);
|
||||
}
|
||||
} catch (err) {
|
||||
log('WARN', 'edge-tts', `clock sync failed: ${(err as Error).message}`);
|
||||
}
|
||||
}
|
||||
|
||||
function synthesizeOnce(req: EdgeSynthesizeRequest): Promise<Uint8Array> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const chunks: Uint8Array[] = [];
|
||||
let settled = false;
|
||||
let ws: WebSocket;
|
||||
try {
|
||||
// Node's global WebSocket accepts extra handshake headers (undici), which the service checks.
|
||||
ws = new (WebSocket as unknown as new (url: string, options: { headers: Record<string, string> }) => WebSocket)(signedUrl(), { headers: headers() });
|
||||
} catch (err) {
|
||||
reject(err as Error);
|
||||
return;
|
||||
}
|
||||
ws.binaryType = 'arraybuffer';
|
||||
const finish = (err?: Error) => {
|
||||
if (settled) return;
|
||||
settled = true;
|
||||
clearTimeout(timer);
|
||||
try {
|
||||
ws.close();
|
||||
} catch {
|
||||
/* already closed */
|
||||
}
|
||||
if (err) reject(err);
|
||||
else {
|
||||
const total = chunks.reduce((n, c) => n + c.byteLength, 0);
|
||||
const out = new Uint8Array(total);
|
||||
let o = 0;
|
||||
for (const c of chunks) {
|
||||
out.set(c, o);
|
||||
o += c.byteLength;
|
||||
}
|
||||
resolve(out);
|
||||
}
|
||||
};
|
||||
const timer = setTimeout(() => finish(new Error('Edge TTS : délai dépassé')), SYNTH_TIMEOUT_MS);
|
||||
ws.onopen = () => {
|
||||
ws.send(edgeConfigMessage());
|
||||
ws.send(edgeSsmlMessage(edgeConnectionId(randomUUID()), edgeSsml(req.text, req.voice, req.speed)));
|
||||
};
|
||||
ws.onmessage = (event: MessageEvent) => {
|
||||
if (typeof event.data === 'string') {
|
||||
if (edgeTextFramePath(event.data) === 'turn.end') finish();
|
||||
return;
|
||||
}
|
||||
const frame = new Uint8Array(event.data as ArrayBuffer);
|
||||
const { path, payload } = parseEdgeBinaryFrame(frame);
|
||||
if (path === 'audio' && payload.byteLength) chunks.push(payload);
|
||||
};
|
||||
ws.onerror = (event: Event) => finish(new Error(`Edge TTS : connexion refusée (${(event as { message?: string }).message ?? 'erreur réseau'})`));
|
||||
ws.onclose = (event: CloseEvent) => {
|
||||
if (!settled) finish(chunks.length ? undefined : new Error(`Edge TTS : connexion fermée (${event.code}${event.reason ? ' ' + event.reason : ''})`));
|
||||
};
|
||||
});
|
||||
}
|
||||
|
||||
export async function edgeSynthesize(req: EdgeSynthesizeRequest): Promise<EdgeSynthesizeResult> {
|
||||
const started = Date.now();
|
||||
let mp3: Uint8Array;
|
||||
try {
|
||||
mp3 = await synthesizeOnce(req);
|
||||
} catch (err) {
|
||||
// A refused handshake is almost always a stale signature: resync the clock and retry once.
|
||||
log('WARN', 'edge-tts', `first attempt failed (${(err as Error).message}), resyncing clock`);
|
||||
await syncClock();
|
||||
mp3 = await synthesizeOnce(req);
|
||||
}
|
||||
if (!mp3.byteLength) throw new Error('Edge TTS : aucun audio reçu');
|
||||
return { mp3, durationMs: Date.now() - started };
|
||||
}
|
||||
|
||||
/** Voice list from the service (cached six hours); falls back to an empty list offline. */
|
||||
export async function edgeVoices(): Promise<EdgeVoice[]> {
|
||||
if (voicesCache && Date.now() - voicesCache.at < VOICES_TTL_MS) return voicesCache.voices;
|
||||
const url = `${EDGE_VOICES_URL}&Sec-MS-GEC=${token()}&Sec-MS-GEC-Version=1-${EDGE_CHROMIUM_VERSION}`;
|
||||
const res = await fetch(url, { headers: headers() });
|
||||
if (!res.ok) throw new Error(`Edge TTS : liste des voix HTTP ${res.status}`);
|
||||
const raw = (await res.json()) as Array<{ ShortName?: string; FriendlyName?: string; Locale?: string; Gender?: string }>;
|
||||
const voices: EdgeVoice[] = raw
|
||||
.filter((v) => typeof v.ShortName === 'string' && typeof v.Locale === 'string')
|
||||
.map((v) => ({
|
||||
shortName: v.ShortName!,
|
||||
name: v.ShortName!.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') || v.FriendlyName || v.ShortName!,
|
||||
locale: v.Locale!,
|
||||
gender: v.Gender === 'Male' ? ('m' as const) : ('f' as const)
|
||||
}))
|
||||
.sort((a, b) => a.locale.localeCompare(b.locale) || a.name.localeCompare(b.name));
|
||||
voicesCache = { at: Date.now(), voices };
|
||||
return voices;
|
||||
}
|
||||
@@ -134,7 +134,7 @@ export function transcribe(req: TranscribeRequest): Promise<TranscribeResult> {
|
||||
|
||||
export async function synthesize(req: SynthesizeRequest): Promise<SynthesizeResult> {
|
||||
const result = await request<Omit<SynthesizeResult, 'wav'> & { wav: string }>(
|
||||
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed },
|
||||
{ type: 'synthesize', model: modelRef(req.modelId), text: req.text, speaker: req.speaker, speed: req.speed, language: req.language },
|
||||
180_000
|
||||
);
|
||||
const buffer = Buffer.from(result.wav, 'base64');
|
||||
|
||||
+14
-2
@@ -1,5 +1,6 @@
|
||||
import { ipcMain } from 'electron';
|
||||
import { VOICE_IPC, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice';
|
||||
import { VOICE_IPC, type EdgeSynthesizeRequest, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice';
|
||||
import { edgeSynthesize, edgeVoices } from './edgeTts';
|
||||
import { engineStatus, kwsFeed, kwsStart, kwsStop, synthesize, transcribe, unload, vadFeed, vadStart, vadStop } from './engine';
|
||||
import { cancelDownload, downloadModel, listModels, removeModel } from './models';
|
||||
|
||||
@@ -24,8 +25,19 @@ export function registerVoiceIpc(): void {
|
||||
ipcMain.handle(VOICE_IPC.synthesize, (_e, req: SynthesizeRequest) => {
|
||||
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
|
||||
if (typeof req.modelId !== 'string') throw new Error('Modèle invalide');
|
||||
return synthesize({ ...req, speaker: Number.isFinite(req.speaker) ? req.speaker : 0, speed: Number.isFinite(req.speed) ? req.speed : 1 });
|
||||
return synthesize({
|
||||
...req,
|
||||
speaker: Number.isFinite(req.speaker) ? req.speaker : 0,
|
||||
speed: Number.isFinite(req.speed) ? req.speed : 1,
|
||||
language: typeof req.language === 'string' ? req.language.slice(0, 8) : undefined
|
||||
});
|
||||
});
|
||||
ipcMain.handle(VOICE_IPC.edgeSynthesize, (_e, req: EdgeSynthesizeRequest) => {
|
||||
if (!req || typeof req.text !== 'string' || !req.text.trim() || req.text.length > 5000) throw new Error('Texte invalide');
|
||||
if (typeof req.voice !== 'string' || !/^[a-z]{2,3}-[A-Za-z]{2,4}-[A-Za-z0-9]+$/.test(req.voice)) throw new Error('Voix Edge invalide');
|
||||
return edgeSynthesize({ text: req.text, voice: req.voice, speed: Number.isFinite(req.speed) ? req.speed : 1 });
|
||||
});
|
||||
ipcMain.handle(VOICE_IPC.edgeVoices, () => edgeVoices());
|
||||
ipcMain.handle(VOICE_IPC.unload, (_e, id?: string) => unload(id));
|
||||
ipcMain.handle(VOICE_IPC.kwsStart, (event, req: KwsStartRequest) => {
|
||||
if (!req || !Array.isArray(req.keywords) || typeof req.modelId !== 'string') throw new Error('Requête invalide');
|
||||
|
||||
@@ -6,6 +6,7 @@
|
||||
import os from 'node:os';
|
||||
import path from 'node:path';
|
||||
import type { VoiceEngineKind } from '../../shared/voice';
|
||||
import { SUPERTONIC_LANGS } from './catalog';
|
||||
|
||||
interface ModelRef {
|
||||
id: string;
|
||||
@@ -17,7 +18,7 @@ interface ModelRef {
|
||||
type Request =
|
||||
| { id: number; type: 'status' }
|
||||
| { id: number; type: 'transcribe'; model: ModelRef; wav: Uint8Array | string; language: string }
|
||||
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number }
|
||||
| { id: number; type: 'synthesize'; model: ModelRef; text: string; speaker: number; speed: number; language?: string }
|
||||
| { id: number; type: 'unload'; modelId?: string }
|
||||
| { id: number; type: 'kws.start'; model: ModelRef; keywordsFile: string; threshold: number; score: number }
|
||||
| { id: number; type: 'kws.audio'; pcm: string; sampleRate: number }
|
||||
@@ -55,8 +56,10 @@ type Sherpa = {
|
||||
OfflineTts: new (config: unknown) => {
|
||||
numSpeakers: number;
|
||||
sampleRate: number;
|
||||
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean }) => { samples: Float32Array; sampleRate: number };
|
||||
generate: (req: { text: string; sid: number; speed: number; enableExternalBuffer?: boolean; generationConfig?: unknown }) => { samples: Float32Array; sampleRate: number };
|
||||
};
|
||||
/** Per-request options for the newer engines (Supertonic reads `extra.lang`). */
|
||||
GenerationConfig: new (opts: { sid: number; speed: number; numSteps?: number; extra?: Record<string, string | number> }) => unknown;
|
||||
version: string;
|
||||
};
|
||||
|
||||
@@ -152,6 +155,19 @@ function getSynthesizer(model: ModelRef) {
|
||||
ttsModel = { vits: { model: p(onnx), tokens: p('tokens.txt'), dataDir: p('espeak-ng-data') } };
|
||||
break;
|
||||
}
|
||||
case 'supertonic':
|
||||
ttsModel = {
|
||||
supertonic: {
|
||||
durationPredictor: p('duration_predictor.int8.onnx'),
|
||||
textEncoder: p('text_encoder.int8.onnx'),
|
||||
vectorEstimator: p('vector_estimator.int8.onnx'),
|
||||
vocoder: p('vocoder.int8.onnx'),
|
||||
ttsJson: p('tts.json'),
|
||||
unicodeIndexer: p('unicode_indexer.bin'),
|
||||
voiceStyle: p('voice.bin')
|
||||
}
|
||||
};
|
||||
break;
|
||||
default:
|
||||
throw new Error(`Moteur TTS non supporté : ${model.engine}`);
|
||||
}
|
||||
@@ -160,6 +176,12 @@ function getSynthesizer(model: ModelRef) {
|
||||
return tts;
|
||||
}
|
||||
|
||||
/** 2-letter code accepted by Supertonic 3 ("fr-FR" → "fr"); English when unknown, as upstream does. */
|
||||
function supertonicLang(language: string | undefined): string {
|
||||
const code = (language ?? '').toLowerCase().split(/[-_]/)[0];
|
||||
return SUPERTONIC_LANGS.has(code) ? code : 'en';
|
||||
}
|
||||
|
||||
// ── audio helpers ──────────────────────────────────────────────────────────
|
||||
function decodeWav(bytes: Uint8Array): { samples: Float32Array; sampleRate: number } {
|
||||
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
|
||||
@@ -397,8 +419,19 @@ function handle(req: Request): unknown {
|
||||
const started = Date.now();
|
||||
const tts = getSynthesizer(req.model);
|
||||
const sid = Math.max(0, Math.min(tts.numSpeakers - 1, Math.floor(req.speaker)));
|
||||
const speed = Math.max(0.5, Math.min(2, req.speed || 1));
|
||||
// Electron forbids N-API external buffers: ask sherpa-onnx to copy the samples into a V8 buffer.
|
||||
const audio = tts.generate({ text: req.text, sid, speed: Math.max(0.5, Math.min(2, req.speed || 1)), enableExternalBuffer: false });
|
||||
const audio =
|
||||
req.model.engine === 'supertonic'
|
||||
? tts.generate({
|
||||
text: req.text,
|
||||
sid,
|
||||
speed,
|
||||
enableExternalBuffer: false,
|
||||
// Supertonic needs the language of the text; 5 denoising steps is the quality/speed sweet spot.
|
||||
generationConfig: new (loadSherpa().GenerationConfig)({ sid, speed, numSteps: 5, extra: { lang: supertonicLang(req.language) } })
|
||||
})
|
||||
: tts.generate({ text: req.text, sid, speed, enableExternalBuffer: false });
|
||||
const wav = encodeWav(audio.samples, audio.sampleRate);
|
||||
return {
|
||||
wav: Buffer.from(wav.buffer, wav.byteOffset, wav.byteLength).toString('base64'),
|
||||
|
||||
+2
-2
@@ -1,7 +1,7 @@
|
||||
{
|
||||
"name": "eveflow",
|
||||
"version": "2.4.1",
|
||||
"releaseVersion": "2.4.1",
|
||||
"version": "2.5.0",
|
||||
"releaseVersion": "2.5.0",
|
||||
"description": "JARVIS-style desktop HUD for Hermes Agent: voice, streaming runs, scheduled jobs, skills and telemetry",
|
||||
"main": "dist-electron/main.js",
|
||||
"private": true,
|
||||
|
||||
@@ -13,6 +13,9 @@ import type {
|
||||
} from './ipc';
|
||||
import type { McpToolRequest, McpToolResponse } from './ipc';
|
||||
import type {
|
||||
EdgeSynthesizeRequest,
|
||||
EdgeSynthesizeResult,
|
||||
EdgeVoice,
|
||||
KwsDetection,
|
||||
KwsStartRequest,
|
||||
VadEvent,
|
||||
@@ -75,6 +78,9 @@ export interface EveFlowBridge {
|
||||
onProgress: (cb: (progress: VoiceDownloadProgress) => void) => Unsubscribe;
|
||||
transcribe: (req: TranscribeRequest) => Promise<TranscribeResult>;
|
||||
synthesize: (req: SynthesizeRequest) => Promise<SynthesizeResult>;
|
||||
/** Microsoft Edge neural voices (online, free): MP3 for one sentence. */
|
||||
edgeSynthesize: (req: EdgeSynthesizeRequest) => Promise<EdgeSynthesizeResult>;
|
||||
edgeVoices: () => Promise<EdgeVoice[]>;
|
||||
unload: (id?: string) => Promise<unknown>;
|
||||
kwsStart: (req: KwsStartRequest) => Promise<{ accepted: string[]; rejected: string[] }>;
|
||||
kwsStop: () => Promise<void>;
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
/**
|
||||
* Pure helpers for the Microsoft Edge "Read aloud" speech service (the endpoint used by the
|
||||
* Edge browser, no API key): request signing input, SSML building, frame parsing and the
|
||||
* default French voices. No Node or DOM dependency so it is shared by main and tests.
|
||||
*/
|
||||
|
||||
export const EDGE_TRUSTED_CLIENT_TOKEN = '6A5AA1D4EAFF4E9FB37E23D68491D6F4';
|
||||
export const EDGE_CHROMIUM_VERSION = '143.0.3650.75';
|
||||
export const EDGE_WSS_URL = `wss://speech.platform.bing.com/consumer/speech/synthesize/readaloud/edge/v1?TrustedClientToken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
|
||||
export const EDGE_VOICES_URL = `https://speech.platform.bing.com/consumer/speech/synthesize/readaloud/voices/list?trustedclienttoken=${EDGE_TRUSTED_CLIENT_TOKEN}`;
|
||||
export const EDGE_OUTPUT_FORMAT = 'audio-24khz-48kbitrate-mono-mp3';
|
||||
|
||||
const WIN_EPOCH_SEC = 11644473600;
|
||||
|
||||
/**
|
||||
* String whose SHA-256 (upper-case hex) is the `Sec-MS-GEC` value: Windows file time (100 ns
|
||||
* ticks since 1601) rounded down to 5 minutes, followed by the trusted client token.
|
||||
* `nowMs` is the client clock, `skewSec` the correction learned from the server's Date header.
|
||||
*/
|
||||
export function edgeTokenInput(nowMs: number, skewSec = 0): string {
|
||||
let seconds = Math.floor(nowMs / 1000 + skewSec) + WIN_EPOCH_SEC;
|
||||
seconds -= seconds % 300;
|
||||
// 10 million ticks per second; BigInt keeps the 18-digit value exact.
|
||||
const ticks = BigInt(seconds) * 10_000_000n;
|
||||
return `${ticks}${EDGE_TRUSTED_CLIENT_TOKEN}`;
|
||||
}
|
||||
|
||||
export function edgeHeaders(): Record<string, string> {
|
||||
const major = EDGE_CHROMIUM_VERSION.split('.')[0];
|
||||
return {
|
||||
'User-Agent': `Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/${major}.0.0.0 Safari/537.36 Edg/${major}.0.0.0`,
|
||||
'Accept-Language': 'en-US,en;q=0.9',
|
||||
Pragma: 'no-cache',
|
||||
'Cache-Control': 'no-cache',
|
||||
Origin: 'chrome-extension://jdiccldimpdaibmpdkjnbmckianbfold'
|
||||
};
|
||||
}
|
||||
|
||||
export function escapeXml(text: string): string {
|
||||
return text.replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"').replace(/'/g, ''');
|
||||
}
|
||||
|
||||
/** Prosody rate attribute for a playback speed multiplier (1 = "+0%", 1.25 = "+25%"). */
|
||||
export function edgeRate(speed: number): string {
|
||||
const pct = Math.round((Math.max(0.5, Math.min(2, speed || 1)) - 1) * 100);
|
||||
return `${pct >= 0 ? '+' : ''}${pct}%`;
|
||||
}
|
||||
|
||||
export function edgeSsml(text: string, voice: string, speed: number): string {
|
||||
const lang = voice.split('-').slice(0, 2).join('-') || 'fr-FR';
|
||||
return (
|
||||
`<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xml:lang='${lang}'>` +
|
||||
`<voice name='${escapeXml(voice)}'><prosody pitch='+0Hz' rate='${edgeRate(speed)}' volume='+0%'>${escapeXml(text)}</prosody></voice></speak>`
|
||||
);
|
||||
}
|
||||
|
||||
/** Timestamp header the service expects ("JavaScript date string"). */
|
||||
export function edgeTimestamp(date = new Date()): string {
|
||||
return date.toUTCString().replace('GMT', 'GMT+0000 (Coordinated Universal Time)');
|
||||
}
|
||||
|
||||
export function edgeConfigMessage(date = new Date()): string {
|
||||
return (
|
||||
`X-Timestamp:${edgeTimestamp(date)}\r\nContent-Type:application/json; charset=utf-8\r\nPath:speech.config\r\n\r\n` +
|
||||
`{"context":{"synthesis":{"audio":{"metadataoptions":{"sentenceBoundaryEnabled":"false","wordBoundaryEnabled":"false"},"outputFormat":"${EDGE_OUTPUT_FORMAT}"}}}}`
|
||||
);
|
||||
}
|
||||
|
||||
export function edgeSsmlMessage(requestId: string, ssml: string, date = new Date()): string {
|
||||
return `X-RequestId:${requestId}\r\nContent-Type:application/ssml+xml\r\nX-Timestamp:${edgeTimestamp(date)}Z\r\nPath:ssml\r\n\r\n${ssml}`;
|
||||
}
|
||||
|
||||
/** Split a binary frame into its text headers and payload (2-byte big-endian header length). */
|
||||
export function parseEdgeBinaryFrame(frame: Uint8Array): { path: string; payload: Uint8Array } {
|
||||
if (frame.byteLength < 2) return { path: '', payload: new Uint8Array(0) };
|
||||
const headerLength = (frame[0] << 8) | frame[1];
|
||||
const end = Math.min(frame.byteLength, 2 + headerLength);
|
||||
let header = '';
|
||||
for (let i = 2; i < end; i++) header += String.fromCharCode(frame[i]);
|
||||
const path = /Path:\s*([^\r\n]+)/i.exec(header)?.[1]?.trim() ?? '';
|
||||
return { path, payload: frame.subarray(end) };
|
||||
}
|
||||
|
||||
/** Path of a text frame ("turn.start", "response", "audio.metadata", "turn.end"). */
|
||||
export function edgeTextFramePath(message: string): string {
|
||||
return /Path:\s*([^\r\n]+)/i.exec(message)?.[1]?.trim() ?? '';
|
||||
}
|
||||
|
||||
export function edgeConnectionId(hex32: string): string {
|
||||
return hex32.replace(/-/g, '').toLowerCase();
|
||||
}
|
||||
|
||||
/** Well-known voices per language so the choice works before the voice list is fetched. */
|
||||
export const EDGE_DEFAULT_VOICES: Record<string, { male: string; female: string }> = {
|
||||
fr: { male: 'fr-FR-HenriNeural', female: 'fr-FR-DeniseNeural' },
|
||||
en: { male: 'en-US-AndrewMultilingualNeural', female: 'en-US-AvaMultilingualNeural' },
|
||||
de: { male: 'de-DE-ConradNeural', female: 'de-DE-KatjaNeural' },
|
||||
es: { male: 'es-ES-AlvaroNeural', female: 'es-ES-ElviraNeural' },
|
||||
it: { male: 'it-IT-DiegoNeural', female: 'it-IT-ElsaNeural' },
|
||||
pt: { male: 'pt-BR-AntonioNeural', female: 'pt-BR-FranciscaNeural' }
|
||||
};
|
||||
|
||||
/** Default Edge voice for a language ("fr-FR", "fr", "en-GB") and gender; French when unknown. */
|
||||
export function defaultEdgeVoice(language: string, gender: 'male' | 'female'): string {
|
||||
const lang = (language || 'fr').toLowerCase().split(/[-_]/)[0];
|
||||
return (EDGE_DEFAULT_VOICES[lang] ?? EDGE_DEFAULT_VOICES.fr)[gender];
|
||||
}
|
||||
|
||||
/** Gender of an Edge voice from its short name, using the built-in table then common first names. */
|
||||
export function edgeVoiceGender(shortName: string): 'male' | 'female' | undefined {
|
||||
for (const pair of Object.values(EDGE_DEFAULT_VOICES)) {
|
||||
if (pair.male === shortName) return 'male';
|
||||
if (pair.female === shortName) return 'female';
|
||||
}
|
||||
const name = shortName.split('-')[2]?.replace(/(Multilingual)?Neural$/i, '') ?? '';
|
||||
if (/^(Henri|Remy|Rémy|Gerard|Antoine|Jean|Thierry|Fabrice|Claude|Andrew|Brian|Guy|Christopher|Eric|Roger|Steffan|Ryan|Thomas|Conrad|Alvaro|Diego|Antonio)$/i.test(name)) return 'male';
|
||||
if (/^(Denise|Eloise|Vivienne|Charline|Sylvie|Ariane|Ava|Emma|Jenny|Aria|Michelle|Ana|Sonia|Libby|Katja|Elvira|Elsa|Francisca)$/i.test(name)) return 'female';
|
||||
return undefined;
|
||||
}
|
||||
+28
-1
@@ -1,7 +1,7 @@
|
||||
/** Local voice engine contract (sherpa-onnx in a utility process). Shared by main and renderer. */
|
||||
|
||||
export type VoiceModelKind = 'stt' | 'tts' | 'kws' | 'vad';
|
||||
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'kws-transducer' | 'silero';
|
||||
export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'supertonic' | 'kws-transducer' | 'silero';
|
||||
|
||||
export interface VoiceSpeaker {
|
||||
id: number;
|
||||
@@ -62,6 +62,31 @@ export interface SynthesizeRequest {
|
||||
text: string;
|
||||
speaker: number;
|
||||
speed: number;
|
||||
/** BCP-47 tag or 2-letter code of the text (multilingual engines such as Supertonic need it). */
|
||||
language?: string;
|
||||
}
|
||||
|
||||
/** Microsoft Edge "Read aloud" neural voices (online, no key). */
|
||||
export interface EdgeVoice {
|
||||
/** e.g. fr-FR-HenriNeural */
|
||||
shortName: string;
|
||||
/** Display name, e.g. Henri */
|
||||
name: string;
|
||||
locale: string;
|
||||
gender: 'm' | 'f';
|
||||
}
|
||||
|
||||
export interface EdgeSynthesizeRequest {
|
||||
text: string;
|
||||
voice: string;
|
||||
/** Playback speed multiplier (0.5 .. 2). */
|
||||
speed: number;
|
||||
}
|
||||
|
||||
export interface EdgeSynthesizeResult {
|
||||
/** MP3, 24 kHz mono 48 kbit/s. */
|
||||
mp3: Uint8Array;
|
||||
durationMs: number;
|
||||
}
|
||||
|
||||
export interface SynthesizeResult {
|
||||
@@ -115,6 +140,8 @@ export const VOICE_IPC = {
|
||||
modelsProgress: 'voice:models:progress',
|
||||
transcribe: 'voice:transcribe',
|
||||
synthesize: 'voice:synthesize',
|
||||
edgeSynthesize: 'voice:edge:synthesize',
|
||||
edgeVoices: 'voice:edge:voices',
|
||||
unload: 'voice:unload',
|
||||
kwsStart: 'voice:kws:start',
|
||||
kwsStop: 'voice:kws:stop',
|
||||
|
||||
@@ -3,6 +3,7 @@ import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, Re
|
||||
import type { VoiceModelStatus } from '../../../shared/voice';
|
||||
import { bridge } from '../../lib/bridge';
|
||||
import { useShallow } from 'zustand/react/shallow';
|
||||
import { pickSpeaker } from '../../lib/voicePreference';
|
||||
import { useVoiceModels } from '../../state/voiceModels';
|
||||
import { useSettings } from '../../state/settings';
|
||||
|
||||
@@ -24,7 +25,10 @@ function ModelRow({ model }: { model: VoiceModelStatus }) {
|
||||
if (model.kind === 'stt') update({ voice: { localModel: model.id, provider: 'local' } });
|
||||
else if (model.kind === 'kws') update({ voice: { wakeMode: 'kws' } });
|
||||
else if (model.kind === 'vad') update({ voice: { neuralVad: true } });
|
||||
else update({ speech: { localModel: model.id, provider: 'local', localSpeaker: model.speakers?.[0]?.id ?? 0 } });
|
||||
else {
|
||||
const { speech } = useSettings.getState().settings;
|
||||
update({ speech: { localModel: model.id, provider: 'local', localSpeaker: pickSpeaker(model, speech.language || 'fr-FR', speech.voiceGender ?? 'male') } });
|
||||
}
|
||||
};
|
||||
const neuralVad = useSettings((s) => s.settings.voice.neuralVad);
|
||||
const activeNow = model.kind === 'kws' ? wakeMode === 'kws' : model.kind === 'vad' ? neuralVad : isActive;
|
||||
|
||||
@@ -5,6 +5,8 @@ import { HermesClient, discoverHermesUrl, hermesUrlCandidates } from '../../serv
|
||||
import { listSystemVoices } from '../../services/voice/tts';
|
||||
import { speech } from '../../services/voice/speech';
|
||||
import { ensurePreferredVoice } from '../../services/voice/voicePreference';
|
||||
import { pickSpeaker, resolveEdgeVoice } from '../../lib/voicePreference';
|
||||
import type { EdgeVoice } from '../../../shared/voice';
|
||||
import { listMicrophones } from '../../services/voice/capture';
|
||||
import { transcribeWav } from '../../services/voice/stt';
|
||||
import { encodeWav } from '../../services/voice/wav';
|
||||
@@ -86,8 +88,18 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
const [sttTest, setSttTest] = useState<TestState>({ status: 'idle', message: '' });
|
||||
const [ttsTest, setTtsTest] = useState<TestState>({ status: 'idle', message: '' });
|
||||
const [voices, setVoices] = useState(listSystemVoices());
|
||||
const [edgeVoices, setEdgeVoices] = useState<EdgeVoice[]>([]);
|
||||
const [webhookSecretVisible, setWebhookSecretVisible] = useState(false);
|
||||
|
||||
const speechProvider = settings.speech.provider;
|
||||
useEffect(() => {
|
||||
if (speechProvider !== 'edge' || edgeVoices.length) return;
|
||||
bridge()
|
||||
?.voice.edgeVoices()
|
||||
.then(setEdgeVoices)
|
||||
.catch((err: Error) => setTtsTest({ status: 'fail', message: `Liste des voix Edge indisponible : ${err.message}` }));
|
||||
}, [speechProvider, edgeVoices.length]);
|
||||
|
||||
useEffect(() => {
|
||||
const refresh = () => setVoices(listSystemVoices());
|
||||
refresh();
|
||||
@@ -359,6 +371,12 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
</>
|
||||
)}
|
||||
<Toggle on={settings.voice.localCommands} onChange={(v) => update({ voice: { localCommands: v } })} label="Commandes locales instantanées" hint="« Verrouille la session », « monte le son », « ouvre Spotify », « regarde mon écran »… exécutées sur ce PC sans passer par Hermes." />
|
||||
<Toggle
|
||||
on={settings.voice.micProcessing ?? true}
|
||||
onChange={(v) => update({ voice: { micProcessing: v } })}
|
||||
label="Traitement du micro par Chromium (écho, bruit, gain automatique)"
|
||||
hint="Désactivez-le si les transcriptions sont approximatives avec un casque ou un bon micro : ces filtres déforment la voix avant la reconnaissance. Gardez-le activé avec des haut-parleurs (sinon la voix de l’assistant est réentendue par le micro)."
|
||||
/>
|
||||
<div className="row" style={{ marginTop: 10 }}>
|
||||
<button className="btn small" onClick={() => void testStt()} disabled={settings.voice.provider === 'browser'}><Mic size={13} /> Tester la reconnaissance</button>
|
||||
</div>
|
||||
@@ -374,7 +392,11 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
<button className={(settings.speech.voiceGender ?? 'male') === 'male' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'male' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Masculine</button>
|
||||
<button className={settings.speech.voiceGender === 'female' ? 'active' : ''} onClick={() => { update({ speech: { voiceGender: 'female' } }); void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m })); }}>Féminine</button>
|
||||
</div>
|
||||
<span className="hint">S’applique à tous les moteurs : voix locale (Piper Tom ou Pierre pour le masculin, téléchargée automatiquement si besoin), voix OpenAI (onyx / nova), voix système Windows (Paul / Hortense). Kokoro n’a pas de voix française masculine.</span>
|
||||
<span className="hint">
|
||||
S’applique à tous les moteurs : Edge (Henri / Denise), voix locale (Supertonic 3, cinq voix de chaque genre, téléchargé automatiquement si besoin), API OpenAI (onyx / nova), voix système Windows (Paul / Hortense).
|
||||
{settings.speech.provider === 'google-free' && ' Google Translate n’a qu’une voix féminine : ce choix n’a pas d’effet avec ce moteur.'}
|
||||
{settings.speech.provider === 'local' && settings.speech.localModel === 'kokoro-v1' && ' Kokoro n’a qu’une voix française, féminine et avec accent : la voix masculine bascule sur Supertonic 3.'}
|
||||
</span>
|
||||
</div>
|
||||
<div className="field">
|
||||
<label>Timbre</label>
|
||||
@@ -388,29 +410,62 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
<button
|
||||
className="btn small primary"
|
||||
onClick={() => {
|
||||
update({ speech: { provider: 'local', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
|
||||
setTtsTest({ status: 'running', message: 'préparation de la voix JARVIS (téléchargement de Piper Tom si nécessaire)…' });
|
||||
update({ speech: { provider: 'edge', edgeVoice: '', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
|
||||
void ensurePreferredVoice().then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
|
||||
}}
|
||||
>
|
||||
<Volume2 size={13} /> Préréglage voix JARVIS (français, masculine, locale)
|
||||
<Volume2 size={13} /> Voix JARVIS en ligne (Edge Henri, masculine)
|
||||
</button>
|
||||
<button
|
||||
className="btn small"
|
||||
onClick={() => {
|
||||
update({ speech: { provider: 'local', voiceGender: 'male', timbre: 'jarvis', speed: 0.97, autoSpeak: true } });
|
||||
setTtsTest({ status: 'running', message: 'préparation de la voix JARVIS locale (téléchargement de Supertonic 3, 129 Mo, si nécessaire)…' });
|
||||
void ensurePreferredVoice({ upgrade: true }).then((m) => setTtsTest({ status: 'ok', message: m || 'Voix JARVIS locale prête.' })).catch((e: Error) => setTtsTest({ status: 'fail', message: e.message }));
|
||||
}}
|
||||
>
|
||||
<Volume2 size={13} /> Voix JARVIS hors ligne (Supertonic 3, masculine)
|
||||
</button>
|
||||
<span className="status-pill">
|
||||
{settings.speech.provider === 'local'
|
||||
? `voix active : ${voiceModels.find((m) => m.id === settings.speech.localModel)?.speakers?.find((s) => s.id === settings.speech.localSpeaker)?.name ?? settings.speech.localModel}`
|
||||
: `moteur : ${settings.speech.provider}`}
|
||||
: settings.speech.provider === 'edge'
|
||||
? `voix active : ${resolveEdgeVoice(settings.speech.edgeVoice ?? '', settings.speech.language, settings.speech.voiceGender ?? 'male')}`
|
||||
: `moteur : ${settings.speech.provider}`}
|
||||
</span>
|
||||
</div>
|
||||
<div className="field">
|
||||
<label>Moteur de synthèse</label>
|
||||
<select className="select" value={settings.speech.provider} onChange={(e) => update({ speech: { provider: e.target.value as typeof settings.speech.provider } })}>
|
||||
<option value="local">Local dans l’application (Kokoro / Piper via sherpa-onnx, hors ligne)</option>
|
||||
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Kokoro, Piper, OpenAI, LocalAI…)</option>
|
||||
<select
|
||||
className="select"
|
||||
value={settings.speech.provider}
|
||||
onChange={(e) => {
|
||||
update({ speech: { provider: e.target.value as typeof settings.speech.provider } });
|
||||
void ensurePreferredVoice().then((m) => m && setTtsTest({ status: 'ok', message: m }));
|
||||
}}
|
||||
>
|
||||
<option value="edge">Microsoft Edge (voix neuronales, gratuit, en ligne, sans clé) — recommandé</option>
|
||||
<option value="local">Local dans l’application (Supertonic 3 / Kokoro / Piper via sherpa-onnx, hors ligne)</option>
|
||||
<option value="openai-compatible">API compatible OpenAI /v1/audio/speech (Qwen3-TTS, Kokoro, OpenAI, LocalAI…)</option>
|
||||
<option value="system">Voix système Windows</option>
|
||||
<option value="google-free">Google Translate (gratuit, en ligne)</option>
|
||||
<option value="google-free">Google Translate (gratuit, en ligne, voix féminine uniquement)</option>
|
||||
<option value="off">Désactivée</option>
|
||||
</select>
|
||||
</div>
|
||||
{settings.speech.provider === 'edge' && (
|
||||
<div className="field">
|
||||
<label>Voix Edge</label>
|
||||
<select className="select" value={settings.speech.edgeVoice ?? ''} onChange={(e) => update({ speech: { edgeVoice: e.target.value } })}>
|
||||
<option value="">Automatique ({resolveEdgeVoice('', settings.speech.language, settings.speech.voiceGender ?? 'male')})</option>
|
||||
{edgeVoices
|
||||
.filter((v) => v.locale.toLowerCase().startsWith(settings.speech.language.toLowerCase().split('-')[0]))
|
||||
.map((v) => (
|
||||
<option key={v.shortName} value={v.shortName}>{v.name} · {v.locale} · {v.gender === 'm' ? 'homme' : 'femme'}</option>
|
||||
))}
|
||||
</select>
|
||||
<span className="hint">Mêmes voix que la lecture à voix haute d’Edge : Henri, Denise, Rémy, Vivienne, Éloise (fr-FR), plus les voix canadiennes, suisses et belges. Aucune donnée locale ; chaque phrase est synthétisée en ligne.</span>
|
||||
</div>
|
||||
)}
|
||||
{settings.speech.provider === 'openai-compatible' && (
|
||||
<>
|
||||
<div className="field">
|
||||
@@ -424,7 +479,7 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
</div>
|
||||
<div className="field">
|
||||
<label>Voix</label>
|
||||
<input className="input" value={settings.speech.voice} placeholder="alloy, onyx, af_heart…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
|
||||
<input className="input" value={settings.speech.voice} placeholder="onyx, nova, af_heart, Ryan (Qwen3-TTS)…" onChange={(e) => update({ speech: { voice: e.target.value } })} />
|
||||
</div>
|
||||
<div className="field">
|
||||
<label>Clé API</label>
|
||||
@@ -444,7 +499,7 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
{settings.speech.provider === 'local' && (
|
||||
installedModels(voiceModels, 'tts').length === 0 ? (
|
||||
<div className="field">
|
||||
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Kokoro dans « Modèles locaux »</a>.</span>
|
||||
<span className="hint">Aucun modèle de voix installé. <a href="#" onClick={(e) => { e.preventDefault(); setSection('models'); }}>Téléchargez Supertonic 3 dans « Modèles locaux »</a>.</span>
|
||||
</div>
|
||||
) : (
|
||||
<div className="grid-2">
|
||||
@@ -452,7 +507,8 @@ export function SettingsDrawer({ onClose }: Props) {
|
||||
<label>Modèle local</label>
|
||||
<select className="select" value={settings.speech.localModel} onChange={(e) => {
|
||||
const m = voiceModels.find((x) => x.id === e.target.value);
|
||||
update({ speech: { localModel: e.target.value, localSpeaker: m?.speakers?.[0]?.id ?? 0 } });
|
||||
// Keep the preferred gender when switching models (Kokoro has no masculine French voice: it falls back to Siwis).
|
||||
update({ speech: { localModel: e.target.value, localSpeaker: m ? pickSpeaker(m, settings.speech.language, settings.speech.voiceGender ?? 'male') : 0 } });
|
||||
}}>
|
||||
{installedModels(voiceModels, 'tts').map((m) => <option key={m.id} value={m.id}>{m.name}</option>)}
|
||||
</select>
|
||||
|
||||
+55
-11
@@ -1,13 +1,18 @@
|
||||
/** Voice gender preference applied to every TTS provider (pure helpers, unit tested). */
|
||||
import type { VoiceModelStatus } from '../../shared/voice';
|
||||
import type { VoiceModelStatus, VoiceSpeaker } from '../../shared/voice';
|
||||
import { defaultEdgeVoice, edgeVoiceGender } from '../../shared/edgeTts';
|
||||
|
||||
export type VoiceGender = 'male' | 'female';
|
||||
|
||||
const MALE_HINTS = /\b(homme|male|masculin|paul|thomas|claude|henri|guillaume|mathieu|antoine|nicolas|denis|pierre|tom|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|david|mark|richard|james|ryan|guy)\b/i;
|
||||
const FEMALE_HINTS = /\b(femme|female|f[ée]minin|hortense|julie|denise|eloise|am[ée]lie|audrey|siwis|jessica|heart|bella|sarah|nicole|sky|alloy|nova|shimmer|zira|aria|jenny|emma|isabella|sophie|charlotte|vivienne|coral|sage)\b/i;
|
||||
const MALE_HINTS =
|
||||
/\b(homme|male|masculin|paul|thomas|claude|henri|remy|rémy|gerard|guillaume|mathieu|antoine|nicolas|denis|pierre|tom|adam|michael|eric|liam|george|lewis|daniel|fenrir|puck|onyx|echo|david|mark|richard|james|ryan|guy|dylan|aiden|uncle fu|andrew|brian|fabrice|jean|thierry)\b/i;
|
||||
const FEMALE_HINTS =
|
||||
/\b(femme|female|f[ée]minin|hortense|julie|denise|eloise|vivienne|charline|sylvie|ariane|am[ée]lie|audrey|siwis|jessica|heart|bella|sarah|nicole|sky|alloy|nova|shimmer|zira|aria|jenny|emma|ava|isabella|sophie|charlotte|coral|sage|vivian|serena|sohee|ono anna)\b/i;
|
||||
|
||||
/** Best-effort gender from a voice or speaker label ("Piper Tom (homme, français)", "Microsoft Paul", "am_adam"). */
|
||||
/** Best-effort gender from a voice or speaker label ("Piper Tom (homme, français)", "Microsoft Paul", "am_adam", "fr-FR-HenriNeural"). */
|
||||
export function inferGender(name: string): VoiceGender | undefined {
|
||||
const edge = edgeVoiceGender(name);
|
||||
if (edge) return edge;
|
||||
const n = name.replace(/_/g, ' ');
|
||||
if (/^(am|bm|em|hm|im|jm|pm|zm)\b/i.test(n) || MALE_HINTS.test(n)) return 'male';
|
||||
if (/^(af|bf|ef|ff|hf|if|jf|pf|zf)\b/i.test(n) || FEMALE_HINTS.test(n)) return 'female';
|
||||
@@ -19,6 +24,13 @@ export function defaultOpenAiVoice(gender: VoiceGender): string {
|
||||
return gender === 'male' ? 'onyx' : 'nova';
|
||||
}
|
||||
|
||||
/** Edge voice to use: the explicit choice when it matches the gender, otherwise the language default. */
|
||||
export function resolveEdgeVoice(explicit: string, language: string, gender: VoiceGender): string {
|
||||
const chosen = explicit.trim();
|
||||
if (chosen && (edgeVoiceGender(chosen) ?? gender) === gender) return chosen;
|
||||
return defaultEdgeVoice(language, gender);
|
||||
}
|
||||
|
||||
/** Rank system voices: language first, then gender, then quality hints. */
|
||||
export function rankSystemVoice(v: { name: string; lang: string; localService?: boolean }, lang: string, gender: VoiceGender): number {
|
||||
const name = v.name.toLowerCase();
|
||||
@@ -38,16 +50,39 @@ export interface LocalVoiceChoice {
|
||||
speaker: number;
|
||||
}
|
||||
|
||||
/** Installed local voice matching language + gender, or null. */
|
||||
/** Local models from best to worst French rendering; unknown models come last. */
|
||||
const LOCAL_QUALITY: Record<string, number> = { 'supertonic-3': 0, 'kokoro-v1': 1, 'piper-fr-upmc': 2, 'piper-fr-tom': 3, 'piper-fr-siwis': 4 };
|
||||
|
||||
export function localQualityRank(modelId: string): number {
|
||||
return LOCAL_QUALITY[modelId] ?? 9;
|
||||
}
|
||||
|
||||
function speakerGender(sp: VoiceSpeaker): VoiceGender | undefined {
|
||||
return sp.gender === 'm' ? 'male' : sp.gender === 'f' ? 'female' : inferGender(sp.name);
|
||||
}
|
||||
|
||||
function speakerSpeaks(sp: VoiceSpeaker, lang: string): boolean {
|
||||
const spLang = (sp.lang || '').toLowerCase();
|
||||
return !spLang || spLang === lang || spLang === 'multi';
|
||||
}
|
||||
|
||||
/** First speaker of a model matching language + gender; falls back to any speaker of the language, then the first one. */
|
||||
export function pickSpeaker(model: Pick<VoiceModelStatus, 'speakers'>, lang: string, gender: VoiceGender): number {
|
||||
const l = lang.toLowerCase().split('-')[0];
|
||||
const speakers = model.speakers ?? [];
|
||||
const exact = speakers.find((sp) => speakerSpeaks(sp, l) && speakerGender(sp) === gender);
|
||||
const sameLang = speakers.find((sp) => speakerSpeaks(sp, l));
|
||||
return (exact ?? sameLang ?? speakers[0])?.id ?? 0;
|
||||
}
|
||||
|
||||
/** Installed local voice matching language + gender (best model first), or null. */
|
||||
export function findLocalVoice(models: VoiceModelStatus[], lang: string, gender: VoiceGender): LocalVoiceChoice | null {
|
||||
const l = lang.toLowerCase().split('-')[0];
|
||||
const installed = models.filter((m) => m.kind === 'tts' && m.installed);
|
||||
const installed = models.filter((m) => m.kind === 'tts' && m.installed).sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id));
|
||||
for (const model of installed) {
|
||||
for (const sp of model.speakers ?? []) {
|
||||
const spLang = (sp.lang || '').toLowerCase();
|
||||
if (spLang && spLang !== l && spLang !== 'multi') continue;
|
||||
const g = (sp as { gender?: string }).gender === 'm' ? 'male' : (sp as { gender?: string }).gender === 'f' ? 'female' : inferGender(sp.name);
|
||||
if (g === gender) return { modelId: model.id, speaker: sp.id };
|
||||
if (!speakerSpeaks(sp, l)) continue;
|
||||
if (speakerGender(sp) === gender) return { modelId: model.id, speaker: sp.id };
|
||||
}
|
||||
}
|
||||
return null;
|
||||
@@ -57,10 +92,19 @@ export function findLocalVoice(models: VoiceModelStatus[], lang: string, gender:
|
||||
export function suggestedDownload(models: VoiceModelStatus[], lang: string, gender: VoiceGender): string | null {
|
||||
const l = lang.toLowerCase().split('-')[0];
|
||||
if (l !== 'fr') return null;
|
||||
const wanted = gender === 'male' ? ['piper-fr-tom', 'piper-fr-upmc'] : ['kokoro-v1', 'piper-fr-siwis'];
|
||||
const wanted = gender === 'male' ? ['supertonic-3', 'piper-fr-tom', 'piper-fr-upmc'] : ['supertonic-3', 'kokoro-v1', 'piper-fr-siwis'];
|
||||
for (const id of wanted) {
|
||||
const m = models.find((x) => x.id === id);
|
||||
if (m && !m.installed) return id;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Best local model in the catalog for the language that is not installed yet (null when the best is already there). */
|
||||
export function bestLocalUpgrade(models: VoiceModelStatus[], lang: string): string | null {
|
||||
const l = lang.toLowerCase().split('-')[0];
|
||||
const best = models
|
||||
.filter((m) => m.kind === 'tts' && (m.languages.includes(l) || m.languages.includes('multi')))
|
||||
.sort((a, b) => localQualityRank(a.id) - localQualityRank(b.id))[0];
|
||||
return best && !best.installed ? best.id : null;
|
||||
}
|
||||
@@ -49,6 +49,8 @@ export interface CaptureCallbacks {
|
||||
|
||||
export interface CaptureOptions {
|
||||
deviceId?: string;
|
||||
/** Chromium's echo cancellation / noise suppression / auto gain (default on). Off keeps the raw signal for the recogniser. */
|
||||
micProcessing?: boolean;
|
||||
mode: CaptureMode;
|
||||
vad?: Partial<VadOptions>;
|
||||
callbacks: CaptureCallbacks;
|
||||
@@ -124,9 +126,9 @@ export class MicCapture {
|
||||
const constraints: MediaStreamConstraints = {
|
||||
audio: {
|
||||
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
|
||||
echoCancellation: true,
|
||||
noiseSuppression: true,
|
||||
autoGainControl: true,
|
||||
echoCancellation: options.micProcessing !== false,
|
||||
noiseSuppression: options.micProcessing !== false,
|
||||
autoGainControl: options.micProcessing !== false,
|
||||
channelCount: 1
|
||||
}
|
||||
};
|
||||
|
||||
@@ -1,16 +1,17 @@
|
||||
/**
|
||||
* Text-to-speech engine with a sentence queue, bounded prefetching and Web Audio playback so
|
||||
* the HUD reacts to the actual waveform. Providers: in-app sherpa-onnx (Kokoro / Piper),
|
||||
* OpenAI-compatible /v1/audio/speech, system voices, and the legacy Google Translate endpoint.
|
||||
* the HUD reacts to the actual waveform. Providers: Microsoft Edge neural voices (online, free),
|
||||
* in-app sherpa-onnx (Supertonic / Kokoro / Piper), OpenAI-compatible /v1/audio/speech, system
|
||||
* voices, and the legacy Google Translate endpoint (one feminine voice per language).
|
||||
*/
|
||||
import { Log } from '../../lib/log';
|
||||
import { defaultOpenAiVoice, rankSystemVoice } from '../../lib/voicePreference';
|
||||
import { defaultOpenAiVoice, rankSystemVoice, resolveEdgeVoice } from '../../lib/voicePreference';
|
||||
import { bridge } from '../../lib/bridge';
|
||||
import { chunkForSpeech, cleanForSpeech, extractSentences } from '../../lib/text';
|
||||
import { httpFetch } from '../../lib/transport';
|
||||
import { audioBus } from './audioBus';
|
||||
|
||||
export type TtsProvider = 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
|
||||
export type TtsProvider = 'edge' | 'openai-compatible' | 'system' | 'google-free' | 'local' | 'off';
|
||||
|
||||
export interface TtsConfig {
|
||||
provider: TtsProvider;
|
||||
@@ -27,6 +28,8 @@ export interface TtsConfig {
|
||||
localModel: string;
|
||||
/** Speaker id inside the local model. */
|
||||
localSpeaker: number;
|
||||
/** Microsoft Edge voice short name (provider 'edge'); empty = automatic from language + gender. */
|
||||
edgeVoice?: string;
|
||||
/** Preferred voice gender, applied to every provider's default voice. */
|
||||
voiceGender?: 'male' | 'female';
|
||||
/** 'jarvis' adds a subtle AI timbre: slightly lower pitch, warm/crisp EQ, short room reverb. */
|
||||
@@ -179,8 +182,9 @@ export class TtsEngine {
|
||||
}
|
||||
|
||||
private prefetch(text: string, gen: number): Promise<AudioBuffer | null> {
|
||||
const provider = this.config.provider;
|
||||
const task =
|
||||
this.config.provider === 'google-free' ? this.fetchGoogle(text) : this.config.provider === 'local' ? this.fetchLocal(text) : this.fetchOpenAi(text);
|
||||
provider === 'google-free' ? this.fetchGoogle(text) : provider === 'local' ? this.fetchLocal(text) : provider === 'edge' ? this.fetchEdge(text) : this.fetchOpenAi(text);
|
||||
return task
|
||||
.then(async (bytes) => {
|
||||
if (gen !== this.generation || !bytes) return null;
|
||||
@@ -228,12 +232,23 @@ export class TtsEngine {
|
||||
modelId: this.config.localModel,
|
||||
text,
|
||||
speaker: this.config.localSpeaker,
|
||||
speed: Math.max(0.5, Math.min(2, this.config.speed || 1))
|
||||
speed: Math.max(0.5, Math.min(2, this.config.speed || 1)),
|
||||
language: this.config.language || 'fr-FR'
|
||||
});
|
||||
Log.debug('tts', `local synthesis ${result.durationMs} ms for ${result.audioSec.toFixed(1)}s of audio`);
|
||||
return result.wav;
|
||||
}
|
||||
|
||||
/** Microsoft Edge neural voices through the main process (no key, online). */
|
||||
private async fetchEdge(text: string): Promise<Uint8Array> {
|
||||
const api = bridge();
|
||||
if (!api) throw new Error('Les voix Edge nécessitent l’application Electron.');
|
||||
const voice = resolveEdgeVoice(this.config.edgeVoice ?? '', this.config.language || 'fr-FR', this.config.voiceGender ?? 'male');
|
||||
const result = await api.voice.edgeSynthesize({ text, voice, speed: Math.max(0.5, Math.min(2, this.config.speed || 1)) });
|
||||
Log.debug('tts', `edge synthesis (${voice}) ${result.durationMs} ms, ${result.mp3.byteLength} bytes`);
|
||||
return result.mp3;
|
||||
}
|
||||
|
||||
private async fetchGoogle(text: string): Promise<Uint8Array> {
|
||||
const url = `https://translate.google.com/translate_tts?ie=UTF-8&tl=${encodeURIComponent(this.config.language.split('-')[0] || 'fr')}&client=tw-ob&q=${encodeURIComponent(text.slice(0, 200))}`;
|
||||
const res = await httpFetch({
|
||||
|
||||
@@ -128,6 +128,7 @@ class VoiceController {
|
||||
}
|
||||
await this.wake.start({
|
||||
deviceId: settings.micDeviceId || undefined,
|
||||
micProcessing: settings.micProcessing ?? true,
|
||||
neuralVad,
|
||||
vad: { silenceMs: settings.silenceMs, speechRatio: SENSITIVITY_RATIO[sensitivity], minRms: SENSITIVITY_MIN_RMS[sensitivity] },
|
||||
callbacks: {
|
||||
@@ -249,6 +250,7 @@ class VoiceController {
|
||||
await audioBus.resume();
|
||||
await this.capture.start({
|
||||
deviceId: settings.micDeviceId || undefined,
|
||||
micProcessing: settings.micProcessing ?? true,
|
||||
mode: settings.captureMode,
|
||||
vad: {
|
||||
silenceMs: settings.silenceMs,
|
||||
|
||||
@@ -1,44 +1,64 @@
|
||||
/**
|
||||
* Applies the preferred voice gender to the active TTS provider: switches the local model/speaker,
|
||||
* downloads a matching French voice when none is installed, and logs what it did.
|
||||
* downloads a matching French voice when none is installed, resets an Edge voice of the other
|
||||
* gender, and explains providers that cannot honour the choice (Google Translate).
|
||||
*/
|
||||
import { Log } from '../../lib/log';
|
||||
import { findLocalVoice, inferGender, suggestedDownload } from '../../lib/voicePreference';
|
||||
import { bestLocalUpgrade, findLocalVoice, resolveEdgeVoice, suggestedDownload } from '../../lib/voicePreference';
|
||||
import { useSettings } from '../../state/settings';
|
||||
import { useVoiceModels } from '../../state/voiceModels';
|
||||
|
||||
let downloading: string | null = null;
|
||||
|
||||
/** Make the local voice match `speech.voiceGender`. Returns a short status message. */
|
||||
export async function ensurePreferredVoice(): Promise<string> {
|
||||
const genderLabel = (g: 'male' | 'female') => (g === 'male' ? 'masculine' : 'féminine');
|
||||
|
||||
/**
|
||||
* Make the active provider match `speech.voiceGender`. Returns a short status message.
|
||||
* With `upgrade`, the best local model for the language is downloaded even if a lesser voice exists.
|
||||
*/
|
||||
export async function ensurePreferredVoice(options: { upgrade?: boolean } = {}): Promise<string> {
|
||||
const { settings, update } = useSettings.getState();
|
||||
const { speech } = settings;
|
||||
const gender = speech.voiceGender ?? 'male';
|
||||
const lang = speech.language || 'fr-FR';
|
||||
|
||||
if (speech.provider === 'edge') {
|
||||
const voice = resolveEdgeVoice(speech.edgeVoice ?? '', lang, gender);
|
||||
if (voice !== (speech.edgeVoice ?? '').trim() && speech.edgeVoice) update({ speech: { edgeVoice: '' } });
|
||||
return `Voix ${genderLabel(gender)} : ${voice.split('-')[2]?.replace(/(Multilingual)?Neural$/, '') ?? voice} (Edge)`;
|
||||
}
|
||||
if (speech.provider === 'google-free') {
|
||||
return gender === 'male'
|
||||
? 'Google Translate n’a qu’une voix féminine par langue : choisissez « Microsoft Edge » ou une voix locale pour une voix masculine.'
|
||||
: '';
|
||||
}
|
||||
if (speech.provider !== 'local') return '';
|
||||
|
||||
const models = useVoiceModels.getState().models;
|
||||
if (!models.length) return '';
|
||||
const lang = speech.language || 'fr-FR';
|
||||
const current = models.find((m) => m.id === speech.localModel);
|
||||
const currentSpeaker = current?.speakers?.find((s) => s.id === speech.localSpeaker);
|
||||
const currentGender = currentSpeaker ? (currentSpeaker as { gender?: string }).gender === 'm' ? 'male' : (currentSpeaker as { gender?: string }).gender === 'f' ? 'female' : inferGender(currentSpeaker.name) : undefined;
|
||||
if (current?.installed && currentGender === gender) return '';
|
||||
const currentGender = currentSpeaker ? (currentSpeaker.gender === 'm' ? 'male' : currentSpeaker.gender === 'f' ? 'female' : undefined) : undefined;
|
||||
const upgrade = options.upgrade ? bestLocalUpgrade(models, lang) : null;
|
||||
if (current?.installed && currentGender === gender && !upgrade) return '';
|
||||
|
||||
const choice = findLocalVoice(models, lang, gender);
|
||||
const choice = upgrade ? null : findLocalVoice(models, lang, gender);
|
||||
if (choice) {
|
||||
update({ speech: { localModel: choice.modelId, localSpeaker: choice.speaker } });
|
||||
const name = models.find((m) => m.id === choice.modelId)?.speakers?.find((s) => s.id === choice.speaker)?.name ?? choice.modelId;
|
||||
Log.info('tts', `voice preference ${gender}: ${name}`);
|
||||
return `Voix ${gender === 'male' ? 'masculine' : 'féminine'} : ${name}`;
|
||||
return `Voix ${genderLabel(gender)} : ${name}`;
|
||||
}
|
||||
const download = suggestedDownload(models, lang, gender);
|
||||
const download = upgrade ?? suggestedDownload(models, lang, gender);
|
||||
if (!download || downloading === download) return download ? 'Téléchargement de la voix en cours…' : '';
|
||||
downloading = download;
|
||||
Log.info('tts', `no ${gender} voice installed, downloading ${download}`);
|
||||
Log.info('tts', `${upgrade ? 'upgrading local voice' : `no ${gender} voice installed`}, downloading ${download}`);
|
||||
try {
|
||||
await useVoiceModels.getState().download(download);
|
||||
const after = findLocalVoice(useVoiceModels.getState().models, lang, gender);
|
||||
if (after) update({ speech: { localModel: after.modelId, localSpeaker: after.speaker } });
|
||||
return after ? `Voix ${gender === 'male' ? 'masculine' : 'féminine'} installée.` : 'Voix téléchargée.';
|
||||
const name = after ? useVoiceModels.getState().models.find((m) => m.id === after.modelId)?.name : undefined;
|
||||
return after ? `Voix ${genderLabel(gender)} installée : ${name ?? after.modelId}.` : 'Voix téléchargée.';
|
||||
} catch (err) {
|
||||
Log.warn('tts', `voice download failed: ${(err as Error).message}`);
|
||||
return `Téléchargement impossible : ${(err as Error).message}`;
|
||||
|
||||
@@ -74,7 +74,7 @@ export class WakeListener {
|
||||
return this.phase;
|
||||
}
|
||||
|
||||
async start(options: { deviceId?: string; vad?: Partial<VadOptions>; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise<void> {
|
||||
async start(options: { deviceId?: string; micProcessing?: boolean; vad?: Partial<VadOptions>; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise<void> {
|
||||
if (this.phase !== 'off') return;
|
||||
this.callbacks = options.callbacks;
|
||||
this.vadOptions = options.vad ?? {};
|
||||
@@ -86,9 +86,9 @@ export class WakeListener {
|
||||
this.stream = await navigator.mediaDevices.getUserMedia({
|
||||
audio: {
|
||||
deviceId: options.deviceId ? { exact: options.deviceId } : undefined,
|
||||
echoCancellation: true,
|
||||
noiseSuppression: true,
|
||||
autoGainControl: true,
|
||||
echoCancellation: options.micProcessing !== false,
|
||||
noiseSuppression: options.micProcessing !== false,
|
||||
autoGainControl: options.micProcessing !== false,
|
||||
channelCount: 1
|
||||
}
|
||||
});
|
||||
|
||||
@@ -26,6 +26,8 @@ export interface VoiceSettings extends SttConfig {
|
||||
neuralVad: boolean;
|
||||
/** Execute short system intents locally (lock, volume, open app…) instead of asking Hermes. */
|
||||
localCommands: boolean;
|
||||
/** Chromium mic processing (echo cancellation, noise suppression, auto gain). Off often transcribes better on a headset. */
|
||||
micProcessing: boolean;
|
||||
}
|
||||
|
||||
export interface SpeechSettings extends TtsConfig {
|
||||
@@ -110,10 +112,12 @@ export const DEFAULT_SETTINGS: Settings = {
|
||||
kwsSensitivity: 3,
|
||||
neuralVad: true,
|
||||
localCommands: true,
|
||||
micProcessing: true,
|
||||
localModel: 'whisper-base'
|
||||
},
|
||||
speech: {
|
||||
provider: 'openai-compatible',
|
||||
// Edge neural voices speak out of the box (no server, no key) with a real masculine/feminine choice.
|
||||
provider: 'edge',
|
||||
apiUrl: 'http://127.0.0.1:8000/v1',
|
||||
apiKey: '',
|
||||
model: 'tts-1',
|
||||
@@ -123,8 +127,9 @@ export const DEFAULT_SETTINGS: Settings = {
|
||||
systemVoice: '',
|
||||
language: 'fr-FR',
|
||||
volume: 1,
|
||||
localModel: 'kokoro-v1',
|
||||
localSpeaker: 30,
|
||||
localModel: 'supertonic-3',
|
||||
localSpeaker: 6,
|
||||
edgeVoice: '',
|
||||
voiceGender: 'male',
|
||||
timbre: 'jarvis',
|
||||
autoSpeak: true,
|
||||
|
||||
@@ -0,0 +1,90 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import {
|
||||
defaultEdgeVoice,
|
||||
edgeConfigMessage,
|
||||
edgeRate,
|
||||
edgeSsml,
|
||||
edgeSsmlMessage,
|
||||
edgeTextFramePath,
|
||||
edgeTokenInput,
|
||||
edgeVoiceGender,
|
||||
escapeXml,
|
||||
parseEdgeBinaryFrame
|
||||
} from '../shared/edgeTts';
|
||||
|
||||
describe('edge token input', () => {
|
||||
it('uses Windows file time rounded down to 5 minutes followed by the client token', () => {
|
||||
// 2026-09-04T15:23:47Z → 15:20:00Z = 1788535200 s since 1970 → +11644473600 = 13433008800 s → ×1e7 ticks.
|
||||
const nowMs = Date.UTC(2026, 8, 4, 15, 23, 47);
|
||||
expect(edgeTokenInput(nowMs)).toBe('1343300880000000006A5AA1D4EAFF4E9FB37E23D68491D6F4');
|
||||
});
|
||||
it('is stable inside a 5-minute window and applies the clock skew', () => {
|
||||
const a = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1));
|
||||
const b = edgeTokenInput(Date.UTC(2026, 8, 4, 15, 24, 59));
|
||||
expect(a).toBe(b);
|
||||
expect(edgeTokenInput(Date.UTC(2026, 8, 4, 15, 20, 1), 300)).not.toBe(a);
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge ssml', () => {
|
||||
it('escapes text and derives the language from the voice', () => {
|
||||
const ssml = edgeSsml('Tom & Jerry <3 "ok"', 'fr-FR-HenriNeural', 1.25);
|
||||
expect(ssml).toContain("xml:lang='fr-FR'");
|
||||
expect(ssml).toContain("<voice name='fr-FR-HenriNeural'>");
|
||||
expect(ssml).toContain("rate='+25%'");
|
||||
expect(ssml).toContain('Tom & Jerry <3 "ok"');
|
||||
expect(escapeXml("l'été")).toBe('l'été');
|
||||
});
|
||||
it('formats rates with a sign', () => {
|
||||
expect(edgeRate(1)).toBe('+0%');
|
||||
expect(edgeRate(0.8)).toBe('-20%');
|
||||
expect(edgeRate(3)).toBe('+100%');
|
||||
});
|
||||
it('builds the config and ssml messages with the expected headers', () => {
|
||||
const date = new Date(Date.UTC(2026, 8, 4, 15, 23, 47));
|
||||
const config = edgeConfigMessage(date);
|
||||
expect(config.startsWith('X-Timestamp:Fri, 04 Sep 2026 15:23:47 GMT+0000 (Coordinated Universal Time)\r\n')).toBe(true);
|
||||
expect(config).toContain('Path:speech.config\r\n\r\n{');
|
||||
expect(config).toContain('audio-24khz-48kbitrate-mono-mp3');
|
||||
const ssml = edgeSsmlMessage('abc123', '<speak/>', date);
|
||||
expect(ssml).toContain('X-RequestId:abc123\r\n');
|
||||
expect(ssml).toContain('Content-Type:application/ssml+xml\r\n');
|
||||
expect(ssml.endsWith('Path:ssml\r\n\r\n<speak/>')).toBe(true);
|
||||
expect(edgeTextFramePath('X-RequestId:1\r\nContent-Type:application/json\r\nPath:turn.end\r\n\r\n{}')).toBe('turn.end');
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge binary frames', () => {
|
||||
it('splits the header (2-byte big-endian length) from the audio payload', () => {
|
||||
const header = 'X-RequestId:1\r\nContent-Type:audio/mpeg\r\nX-StreamId:2\r\nPath:audio\r\n';
|
||||
const payload = new Uint8Array([0xff, 0xfb, 0x90, 0x00]);
|
||||
const frame = new Uint8Array(2 + header.length + payload.length);
|
||||
frame[0] = header.length >> 8;
|
||||
frame[1] = header.length & 0xff;
|
||||
for (let i = 0; i < header.length; i++) frame[2 + i] = header.charCodeAt(i);
|
||||
frame.set(payload, 2 + header.length);
|
||||
const parsed = parseEdgeBinaryFrame(frame);
|
||||
expect(parsed.path).toBe('audio');
|
||||
expect([...parsed.payload]).toEqual([...payload]);
|
||||
});
|
||||
it('tolerates truncated frames', () => {
|
||||
expect(parseEdgeBinaryFrame(new Uint8Array([0x00])).payload.byteLength).toBe(0);
|
||||
expect(parseEdgeBinaryFrame(new Uint8Array([0x10, 0x00, 0x41])).path).toBe('');
|
||||
});
|
||||
});
|
||||
|
||||
describe('edge voices', () => {
|
||||
it('picks Henri / Denise for French and falls back to French for unknown languages', () => {
|
||||
expect(defaultEdgeVoice('fr-FR', 'male')).toBe('fr-FR-HenriNeural');
|
||||
expect(defaultEdgeVoice('fr', 'female')).toBe('fr-FR-DeniseNeural');
|
||||
expect(defaultEdgeVoice('en-GB', 'male')).toBe('en-US-AndrewMultilingualNeural');
|
||||
expect(defaultEdgeVoice('xx', 'female')).toBe('fr-FR-DeniseNeural');
|
||||
});
|
||||
it('knows the gender of the common French voices', () => {
|
||||
expect(edgeVoiceGender('fr-FR-HenriNeural')).toBe('male');
|
||||
expect(edgeVoiceGender('fr-FR-RemyMultilingualNeural')).toBe('male');
|
||||
expect(edgeVoiceGender('fr-FR-VivienneMultilingualNeural')).toBe('female');
|
||||
expect(edgeVoiceGender('fr-CA-SylvieNeural')).toBe('female');
|
||||
expect(edgeVoiceGender('zz-ZZ-NobodyNeural')).toBeUndefined();
|
||||
});
|
||||
});
|
||||
@@ -1,36 +1,61 @@
|
||||
import { describe, expect, it } from 'vitest';
|
||||
import { defaultOpenAiVoice, findLocalVoice, inferGender, rankSystemVoice, suggestedDownload } from '../src/lib/voicePreference';
|
||||
import type { VoiceModelStatus } from '../shared/voice';
|
||||
import { bestLocalUpgrade, defaultOpenAiVoice, findLocalVoice, inferGender, pickSpeaker, rankSystemVoice, resolveEdgeVoice, suggestedDownload } from '../src/lib/voicePreference';
|
||||
import type { VoiceModelStatus, VoiceSpeaker } from '../shared/voice';
|
||||
|
||||
const model = (id: string, installed: boolean, speakers: Array<{ id: number; name: string; lang: string }>): VoiceModelStatus =>
|
||||
({ id, kind: 'tts', engine: 'piper', name: id, description: '', languages: ['fr'], sizeMb: 1, url: '', dir: id, files: [], speakers, installed, installedBytes: 0 }) as unknown as VoiceModelStatus;
|
||||
const model = (id: string, installed: boolean, speakers: VoiceSpeaker[], languages = ['fr']): VoiceModelStatus =>
|
||||
({ id, kind: 'tts', engine: 'piper', name: id, description: '', languages, sizeMb: 1, url: '', dir: id, files: [], speakers, installed, installedBytes: 0 }) as unknown as VoiceModelStatus;
|
||||
|
||||
describe('inferGender', () => {
|
||||
it('reads catalog labels, system voices and kokoro ids', () => {
|
||||
it('reads catalog labels, system voices, kokoro ids and Edge short names', () => {
|
||||
expect(inferGender('Piper Tom (homme, français)')).toBe('male');
|
||||
expect(inferGender('Siwis (femme, français)')).toBe('female');
|
||||
expect(inferGender('Microsoft Paul - French (France)')).toBe('male');
|
||||
expect(inferGender('Microsoft Hortense - French (France)')).toBe('female');
|
||||
expect(inferGender('am_adam')).toBe('male');
|
||||
expect(inferGender('fr-FR-HenriNeural')).toBe('male');
|
||||
expect(inferGender('fr-FR-DeniseNeural')).toBe('female');
|
||||
expect(inferGender('Voix 3')).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe('local voice selection', () => {
|
||||
const supertonic: VoiceSpeaker[] = [
|
||||
{ id: 6, name: 'Homme 2 (grave, posé)', lang: 'multi', gender: 'm' },
|
||||
{ id: 0, name: 'Femme 1', lang: 'multi', gender: 'f' }
|
||||
];
|
||||
const models = [
|
||||
model('kokoro-v1', true, [{ id: 30, name: 'Siwis (femme, français)', lang: 'fr' }, { id: 4, name: 'Adam (homme, anglais US)', lang: 'en' }]),
|
||||
model('kokoro-v1', true, [{ id: 30, name: 'Siwis (femme, français)', lang: 'fr' }, { id: 4, name: 'Adam (homme, anglais US)', lang: 'en' }], ['en', 'fr', 'multi']),
|
||||
model('piper-fr-tom', false, [{ id: 0, name: 'Tom', lang: 'fr' }]),
|
||||
model('piper-fr-upmc', true, [{ id: 0, name: 'Jessica (femme)', lang: 'fr' }, { id: 1, name: 'Pierre (homme)', lang: 'fr' }])
|
||||
model('piper-fr-upmc', true, [{ id: 0, name: 'Jessica (femme)', lang: 'fr' }, { id: 1, name: 'Pierre (homme)', lang: 'fr' }]),
|
||||
model('supertonic-3', false, supertonic, ['fr', 'en', 'multi'])
|
||||
];
|
||||
it('prefers an installed speaker of the wanted gender in the right language', () => {
|
||||
expect(findLocalVoice(models, 'fr-FR', 'male')).toEqual({ modelId: 'piper-fr-upmc', speaker: 1 });
|
||||
expect(findLocalVoice(models, 'fr-FR', 'female')).toEqual({ modelId: 'kokoro-v1', speaker: 30 });
|
||||
expect(findLocalVoice(models.slice(0, 2), 'fr', 'male')).toBeNull();
|
||||
});
|
||||
it('suggests a French download when nothing matches', () => {
|
||||
expect(suggestedDownload(models, 'fr', 'male')).toBe('piper-fr-tom');
|
||||
it('ranks Supertonic above Kokoro and Piper once installed', () => {
|
||||
const installed = models.map((m) => (m.id === 'supertonic-3' ? { ...m, installed: true } : m));
|
||||
expect(findLocalVoice(installed, 'fr-FR', 'male')).toEqual({ modelId: 'supertonic-3', speaker: 6 });
|
||||
expect(findLocalVoice(installed, 'fr-FR', 'female')).toEqual({ modelId: 'supertonic-3', speaker: 0 });
|
||||
});
|
||||
it('suggests Supertonic first, then the Piper voices, for a French download', () => {
|
||||
expect(suggestedDownload(models, 'fr', 'male')).toBe('supertonic-3');
|
||||
expect(suggestedDownload(models.filter((m) => m.id !== 'supertonic-3'), 'fr', 'male')).toBe('piper-fr-tom');
|
||||
expect(suggestedDownload(models, 'en', 'male')).toBeNull();
|
||||
});
|
||||
it('reports the best local model still to download', () => {
|
||||
expect(bestLocalUpgrade(models, 'fr-FR')).toBe('supertonic-3');
|
||||
expect(bestLocalUpgrade(models.map((m) => ({ ...m, installed: true })), 'fr-FR')).toBeNull();
|
||||
});
|
||||
it('keeps the gender when switching models and falls back to the language', () => {
|
||||
expect(pickSpeaker(models[3], 'fr-FR', 'male')).toBe(6);
|
||||
expect(pickSpeaker(models[3], 'fr-FR', 'female')).toBe(0);
|
||||
expect(pickSpeaker(models[2], 'fr', 'male')).toBe(1);
|
||||
// Kokoro: no masculine French voice → the French voice, not an English one.
|
||||
expect(pickSpeaker(models[0], 'fr', 'male')).toBe(30);
|
||||
expect(pickSpeaker({ speakers: [] }, 'fr', 'male')).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
describe('provider defaults', () => {
|
||||
@@ -38,6 +63,12 @@ describe('provider defaults', () => {
|
||||
expect(defaultOpenAiVoice('male')).toBe('onyx');
|
||||
expect(defaultOpenAiVoice('female')).toBe('nova');
|
||||
});
|
||||
it('keeps an explicit Edge voice only when it matches the gender', () => {
|
||||
expect(resolveEdgeVoice('', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
|
||||
expect(resolveEdgeVoice('fr-FR-RemyMultilingualNeural', 'fr-FR', 'male')).toBe('fr-FR-RemyMultilingualNeural');
|
||||
expect(resolveEdgeVoice('fr-FR-DeniseNeural', 'fr-FR', 'male')).toBe('fr-FR-HenriNeural');
|
||||
expect(resolveEdgeVoice('fr-CH-ArianeNeural', 'fr-FR', 'female')).toBe('fr-CH-ArianeNeural');
|
||||
});
|
||||
it('ranks system voices by language then gender', () => {
|
||||
const paul = { name: 'Microsoft Paul - French (France)', lang: 'fr-FR', localService: true };
|
||||
const hortense = { name: 'Microsoft Hortense - French (France)', lang: 'fr-FR', localService: true };
|
||||
|
||||
Reference in new issue
Block a user