diff --git a/README.md b/README.md index a0ca7dd..54d816a 100644 --- a/README.md +++ b/README.md @@ -1,7 +1,7 @@ # EveFlow 2 — Interface vocale JARVIS pour Hermes Agent [![Build](https://img.shields.io/github/actions/workflow/status/R0m1k3/EveFlow/windows-release.yml?style=flat-square)](https://github.com/R0m1k3/EveFlow/actions) -[![Version](https://img.shields.io/badge/version-2.2.0-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases) +[![Version](https://img.shields.io/badge/version-2.3.0-brightgreen.svg?style=flat-square)](https://github.com/R0m1k3/EveFlow/releases) [![License](https://img.shields.io/badge/license-MIT-lightgrey.svg?style=flat-square)](LICENSE) **EveFlow** est un compagnon de bureau Windows qui transforme [Hermes Agent](https://hermes-agent.nousresearch.com/) en assistant vocal à la JARVIS : un noyau holographique réactif au son, une conversation en streaming, les outils, sous-agents, approbations, crons, skills et sessions d'Hermes pilotés depuis un seul HUD. @@ -23,6 +23,9 @@ La version 2 est une réécriture complète : plus de robot 3D, un pipeline voca * **Détection d'activité vocale** (seuil adaptatif, sensibilité et silence de fin réglables) : l'enregistrement s'arrête tout seul quand vous avez fini de parler. * **Mains libres** : le micro se réactive après chaque réponse. * **Modèles intégrés, hors ligne** (sherpa-onnx dans un processus séparé) : reconnaissance Whisper (base, small, large-v3 turbo) ou SenseVoice, synthèse Kokoro v1.0 (voix française Siwis et voix anglaises) ou Piper (Siwis, Tom, UPMC). Les modèles se téléchargent depuis **Paramètres → Modèles locaux** et tournent sur le processeur. +* **Fin de phrase neuronale** : en écoute permanente, Silero VAD (0,6 Mo, sherpa-onnx) décide du début et de la fin de la commande à la place du seuil d'énergie ; moins de faux départs sur le bruit, coupure plus nette. Repli automatique sur le VAD énergétique si le modèle n'est pas installé. +* **Vision d'écran** : « Jarvis, regarde mon écran » (ou le bouton de la barre de commande) joint une capture de l'écran principal à la question envoyée à Hermes. +* **Actions locales instantanées** : « verrouille la session », « monte le son », « coupe le son », « piste suivante », « ouvre Spotify », « ouvre github.com »… exécutées sur le PC sans passer par Hermes, résultat lu à voix haute. Liste blanche d'actions dans le processus principal, désactivable dans les paramètres. * **Écoute permanente** : un détecteur de mot-clé de 3 Mo (sherpa-onnx, keyword spotting) tourne en continu sur le micro, quasi gratuit en CPU. « Jarvis » (ou n'importe quel mot-clé) ouvre l'écoute, « Jarvis, allume… » envoie directement la commande, et le mot coupe la voix en cours. Alternative : filtre du mot après transcription en mains libres. * **STT externe** : n'importe quelle API `/v1/audio/transcriptions` compatible OpenAI (Qwen3-ASR, Whisper, Speaches, faster-whisper-server, LocalAI, OpenAI). Repli sur la reconnaissance Chromium. * **TTS externe** : API `/v1/audio/speech` compatible OpenAI, voix système Windows ou Google Translate. Lecture phrase par phrase pendant le streaming, préchargement du segment suivant, coupure instantanée. diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 366ac83..44621b9 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -11,7 +11,9 @@ | Synthèse vocale locale | Fait | Kokoro v1.0 (voix française Siwis) et Piper fr | | Mot d'activation permanent | Fait (2.2.0) | Keyword spotting sherpa-onnx en continu ; mot-clé libre encodé en BPE ; validé sur audio réel (détection, zéro faux positif sur le test anglais) | | Mot d'activation après transcription | Fait | Filtre « Jarvis … » en mains libres, tolérant aux erreurs de transcription | -| Détection de fin de phrase | Fait | VAD énergétique adaptatif, pré-roll 400 ms | +| Détection de fin de phrase | Fait (2.3.0) | Silero VAD neuronal dans le worker (segment renvoyé au renderer), repli sur le VAD énergétique si le modèle manque | +| Vision d'écran | Fait (2.3.0) | Capture `desktopCapturer` jointe à la requête Hermes (bouton, ou « regarde mon écran ») | +| Actions système locales | Fait (2.3.0) | Verrouillage, volume et touches média, ouvrir une application ou une URL, presse-papiers, recherche de fichiers ; intentions courtes exécutées sans passer par Hermes | | Hermes : runs, sessions, chat completions | Fait | Transport choisi selon `/v1/capabilities` | | Approbations, steer, stop | Fait | Modales, injection de consigne en cours de run | | Crons, skills, toolsets, sessions | Fait | Panneau Hermes Ops | @@ -23,10 +25,10 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai), [JarvisAi](https://github.com/PanPenek/JarvisAi), [bertrandmbanwi/Jarvis](https://github.com/bertrandmbanwi/Jarvis), [InterGenJLU/jarvis](https://github.com/InterGenJLU/jarvis), [livekit-wakeword](https://livekit.com/blog/livekit-wakeword), [sherpa-onnx keyword spotting](https://k2-fsa.github.io/sherpa/onnx/kws/index.html), [Hermes Agent features](https://hermes-agent.nousresearch.com/docs/user-guide/features/overview). 1. **Mot d'activation permanent, quasi gratuit en CPU.** Les projets de référence utilisent openWakeWord (« hey jarvis ») ou un modèle de keyword spotting qui écoute en continu, au lieu de transcrire chaque phrase. sherpa-onnx fournit un modèle KWS anglais de 3,3 Mo qui accepte n'importe quel mot-clé sans réentraînement ; « JARVIS » s'encode `▁JA R VI S` avec son modèle BPE (vérifié). Livré en 2.2.0 (voir l'étape 1 ci-dessous). -2. **VAD neuronal (Silero) au lieu du seuil d'énergie.** Fin de phrase plus nette (environ 500 ms gagnés) et beaucoup moins de faux départs sur le bruit ambiant. Silero est déjà livré dans sherpa-onnx (`silero_vad.onnx`, 0,6 Mo). +2. **VAD neuronal (Silero) au lieu du seuil d'énergie.** Fin de phrase plus nette (environ 500 ms gagnés) et beaucoup moins de faux départs sur le bruit ambiant. Silero est déjà livré dans sherpa-onnx (`silero_vad.onnx`, 0,6 Mo). Livré en 2.3.0. 3. **Latence perçue sous la seconde.** Les références visent 1 s entre la fin de parole et le premier mot prononcé : STT rapide, premier token en streaming, TTS phrase par phrase (déjà en place), et un modèle Hermes rapide pour la conversation courante. -4. **Vision d'écran.** Capture d'écran à la demande (« Jarvis, qu'est-ce que je regarde ? ») envoyée à Hermes comme image, ou lecture d'une fenêtre. Hermes accepte déjà les images inline. -5. **Actions système locales.** Ouvrir une application, régler le volume, verrouiller la session, chercher un fichier ; ce sont des outils EveFlow côté client à exposer à Hermes (mode chat completions) ou un petit serveur MCP local que Hermes appelle. +4. **Vision d'écran.** Capture d'écran à la demande (« Jarvis, qu'est-ce que je regarde ? ») envoyée à Hermes comme image, ou lecture d'une fenêtre. Hermes accepte déjà les images inline. Livré en 2.3.0 (transport chat completions pour les images). +5. **Actions système locales.** Ouvrir une application, régler le volume, verrouiller la session, chercher un fichier. Livré en 2.3.0 sous forme d'intentions courtes exécutées localement avant Hermes ; reste à exposer les mêmes actions à Hermes via un serveur MCP local. 6. **Proactivité.** Notifications parlées à l'arrivée d'un cron, rappel, événement webhook, avec un résumé plutôt que la lecture intégrale ; c'est en partie fait via le webhook, à enrichir avec des règles (heures calmes, priorité). 7. **Mémoire et personnalisation.** Hermes gère la mémoire longue durée (`X-Hermes-Session-Key`) ; côté EveFlow, un profil (nom, préférences de voix, style de réponse) déjà transmis dans les instructions. @@ -36,11 +38,13 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai) - Le renderer garde un seul flux micro (AudioWorklet 16 kHz) et envoie des blocs de 256 ms au processus principal, qui alimente le `KeywordSpotter` sherpa-onnx dans le worker. - Mots-clés encodés en BPE (table SentencePiece pour les mots courants, repli glouton sur le vocabulaire du modèle), sensibilité réglable (seuil 0,45 → 0,12). - À la détection : chime, capture de la commande sur le même flux (VAD énergétique, pré-roll 400 ms), transcription locale ou API, envoi à Hermes ; retour automatique à l'écoute. -- Reste à faire : remplacer le VAD énergétique par Silero pour la fin de phrase. +- Fin de phrase Silero livrée en 2.3.0 : le renderer envoie des trames de 128 ms au worker pendant la commande, le worker renvoie le segment WAV complet ; VAD énergétique en repli. -### Étape 2 : vision et actions locales -- Outil `capture_screen` (Electron `desktopCapturer`) qui joint une capture à la requête Hermes. -- Outils système : `open_app`, `set_volume`, `lock_session`, `find_file`, `clipboard` ; exposés en chat completions et via un serveur MCP local pour les transports runs/sessions. +### Étape 2 : vision et actions locales — livrée en 2.3.0 +- Capture d'écran (`desktopCapturer`, JPEG 1600 px) jointe à la requête Hermes : bouton dans la barre de commande, ou phrase « regarde mon écran… » à l'oral comme à l'écrit. +- Actions locales (processus principal, liste blanche) : verrouiller, volume/mute/lecture/piste, ouvrir une application connue ou une URL http(s), presse-papiers, recherche de fichiers dans Documents/Bureau/Téléchargements/Images. +- Routeur d'intentions FR/EN (`src/services/localCommands.ts`) exécuté avant l'envoi à Hermes ; désactivable dans Paramètres → Micro. Les phrases composées (« ouvre X puis… ») partent à Hermes. +- Reste à faire : exposer ces actions à Hermes lui-même (serveur MCP local) pour qu'un run puisse les enchaîner. ### Étape 3 : conversation plus naturelle - Barge-in réel : couper la voix dès que l'utilisateur parle (déjà préparé, à valider avec l'annulation d'écho Windows). @@ -61,5 +65,8 @@ Sources : [jarvis-desktop-ai](https://github.com/ccarloshenri/jarvis-desktop-ai) | Kokoro (fr) → Whisper base, phrase courte de 2 s | synthèse 2,9 s, transcription 2,5 s, texte approximatif | | Kokoro (fr) → Whisper base, phrase de 7 s | transcription correcte à un mot près | | Piper (fr) → Whisper base | transcription exacte | +| Silero VAD sur phrase Kokoro de 2,3 s (2.3.0) | un seul segment, début et fin détectés, 6,7 s d'audio traités en 190 ms | +| Capture d'écran → Hermes (2.3.0) | JPEG de 107 ko reçu côté Hermes (mock chat completions) | +| Intention locale « coupe le son » (2.3.0) | traitée sans Hermes, résultat affiché dans le fil | Sur un PC à 28 cœurs les temps sont nettement plus courts. Whisper small est maintenant recommandé pour le français. diff --git a/electron/ipc/system.ts b/electron/ipc/system.ts new file mode 100644 index 0000000..b0976f2 --- /dev/null +++ b/electron/ipc/system.ts @@ -0,0 +1,184 @@ +/** + * Local "JARVIS" actions: screenshot for vision, and an allow-list of system actions + * (lock, open an application or URL, media keys, clipboard, file search). + */ +import { app, clipboard, desktopCapturer, ipcMain, screen, shell } from 'electron'; +import { execFile } from 'node:child_process'; +import fs from 'node:fs'; +import os from 'node:os'; +import path from 'node:path'; +import { IPC } from '../../shared/ipc'; +import type { SystemAction, SystemActionResult } from '../../shared/bridge'; +import { log } from '../logger'; + +const MEDIA_KEYS: Record = { + 'volume-up': 175, + 'volume-down': 174, + mute: 173, + 'play-pause': 179, + next: 176, + previous: 177 +}; + +/** Applications the assistant may launch by name (Windows aliases + common Linux/macOS names). */ +const APP_ALIASES: Record = { + 'bloc-notes': ['notepad.exe', 'gedit', 'TextEdit'], + notepad: ['notepad.exe', 'gedit', 'TextEdit'], + calculatrice: ['calc.exe', 'gnome-calculator', 'Calculator'], + calc: ['calc.exe', 'gnome-calculator', 'Calculator'], + explorateur: ['explorer.exe', 'nautilus', 'Finder'], + explorer: ['explorer.exe', 'nautilus', 'Finder'], + terminal: ['wt.exe', 'cmd.exe', 'gnome-terminal', 'Terminal'], + cmd: ['cmd.exe'], + powershell: ['powershell.exe'], + paint: ['mspaint.exe'], + chrome: ['chrome', 'google-chrome', 'Google Chrome'], + edge: ['msedge', 'microsoft-edge', 'Microsoft Edge'], + firefox: ['firefox', 'Firefox'], + vscode: ['code', 'Visual Studio Code'], + code: ['code', 'Visual Studio Code'], + spotify: ['spotify', 'Spotify'], + discord: ['discord', 'Discord'], + steam: ['steam', 'Steam'], + word: ['winword', 'Microsoft Word'], + excel: ['excel', 'Microsoft Excel'], + outlook: ['outlook', 'Microsoft Outlook'], + teams: ['ms-teams', 'Microsoft Teams'], + 'task manager': ['taskmgr.exe'], + 'gestionnaire des tâches': ['taskmgr.exe'], + paramètres: ['ms-settings:'], + settings: ['ms-settings:'] +}; + +function run(cmd: string, args: string[], timeoutMs = 8000): Promise<{ code: number; out: string }> { + return new Promise((resolve) => { + execFile(cmd, args, { timeout: timeoutMs, windowsHide: true }, (err, stdout, stderr) => { + resolve({ code: err ? 1 : 0, out: `${stdout}${stderr}`.trim() }); + }); + }); +} + +async function openApp(name: string): Promise { + const key = name.trim().toLowerCase(); + if (!key || key.length > 40) return { ok: false, message: 'Nom d’application invalide' }; + const candidates = APP_ALIASES[key] ?? [key.replace(/[^a-z0-9 ._-]/gi, '')]; + if (process.platform === 'win32') { + for (const candidate of candidates) { + if (candidate.endsWith(':')) { + await shell.openExternal(candidate); + return { ok: true, message: `${name} ouvert` }; + } + // `start` resolves App Paths, PATH and Start Menu names. + const result = await run('cmd.exe', ['/c', 'start', '', candidate]); + if (result.code === 0) return { ok: true, message: `${name} lancé` }; + } + return { ok: false, message: `Impossible de lancer ${name}` }; + } + for (const candidate of candidates) { + const result = process.platform === 'darwin' ? await run('open', ['-a', candidate]) : await run('sh', ['-c', `command -v ${JSON.stringify(candidate)} >/dev/null && (nohup ${JSON.stringify(candidate)} >/dev/null 2>&1 &)`]); + if (result.code === 0) return { ok: true, message: `${name} lancé` }; + } + return { ok: false, message: `Application introuvable : ${name}` }; +} + +async function lockSession(): Promise { + if (process.platform === 'win32') { + const r = await run('rundll32.exe', ['user32.dll,LockWorkStation']); + return { ok: r.code === 0, message: r.code === 0 ? 'Session verrouillée' : r.out }; + } + if (process.platform === 'darwin') { + const r = await run('osascript', ['-e', 'tell application "System Events" to keystroke "q" using {command down, control down}']); + return { ok: r.code === 0, message: r.out }; + } + const r = await run('sh', ['-c', 'loginctl lock-session || xdg-screensaver lock || gnome-screensaver-command -l']); + return { ok: r.code === 0, message: r.code === 0 ? 'Session verrouillée' : r.out }; +} + +async function mediaKey(key: string): Promise { + const code = MEDIA_KEYS[key]; + if (!code) return { ok: false, message: 'Touche inconnue' }; + if (process.platform === 'win32') { + const script = `$s=Add-Type -MemberDefinition '[DllImport("user32.dll")] public static extern void keybd_event(byte b,byte s,uint f,UIntPtr e);' -Name K -Namespace W -PassThru; $s::keybd_event(${code},0,0,[UIntPtr]::Zero); $s::keybd_event(${code},0,2,[UIntPtr]::Zero)`; + const r = await run('powershell.exe', ['-NoProfile', '-NonInteractive', '-Command', script]); + return { ok: r.code === 0, message: r.code === 0 ? key : r.out }; + } + const xdo: Record = { 'volume-up': 'XF86AudioRaiseVolume', 'volume-down': 'XF86AudioLowerVolume', mute: 'XF86AudioMute', 'play-pause': 'XF86AudioPlay', next: 'XF86AudioNext', previous: 'XF86AudioPrev' }; + const r = await run('sh', ['-c', `xdotool key ${xdo[key]}`]); + return { ok: r.code === 0, message: r.code === 0 ? key : 'xdotool indisponible' }; +} + +async function findFiles(query: string): Promise { + const needle = query.trim().toLowerCase(); + if (needle.length < 2) return { ok: false, message: 'Recherche trop courte' }; + const roots = ['Documents', 'Desktop', 'Downloads', 'Pictures'].map((d) => path.join(os.homedir(), d)); + const hits: string[] = []; + const deadline = Date.now() + 4000; + const walk = async (dir: string, depth: number) => { + if (depth > 4 || hits.length >= 25 || Date.now() > deadline) return; + let entries: fs.Dirent[] = []; + try { + entries = await fs.promises.readdir(dir, { withFileTypes: true }); + } catch { + return; + } + for (const entry of entries) { + if (entry.name.startsWith('.') || entry.name === 'node_modules') continue; + const full = path.join(dir, entry.name); + if (entry.name.toLowerCase().includes(needle)) hits.push(full); + if (entry.isDirectory()) await walk(full, depth + 1); + if (hits.length >= 25) return; + } + }; + for (const root of roots) await walk(root, 0); + return { ok: true, message: `${hits.length} résultat(s)`, data: hits }; +} + +export async function captureScreen(maxWidth = 1600): Promise { + const display = screen.getPrimaryDisplay(); + const scale = Math.min(1, maxWidth / display.size.width); + const sources = await desktopCapturer.getSources({ + types: ['screen'], + thumbnailSize: { width: Math.round(display.size.width * scale), height: Math.round(display.size.height * scale) } + }); + const primary = sources.find((s) => s.display_id === String(display.id)) ?? sources[0]; + if (!primary) throw new Error('Aucun écran capturable'); + return primary.thumbnail.toJPEG(82).toString('base64').replace(/^/, 'data:image/jpeg;base64,'); +} + +export async function runSystemAction(action: SystemAction): Promise { + switch (action.type) { + case 'lock': + return lockSession(); + case 'open-app': + return openApp(String(action.name ?? '')); + case 'open-url': { + const url = String(action.url ?? ''); + if (!/^https?:\/\//i.test(url) || url.length > 2048) return { ok: false, message: 'URL refusée' }; + await shell.openExternal(url); + return { ok: true, message: 'Ouvert dans le navigateur' }; + } + case 'media': + return mediaKey(String(action.key)); + case 'clipboard-read': + return { ok: true, data: (await clipboard.readText()).slice(0, 20_000) }; + case 'clipboard-write': + await clipboard.writeText(String(action.text ?? '').slice(0, 100_000)); + return { ok: true, message: 'Copié dans le presse-papiers' }; + case 'find-files': + return findFiles(String(action.query ?? '')); + default: + return { ok: false, message: 'Action inconnue' }; + } +} + +export function registerSystemIpc(): void { + ipcMain.handle(IPC.screenCapture, (_e, maxWidth?: number) => captureScreen(Number(maxWidth) || 1600)); + ipcMain.handle(IPC.systemAction, async (_e, action: SystemAction) => { + if (!action || typeof action.type !== 'string') throw new Error('Action invalide'); + log('INFO', 'system', `action ${action.type}`); + const result = await runSystemAction(action); + if (!result.ok) log('WARN', 'system', `action ${action.type} failed: ${result.message}`); + return result; + }); + void app; +} diff --git a/electron/main.ts b/electron/main.ts index b1b5b09..9fb41b6 100644 --- a/electron/main.ts +++ b/electron/main.ts @@ -9,6 +9,7 @@ import { getSharedDirectory, registerFilesIpc } from './ipc/files'; import { createMainWindow, getMainWindow, getWindowMode, setWindowMode, toggleWindowVisibility, windowEvents } from './window'; import { getWebhookStatus, startWebhookServer, stopWebhookServer } from './webhook'; import { registerVoiceIpc } from './voice/ipc'; +import { registerSystemIpc } from './ipc/system'; import { stopEngine } from './voice/engine'; // Audio playback must never be blocked behind a user gesture (TTS starts on incoming events). @@ -159,6 +160,7 @@ async function boot(): Promise { registerFilesIpc(); registerCoreIpc(); registerVoiceIpc(); + registerSystemIpc(); createMainWindow(); createTray(); registerShortcuts(); diff --git a/electron/preload.ts b/electron/preload.ts index d48c907..dbc8c14 100644 --- a/electron/preload.ts +++ b/electron/preload.ts @@ -13,8 +13,8 @@ import { type WebhookStatus, type WindowMode } from '../shared/ipc'; -import type { EveFlowBridge, Unsubscribe } from '../shared/bridge'; -import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice'; +import type { EveFlowBridge, SystemAction, SystemActionResult, Unsubscribe } from '../shared/bridge'; +import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type VadEvent, type VadStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceDownloadProgress, type VoiceEngineStatus, type VoiceModelStatus } from '../shared/voice'; function subscribe(channel: string, callback: (payload: T) => void): Unsubscribe { const listener = (_event: Electron.IpcRendererEvent, payload: T) => callback(payload); @@ -43,7 +43,9 @@ const api: EveFlowBridge = { }, system: { metrics: () => ipcRenderer.invoke(IPC.metrics) as Promise, - appInfo: () => ipcRenderer.invoke(IPC.appInfo) as Promise + appInfo: () => ipcRenderer.invoke(IPC.appInfo) as Promise, + captureScreen: (maxWidth?: number) => ipcRenderer.invoke(IPC.screenCapture, maxWidth) as Promise, + action: (action: SystemAction) => ipcRenderer.invoke(IPC.systemAction, action) as Promise }, files: { readLocal: (filePath: string) => ipcRenderer.invoke(IPC.readLocalFile, filePath) as Promise, @@ -73,7 +75,11 @@ const api: EveFlowBridge = { kwsStart: (req: KwsStartRequest) => ipcRenderer.invoke(VOICE_IPC.kwsStart, req) as Promise<{ accepted: string[]; rejected: string[] }>, kwsStop: () => ipcRenderer.invoke(VOICE_IPC.kwsStop) as Promise, kwsAudio: (pcm: Uint8Array, sampleRate: number) => ipcRenderer.send(VOICE_IPC.kwsAudio, pcm, sampleRate), - onKwsDetected: (cb: (detection: KwsDetection) => void) => subscribe(VOICE_IPC.kwsDetected, cb) + onKwsDetected: (cb: (detection: KwsDetection) => void) => subscribe(VOICE_IPC.kwsDetected, cb), + vadStart: (req: VadStartRequest) => ipcRenderer.invoke(VOICE_IPC.vadStart, req) as Promise, + vadStop: () => ipcRenderer.invoke(VOICE_IPC.vadStop) as Promise, + vadAudio: (pcm: Uint8Array, sampleRate: number) => ipcRenderer.send(VOICE_IPC.vadAudio, pcm, sampleRate), + onVadEvent: (cb: (event: VadEvent) => void) => subscribe(VOICE_IPC.vadEvent, cb) } }; diff --git a/electron/voice/catalog.ts b/electron/voice/catalog.ts index b2d9569..62f7fa1 100644 --- a/electron/voice/catalog.ts +++ b/electron/voice/catalog.ts @@ -86,6 +86,19 @@ export const VOICE_CATALOG: VoiceModelSpec[] = [ ], recommended: true }, + { + id: 'silero-vad', + kind: 'vad', + engine: 'silero', + name: 'Silero VAD (fin de phrase neuronale, 0,6 Mo)', + description: 'Détecte précisément le début et la fin de la parole pendant l’écoute permanente ; moins de faux départs sur le bruit.', + languages: ['multi'], + sizeMb: 1, + url: `${ASR}/silero_vad.onnx`, + dir: 'silero-vad', + files: ['silero_vad.onnx'], + recommended: true + }, { id: 'kokoro-v1', kind: 'tts', diff --git a/electron/voice/engine.ts b/electron/voice/engine.ts index a3a0b00..c880963 100644 --- a/electron/voice/engine.ts +++ b/electron/voice/engine.ts @@ -6,7 +6,7 @@ import { utilityProcess, type UtilityProcess, type WebContents } from 'electron' import { createHash } from 'node:crypto'; import fs from 'node:fs'; import path from 'node:path'; -import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VoiceEngineStatus } from '../../shared/voice'; +import { VOICE_IPC, type KwsDetection, type KwsStartRequest, type SynthesizeRequest, type SynthesizeResult, type TranscribeRequest, type TranscribeResult, type VadEvent, type VadStartRequest, type VoiceEngineStatus } from '../../shared/voice'; import { buildKeywordsFile, parseTokens } from '../../shared/keywords'; import { findModel } from './catalog'; import { isInstalled, modelDir, modelsDir } from './models'; @@ -16,6 +16,8 @@ import { log } from '../logger'; let kwsSubscriber: WebContents | null = null; let kwsActive = false; let kwsRequest: KwsStartRequest | null = null; +let vadSubscriber: WebContents | null = null; +let vadActive = false; interface Pending { resolve: (value: unknown) => void; @@ -42,7 +44,20 @@ function spawn(): UtilityProcess { child = proc; proc.stdout?.on('data', (d: Buffer) => log('DEBUG', 'voice-worker', d.toString().trim())); proc.stderr?.on('data', (d: Buffer) => log('WARN', 'voice-worker', d.toString().trim())); - proc.on('message', (msg: { id?: number; type?: string; ok?: boolean; result?: unknown; error?: string; keyword?: string; at?: number }) => { + proc.on('message', (msg: { id?: number; type?: string; ok?: boolean; result?: unknown; error?: string; keyword?: string; at?: number; event?: { type: string; wav?: string; durationSec?: number } }) => { + if (msg.type === 'vad.event' && msg.event) { + if (vadSubscriber && !vadSubscriber.isDestroyed()) { + const ev = msg.event; + const payload: VadEvent = + ev.type === 'segment' && ev.wav + ? { type: 'segment', wav: new Uint8Array(Buffer.from(ev.wav, 'base64')), durationSec: ev.durationSec ?? 0 } + : ev.type === 'speech-start' + ? { type: 'speech-start' } + : { type: 'error', message: 'événement VAD inconnu' }; + vadSubscriber.send(VOICE_IPC.vadEvent, payload); + } + return; + } if (msg.type === 'kws.detected') { if (kwsSubscriber && !kwsSubscriber.isDestroyed()) { kwsSubscriber.send(VOICE_IPC.kwsDetected, { keyword: msg.keyword ?? '', at: msg.at ?? Date.now() } satisfies KwsDetection); @@ -180,6 +195,35 @@ export function kwsFeed(pcm: Uint8Array, sampleRate: number): void { } } +/** Neural end-of-speech detection (Silero) on frames streamed by the renderer. */ +export async function vadStart(req: VadStartRequest, sender: WebContents): Promise { + const spec = findModel(req.modelId); + if (!spec || spec.kind !== 'vad') throw new Error('Modèle VAD introuvable'); + if (!isInstalled(spec)) throw new Error('Silero VAD non installé (Paramètres → Modèles locaux).'); + vadSubscriber = sender; + await request( + { type: 'vad.start', model: { id: spec.id, engine: spec.engine, dir: modelDir(spec), files: spec.files }, silenceMs: req.silenceMs, threshold: req.threshold, maxUtteranceSec: req.maxUtteranceSec }, + 30_000 + ); + vadActive = true; +} + +export async function vadStop(): Promise { + vadActive = false; + vadSubscriber = null; + if (child) await request({ type: 'vad.stop' }, 10_000).catch(() => undefined); +} + +export function vadFeed(pcm: Uint8Array, sampleRate: number): void { + if (!vadActive || !child) return; + const b64 = Buffer.from(pcm.buffer, pcm.byteOffset, pcm.byteLength).toString('base64'); + try { + child.postMessage({ id: -1, type: 'vad.audio', pcm: b64, sampleRate }); + } catch (err) { + log('WARN', 'voice', `vad feed failed: ${(err as Error).message}`); + } +} + /** Native model memory is only released deterministically by restarting the worker. */ export function unload(_modelId?: string): Promise { stopEngine(); diff --git a/electron/voice/ipc.ts b/electron/voice/ipc.ts index a89300f..831871a 100644 --- a/electron/voice/ipc.ts +++ b/electron/voice/ipc.ts @@ -1,6 +1,6 @@ import { ipcMain } from 'electron'; -import { VOICE_IPC, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest } from '../../shared/voice'; -import { engineStatus, kwsFeed, kwsStart, kwsStop, synthesize, transcribe, unload } from './engine'; +import { VOICE_IPC, type KwsStartRequest, type SynthesizeRequest, type TranscribeRequest, type VadStartRequest } from '../../shared/voice'; +import { engineStatus, kwsFeed, kwsStart, kwsStop, synthesize, transcribe, unload, vadFeed, vadStart, vadStop } from './engine'; import { cancelDownload, downloadModel, listModels, removeModel } from './models'; export function registerVoiceIpc(): void { @@ -32,6 +32,15 @@ export function registerVoiceIpc(): void { return kwsStart({ ...req, keywords: req.keywords.filter((k) => typeof k === 'string'), sensitivity: Number(req.sensitivity) || 3 }, event.sender); }); ipcMain.handle(VOICE_IPC.kwsStop, () => kwsStop()); + ipcMain.handle(VOICE_IPC.vadStart, (event, req: VadStartRequest) => { + if (!req || typeof req.modelId !== 'string') throw new Error('Requête invalide'); + return vadStart({ modelId: req.modelId, silenceMs: Number(req.silenceMs) || 700, threshold: Number(req.threshold) || 0.5, maxUtteranceSec: Number(req.maxUtteranceSec) || 30 }, event.sender); + }); + ipcMain.handle(VOICE_IPC.vadStop, () => vadStop()); + ipcMain.on(VOICE_IPC.vadAudio, (_e, pcm: unknown, sampleRate: unknown) => { + if (!(pcm instanceof Uint8Array) || pcm.byteLength === 0 || pcm.byteLength > 1024 * 1024) return; + vadFeed(pcm, typeof sampleRate === 'number' && sampleRate > 0 ? sampleRate : 16000); + }); ipcMain.on(VOICE_IPC.kwsAudio, (_e, pcm: unknown, sampleRate: unknown) => { if (!(pcm instanceof Uint8Array) || pcm.byteLength === 0 || pcm.byteLength > 1024 * 1024) return; kwsFeed(pcm, typeof sampleRate === 'number' && sampleRate > 0 ? sampleRate : 16000); diff --git a/electron/voice/worker.ts b/electron/voice/worker.ts index 97c7ba1..1acaaab 100644 --- a/electron/voice/worker.ts +++ b/electron/voice/worker.ts @@ -21,7 +21,10 @@ type Request = | { id: number; type: 'unload'; modelId?: string } | { id: number; type: 'kws.start'; model: ModelRef; keywordsFile: string; threshold: number; score: number } | { id: number; type: 'kws.audio'; pcm: string; sampleRate: number } - | { id: number; type: 'kws.stop' }; + | { id: number; type: 'kws.stop' } + | { id: number; type: 'vad.start'; model: ModelRef; silenceMs: number; threshold: number; maxUtteranceSec: number } + | { id: number; type: 'vad.audio'; pcm: string; sampleRate: number } + | { id: number; type: 'vad.stop' }; type Response = { id: number; ok: true; result: unknown } | { id: number; ok: false; error: string }; @@ -32,6 +35,16 @@ type Sherpa = { decode: (s: unknown) => void; getResult: (s: unknown) => { text: string; lang?: string }; }; + Vad: new (config: unknown, bufferSizeInSeconds: number) => { + acceptWaveform: (samples: Float32Array) => void; + isEmpty: () => boolean; + isDetected: () => boolean; + pop: () => void; + clear: () => void; + front: (enableExternalBuffer?: boolean) => { start: number; samples: Float32Array }; + reset: () => void; + flush: () => void; + }; KeywordSpotter: new (config: unknown) => { createStream: () => KwsStream; isReady: (s: KwsStream) => boolean; @@ -50,6 +63,7 @@ type Sherpa = { type KwsStream = { acceptWaveform: (w: { sampleRate: number; samples: Float32Array }) => void }; let sherpa: Sherpa | null = null; let kws: { spotter: InstanceType; stream: KwsStream } | null = null; +let vad: { detector: InstanceType; speaking: boolean; windowSize: number; carry: Float32Array } | null = null; let notify: ((message: unknown) => void) | null = null; let loadError: string | null = null; @@ -275,9 +289,74 @@ function feedKws(pcmBase64: string, sampleRate: number): void { } } +function startVad(model: ModelRef, silenceMs: number, threshold: number, maxUtteranceSec: number): void { + const s = loadSherpa(); + const windowSize = 512; + const detector = new s.Vad( + { + sileroVad: { + model: path.join(model.dir, 'silero_vad.onnx'), + threshold: Math.max(0.1, Math.min(0.95, threshold)), + minSilenceDuration: Math.max(0.15, silenceMs / 1000), + minSpeechDuration: 0.2, + windowSize, + maxSpeechDuration: Math.max(3, maxUtteranceSec) + }, + sampleRate: 16000, + numThreads: 1, + provider: 'cpu', + debug: 0 + }, + Math.max(10, maxUtteranceSec + 5) + ); + vad = { detector, speaking: false, windowSize, carry: new Float32Array(0) }; +} + +function feedVad(pcmBase64: string, sampleRate: number): void { + if (!vad) return; + const bytes = Buffer.from(pcmBase64, 'base64'); + const int16 = new Int16Array(bytes.buffer, bytes.byteOffset, Math.floor(bytes.byteLength / 2)); + let samples: Float32Array = new Float32Array(int16.length); + for (let i = 0; i < int16.length; i++) samples[i] = int16[i] / 32768; + if (sampleRate !== 16000) samples = resampleTo16k(samples, sampleRate); + // Silero expects fixed windows: keep the remainder for the next frame. + const merged = new Float32Array(vad.carry.length + samples.length); + merged.set(vad.carry, 0); + merged.set(samples, vad.carry.length); + const usable = merged.length - (merged.length % vad.windowSize); + for (let i = 0; i < usable; i += vad.windowSize) { + vad.detector.acceptWaveform(merged.subarray(i, i + vad.windowSize)); + const detected = vad.detector.isDetected(); + if (detected && !vad.speaking) { + vad.speaking = true; + notify?.({ type: 'vad.event', event: { type: 'speech-start' } }); + } + while (!vad.detector.isEmpty()) { + const segment = vad.detector.front(false); + vad.detector.pop(); + vad.speaking = false; + const wav = encodeWav(segment.samples, 16000); + notify?.({ + type: 'vad.event', + event: { type: 'segment', wav: Buffer.from(wav.buffer, wav.byteOffset, wav.byteLength).toString('base64'), durationSec: segment.samples.length / 16000 } + }); + } + } + vad.carry = merged.slice(usable); +} + // ── request handling ─────────────────────────────────────────────────────── function handle(req: Request): unknown { switch (req.type) { + case 'vad.start': + startVad(req.model, req.silenceMs, req.threshold, req.maxUtteranceSec); + return { ok: true }; + case 'vad.audio': + feedVad(req.pcm, req.sampleRate); + return { ok: true }; + case 'vad.stop': + vad = null; + return { ok: true }; case 'kws.start': startKws(req.model, req.keywordsFile, req.threshold, req.score); return { ok: true }; @@ -331,6 +410,7 @@ function handle(req: Request): unknown { recognizers.clear(); synthesizers.clear(); kws = null; + vad = null; } return { ok: true }; } diff --git a/package.json b/package.json index 7d48ea2..2265949 100644 --- a/package.json +++ b/package.json @@ -1,7 +1,7 @@ { "name": "eveflow", - "version": "2.2.0", - "releaseVersion": "2.2.0", + "version": "2.3.0", + "releaseVersion": "2.3.0", "description": "JARVIS-style desktop HUD for Hermes Agent: voice, streaming runs, scheduled jobs, skills and telemetry", "main": "dist-electron/main.js", "private": true, diff --git a/shared/bridge.ts b/shared/bridge.ts index b99298e..d25d7a8 100644 --- a/shared/bridge.ts +++ b/shared/bridge.ts @@ -14,6 +14,8 @@ import type { import type { KwsDetection, KwsStartRequest, + VadEvent, + VadStartRequest, SynthesizeRequest, SynthesizeResult, TranscribeRequest, @@ -46,10 +48,6 @@ export interface EveFlowBridge { streamAbort: (id: string) => void; onStreamEvent: (cb: (event: HttpStreamEvent) => void) => Unsubscribe; }; - system: { - metrics: () => Promise; - appInfo: () => Promise; - }; files: { readLocal: (filePath: string) => Promise; writeShared: (filename: string, content: string, isBase64?: boolean) => Promise<{ path: string; url: string }>; @@ -79,5 +77,32 @@ export interface EveFlowBridge { /** Fire-and-forget 16-bit PCM frames for the keyword spotter. */ kwsAudio: (pcm: Uint8Array, sampleRate: number) => void; onKwsDetected: (cb: (detection: KwsDetection) => void) => Unsubscribe; + vadStart: (req: VadStartRequest) => Promise; + vadStop: () => Promise; + vadAudio: (pcm: Uint8Array, sampleRate: number) => void; + onVadEvent: (cb: (event: VadEvent) => void) => Unsubscribe; + }; + system: { + metrics: () => Promise; + appInfo: () => Promise; + /** Screenshot of the primary display as a JPEG data URL. */ + captureScreen: (maxWidth?: number) => Promise; + /** Allow-listed local actions (lock session, open app/url, media keys, clipboard). */ + action: (action: SystemAction) => Promise; }; } + +export type SystemAction = + | { type: 'lock' } + | { type: 'open-app'; name: string } + | { type: 'open-url'; url: string } + | { type: 'media'; key: 'volume-up' | 'volume-down' | 'mute' | 'play-pause' | 'next' | 'previous' } + | { type: 'clipboard-read' } + | { type: 'clipboard-write'; text: string } + | { type: 'find-files'; query: string }; + +export interface SystemActionResult { + ok: boolean; + message?: string; + data?: unknown; +} diff --git a/shared/ipc.ts b/shared/ipc.ts index f498ee5..2ac8ba1 100644 --- a/shared/ipc.ts +++ b/shared/ipc.ts @@ -111,6 +111,8 @@ export const IPC = { httpStreamEvent: 'http:stream:event', metrics: 'system:metrics', appInfo: 'app:info', + screenCapture: 'system:screen-capture', + systemAction: 'system:action', readLocalFile: 'files:read-local', writeSharedFile: 'files:write-shared', openPath: 'files:open-path', diff --git a/shared/voice.ts b/shared/voice.ts index f95e40a..7485b0c 100644 --- a/shared/voice.ts +++ b/shared/voice.ts @@ -1,7 +1,7 @@ /** Local voice engine contract (sherpa-onnx in a utility process). Shared by main and renderer. */ -export type VoiceModelKind = 'stt' | 'tts' | 'kws'; -export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'kws-transducer'; +export type VoiceModelKind = 'stt' | 'tts' | 'kws' | 'vad'; +export type VoiceEngineKind = 'whisper' | 'sense-voice' | 'nemo-transducer' | 'kokoro' | 'piper' | 'kws-transducer' | 'silero'; export interface VoiceSpeaker { id: number; @@ -82,6 +82,20 @@ export interface KwsDetection { at: number; } +export interface VadStartRequest { + modelId: string; + /** Silence that ends an utterance, in ms. */ + silenceMs: number; + /** Detection threshold 0..1 (0.5 default). */ + threshold: number; + maxUtteranceSec: number; +} + +export type VadEvent = + | { type: 'speech-start' } + | { type: 'segment'; wav: Uint8Array; durationSec: number } + | { type: 'error'; message: string }; + export interface VoiceEngineStatus { available: boolean; error?: string; @@ -103,5 +117,9 @@ export const VOICE_IPC = { kwsStart: 'voice:kws:start', kwsStop: 'voice:kws:stop', kwsAudio: 'voice:kws:audio', - kwsDetected: 'voice:kws:detected' + kwsDetected: 'voice:kws:detected', + vadStart: 'voice:vad:start', + vadStop: 'voice:vad:stop', + vadAudio: 'voice:vad:audio', + vadEvent: 'voice:vad:event' } as const; diff --git a/src/components/chat/CommandBar.tsx b/src/components/chat/CommandBar.tsx index b8f3879..c08826c 100644 --- a/src/components/chat/CommandBar.tsx +++ b/src/components/chat/CommandBar.tsx @@ -1,6 +1,6 @@ import { useEffect, useRef, useState } from 'react'; -import { Mic, Paperclip, Send, Square, X, Radio, Loader2, Volume2, VolumeX, Navigation } from 'lucide-react'; -import { sendMessage, steer, stopGeneration } from '../../services/conversation'; +import { Mic, Paperclip, Send, Square, X, Radio, Loader2, Volume2, VolumeX, Navigation, ScanEye } from 'lucide-react'; +import { sendMessage, steer, stopGeneration , captureScreen } from '../../services/conversation'; import { speech } from '../../services/voice/speech'; import { voiceController } from '../../services/voice/voiceController'; import { useChat } from '../../state/chat'; @@ -167,9 +167,20 @@ export function CommandBar({ compact }: Props) { spellCheck={false} /> {!compact && ( - + <> + + + )} void onFiles(e.target.files)} /> {isSending && !steerMode ? ( diff --git a/src/components/settings/ModelsSection.tsx b/src/components/settings/ModelsSection.tsx index ba02ee1..9c38e31 100644 --- a/src/components/settings/ModelsSection.tsx +++ b/src/components/settings/ModelsSection.tsx @@ -1,5 +1,5 @@ import { useEffect } from 'react'; -import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, RefreshCw, Ear } from 'lucide-react'; +import { Download, Trash2, X, CheckCircle2, Cpu, Mic, Volume2, AlertTriangle, RefreshCw, Ear, Activity } from 'lucide-react'; import type { VoiceModelStatus } from '../../../shared/voice'; import { bridge } from '../../lib/bridge'; import { useShallow } from 'zustand/react/shallow'; @@ -23,9 +23,11 @@ function ModelRow({ model }: { model: VoiceModelStatus }) { const activate = () => { if (model.kind === 'stt') update({ voice: { localModel: model.id, provider: 'local' } }); else if (model.kind === 'kws') update({ voice: { wakeMode: 'kws' } }); + else if (model.kind === 'vad') update({ voice: { neuralVad: true } }); else update({ speech: { localModel: model.id, provider: 'local', localSpeaker: model.speakers?.[0]?.id ?? 0 } }); }; - const activeNow = model.kind === 'kws' ? wakeMode === 'kws' : isActive; + const neuralVad = useSettings((s) => s.settings.voice.neuralVad); + const activeNow = model.kind === 'kws' ? wakeMode === 'kws' : model.kind === 'vad' ? neuralVad : isActive; return (
@@ -76,6 +78,7 @@ export function ModelsSection() { const stt = models.filter((m) => m.kind === 'stt'); const tts = models.filter((m) => m.kind === 'tts'); const kws = models.filter((m) => m.kind === 'kws'); + const vad = models.filter((m) => m.kind === 'vad'); return ( <>
@@ -99,6 +102,11 @@ export function ModelsSection() {
Mot d’activation
{kws.map((m) => )}
+
+
Fin de phrase
+
{vad.map((m) => )}
+ Utilisé automatiquement en écoute permanente dès qu’il est installé ; sinon le VAD énergétique prend le relais. +
Reconnaissance vocale
{stt.map((m) => )}
diff --git a/src/components/settings/SettingsDrawer.tsx b/src/components/settings/SettingsDrawer.tsx index 5d2251e..c8b4d86 100644 --- a/src/components/settings/SettingsDrawer.tsx +++ b/src/components/settings/SettingsDrawer.tsx @@ -336,8 +336,12 @@ export function SettingsDrawer({ onClose }: Props) {
)} {settings.voice.wakeMode === 'kws' && ( - + <> + + update({ voice: { neuralVad: v } })} label="Fin de phrase neuronale (Silero)" hint="Coupe l’écoute au bon moment, même avec du bruit de fond. Nécessite le modèle Silero VAD (0,6 Mo) dans Modèles locaux." /> + )} + update({ voice: { localCommands: v } })} label="Commandes locales instantanées" hint="« Verrouille la session », « monte le son », « ouvre Spotify », « regarde mon écran »… exécutées sur ce PC sans passer par Hermes." />
diff --git a/src/services/conversation.ts b/src/services/conversation.ts index 439d136..a010d63 100644 --- a/src/services/conversation.ts +++ b/src/services/conversation.ts @@ -12,6 +12,8 @@ import { useSettings } from '../state/settings'; import { executeLocalTool, LOCAL_TOOL_DEFINITIONS } from './hermes/localTools'; import type { HermesStreamEvent, SendHandle } from './hermes/types'; import { speech } from './voice/speech'; +import { bridge } from '../lib/bridge'; +import { parseLocalIntent, runLocalIntent } from './localCommands'; let active: SendHandle | null = null; let activeMessageId: string | null = null; @@ -165,9 +167,49 @@ function addPending(request: Omit): void { speech.say(request.kind === 'approval' ? 'Autorisation requise.' : request.kind === 'clarify' ? request.description : 'Saisie requise.', { interrupt: false }); } +/** Screenshot of the primary display as a data URL (Electron only). */ +export async function captureScreen(): Promise { + const api = bridge(); + if (!api) return null; + try { + return await api.system.captureScreen(1600); + } catch (err) { + useChat.getState().setError(`Capture d’écran : ${(err as Error).message}`); + return null; + } +} + +/** Short system intents handled on the machine without a Hermes round trip. Returns true when consumed. */ +async function tryLocalIntent(text: string, source: string): Promise<{ handled: boolean; images?: string[]; text?: string }> { + const settings = useSettings.getState().settings; + if (!settings.voice.localCommands || !bridge()) return { handled: false }; + const intent = parseLocalIntent(text); + if (!intent) return { handled: false }; + if (intent.kind === 'screenshot') { + const shot = await captureScreen(); + if (!shot) return { handled: false }; + return { handled: false, images: [shot], text: intent.question || text }; + } + const chat = useChat.getState(); + chat.addMessage({ role: 'user', content: text, source, status: 'done' }); + const result = await runLocalIntent(intent); + chat.addMessage({ role: 'assistant', content: result.message, source: 'local', status: result.ok ? 'done' : 'error' }); + chat.setHud(result.ok ? 'idle' : 'error'); + if (settings.speech.autoSpeak) speech.say(result.message, { interrupt: true }); + return { handled: true }; +} + export async function sendMessage(text: string, images: string[] = [], source = 'eveflow'): Promise { - const trimmed = text.trim(); + let trimmed = text.trim(); if ((!trimmed && images.length === 0) || active) return; + if (trimmed && images.length === 0) { + const local = await tryLocalIntent(trimmed, source); + if (local.handled) return; + if (local.images) { + images = local.images; + trimmed = (local.text ?? trimmed).trim(); + } + } const chat = useChat.getState(); const hermes = useHermes.getState(); const settings = useSettings.getState().settings; diff --git a/src/services/hermes/client.ts b/src/services/hermes/client.ts index ccfcc4b..c79aa07 100644 --- a/src/services/hermes/client.ts +++ b/src/services/hermes/client.ts @@ -467,6 +467,8 @@ export class HermesClient { throw new HttpError(handle.start.status, detail); } ready = true; + // Chunks that raced ahead of the start event were buffered as potential error bodies: replay them. + for (const chunk of errorChunks.splice(0)) parser.feed(chunk); const rotated = handle.start.headers['x-hermes-session-id']; if (rotated && rotated !== options.sessionId) onEvent({ kind: 'session', sessionId: rotated }); diff --git a/src/services/localCommands.ts b/src/services/localCommands.ts new file mode 100644 index 0000000..b0b6f65 --- /dev/null +++ b/src/services/localCommands.ts @@ -0,0 +1,75 @@ +/** + * Instant local commands: short French/English intents handled on the machine without a + * round trip to Hermes (lock, open app/url, volume, media, screenshot to Hermes). + * Anything unmatched goes to Hermes as usual. + */ +import type { SystemAction } from '../../shared/bridge'; +import { bridge } from '../lib/bridge'; +import { Log } from '../lib/log'; + +export interface LocalIntent { + kind: 'action' | 'screenshot'; + action?: SystemAction; + /** Spoken confirmation. */ + reply: string; + /** For screenshot: the question to send to Hermes with the image. */ + question?: string; +} + +const norm = (t: string) => + t + .normalize('NFD') + .replace(/[̀-ͯ]/g, '') + .toLowerCase() + .replace(/[’']/g, ' ') + .replace(/[^a-z0-9 :/._-]+/g, ' ') + .replace(/\s+/g, ' ') + .trim(); + +const LOCK = /^(verrouille|verrouiller|bloque|lock)( (la |ma )?(session|le pc|l ordinateur|the (pc|computer|screen)))?$/; +const VOL_UP = /^(monte|augmente|hausse|plus fort|volume plus|turn up|raise)( (le|the)? ?(son|volume))?( de \d+)?$/; +const VOL_DOWN = /^(baisse|diminue|moins fort|volume moins|turn down|lower)( (le|the)? ?(son|volume))?( de \d+)?$/; +const MUTE = /^(coupe|couper|mute|silence|desactive)( (le|the)? ?(son|volume|audio))?$/; +const PLAY = /^(pause|play|lecture|reprends|reprendre|mets en pause|met en pause|stop la musique|arrete la musique)$/; +const NEXT = /^(suivant|suivante|piste suivante|musique suivante|next|skip)$/; +const PREV = /^(precedent|precedente|piste precedente|previous)$/; +const OPEN_APP = /^(ouvre|ouvrir|lance|lancer|demarre|open|launch|start) ((l application|l appli|le logiciel|le programme|the app|le|la|les|l|un|une|the|moi) )*(.+)$/; +const OPEN_URL = /^(ouvre|ouvrir|va sur|open|go to) (https?:\/\/\S+|(www\.)?[a-z0-9-]+\.[a-z]{2,}(\/\S*)?)$/; +const SCREEN = /(regarde|regardes|analyse|decris|decrit|lis|explique|qu est ce qu il y a sur|que vois tu sur|what is on|look at|read) (mon |l |the )?(ecran|screen)|capture (d )?ecran|screenshot/; + +export function parseLocalIntent(text: string): LocalIntent | null { + const t = norm(text).replace(/^(jarvis|hey jarvis|ok jarvis)[ ,]*/, ''); + if (!t) return null; + if (LOCK.test(t)) return { kind: 'action', action: { type: 'lock' }, reply: 'Session verrouillée.' }; + if (MUTE.test(t)) return { kind: 'action', action: { type: 'media', key: 'mute' }, reply: 'Son coupé.' }; + if (VOL_UP.test(t)) return { kind: 'action', action: { type: 'media', key: 'volume-up' }, reply: 'Volume augmenté.' }; + if (VOL_DOWN.test(t)) return { kind: 'action', action: { type: 'media', key: 'volume-down' }, reply: 'Volume baissé.' }; + if (PLAY.test(t)) return { kind: 'action', action: { type: 'media', key: 'play-pause' }, reply: 'Lecture.' }; + if (NEXT.test(t)) return { kind: 'action', action: { type: 'media', key: 'next' }, reply: 'Piste suivante.' }; + if (PREV.test(t)) return { kind: 'action', action: { type: 'media', key: 'previous' }, reply: 'Piste précédente.' }; + if (SCREEN.test(t)) return { kind: 'screenshot', reply: 'Je regarde votre écran.', question: text.trim() }; + const url = OPEN_URL.exec(t); + if (url) { + const target = url[2].startsWith('http') ? url[2] : `https://${url[2]}`; + return { kind: 'action', action: { type: 'open-url', url: target }, reply: `J’ouvre ${url[2]}.` }; + } + const app = OPEN_APP.exec(t); + if (app) { + const name = app[app.length - 1].trim(); + if (name.length <= 40 && !/\s(et|puis|and)\s/.test(name)) return { kind: 'action', action: { type: 'open-app', name }, reply: `J’ouvre ${name}.` }; + } + return null; +} + +/** Execute an action intent through the bridge. Returns the spoken outcome. */ +export async function runLocalIntent(intent: LocalIntent): Promise<{ ok: boolean; message: string }> { + const api = bridge(); + if (!api || !intent.action) return { ok: false, message: 'Actions locales indisponibles hors Electron.' }; + try { + const result = await api.system.action(intent.action); + Log.info('local', `${intent.action.type}: ${result.ok ? 'ok' : result.message}`); + return { ok: result.ok, message: result.ok ? intent.reply : result.message || 'Échec de la commande locale.' }; + } catch (err) { + return { ok: false, message: (err as Error).message }; + } +} diff --git a/src/services/voice/voiceController.ts b/src/services/voice/voiceController.ts index 9307061..06e0013 100644 --- a/src/services/voice/voiceController.ts +++ b/src/services/voice/voiceController.ts @@ -113,8 +113,21 @@ class VoiceController { this.unsubscribeKws?.(); this.unsubscribeKws = api.voice.onKwsDetected((d) => this.onWakeDetected(d.keyword)); const sensitivity = Math.min(5, Math.max(1, Math.round(settings.sensitivity))) as 1 | 2 | 3 | 4 | 5; + let neuralVad = false; + if (settings.neuralVad) { + try { + const installed = (await api.voice.listModels()).some((m) => m.id === 'silero-vad' && m.installed); + if (installed) { + await api.voice.vadStart({ modelId: 'silero-vad', silenceMs: Math.max(300, Math.min(1500, settings.silenceMs - 200)), threshold: 0.5, maxUtteranceSec: 25 }); + neuralVad = true; + } + } catch (err) { + Log.warn('voice', `silero unavailable, energy VAD fallback: ${(err as Error).message}`); + } + } await this.wake.start({ deviceId: settings.micDeviceId || undefined, + neuralVad, vad: { silenceMs: settings.silenceMs, speechRatio: SENSITIVITY_RATIO[sensitivity], minRms: SENSITIVITY_MIN_RMS[sensitivity] }, callbacks: { onPhase: (phase) => { @@ -141,7 +154,8 @@ class VoiceController { } }); voice.setWake('spotting', result.accepted.map((k) => k.replace(/_/g, ' '))); - Log.info('voice', `wake mode on: ${result.accepted.join(', ')}${result.rejected.length ? ` (rejetés : ${result.rejected.join(', ')})` : ''}`); + voice.setNeuralVad(neuralVad); + Log.info('voice', `wake mode on: ${result.accepted.join(', ')}${result.rejected.length ? ` (rejetés : ${result.rejected.join(', ')})` : ''}${neuralVad ? ' · Silero' : ''}`); } catch (err) { const message = (err as Error).message; voice.setWake('error'); @@ -149,6 +163,7 @@ class VoiceController { useChat.getState().setError(`Mot d’activation : ${message}`); Log.error('voice', `wake mode failed: ${message}`); await api.voice.kwsStop().catch(() => undefined); + await api.voice.vadStop().catch(() => undefined); } } @@ -157,7 +172,10 @@ class VoiceController { this.unsubscribeKws = null; this.wake.stop(); useVoice.getState().setWake('off'); - await bridge()?.voice.kwsStop().catch(() => undefined); + useVoice.getState().setNeuralVad(false); + const api = bridge(); + await api?.voice.kwsStop().catch(() => undefined); + await api?.voice.vadStop().catch(() => undefined); } /** Restart spotting with the current settings (wake word, sensitivity, microphone). */ diff --git a/src/services/voice/wakeListener.ts b/src/services/voice/wakeListener.ts index ba5a52d..a429ee2 100644 --- a/src/services/voice/wakeListener.ts +++ b/src/services/voice/wakeListener.ts @@ -9,6 +9,7 @@ import { Log } from '../../lib/log'; import { audioBus } from './audioBus'; import { DEFAULT_VAD, EnergyVad, type VadOptions } from './vad'; import { buildWav16k, rms, type WavResult } from './wav'; +import type { VadEvent } from '../../../shared/voice'; const WORKLET_SOURCE = ` class EveFlowWakeProcessor extends AudioWorkletProcessor { @@ -58,6 +59,11 @@ export class WakeListener { private callbacks: WakeCallbacks | null = null; private vadOptions: Partial = {}; private noSpeechTimer: ReturnType | null = null; + /** Neural end-of-speech (Silero in the main process) instead of the energy VAD. */ + private neural = false; + private unsubscribeVad: (() => void) | null = null; + private vadPending: Float32Array[] = []; + private vadPendingSamples = 0; get isActive(): boolean { return this.phase !== 'off'; @@ -67,10 +73,15 @@ export class WakeListener { return this.phase; } - async start(options: { deviceId?: string; vad?: Partial; callbacks: WakeCallbacks }): Promise { + async start(options: { deviceId?: string; vad?: Partial; neuralVad?: boolean; callbacks: WakeCallbacks }): Promise { if (this.phase !== 'off') return; this.callbacks = options.callbacks; this.vadOptions = options.vad ?? {}; + this.neural = !!options.neuralVad; + if (this.neural) { + this.unsubscribeVad?.(); + this.unsubscribeVad = bridge()?.voice.onVadEvent((event) => this.onVadEvent(event)) ?? null; + } this.stream = await navigator.mediaDevices.getUserMedia({ audio: { deviceId: options.deviceId ? { exact: options.deviceId } : undefined, @@ -98,13 +109,34 @@ export class WakeListener { this.node.port.onmessage = (event: MessageEvent) => this.onSamples(event.data); source.connect(this.node); this.setPhase('spotting'); - Log.info('wake', `listener started (${this.sampleRate} Hz)`); + Log.info('wake', `listener started (${this.sampleRate} Hz, fin de phrase ${this.neural ? 'Silero' : 'énergie'})`); + } + + private onVadEvent(event: VadEvent): void { + if (!this.neural || (this.phase !== 'command' && this.phase !== 'speech')) return; + if (event.type === 'speech-start') { + if (this.phase === 'command') { + this.setPhase('speech'); + if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer); + this.noSpeechTimer = null; + } + } else if (event.type === 'segment') { + const wav: WavResult = { bytes: event.wav, sampleRate: 16_000, durationSec: event.durationSec }; + this.resumeSpotting(); + if (wav.durationSec >= 0.25) this.callbacks?.onUtterance(wav); + else this.callbacks?.onNoSpeech(); + } else if (event.type === 'error') { + this.callbacks?.onError(event.message); + this.resumeSpotting(); + } } /** Called by the controller when the main process reports the wake word (or on manual trigger). */ beginCommand(): void { if (this.phase === 'off' || this.phase === 'command' || this.phase === 'speech') return; - this.vad = new EnergyVad({ ...DEFAULT_VAD, ...this.vadOptions, noSpeechTimeoutMs: Number.POSITIVE_INFINITY }); + this.vad = this.neural ? null : new EnergyVad({ ...DEFAULT_VAD, ...this.vadOptions, noSpeechTimeoutMs: Number.POSITIVE_INFINITY }); + this.vadPending = []; + this.vadPendingSamples = 0; this.chunks = []; this.totalSamples = 0; this.setPhase('command'); @@ -126,6 +158,8 @@ export class WakeListener { if (this.phase === 'off') return; if (this.noSpeechTimer) clearTimeout(this.noSpeechTimer); this.noSpeechTimer = null; + this.unsubscribeVad?.(); + this.unsubscribeVad = null; audioBus.setInputAnalyser(null); if (this.node) { this.node.port.onmessage = null; @@ -139,6 +173,8 @@ export class WakeListener { this.chunks = []; this.pending = []; this.pendingSamples = 0; + this.vadPending = []; + this.vadPendingSamples = 0; this.setPhase('off'); Log.info('wake', 'listener stopped'); } @@ -170,6 +206,14 @@ export class WakeListener { return; } + if (this.neural) { + // command / speech with Silero: stream ~128 ms frames to the main process, which returns the segment. + this.vadPending.push(samples); + this.vadPendingSamples += samples.length; + if (this.vadPendingSamples >= this.sampleRate * 0.128) this.flushToVad(); + return; + } + // command / speech: collect the utterance this.chunks.push(samples); this.totalSamples += samples.length; @@ -196,21 +240,34 @@ export class WakeListener { } } - private flushToSpotter(): void { - const api = bridge(); - if (!api) return; - const total = this.pendingSamples; + private static toInt16(chunks: Float32Array[], total: number): Uint8Array { const merged = new Int16Array(total); let offset = 0; - for (const chunk of this.pending) { + for (const chunk of chunks) { for (let i = 0; i < chunk.length; i++) { const s = Math.max(-1, Math.min(1, chunk[i])); merged[offset + i] = s < 0 ? s * 0x8000 : s * 0x7fff; } offset += chunk.length; } + return new Uint8Array(merged.buffer); + } + + private flushToSpotter(): void { + const api = bridge(); + if (!api) return; + const bytes = WakeListener.toInt16(this.pending, this.pendingSamples); this.pending = []; this.pendingSamples = 0; - api.voice.kwsAudio(new Uint8Array(merged.buffer), this.sampleRate); + api.voice.kwsAudio(bytes, this.sampleRate); + } + + private flushToVad(): void { + const api = bridge(); + if (!api) return; + const bytes = WakeListener.toInt16(this.vadPending, this.vadPendingSamples); + this.vadPending = []; + this.vadPendingSamples = 0; + api.voice.vadAudio(bytes, this.sampleRate); } } diff --git a/src/state/settings.ts b/src/state/settings.ts index ade09a2..f957a33 100644 --- a/src/state/settings.ts +++ b/src/state/settings.ts @@ -22,6 +22,10 @@ export interface VoiceSettings extends SttConfig { /** off = push-to-talk / hands-free; transcript = filter after transcription; kws = always-on keyword spotting. */ wakeMode: 'off' | 'transcript' | 'kws'; kwsSensitivity: number; // 1..5 + /** Use Silero VAD for end-of-speech in always-on mode when the model is installed. */ + neuralVad: boolean; + /** Execute short system intents locally (lock, volume, open app…) instead of asking Hermes. */ + localCommands: boolean; } export interface SpeechSettings extends TtsConfig { @@ -88,6 +92,8 @@ export const DEFAULT_SETTINGS: Settings = { wakeWord: 'jarvis', wakeMode: 'off', kwsSensitivity: 3, + neuralVad: true, + localCommands: true, localModel: 'whisper-base' }, speech: { diff --git a/src/state/voice.ts b/src/state/voice.ts index d4a6c57..f1989ad 100644 --- a/src/state/voice.ts +++ b/src/state/voice.ts @@ -15,6 +15,8 @@ interface VoiceStore { micDevices: Array<{ deviceId: string; label: string }>; wake: WakeState; wakeKeywords: string[]; + neuralVad: boolean; + setNeuralVad: (on: boolean) => void; setWake: (state: WakeState, keywords?: string[]) => void; setPhase: (phase: ListenPhase) => void; setInputLevel: (level: number) => void; @@ -37,6 +39,8 @@ export const useVoice = create((set) => ({ micDevices: [], wake: 'off', wakeKeywords: [], + neuralVad: false, + setNeuralVad: (neuralVad) => set({ neuralVad }), setWake: (wake, wakeKeywords) => set(wakeKeywords ? { wake, wakeKeywords } : { wake }), setPhase: (phase) => set({ phase }), setInputLevel: (inputLevel) => set({ inputLevel }), diff --git a/tests/localCommands.test.ts b/tests/localCommands.test.ts new file mode 100644 index 0000000..ee74b55 --- /dev/null +++ b/tests/localCommands.test.ts @@ -0,0 +1,24 @@ +import { describe, expect, it } from 'vitest'; +import { parseLocalIntent } from '../src/services/localCommands'; + +describe('parseLocalIntent', () => { + it('recognises system actions in French and English', () => { + expect(parseLocalIntent('Jarvis, verrouille la session')?.action).toEqual({ type: 'lock' }); + expect(parseLocalIntent('monte le son')?.action).toEqual({ type: 'media', key: 'volume-up' }); + expect(parseLocalIntent('Baisse le volume')?.action).toEqual({ type: 'media', key: 'volume-down' }); + expect(parseLocalIntent('coupe le son')?.action).toEqual({ type: 'media', key: 'mute' }); + expect(parseLocalIntent('piste suivante')?.action).toEqual({ type: 'media', key: 'next' }); + expect(parseLocalIntent('ouvre le bloc-notes')?.action).toEqual({ type: 'open-app', name: 'bloc-notes' }); + expect(parseLocalIntent('open spotify')?.action).toEqual({ type: 'open-app', name: 'spotify' }); + expect(parseLocalIntent('ouvre github.com')?.action).toEqual({ type: 'open-url', url: 'https://github.com' }); + }); + it('detects screen questions', () => { + expect(parseLocalIntent('Jarvis, regarde mon écran et dis-moi ce que tu vois')?.kind).toBe('screenshot'); + expect(parseLocalIntent('fais une capture d’écran')?.kind).toBe('screenshot'); + }); + it('leaves everything else to Hermes', () => { + expect(parseLocalIntent('Quelle est la météo à Paris demain ?')).toBeNull(); + expect(parseLocalIntent('ouvre le fichier puis envoie-le à Marc')).toBeNull(); + expect(parseLocalIntent('')).toBeNull(); + }); +});