Commit Graph
86 Commits
Author SHA1 Message Date
MichaelandClaude Opus 4.7 c008b41eb7 feat(ffmpeg): Gemini TTS as primary engine, Edge TTS as fallback
Default tts_engine changed from "edge" to "gemini".
Both /process-reel and /preview-tts now try Gemini first,
fall back to Edge TTS on failure.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 16:42:45 +02:00
Claude 89bc1bf88b Decode Gemini TTS raw PCM with correct format flags
Gemini returns raw PCM (audio/L16;codec=pcm;rate=24000), not a real WAV
file with a header, so FFmpeg fails to auto-detect the format. Pass
-f s16le, -ar (from mime type), and -ac 1 explicitly so FFmpeg can
decode the byte stream. Also surface FFmpeg stderr on failure instead
of swallowing it.

https://claude.ai/code/session_01QBvwHAMZzVYfWvy1U5paat
2026-05-19 13:31:15 +00:00
Claude 46c31a1447 Switch Gemini TTS from Google Cloud TTS to Gemini native API
Replace the texttospeech.googleapis.com call (which requires a separate
GCP project with Cloud TTS enabled and billing) with the Gemini native
TTS endpoint (gemini-2.5-flash-preview-tts). This uses the same Gemini
API key already configured in the app, with no extra GCP setup needed.

Voice mapping: fr-FR-Standard-A/C -> Kore (female), B/D -> Charon (male).
Audio is returned as WAV and converted to MP3 via FFmpeg. Subtitle sync
uses ffsubsync as Gemini TTS does not return word boundaries.

https://claude.ai/code/session_01QBvwHAMZzVYfWvy1U5paat
2026-05-19 13:15:30 +00:00
Claude a146705d15 Fix Gemini TTS crash from surrogate escape sequences in log strings
Replace \uXXXX surrogate-pair escapes with the actual emoji characters
in print() calls. Lone surrogates are invalid in UTF-8, causing
UnicodeEncodeError on stdout encode and aborting TTS generation before
the Gemini API call.

https://claude.ai/code/session_01QBvwHAMZzVYfWvy1U5paat
2026-05-19 12:39:06 +00:00
Michael 6ff8f6b283 Debug Gemini TTS timepoints and add verbose logging
- Add response structure logging to see actual keys returned by Gemini API
- Log timepoints count and first item to diagnose word boundary parsing
- Add 'unexpected timepoint format' warning for easier debugging
- Convert emoji to ASCII-safe versions for container logs
2026-05-19 13:55:47 +02:00
Michael 64fe3810b5 Fix Edge TTS voice fallback and Gemini voice validation
- Edge TTS fallback list now includes fr-FR-RemyMultilingualNeural
- Removed 'Neural not in voice' check that rejected Gemini voices (fr-FR-Standard-A etc) before falling back - these are valid voices and Edge-tts handles them properly
- Improved is_male check to match 'Remy' not 'Remi'
2026-05-19 12:56:21 +02:00
Michael 72b13b3e42 Add Google Gemini TTS as alternative to Edge TTS
- Store Gemini API key globally in appConfig (single key for all users)
- Add /api/settings/gemini GET/POST/DELETE routes for API key management
- Backend: add tts_engine parameter ("edge" or "gemini") to FFmpeg service
- Backend: add generate_tts_gemini() using Google Cloud TTS REST API
- Frontend: Settings page shows Google Gemini API key input card
- Frontend: new-reel, mobile/new-reel, remotion-video, mobile/remotion-video
  pages now have Edge/Gemini engine toggle and French voice selector
- Fix tts-preview route to extract ttsVoice from req.body instead of
  undefined voice variable
2026-05-19 12:17:42 +02:00
Michael 59bc56aedc refactor: remove Piper TTS, keep edge_tts only
- Remove piper_url from ReelRequest model and all function signatures
- Remove generate_tts_piper() function and Piper branch in generate_tts_with_subs()
- Remove /api/piper/config routes from server/routes.ts
- Remove piperConfig table, schemas and storage methods
- Remove piper_config migration
- Remove Piper TTS settings UI card
- Update "Piper TTS" labels to "TTS — voix activée"

edge_tts is now the only TTS engine, using precise word-boundary timing
2026-05-19 10:55:01 +02:00
Michael 8736e4f6a0 fix: pin edge-tts to 6.1.12 for reproducible builds 2026-05-19 10:30:02 +02:00
MichaelandClaude Sonnet 4.6 797e12dde5 refactor(tts): replace Minimax+Freesound with Piper TTS
Remove Minimax Speech API and Freesound integrations entirely.
Add Piper TTS as the sole TTS provider via configurable HTTP URL.

- Add piper_config DB table (url field, per-user)
- Add GET/POST /api/piper/config routes
- Python: replace generate_tts_minimax with generate_tts_piper (GET ?text=, WAV→MP3)
- Simplify generate_tts_with_subs: piper_url param replaces tts_provider+minimax fields
- Drop ttsVoice/ttsProvider from all UI, API, and background job params
- Settings page: replace Minimax+Freesound cards with single Piper URL input
- Remove freeSoundService init from server startup and CSP headers

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-19 09:15:26 +02:00
Michael 0209c3eacf feat(minimax): add Group ID support to fix rate limit quota
Pass GroupId query param to Minimax T2A v2 API — required for paid
plan quota allocation. Without it, Minimax defaults to (0/0 used).

- shared/schema.ts: groupId field on minimaxConfig table
- server/migrate.ts: ADD COLUMN IF NOT EXISTS group_id
- server/routes/reels.ts: fetch and pass groupId alongside apiKey
- server/services/ffmpeg.ts: minimax_group_id in request interface/body
- ffmpeg-service/main.py: GroupId in URL, threaded through all call sites
- client/src/pages/settings.tsx: Group ID input in Minimax settings card
2026-05-18 13:50:41 +02:00
Michael cf20b0a264 fix(tts): surface TTS errors from Python to Node.js logs and DB
- Python: capture tts_error_msg on exception, return in response
- Python: log tts_provider, minimax_api_key presence before call
- ffmpeg.ts: read tts_error from response, log and return it
- reels.ts: store TTS error in post.generationError for visibility
2026-05-18 13:26:13 +02:00
Michael 211b7dc272 fix(minimax): correct API payload and add TTS debug logging
- Fix model name: speech-2.8-hd -> speech-02-hd
- Fix bitrate: 128 -> 128000 bps
- Add channel, speed, pitch, vol to voice/audio settings
- Handle binary audio response (Content-Type: audio/*)
- Fallback audio field lookup: data.data.audio || data.audio
- Add background job log: ttsProvider + hasMinimaxKey
- Log full FFmpeg request body (text truncated, key masked)
2026-05-18 12:50:59 +02:00
Michael b7ef4940f7 feat: add Minimax Speech as alternative TTS provider for Reels
- Add minimax_config table (migration + schema + storage CRUD)
- Add GET/POST /api/minimax/config routes
- Pass tts_provider + minimax_api_key through ffmpeg service
- Add generate_tts_minimax() in Python using Minimax T2A v2 API
- Fall back to ffsubsync for subtitle sync (no WordBoundary events)
- Add Minimax config card in Settings page
- Add provider toggle (Edge TTS / Minimax) + French voices in new-reel
2026-05-18 11:55:07 +02:00
MichaelandClaude Opus 4.7 9e6fe9fe31 fix(ffmpeg): fallback to ffsubsync when edge_tts emits no word boundaries
Some voices/texts do not produce WordBoundary events from edge_tts.
When word_boundaries is empty, generate_ass_from_word_boundaries returns
without writing the ASS file, causing FFmpeg to fail with ENOENT.

Fallback to the old ffsubsync path when no boundaries are captured.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-04 14:34:20 +02:00
MichaelandClaude Opus 4.7 37aba33425 fix: precise TTS word-boundary timing and remove wasted sync calculations
- ffmpeg-service/main.py: replace ffsubsync path with exact word-boundary timing from edge_tts for TTS subtitles. Karaoke styling preserved.
- server/services/ttsSync.ts: fix word count to match TTS-cleaned text, remove artificial punctuationPause subtraction, strip punctuation tokens from count.
- server/routes/reels.ts: remove redundant ttsSyncService calls in preview/background; word_duration is ignored by Python, these only wasted TTS generations.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-04 13:13:55 +02:00
Michael 177f4db455 feat: introduce FastAPI FFmpeg service for video reel creation with TTS and synchronized subtitles. 2026-03-24 08:59:28 +01:00
Michael af082eaf2f feat: Add Dockerfile for the ffmpeg-service, including FFmpeg with vidstab support and Python dependencies. 2026-03-24 08:45:35 +01:00
Michael eff432b585 feat: implement initial FFmpeg processing service with FastAPI, TTS generation, and word-level subtitles. 2026-03-24 08:39:45 +01:00
Michael e6a1fdc9da feat: Implement new FFmpeg service with FastAPI for video processing, TTS, and subtitle generation, including font management and debugging endpoints. 2026-03-19 15:00:59 +01:00
Michael 3a5acd425c feat: Implement new Reel creation functionality with dedicated mobile and desktop UIs, server-side API routes, and a new FFmpeg processing service. 2026-03-19 14:51:16 +01:00
Michael 5a67ed2e83 feat: add initial API routes for Reels, including music search, popular tracks, and favorite management. 2026-03-12 16:18:09 +01:00
Michael 7c88527501 feat: Introduce FFmpeg service with TTS, synchronized subtitles, and video stabilization features. 2026-03-03 19:26:31 +01:00
Michael 5a8c548a62 feat: Implement FastAPI service for video processing with advanced TTS generation and word-level synchronized subtitles. 2026-03-03 19:13:03 +01:00
Michael 64969b17b5 feat: Create FastAPI service for generating TTS with word-level synchronized subtitles, robust font handling, and FFmpeg diagnostics. 2026-03-03 18:19:52 +01:00
Michael 0fd7812b05 feat: Implement mobile new reel creation page with multi-step workflow for video upload, music selection, text overlay, and scheduling. 2026-03-03 16:27:48 +01:00
Michael 108f020dc5 test: add sample ASS subtitle files for subtitle generation and styling tests. 2026-03-03 14:41:09 +01:00
Michael c3356c2140 feat: Introduce ffmpeg-service using FastAPI to process video, generate TTS with synchronized subtitles, and manage fonts. 2026-03-03 14:28:01 +01:00
Michael 101e27a2fd feat: Implement new FFmpeg service for video processing, including TTS with synchronized subtitles and font management. 2026-03-03 13:59:17 +01:00
Michael b413c1ca41 feat: implement an FFmpeg processing service with FastAPI, supporting TTS, word-level subtitles, and font management. 2026-03-03 13:11:43 +01:00
Michael 9d29e7784a feat: Implement core FastAPI service for video processing with TTS generation, dynamic subtitles, and FFmpeg enhancements. 2026-03-03 12:44:32 +01:00
Michael 4fcc8cf53c feat: Add music search and favorite API routes, and introduce FFmpeg service for Reels. 2026-03-03 12:34:09 +01:00
Michael b4e48a1f0e feat: implement FastAPI service for video processing with TTS, synchronized subtitles, and stabilization. 2026-03-03 12:08:28 +01:00
Michael 40a057e924 feat: Introduce FastAPI service for video processing, supporting TTS, dynamic subtitles, and video stabilization. 2026-03-02 13:39:51 +01:00
Michael 0e9622b5eb feat: Implement internal MP3 management and logo overlay functionality for Reels. 2026-03-02 11:33:22 +01:00
Michael dbddf3a74e feat: implement reel improvements (clean text, deletion, mobile TTS) 2026-02-13 10:11:04 +01:00
Michael 11b5ef0ad2 feat(mobile-reel): suppression bouton caméra, amélioration sync TTS (chunks 3 mots + overlap) et style sous-titres (outline/shadow) 2026-02-12 10:08:46 +01:00
Michael cdabf13e94 feat(ffmpeg): synchronisation texte/voix precise via WordBoundary edge-tts 2026-02-11 15:29:12 +01:00
Michael 58f513d476 feat: Implement precise TTS sync using edge-tts WordBoundary events 2026-02-03 16:24:42 +01:00
Michael 1f7d0d1675 feat: Implement a new FastAPI service for video processing with FFmpeg, TTS, and subtitle generation, including font management and diagnostic endpoints. 2026-02-03 15:24:10 +01:00
Michael 838d168239 fix: remove emojis from display overlay and increase font size to 65 2026-01-29 13:32:11 +01:00
Michael 5c33d7fac0 fix: auto-download Noto Color Emoji and use generic Sans font family for robust emoji support 2026-01-29 13:25:16 +01:00
Michael 70f348553f fix: revert to DejaVu Sans to fix emoji visibility regression 2026-01-29 13:19:17 +01:00
Michael 214ff257a2 fix: emojis color, remove hashtags from overlay, verify stabilization logic 2026-01-29 13:12:12 +01:00
Michael 817bd2a980 feat: emoji support in subtitles and clear stabilization UI 2026-01-29 12:58:53 +01:00
Michael eee6e65c94 fix: tune encoded stabilization for aggressive handheld shake 2026-01-29 12:20:12 +01:00
Michael 15397974e0 fix: optimize video stabilization with adaptive zoom and single-pass encoding 2026-01-29 12:15:09 +01:00
Michael d7377914e7 fix: synchronize text overlay speed with dynamic TTS duration 2026-01-29 12:08:52 +01:00
Michael 456139cf2d feat: optimize facebook reels video quality (1080p, 30fps, 10Mbps) and fix camera capture settings 2026-01-29 11:19:40 +01:00
Jacques c8943accb8 fix(ffmpeg): use ASS subtitles for precise centering and crop-to-fill for full height 2026-01-23 20:44:56 +01:00