Commit Graph
552 Commits
Author SHA1 Message Date
Michael c6b9befc9a feat: introduce BrowserlessService for persistent, auto-reconnecting browser operations and enhanced price extraction 2025-11-30 13:25:13 +01:00
LogiFlow 65e8f20c56 Merge pull request #170 from R0m1k3/main
Merge pull request #169 from R0m1k3/antigravity
2025-11-30 13:00:45 +01:00
LogiFlow ca0eaf67b9 Merge pull request #169 from R0m1k3/antigravity
feat: Add new services for search orchestration, browser automation, …
2025-11-30 12:50:09 +01:00
Michael c93e55519e Merge main: Resolved conflicts by keeping enhanced browserless service with auto-reconnection + price extraction 2025-11-30 12:49:40 +01:00
Michael 2710388faf feat: Implement scheduled item checks with AI price extraction, database updates, and price change notifications. 2025-11-30 12:48:30 +01:00
Michael 93d73dbed8 feat: Implement BrowserlessService for robust browser automation with auto-reconnection, popup handling, and price extraction capabilities. 2025-11-30 12:43:22 +01:00
Michael 23afe4bc6c feat: Add new services for search orchestration, browser automation, and scheduling. 2025-11-30 12:39:26 +01:00
LogiFlow 7df0b26056 Merge pull request #168 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
fix: Keep original product images from search pages instead of missin…
2025-11-30 12:25:34 +01:00
Claude cf1831642b fix: Keep original product images from search pages instead of missing screenshots
- Improved image extraction with multiple strategies (src, data-src, srcset, picture elements)
- Filter out placeholder and 1x1 pixel images
- Don't replace real images with non-existent screenshots
- Screenshots would need FastAPI static serving which is unnecessary
2025-11-30 11:01:55 +00:00
LogiFlow 056e3d7b87 Merge pull request #167 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
feat: Improve Search & Comparateur with persistent browser ScraperSer…
2025-11-30 11:55:54 +01:00
Claude 68db9ad15e feat: Improve Search & Comparateur with persistent browser ScraperService pattern
- Created improved_search_service.py using persistent browser pattern
- Persistent browser connection with auto-reconnect capability
- Proper popup/cookie handling across all sites
- Price extraction with validation (reject unrealistic prices)
- Stock status detection
- Concurrent scraping with semaphore limits (2 sites, 2 products)
- Reuses browser context for better session management
- Screenshot capture for product images
- Updated search router to use improved service
- Added initialization/shutdown in main.py lifespan

Benefits:
✓ More reliable scraping with persistent connections
✓ Better anti-detection (consistent sessions)
✓ Improved price accuracy with validation
✓ Faster performance (reuses browser contexts)
✓ Auto-recovery from connection failures
2025-11-30 10:51:22 +00:00
LogiFlow 2ee9fdfe4e Merge pull request #166 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
fix: Load Amazon homepage first to establish session
2025-11-30 11:44:51 +01:00
Claude 85b19e7680 fix: Load Amazon homepage first to establish session
- CRITICAL FIX: Load amazon.fr homepage in same context before search
- Preserves cookies/session between homepage and search
- Should fix 2065 bytes issue (blocked requests)
- Homepage popups handled before search
2025-11-30 10:43:24 +00:00
LogiFlow 963b6d4ea4 Merge pull request #165 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 11:37:36 +01:00
Claude 426608c029 feat: Initialize Amazon scraper service on app startup
- Auto-initialize persistent browser on startup
- Proper shutdown on app exit
- Browser stays connected across requests
- Better for session/cookie persistence
2025-11-30 10:35:04 +00:00
Claude ea2af8a8f2 feat: Create persistent Amazon scraper service (ScraperService pattern)
- New amazon_scraper_service.py with persistent browser connection
- Shared browser instance across requests (better session management)
- Auto-reconnect if browser connection drops
- Better stealth mode (webdriver undefined, chrome runtime)
- Proper cookie/popup handling for Amazon
- Uses Playwright via Browserless (not Crawl4AI)
- Should fix 503 errors with persistent session

Based on user's ScraperService pattern for reliability
2025-11-30 10:34:41 +00:00
Claude 3c72bd3229 fix: Improve price extraction to avoid fantasy prices on Dashboard
- Added more price selectors (data-testid, price-current, prix-actuel, etc.)
- Added validation to reject unreasonable prices:
  * Reject if <= 0 or < 0.01€ (errors)
  * Reject if > 100,000€ (wrong element)
- Better strikethrough detection to skip old prices
- More detailed logging with price values

Fixes Dashboard showing incorrect/fantasy prices
2025-11-30 10:33:21 +00:00
LogiFlow 1893b5ca73 Merge pull request #164 from R0m1k3/antigravity
feat: Implement Catalogues page for displaying, filtering, searching,…
2025-11-30 11:32:58 +01:00
Michael 606fbf1523 feat: Implement Catalogues page for displaying, filtering, searching, and managing promotional catalogues. 2025-11-30 11:32:28 +01:00
LogiFlow 86c566ceeb Merge pull request #163 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
feat: Load Amazon homepage first to establish session
2025-11-30 11:26:19 +01:00
Claude 29a4635995 feat: Load Amazon homepage first to establish session
- Load amazon.fr homepage before search to get cookies
- Add 2s delay between homepage and search (human-like)
- Trying to bypass 503 errors by establishing session first
- Each request still creates new context (limitation)
2025-11-30 10:21:25 +00:00
LogiFlow 5d70aceb71 Merge pull request #162 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce n…
2025-11-30 11:17:50 +01:00
Michael 99650a992b feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce new Catalogues frontend page. 2025-11-30 11:17:28 +01:00
LogiFlow e641de52b9 Merge pull request #161 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
test: Try Amazon scraping without proxy
2025-11-30 11:11:00 +01:00
Claude 77bc8ececf test: Try Amazon scraping without proxy
- Disabled proxy (use_proxy=False) for testing
- Previous attempt got 503 from Amazon with proxy
- Will test if direct connection works better
- Can re-enable proxy later if needed
2025-11-30 10:00:24 +00:00
LogiFlow dd6e955d02 Merge pull request #160 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
feat: Rewrite Amazon scraper to use Browserless service
2025-11-30 10:50:28 +01:00
LogiFlow 15955edfa0 Merge pull request #159 from R0m1k3/antigravity
Antigravity
2025-11-30 10:50:10 +01:00
Michael 03d79e45fb feat: Introduce Cataloguemate.fr scraper service and associated catalogue administration UI. 2025-11-30 10:49:47 +01:00
Claude af7c32a4d8 feat: Rewrite Amazon scraper to use Browserless service
- Created amazon_scraper_v2.py using browserless_service (Playwright)
- Removed dependency on Crawl4AI which was not loading pages correctly
- Use browserless proxy rotation (use_proxy=True)
- More robust selector fallbacks for title, price, rating
- Fixed 2128 bytes issue - now loads full Amazon pages (>100KB)
- Updated router to use new scraper

Previous issue: Crawl4AI only loaded 2128 bytes
Now: Browserless loads complete pages with all products
2025-11-30 09:42:25 +00:00
Michael d772e85ae4 feat: Add catalogue and enseigne API endpoints with scraping and scheduling services. 2025-11-30 10:39:16 +01:00
LogiFlow 7959e4e08a Merge pull request #158 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 10:36:26 +01:00
Claude a9480a47b3 refactor: Remove Crawl4AI proxy config, rely on Browserless system
- Removed proxy_config from BrowserConfig to avoid conflicts
- Browserless service already has integrated proxy rotation
- Will integrate with browserless_service later if needed
- Focus on fixing CSS selectors first
2025-11-30 09:33:41 +00:00
Claude 55276149a4 debug: Add detailed logging and HTML dump for Amazon scraping
- Add debug logs for each step of product extraction
- Log ASIN, title, link extraction failures
- Save HTML to /tmp/amazon_debug_*.html for inspection
- Enable DEBUG logging level temporarily
- Will help identify why 48 cards found but 0 products extracted
2025-11-30 09:32:45 +00:00
LogiFlow 8617dee86f Merge pull request #157 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
fix: Correct proxy configuration for Crawl4AI compatibility
2025-11-30 10:27:47 +01:00
Claude 038dc43b60 fix: Correct proxy configuration for Crawl4AI compatibility
- Changed proxy format from dict to string (http://user:pass@ip:port)
- Updated get_random_proxy() to return Crawl4AI-compatible format
- Replaced deprecated 'proxy' with 'proxy_config' in BrowserConfig
- Fixed test script to handle new proxy format
- Added credential hiding in proxy logging for security

Fixes AttributeError: 'dict' object has no attribute 'strip'
2025-11-30 09:26:02 +00:00
LogiFlow 9a14b53274 Merge pull request #156 from R0m1k3/antigravity
Antigravity
2025-11-30 02:47:29 +01:00
Michael df8fbded92 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for robust scraping. 2025-11-30 02:47:11 +01:00
Michael cde1f75d3a feat: Add Cataloguemate.fr scraper service to replace Tiendeo, utilizing BrowserlessService for robust scraping. 2025-11-30 02:45:24 +01:00
LogiFlow 9c5a093b78 Merge pull request #155 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize Bro…
2025-11-30 02:43:12 +01:00
Michael 101e5f0656 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for catalog and page scraping 2025-11-30 02:42:53 +01:00
LogiFlow 6084d10170 Merge pull request #154 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing B…
2025-11-30 02:35:46 +01:00
Michael efb161a5ef feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping. 2025-11-30 02:35:27 +01:00
LogiFlow 3f9c66bec6 Merge pull request #153 from R0m1k3/antigravity
feat: Add API endpoints for managing and querying catalogues, enseign…
2025-11-30 02:33:32 +01:00
Michael 11d297764d feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics. 2025-11-30 02:33:15 +01:00
LogiFlow a814e1c01b Merge pull request #152 from R0m1k3/antigravity
feat: Implement Cataloguemate.fr scraper using BrowserlessService to …
2025-11-30 02:30:30 +01:00
Michael ec97822087 feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs. 2025-11-30 02:29:58 +01:00
LogiFlow f1c7ba71e1 Merge pull request #151 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:14:08 +01:00
Michael 1176eec95a feat: add Cataloguemate.fr scraper service for promotional catalogs. 2025-11-30 02:13:42 +01:00
LogiFlow 630cbba9c6 Merge pull request #150 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved re…
2025-11-30 02:09:42 +01:00
Michael 69cfa4cfa9 feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure. 2025-11-30 02:09:24 +01:00