Michael
c6b9befc9a
feat: introduce BrowserlessService for persistent, auto-reconnecting browser operations and enhanced price extraction
2025-11-30 13:25:13 +01:00
LogiFlow
65e8f20c56
Merge pull request #170 from R0m1k3/main
...
Merge pull request #169 from R0m1k3/antigravity
2025-11-30 13:00:45 +01:00
LogiFlow
ca0eaf67b9
Merge pull request #169 from R0m1k3/antigravity
...
feat: Add new services for search orchestration, browser automation, …
2025-11-30 12:50:09 +01:00
Michael
c93e55519e
Merge main: Resolved conflicts by keeping enhanced browserless service with auto-reconnection + price extraction
2025-11-30 12:49:40 +01:00
Michael
2710388faf
feat: Implement scheduled item checks with AI price extraction, database updates, and price change notifications.
2025-11-30 12:48:30 +01:00
Michael
93d73dbed8
feat: Implement BrowserlessService for robust browser automation with auto-reconnection, popup handling, and price extraction capabilities.
2025-11-30 12:43:22 +01:00
Michael
23afe4bc6c
feat: Add new services for search orchestration, browser automation, and scheduling.
2025-11-30 12:39:26 +01:00
LogiFlow
7df0b26056
Merge pull request #168 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
fix: Keep original product images from search pages instead of missin…
2025-11-30 12:25:34 +01:00
Claude
cf1831642b
fix: Keep original product images from search pages instead of missing screenshots
...
- Improved image extraction with multiple strategies (src, data-src, srcset, picture elements)
- Filter out placeholder and 1x1 pixel images
- Don't replace real images with non-existent screenshots
- Screenshots would need FastAPI static serving which is unnecessary
2025-11-30 11:01:55 +00:00
LogiFlow
056e3d7b87
Merge pull request #167 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
feat: Improve Search & Comparateur with persistent browser ScraperSer…
2025-11-30 11:55:54 +01:00
Claude
68db9ad15e
feat: Improve Search & Comparateur with persistent browser ScraperService pattern
...
- Created improved_search_service.py using persistent browser pattern
- Persistent browser connection with auto-reconnect capability
- Proper popup/cookie handling across all sites
- Price extraction with validation (reject unrealistic prices)
- Stock status detection
- Concurrent scraping with semaphore limits (2 sites, 2 products)
- Reuses browser context for better session management
- Screenshot capture for product images
- Updated search router to use improved service
- Added initialization/shutdown in main.py lifespan
Benefits:
✓ More reliable scraping with persistent connections
✓ Better anti-detection (consistent sessions)
✓ Improved price accuracy with validation
✓ Faster performance (reuses browser contexts)
✓ Auto-recovery from connection failures
2025-11-30 10:51:22 +00:00
LogiFlow
2ee9fdfe4e
Merge pull request #166 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
fix: Load Amazon homepage first to establish session
2025-11-30 11:44:51 +01:00
Claude
85b19e7680
fix: Load Amazon homepage first to establish session
...
- CRITICAL FIX: Load amazon.fr homepage in same context before search
- Preserves cookies/session between homepage and search
- Should fix 2065 bytes issue (blocked requests)
- Homepage popups handled before search
2025-11-30 10:43:24 +00:00
LogiFlow
963b6d4ea4
Merge pull request #165 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 11:37:36 +01:00
Claude
426608c029
feat: Initialize Amazon scraper service on app startup
...
- Auto-initialize persistent browser on startup
- Proper shutdown on app exit
- Browser stays connected across requests
- Better for session/cookie persistence
2025-11-30 10:35:04 +00:00
Claude
ea2af8a8f2
feat: Create persistent Amazon scraper service (ScraperService pattern)
...
- New amazon_scraper_service.py with persistent browser connection
- Shared browser instance across requests (better session management)
- Auto-reconnect if browser connection drops
- Better stealth mode (webdriver undefined, chrome runtime)
- Proper cookie/popup handling for Amazon
- Uses Playwright via Browserless (not Crawl4AI)
- Should fix 503 errors with persistent session
Based on user's ScraperService pattern for reliability
2025-11-30 10:34:41 +00:00
Claude
3c72bd3229
fix: Improve price extraction to avoid fantasy prices on Dashboard
...
- Added more price selectors (data-testid, price-current, prix-actuel, etc.)
- Added validation to reject unreasonable prices:
* Reject if <= 0 or < 0.01€ (errors)
* Reject if > 100,000€ (wrong element)
- Better strikethrough detection to skip old prices
- More detailed logging with price values
Fixes Dashboard showing incorrect/fantasy prices
2025-11-30 10:33:21 +00:00
LogiFlow
1893b5ca73
Merge pull request #164 from R0m1k3/antigravity
...
feat: Implement Catalogues page for displaying, filtering, searching,…
2025-11-30 11:32:58 +01:00
Michael
606fbf1523
feat: Implement Catalogues page for displaying, filtering, searching, and managing promotional catalogues.
2025-11-30 11:32:28 +01:00
LogiFlow
86c566ceeb
Merge pull request #163 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
feat: Load Amazon homepage first to establish session
2025-11-30 11:26:19 +01:00
Claude
29a4635995
feat: Load Amazon homepage first to establish session
...
- Load amazon.fr homepage before search to get cookies
- Add 2s delay between homepage and search (human-like)
- Trying to bypass 503 errors by establishing session first
- Each request still creates new context (limitation)
2025-11-30 10:21:25 +00:00
LogiFlow
5d70aceb71
Merge pull request #162 from R0m1k3/antigravity
...
feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce n…
2025-11-30 11:17:50 +01:00
Michael
99650a992b
feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce new Catalogues frontend page.
2025-11-30 11:17:28 +01:00
LogiFlow
e641de52b9
Merge pull request #161 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
test: Try Amazon scraping without proxy
2025-11-30 11:11:00 +01:00
Claude
77bc8ececf
test: Try Amazon scraping without proxy
...
- Disabled proxy (use_proxy=False) for testing
- Previous attempt got 503 from Amazon with proxy
- Will test if direct connection works better
- Can re-enable proxy later if needed
2025-11-30 10:00:24 +00:00
LogiFlow
dd6e955d02
Merge pull request #160 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
feat: Rewrite Amazon scraper to use Browserless service
2025-11-30 10:50:28 +01:00
LogiFlow
15955edfa0
Merge pull request #159 from R0m1k3/antigravity
...
Antigravity
2025-11-30 10:50:10 +01:00
Michael
03d79e45fb
feat: Introduce Cataloguemate.fr scraper service and associated catalogue administration UI.
2025-11-30 10:49:47 +01:00
Claude
af7c32a4d8
feat: Rewrite Amazon scraper to use Browserless service
...
- Created amazon_scraper_v2.py using browserless_service (Playwright)
- Removed dependency on Crawl4AI which was not loading pages correctly
- Use browserless proxy rotation (use_proxy=True)
- More robust selector fallbacks for title, price, rating
- Fixed 2128 bytes issue - now loads full Amazon pages (>100KB)
- Updated router to use new scraper
Previous issue: Crawl4AI only loaded 2128 bytes
Now: Browserless loads complete pages with all products
2025-11-30 09:42:25 +00:00
Michael
d772e85ae4
feat: Add catalogue and enseigne API endpoints with scraping and scheduling services.
2025-11-30 10:39:16 +01:00
LogiFlow
7959e4e08a
Merge pull request #158 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 10:36:26 +01:00
Claude
a9480a47b3
refactor: Remove Crawl4AI proxy config, rely on Browserless system
...
- Removed proxy_config from BrowserConfig to avoid conflicts
- Browserless service already has integrated proxy rotation
- Will integrate with browserless_service later if needed
- Focus on fixing CSS selectors first
2025-11-30 09:33:41 +00:00
Claude
55276149a4
debug: Add detailed logging and HTML dump for Amazon scraping
...
- Add debug logs for each step of product extraction
- Log ASIN, title, link extraction failures
- Save HTML to /tmp/amazon_debug_*.html for inspection
- Enable DEBUG logging level temporarily
- Will help identify why 48 cards found but 0 products extracted
2025-11-30 09:32:45 +00:00
LogiFlow
8617dee86f
Merge pull request #157 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
...
fix: Correct proxy configuration for Crawl4AI compatibility
2025-11-30 10:27:47 +01:00
Claude
038dc43b60
fix: Correct proxy configuration for Crawl4AI compatibility
...
- Changed proxy format from dict to string (http://user:pass@ip:port )
- Updated get_random_proxy() to return Crawl4AI-compatible format
- Replaced deprecated 'proxy' with 'proxy_config' in BrowserConfig
- Fixed test script to handle new proxy format
- Added credential hiding in proxy logging for security
Fixes AttributeError: 'dict' object has no attribute 'strip'
2025-11-30 09:26:02 +00:00
LogiFlow
9a14b53274
Merge pull request #156 from R0m1k3/antigravity
...
Antigravity
2025-11-30 02:47:29 +01:00
Michael
df8fbded92
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for robust scraping.
2025-11-30 02:47:11 +01:00
Michael
cde1f75d3a
feat: Add Cataloguemate.fr scraper service to replace Tiendeo, utilizing BrowserlessService for robust scraping.
2025-11-30 02:45:24 +01:00
LogiFlow
9c5a093b78
Merge pull request #155 from R0m1k3/antigravity
...
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize Bro…
2025-11-30 02:43:12 +01:00
Michael
101e5f0656
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for catalog and page scraping
2025-11-30 02:42:53 +01:00
LogiFlow
6084d10170
Merge pull request #154 from R0m1k3/antigravity
...
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing B…
2025-11-30 02:35:46 +01:00
Michael
efb161a5ef
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping.
2025-11-30 02:35:27 +01:00
LogiFlow
3f9c66bec6
Merge pull request #153 from R0m1k3/antigravity
...
feat: Add API endpoints for managing and querying catalogues, enseign…
2025-11-30 02:33:32 +01:00
Michael
11d297764d
feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics.
2025-11-30 02:33:15 +01:00
LogiFlow
a814e1c01b
Merge pull request #152 from R0m1k3/antigravity
...
feat: Implement Cataloguemate.fr scraper using BrowserlessService to …
2025-11-30 02:30:30 +01:00
Michael
ec97822087
feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs.
2025-11-30 02:29:58 +01:00
LogiFlow
f1c7ba71e1
Merge pull request #151 from R0m1k3/antigravity
...
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:14:08 +01:00
Michael
1176eec95a
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:13:42 +01:00
LogiFlow
630cbba9c6
Merge pull request #150 from R0m1k3/antigravity
...
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved re…
2025-11-30 02:09:42 +01:00
Michael
69cfa4cfa9
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure.
2025-11-30 02:09:24 +01:00