- Removed proxy_config from BrowserConfig to avoid conflicts
- Browserless service already has integrated proxy rotation
- Will integrate with browserless_service later if needed
- Focus on fixing CSS selectors first
- Add debug logs for each step of product extraction
- Log ASIN, title, link extraction failures
- Save HTML to /tmp/amazon_debug_*.html for inspection
- Enable DEBUG logging level temporarily
- Will help identify why 48 cards found but 0 products extracted
- Changed proxy format from dict to string (http://user:pass@ip:port)
- Updated get_random_proxy() to return Crawl4AI-compatible format
- Replaced deprecated 'proxy' with 'proxy_config' in BrowserConfig
- Fixed test script to handle new proxy format
- Added credential hiding in proxy logging for security
Fixes AttributeError: 'dict' object has no attribute 'strip'
- Created Amazon scraper service with advanced anti-bot techniques:
* User-Agent rotation from realistic pool
* Complete browser headers (Accept, Accept-Language, etc.)
* Proxy rotation (10 residential proxies)
* Random delays (1.5-4s) to mimic human behavior
* Crawl4AI browser fingerprint randomization
* NetworkIdle waiting for complete page load
* Cookie acceptance automation
- Added Amazon search API endpoint with SSE streaming
* Real-time progress updates
* Proper error handling
* Health check endpoint
- Created dedicated Amazon France frontend page:
* Modern UI with product cards
* Rating display (stars + review count)
* Price formatting with discount badges
* Prime badge support
* Stock status indicators
* Sponsored product labels
* Direct Amazon links
- Removed store list (ENSEIGNES_DATA cleared)
* Migration from discount stores to Amazon France
* Catalog system kept for future use
- Updated navigation:
* Added "Amazon France" menu item with ShoppingBag icon
* Positioned between Search and Compare
* Available on desktop and mobile
Technical stack:
- Backend: Crawl4AI + BeautifulSoup for scraping
- Frontend: React + Shadcn UI components
- API: FastAPI with SSE streaming