- Disabled proxy (use_proxy=False) for testing
- Previous attempt got 503 from Amazon with proxy
- Will test if direct connection works better
- Can re-enable proxy later if needed
- Created amazon_scraper_v2.py using browserless_service (Playwright)
- Removed dependency on Crawl4AI which was not loading pages correctly
- Use browserless proxy rotation (use_proxy=True)
- More robust selector fallbacks for title, price, rating
- Fixed 2128 bytes issue - now loads full Amazon pages (>100KB)
- Updated router to use new scraper
Previous issue: Crawl4AI only loaded 2128 bytes
Now: Browserless loads complete pages with all products
- Removed proxy_config from BrowserConfig to avoid conflicts
- Browserless service already has integrated proxy rotation
- Will integrate with browserless_service later if needed
- Focus on fixing CSS selectors first
- Add debug logs for each step of product extraction
- Log ASIN, title, link extraction failures
- Save HTML to /tmp/amazon_debug_*.html for inspection
- Enable DEBUG logging level temporarily
- Will help identify why 48 cards found but 0 products extracted
- Changed proxy format from dict to string (http://user:pass@ip:port)
- Updated get_random_proxy() to return Crawl4AI-compatible format
- Replaced deprecated 'proxy' with 'proxy_config' in BrowserConfig
- Fixed test script to handle new proxy format
- Added credential hiding in proxy logging for security
Fixes AttributeError: 'dict' object has no attribute 'strip'
- Test basic search functionality
- Test multiple queries with delays
- Test anti-detection system (proxies, user-agents)
- Detailed logging for debugging
- Success/failure reporting
- Created Amazon scraper service with advanced anti-bot techniques:
* User-Agent rotation from realistic pool
* Complete browser headers (Accept, Accept-Language, etc.)
* Proxy rotation (10 residential proxies)
* Random delays (1.5-4s) to mimic human behavior
* Crawl4AI browser fingerprint randomization
* NetworkIdle waiting for complete page load
* Cookie acceptance automation
- Added Amazon search API endpoint with SSE streaming
* Real-time progress updates
* Proper error handling
* Health check endpoint
- Created dedicated Amazon France frontend page:
* Modern UI with product cards
* Rating display (stars + review count)
* Price formatting with discount badges
* Prime badge support
* Stock status indicators
* Sponsored product labels
* Direct Amazon links
- Removed store list (ENSEIGNES_DATA cleared)
* Migration from discount stores to Amazon France
* Catalog system kept for future use
- Updated navigation:
* Added "Amazon France" menu item with ShoppingBag icon
* Positioned between Search and Compare
* Available on desktop and mobile
Technical stack:
- Backend: Crawl4AI + BeautifulSoup for scraping
- Frontend: React + Shadcn UI components
- API: FastAPI with SSE streaming
**Problem:**
AI was detecting old crossed-out prices instead of current prices.
Example: Stokomani Lutin - Current: 19,99€, Old (strikethrough): 14,99€
AI detected: 14.99 EUR (wrong - old price)
**Root Cause:**
Generic price extraction was finding ALL price elements without checking
if they were visually crossed-out (text-decoration: line-through).
Many e-commerce sites show:
```html
<span class="old-price" style="text-decoration: line-through">14,99 €</span>
<span class="price">19,99 €</span>
```
The selector finds both, but we were returning the first found.
**Solution: Check CSS text-decoration**
Added strikethrough detection (browserless_service.py:182-189):
```python
# Check if element is strikethrough (old price)
text_decoration = await element.evaluate(
"el => window.getComputedStyle(el).textDecoration"
)
if "line-through" in text_decoration:
continue # Skip crossed-out prices
```
**Flow:**
```
Found elements with .price selector: [elem1, elem2, elem3]
↓
For each element:
1. Check visibility ✓
2. Check text-decoration
→ "line-through" → SKIP ✓
→ "none" → CONTINUE
3. Extract price text
↓
Return first non-strikethrough price
```
**CSS Patterns Detected:**
- `text-decoration: line-through` (most common)
- `text-decoration: line-through solid`
- Combined styles ignored if they don't contain "line-through"
**Expected Results:**
- Before: Returns first price found (14.99 if it's first in DOM)
- After: Skips strikethrough prices, returns current price (19.99)
**Note:** Existing selector already excludes common classes:
`:not([class*='old']):not([class*='was']):not([class*='original'])`
This adds runtime CSS check as additional safety layer.
Partial fix for Stokomani Lutin 14.99 vs 19.99 issue.
May need site-specific selectors if problem persists.