Resolved conflict in browserless_service.py by merging Amazon interstitial selectors.
Combined local robust selectors with remote additions for maximum coverage.
- Added Amazon 'Continue' interstitial selectors to BrowserlessService
- Ported robust popup handling from TrackingScraperService to BrowserlessService
- Updated ItemService to fetch dynamic screenshot URLs from PriceHistory
(fixes issue where forced updates didn't show new screenshots)
- Increased default Amazon search results limit from 20 to 50 in backend and frontend
- Improved TrackingScraperService to target main product area on Amazon (avoiding reviews)
- Added specific selectors for Amazon product page (#imgTagWrapperId, #productTitle, etc.)
Amazon detects the homepage-then-search pattern as bot behavior and blocks
with 2065 byte pages. Now navigating directly to search URL to appear more
natural.
This reverts the homepage loading strategy from commit 85b19e7 as Amazon's
bot detection has evolved and now flags this pattern.
ImprovedSearchService and AmazonScraperService were both connecting to
Browserless at startup, causing connection conflicts and Amazon blocking
(2065 byte pages).
Now only AmazonScraperService initializes at startup, keeping Browserless
in continuous connection for Amazon. ImprovedSearchService will initialize
on-demand when needed.
This restores Amazon search to working state as it was at commit 85b19e7.
TrackingScraperService now initializes on-demand instead of at startup.
This prevents connection conflicts with AmazonScraperService and ImprovedSearchService
since all services connect to the same Browserless instance.
Fixes: Amazon scraper blocking issue (page too small - 2065 bytes)
Improves GenericParser to automatically try common e-commerce selectors
when the configured selector doesn't match any products.
**Problem:**
- Auchan and other sites return 0 products because configured selector
doesn't match their current HTML structure
- GenericParser was too rigid: if config selector fails, give up
**Solution:**
Automatic fallback selector cascade:
1. Try configured selector first (as before)
2. If no results, try 11 common e-commerce patterns:
- Product URL patterns: /produit, /product, /p/, /item
- Class patterns: article a, .product a, .product-card a
- Generic patterns: [class*='product'] a
3. Accept first selector that finds >= 3 links
4. Log which selector worked for debugging
**Selector Priority:**
- Most specific first (URL patterns)
- Medium specificity (semantic classes)
- Generic last resort (any product-related class)
**Benefits:**
- ✅ Works on more sites without manual config updates
- ✅ Resilient to site HTML changes
- ✅ Clear logging shows which selector worked
- ✅ Still uses configured selector when available
- ✅ Requires >= 3 matches to avoid false positives
**Expected Impact:**
Sites like Auchan should now find products even if configured
selector is outdated, without needing a specialized parser.
Critical fix for E.Leclerc and similar sites where DB domain differs
from config key by punctuation (. vs -).
**Problem:**
- DB contains: "e.leclerc" (with dot)
- Config key: "e-leclerc.com" (with dash)
- Previous matching failed: no match found
**Solution:**
Enhanced matching with 5-level strategy:
1. **Exact match** - as before
2. **Normalized punctuation match** (NEW)
- Remove all `-` and `.` from both sides
- Compare: "e.leclerc" → "eleclerc", "e-leclerc.com" → "eleclecrcom"
- Use contains logic: "eleclerc" in "eleclecrcom" → ✅ MATCH
3. **Contains match** - with original punctuation
4. **Prefix removal** - now handles both "e-" and "e." (plus "la-", "la.")
5. **Fuzzy core match** - last resort for edge cases
**Impact:**
- ✅ E.Leclerc now matches correctly
- ✅ La Foir'Fouille handles variations
- ✅ Any site with punctuation differences works
- ✅ Backward compatible with existing matches
**Debug logging:**
All strategies log which method matched, making issues easy to diagnose.
Fixes two critical issues preventing some sites from working:
**1. Site Mapping Bug (E.Leclerc, La Foir'Fouille not searching)**
- Previous mapping logic was too strict: `if key in domain or domain in key`
- E.Leclerc in DB as "leclerc.com" didn't match config key "e-leclerc.com"
- La Foir'Fouille similar issues
Solution: Multi-strategy matching
- Normalize domains (remove www., lowercase)
- Try exact match first
- Try substring match
- Try without prefixes (e-, la-)
- Log warnings for unmatched sites
**2. GenericParser Over-Filtering (Stokomani: 48 links → 0 products)**
- Previous filter required at least one query word in title
- Too strict: filtered valid products with different wording
Solution: Very relaxed filtering
- Keep products with matching query words
- ALSO keep products with substantial titles (>15 chars) even without match
- Only filter very short titles (<15 chars) without query match
- Added debug logging to track filtering decisions
**Additional Improvements:**
- Added detailed logging for title extraction
- Better visibility of why products are filtered/kept
- Helps diagnose future parser issues
These changes should significantly improve product discovery
across all sites using GenericParser.
Complete refactoring of the search module for robustness and reliability:
**New Architecture:**
- Modular parser system with specialized parsers for each major site
- Factory pattern for automatic parser selection
- Generic fallback parser for sites without specific implementation
**Specialized Parsers:**
- ✅ Amazon.fr - Robust parser with multiple fallback selectors
- Handles sponsored products, Prime eligibility, ratings, reviews
- Based on proven amazon_scraper_v2.py implementation
- ✅ Cdiscount.com - Optimized for Cdiscount's HTML structure
- ✅ Fnac.com - Handles Fnac's picture elements and article structure
- ✅ Darty.com - Specialized for electronics retailer
- ✅ Boulanger.com - Optimized for Boulanger's product cards
**Key Features:**
- Multiple fallback selectors for each element (title, price, image)
- Proper price parsing for French format (1 234,56 €)
- Rating and review count extraction
- Stock status detection
- Image URL extraction with multiple attribute fallbacks
- Query filtering to ensure relevant results
**Benefits:**
- Much more robust than previous generic parsing
- Easier to maintain (one parser per site)
- Easily extensible (add new parser for new site)
- Better error handling and logging
- Higher success rate for product extraction
**Updated Services:**
- improved_search_service.py now uses ParserFactory
- Legacy _parse_results marked as deprecated
- Conversion layer between ProductResult and SearchResult
This implements the Amazon France model across all major sites
for consistent, reliable product search results.