Commit Graph
288 Commits
Author SHA1 Message Date
Michael 8a8ef02168 feat: Introduce BrowserlessService for persistent web scraping with enhanced Amazon price extraction, alongside initial scheduler service and task tracking. 2025-12-02 13:59:39 +01:00
Michael 3bbab058c6 feat: implement persistent browserless service for robust web scraping and price extraction with auto-reconnection. 2025-12-02 11:46:32 +01:00
Michael e7875f35a7 feat: Add French localization file with comprehensive translations for the application UI. 2025-12-02 09:51:56 +01:00
Michael 86aae8f628 Merge branch 'origin/antigravity' into antigravity
Resolved conflict in browserless_service.py by merging Amazon interstitial selectors.
Combined local robust selectors with remote additions for maximum coverage.
2025-12-02 09:37:31 +01:00
Michael 2b21cd176c feat: add TrackingScraperService for Playwright-based web scraping with popup handling and smart scrolling capabilities. 2025-12-02 09:28:08 +01:00
Michael 5516562cb6 fix: Handle Amazon interstitials and dynamic screenshot updates
- Added Amazon 'Continue' interstitial selectors to BrowserlessService
- Ported robust popup handling from TrackingScraperService to BrowserlessService
- Updated ItemService to fetch dynamic screenshot URLs from PriceHistory
  (fixes issue where forced updates didn't show new screenshots)
2025-12-02 09:27:12 +01:00
Michael 9111aa02b9 feat: Add Playwright script to detect and analyze Amazon "Continuer les achats" popup elements. 2025-12-01 17:26:47 +01:00
Michael 9178512465 feat: implement BrowserlessService for unified, persistent Playwright browser management with auto-reconnection and enhanced price extraction. 2025-12-01 17:14:12 +01:00
Michael 6eb821c40d feat: Add automated item checking scheduler service and new dashboard item display components. 2025-12-01 17:07:50 +01:00
Michael f0e1c6668d feat: Implement new search service for multi-site e-commerce product search and item detail scraping. 2025-12-01 16:56:32 +01:00
Michael f223dc5a29 feat: Increase Amazon limit to 50 and fix screenshot framing
- Increased default Amazon search results limit from 20 to 50 in backend and frontend
- Improved TrackingScraperService to target main product area on Amazon (avoiding reviews)
- Added specific selectors for Amazon product page (#imgTagWrapperId, #productTitle, etc.)
2025-12-01 16:18:35 +01:00
Michael acfc5c6397 feat: Enhance Amazon stealth with comprehensive anti-detection
Added comprehensive stealth techniques to bypass Amazon bot detection:
- Complete HTTP headers (Accept, Accept-Language, Sec-Fetch-*, etc.)
- Extended chrome object with loadTimes, csi, app
- Permissions API override
- Plugins, languages, platform spoofing
- Battery API mocking

This significantly improves bot detection evasion compared to the basic
2-line stealth that was insufficient for Amazon's detection systems.
2025-12-01 13:55:57 +01:00
Michael 2429fa2018 Revert "fix: Remove homepage redirect to avoid Amazon bot detection"
This reverts commit 5d3d181f5e.
2025-12-01 13:26:05 +01:00
Michael 5d3d181f5e fix: Remove homepage redirect to avoid Amazon bot detection
Amazon detects the homepage-then-search pattern as bot behavior and blocks
with 2065 byte pages. Now navigating directly to search URL to appear more
natural.

This reverts the homepage loading strategy from commit 85b19e7 as Amazon's
bot detection has evolved and now flags this pattern.
2025-12-01 13:25:51 +01:00
Michael 95846ec606 fix: Disable ImprovedSearchService auto-init to restore Amazon functionality
ImprovedSearchService and AmazonScraperService were both connecting to
Browserless at startup, causing connection conflicts and Amazon blocking
(2065 byte pages).

Now only AmazonScraperService initializes at startup, keeping Browserless
in continuous connection for Amazon. ImprovedSearchService will initialize
on-demand when needed.

This restores Amazon search to working state as it was at commit 85b19e7.
2025-12-01 13:21:32 +01:00
Michael 53a05f45e4 fix: Remove auto-initialization of TrackingScraperService to fix Amazon scraper
TrackingScraperService now initializes on-demand instead of at startup.
This prevents connection conflicts with AmazonScraperService and ImprovedSearchService
since all services connect to the same Browserless instance.

Fixes: Amazon scraper blocking issue (page too small - 2065 bytes)
2025-12-01 13:05:49 +01:00
Michael e2a874bf6c feat: implement TrackingScraperService for robust web scraping with Playwright, including shared browser management and popup handling. 2025-12-01 12:42:14 +01:00
Michael a792060f08 feat: Add TrackingScraperService with enhanced popup handling for screenshots
- Add 47 popup selectors (vs 12 generic before)
- Support French RGPD platforms (Axeptio, Didomi, OneTrust, TarteAuCitron)
- Amazon-specific detection and handling (6 selectors)
- Retry logic with verification pass
- Double popup cleanup before screenshot
- Increased timeouts for slow animations
- New _verify_no_popups() helper function

This resolves popup visibility issues in the Suivi/Dashboard screenshots.
2025-12-01 12:27:51 +01:00
Michael ffc2f0d632 feat: introduce centralized search configuration for scraping, including site selectors, proxies, and user agents. 2025-12-01 07:40:50 +01:00
Michael 36ae466583 feat: Introduce ImprovedSearchService utilizing Playwright for persistent browser-based e-commerce scraping and price extraction, along with new search configuration. 2025-12-01 01:30:06 +01:00
Michael b35401a87b feat: Add improved search and AI price extraction services, update Docker Compose, and introduce benchmark verification script. 2025-12-01 01:20:39 +01:00
Michael 8303e767f7 feat: implement a new persistent browser-based search service for e-commerce scraping. 2025-12-01 01:11:42 +01:00
Michael 8bad965e29 feat: Add ImprovedSearchService for persistent browser-based e-commerce search scraping. 2025-12-01 01:07:22 +01:00
Michael 759a28b2d4 feat: Add ImprovedSearchService for persistent browser-based e-commerce search and scraping. 2025-12-01 01:05:40 +01:00
Michael 4dbb9f6424 feat: add improved search service with persistent browser connection 2025-12-01 01:03:51 +01:00
Michael 39893e8c1e feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-12-01 01:00:25 +01:00
Michael b5c5961618 feat: Implement AI-powered price extraction using Gemma 2 and an improved search service with persistent browser connections for web scraping. 2025-12-01 00:57:05 +01:00
Michael c159036583 feat: Implement comprehensive search configurations for discount stores and add Gifi-specific search and analysis tools. 2025-12-01 00:36:54 +01:00
Michael 63073d03b5 feat: introduce new search configuration, improved search service, and Carrefour search integration test. 2025-11-30 23:54:48 +01:00
Michael bf97cd22e7 feat: introduce improved search service with persistent browser and modular parsers 2025-11-30 23:49:34 +01:00
Michael 7cb0e2f94b feat: add improved search service with persistent browser connection and modular site-specific parsers 2025-11-30 23:45:51 +01:00
Michael 8bf9ddb2d7 feat: implement improved search service using persistent browser connection and modular parsers, alongside a BeautifulSoup check utility. 2025-11-30 23:43:11 +01:00
Michael 87568866e7 feat: introduce improved search service with persistent browser connection and modular site-specific parsers. 2025-11-30 23:40:22 +01:00
Michael 4dd77bd889 feat: add improved search service with persistent browser, modular parsers, and popup handling. 2025-11-30 23:35:40 +01:00
Michael 9855040f3e feat: Implement improved search service with supporting models, schemas, and API routers, and refactor site verification script. 2025-11-30 23:33:05 +01:00
Michael 4720ed6b93 feat: Implement improved search service with persistent browser and modular site-specific parsers, including Gifi. 2025-11-30 22:21:08 +01:00
Michael 64067347d1 feat: implement new search service with multiple site parsers, search configuration, and supporting inspection/verification scripts. 2025-11-30 22:04:14 +01:00
Michael b44082aeb4 feat: Implement new product parsers for Stokomani, Auchan, Carrefour, and Gifi, and add a centralized search configuration module. 2025-11-30 20:26:19 +01:00
Michael e0d73f4d0e feat: add LaFoirFouille parser to extract product search results from lafoirfouille.fr 2025-11-30 19:09:09 +01:00
Michael fbd06cfb17 feat: Add initial parser for Auchan.fr to extract product search results. 2025-11-30 19:05:38 +01:00
Michael 078060f155 feat: Add abstract base parser with product result dataclass and implement Stokomani parser for search results. 2025-11-30 19:03:26 +01:00
Michael ebe685fdf9 feat: add Stokomani parser to extract product search results from stokomani.fr 2025-11-30 18:17:07 +01:00
Michael b3d5c56707 feat: Add Stokomani search results parser 2025-11-30 18:13:11 +01:00
Michael c2f042beee feat: Implement improved search service with centralized configuration and site-specific parsers for various retailers. 2025-11-30 18:05:34 +01:00
Michael a8bffaec08 feat: Implement initial e-commerce site parsers with a base class and factory for extracting product search results. 2025-11-30 17:42:32 +01:00
Michael ee13a0bccd fix: resolve merge conflict in improved_search_service - unified domain matching logic 2025-11-30 17:37:16 +01:00
Michael 3db7db48fa feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-11-30 17:32:58 +01:00
Claude 43ea314e41 fix: Improve Fnac image extraction and add Shopify fallbacks
Fixes two issues:

**1. Fnac - No Images Displayed**
Problem: FnacParser only searched for images in parent container,
missing images nested directly in the link element.

Solution: Two-priority search
- PRIORITY 1: Search in link element first (picture, img)
- PRIORITY 2: Search in parent container
- Added warning log when no image found
- Ensures parent is always defined before use

**2. L'Incroyable - 0 Products (Shopify sites)**
Problem: GenericParser fallbacks didn't include Shopify patterns.
L'Incroyable uses Shopify structure with /products/ URLs.

Solution: Added Shopify-specific fallbacks
- a[href*='/products/'] (Shopify product URLs)
- a.product-card__link (Shopify class)
- a[href*='/collections/'] (Shopify collections)
- li[class*='product'] a (list-based layouts)

**Fallback Order (now 17 selectors):**
1. Product URL patterns (Shopify, generic)
2. Semantic classes (article, .product-card)
3. Shopify-specific selectors
4. Generic structure patterns
5. Last resort wildcards

**Expected Impact:**
- ✅ Fnac images should now display
- ✅ L'Incroyable should find products (Shopify)
- ✅ Better coverage of Shopify-based stores
2025-11-30 16:25:30 +00:00
Claude ff2615ff81 feat: Add fallback selectors to GenericParser for better site coverage
Improves GenericParser to automatically try common e-commerce selectors
when the configured selector doesn't match any products.

**Problem:**
- Auchan and other sites return 0 products because configured selector
  doesn't match their current HTML structure
- GenericParser was too rigid: if config selector fails, give up

**Solution:**
Automatic fallback selector cascade:

1. Try configured selector first (as before)
2. If no results, try 11 common e-commerce patterns:
   - Product URL patterns: /produit, /product, /p/, /item
   - Class patterns: article a, .product a, .product-card a
   - Generic patterns: [class*='product'] a
3. Accept first selector that finds >= 3 links
4. Log which selector worked for debugging

**Selector Priority:**
- Most specific first (URL patterns)
- Medium specificity (semantic classes)
- Generic last resort (any product-related class)

**Benefits:**
- ✅ Works on more sites without manual config updates
- ✅ Resilient to site HTML changes
- ✅ Clear logging shows which selector worked
- ✅ Still uses configured selector when available
- ✅ Requires >= 3 matches to avoid false positives

**Expected Impact:**
Sites like Auchan should now find products even if configured
selector is outdated, without needing a specialized parser.
2025-11-30 16:20:31 +00:00
Claude aa9e8dc427 fix: Handle punctuation differences in site domain matching
Critical fix for E.Leclerc and similar sites where DB domain differs
from config key by punctuation (. vs -).

**Problem:**
- DB contains: "e.leclerc" (with dot)
- Config key: "e-leclerc.com" (with dash)
- Previous matching failed: no match found

**Solution:**
Enhanced matching with 5-level strategy:

1. **Exact match** - as before
2. **Normalized punctuation match** (NEW)
   - Remove all `-` and `.` from both sides
   - Compare: "e.leclerc" → "eleclerc", "e-leclerc.com" → "eleclecrcom"
   - Use contains logic: "eleclerc" in "eleclecrcom" → ✅ MATCH
3. **Contains match** - with original punctuation
4. **Prefix removal** - now handles both "e-" and "e." (plus "la-", "la.")
5. **Fuzzy core match** - last resort for edge cases

**Impact:**
- ✅ E.Leclerc now matches correctly
- ✅ La Foir'Fouille handles variations
- ✅ Any site with punctuation differences works
- ✅ Backward compatible with existing matches

**Debug logging:**
All strategies log which method matched, making issues easy to diagnose.
2025-11-30 16:19:06 +00:00