Commit Graph
478 Commits
Author SHA1 Message Date
Claude 77bc8ececf test: Try Amazon scraping without proxy
- Disabled proxy (use_proxy=False) for testing
- Previous attempt got 503 from Amazon with proxy
- Will test if direct connection works better
- Can re-enable proxy later if needed
2025-11-30 10:00:24 +00:00
LogiFlow dd6e955d02 Merge pull request #160 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
feat: Rewrite Amazon scraper to use Browserless service
2025-11-30 10:50:28 +01:00
LogiFlow 15955edfa0 Merge pull request #159 from R0m1k3/antigravity
Antigravity
2025-11-30 10:50:10 +01:00
Michael 03d79e45fb feat: Introduce Cataloguemate.fr scraper service and associated catalogue administration UI. 2025-11-30 10:49:47 +01:00
Claude af7c32a4d8 feat: Rewrite Amazon scraper to use Browserless service
- Created amazon_scraper_v2.py using browserless_service (Playwright)
- Removed dependency on Crawl4AI which was not loading pages correctly
- Use browserless proxy rotation (use_proxy=True)
- More robust selector fallbacks for title, price, rating
- Fixed 2128 bytes issue - now loads full Amazon pages (>100KB)
- Updated router to use new scraper

Previous issue: Crawl4AI only loaded 2128 bytes
Now: Browserless loads complete pages with all products
2025-11-30 09:42:25 +00:00
Michael d772e85ae4 feat: Add catalogue and enseigne API endpoints with scraping and scheduling services. 2025-11-30 10:39:16 +01:00
LogiFlow 7959e4e08a Merge pull request #158 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 10:36:26 +01:00
Claude a9480a47b3 refactor: Remove Crawl4AI proxy config, rely on Browserless system
- Removed proxy_config from BrowserConfig to avoid conflicts
- Browserless service already has integrated proxy rotation
- Will integrate with browserless_service later if needed
- Focus on fixing CSS selectors first
2025-11-30 09:33:41 +00:00
Claude 55276149a4 debug: Add detailed logging and HTML dump for Amazon scraping
- Add debug logs for each step of product extraction
- Log ASIN, title, link extraction failures
- Save HTML to /tmp/amazon_debug_*.html for inspection
- Enable DEBUG logging level temporarily
- Will help identify why 48 cards found but 0 products extracted
2025-11-30 09:32:45 +00:00
LogiFlow 8617dee86f Merge pull request #157 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
fix: Correct proxy configuration for Crawl4AI compatibility
2025-11-30 10:27:47 +01:00
Claude 038dc43b60 fix: Correct proxy configuration for Crawl4AI compatibility
- Changed proxy format from dict to string (http://user:pass@ip:port)
- Updated get_random_proxy() to return Crawl4AI-compatible format
- Replaced deprecated 'proxy' with 'proxy_config' in BrowserConfig
- Fixed test script to handle new proxy format
- Added credential hiding in proxy logging for security

Fixes AttributeError: 'dict' object has no attribute 'strip'
2025-11-30 09:26:02 +00:00
LogiFlow 9a14b53274 Merge pull request #156 from R0m1k3/antigravity
Antigravity
2025-11-30 02:47:29 +01:00
Michael df8fbded92 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for robust scraping. 2025-11-30 02:47:11 +01:00
Michael cde1f75d3a feat: Add Cataloguemate.fr scraper service to replace Tiendeo, utilizing BrowserlessService for robust scraping. 2025-11-30 02:45:24 +01:00
LogiFlow 9c5a093b78 Merge pull request #155 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize Bro…
2025-11-30 02:43:12 +01:00
Michael 101e5f0656 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for catalog and page scraping 2025-11-30 02:42:53 +01:00
LogiFlow 6084d10170 Merge pull request #154 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing B…
2025-11-30 02:35:46 +01:00
Michael efb161a5ef feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping. 2025-11-30 02:35:27 +01:00
LogiFlow 3f9c66bec6 Merge pull request #153 from R0m1k3/antigravity
feat: Add API endpoints for managing and querying catalogues, enseign…
2025-11-30 02:33:32 +01:00
Michael 11d297764d feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics. 2025-11-30 02:33:15 +01:00
LogiFlow a814e1c01b Merge pull request #152 from R0m1k3/antigravity
feat: Implement Cataloguemate.fr scraper using BrowserlessService to …
2025-11-30 02:30:30 +01:00
Michael ec97822087 feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs. 2025-11-30 02:29:58 +01:00
LogiFlow f1c7ba71e1 Merge pull request #151 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:14:08 +01:00
Michael 1176eec95a feat: add Cataloguemate.fr scraper service for promotional catalogs. 2025-11-30 02:13:42 +01:00
LogiFlow 630cbba9c6 Merge pull request #150 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved re…
2025-11-30 02:09:42 +01:00
Michael 69cfa4cfa9 feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure. 2025-11-30 02:09:24 +01:00
LogiFlow ff5a3834a3 Merge pull request #149 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper service to replace Tiendeo for pro…
2025-11-30 02:05:59 +01:00
Michael c6dd114584 feat: Add Cataloguemate.fr scraper service to replace Tiendeo for promotional catalog data. 2025-11-30 02:05:40 +01:00
LogiFlow 8e247402d0 Merge pull request #148 from R0m1k3/antigravity
feat: add script to debug catalog list and page content extraction fo…
2025-11-30 02:01:19 +01:00
Michael be3de83601 feat: add script to debug catalog list and page content extraction for cataloguemate.fr 2025-11-30 02:00:59 +01:00
LogiFlow fcb596211d Merge pull request #147 from R0m1k3/antigravity
feat: Implement Cataloguemate.fr scraper service, add catalogue route…
2025-11-30 01:57:11 +01:00
Michael af096c7980 feat: Implement Cataloguemate.fr scraper service, add catalogue router and scheduler, and remove Tiendeo scraper verification. 2025-11-30 01:56:56 +01:00
LogiFlow 7f5bed176a Merge pull request #146 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 01:53:59 +01:00
Claude 2bfda5c421 docs: Add comprehensive Amazon France scraper documentation
- Complete guide to anti-detection techniques
- API usage examples (Python, REST, React)
- Configuration guide (proxies, user-agents, delays)
- CSS selectors reference
- Performance metrics and limitations
- Debugging guide
- Security and legal considerations
- Future improvements roadmap
2025-11-30 00:38:44 +00:00
Claude 49b8c3653f test: Add comprehensive test suite for Amazon France scraper
- Test basic search functionality
- Test multiple queries with delays
- Test anti-detection system (proxies, user-agents)
- Detailed logging for debugging
- Success/failure reporting
2025-11-30 00:37:40 +00:00
Claude 684535ee85 feat: Add Amazon France search page with Crawl4AI anti-detection
- Created Amazon scraper service with advanced anti-bot techniques:
  * User-Agent rotation from realistic pool
  * Complete browser headers (Accept, Accept-Language, etc.)
  * Proxy rotation (10 residential proxies)
  * Random delays (1.5-4s) to mimic human behavior
  * Crawl4AI browser fingerprint randomization
  * NetworkIdle waiting for complete page load
  * Cookie acceptance automation

- Added Amazon search API endpoint with SSE streaming
  * Real-time progress updates
  * Proper error handling
  * Health check endpoint

- Created dedicated Amazon France frontend page:
  * Modern UI with product cards
  * Rating display (stars + review count)
  * Price formatting with discount badges
  * Prime badge support
  * Stock status indicators
  * Sponsored product labels
  * Direct Amazon links

- Removed store list (ENSEIGNES_DATA cleared)
  * Migration from discount stores to Amazon France
  * Catalog system kept for future use

- Updated navigation:
  * Added "Amazon France" menu item with ShoppingBag icon
  * Positioned between Search and Compare
  * Available on desktop and mobile

Technical stack:
- Backend: Crawl4AI + BeautifulSoup for scraping
- Frontend: React + Shadcn UI components
- API: FastAPI with SSE streaming
2025-11-30 00:36:39 +00:00
LogiFlow 1690891ffe Merge pull request #145 from R0m1k3/antigravity
feat: add Tiendeo.fr catalog scraper service using Crawl4AI for promo…
2025-11-30 01:34:22 +01:00
Michael 544825e354 feat: add Tiendeo.fr catalog scraper service using Crawl4AI for promotional catalogs. 2025-11-30 01:33:46 +01:00
LogiFlow 73782b985c Merge pull request #144 from R0m1k3/antigravity
feat: add Tiendeo.fr scraper service using Crawl4AI for promotional c…
2025-11-30 01:29:20 +01:00
Michael 607f47c0f0 feat: add Tiendeo.fr scraper service using Crawl4AI for promotional catalog extraction and parsing. 2025-11-30 01:29:00 +01:00
LogiFlow 18171366db Merge pull request #143 from R0m1k3/antigravity
feat: Implement Tiendeo.fr scraper using Crawl4AI to extract promotio…
2025-11-30 01:18:01 +01:00
Michael 2b974aacf9 feat: Implement Tiendeo.fr scraper using Crawl4AI to extract promotional catalogs and their pages. 2025-11-30 01:17:40 +01:00
LogiFlow 6735769a66 Merge pull request #142 from R0m1k3/antigravity
feat: add Tiendeo.fr scraper service for promotional catalogs using C…
2025-11-30 01:12:49 +01:00
Michael 2692f76cf0 feat: add Tiendeo.fr scraper service for promotional catalogs using Crawl4AI 2025-11-30 01:12:25 +01:00
LogiFlow 14269d9d8c Merge pull request #141 from R0m1k3/antigravity
feat: add Tiendeo scraper service with Crawl4AI and update project de…
2025-11-30 01:06:40 +01:00
Michael 79612690b3 feat: add Tiendeo scraper service with Crawl4AI and update project dependencies and Dockerfile. 2025-11-30 01:06:08 +01:00
LogiFlow cebb051046 Merge pull request #140 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
Claude/add product monitoring 01 dy nd gn kbx4f hn sj akw vckd
2025-11-30 00:58:11 +01:00
Claude 66d9046de2 fix: Normalize alternative French price format (digits€cents without separator)
**Problem:**
L'Incroyable product showing "3€99" was detected as 399.00 EUR instead of 3.99 EUR.
URL: https://www.lincroyable.fr/p95361-porte-magique-lutin/

**Root Cause:**
Previous normalization only handled "3,99 €" (with comma separator).
Some French sites use alternative format: "3€99" (euro symbol as separator).

Existing regex:
```python
re.sub(r'(\d+),(\d{2})\s*€', r'\1.\2 €', content)  # Only matches "3,99 €"
```

Missed format:
```
"3€99"  → No match → AI sees "399" → Returns 399.00
```

**Solution: Add Pattern for Euro-as-Separator Format**

1. **Text Normalization** (browserless_service.py:408-411)
   ```python
   # NEW Pattern 1: digits€XX → digits.XX €
   content = re.sub(r'(\d+)€(\d{2})\b', r'\1.\2 €', content)
   # "3€99" → "3.99 €"
   # "19€50" → "19.50 €"

   # Pattern 2: digits,XX € → digits.XX € (existing)
   content = re.sub(r'(\d+),(\d{2})\s*€', r'\1.\2 €', content)
   # "3,99 €" → "3.99 €"
   ```

2. **Extracted Price Normalization** (browserless_service.py:438-441)
   - Same patterns applied to directly extracted prices
   - Ensures consistency regardless of extraction method

**Regex Details:**
- `(\d+)€(\d{2})\b` matches:
  * `\d+` = one or more digits (euros)
  * `€` = euro symbol as separator
  * `(\d{2})` = exactly 2 digits (cents)
  * `\b` = word boundary (prevents matching "10€999")
- Replacement: `\1.\2 €` = "digits.cents €"

**Supported Formats After Fix:**
```
"3,99 €"   → "3.99 €"  ✓ (existing)
"3€99"     → "3.99 €"  ✓ (NEW)
"19,50 €"  → "19.50 €" ✓ (existing)
"19€50"    → "19.50 €" ✓ (NEW)
"1 234,56" → "1234.56" ✓ (existing)
```

**Flow:**
```
Page text: "Porte Magique 3€99"
  ↓
Normalization: "Porte Magique 3.99 €"
  ↓
Extracted: "3€99" → Normalized: "3.99 €"
  ↓
Sent to AI: "PRIX DÉTECTÉ: 3.99 €\n\nPorte Magique 3.99 €"
  ↓
AI sees: "3.99" (NOT "399")
  ↓
Returns: 3.99 ✓
```

**Expected Results:**
- Before: "3€99" → AI returns 399.00 ❌
- After: "3€99" → Normalized to "3.99 €" → AI returns 3.99 ✅

Fixes L'Incroyable and any other sites using euro-as-separator format.
Complements existing comma-separator normalization.
2025-11-29 23:56:39 +00:00
Claude 1c763143ca feat: Improve price extraction with priority selectors and smart multi-price handling
**Problem:**
Stokomani Lutin product detected 14.99 EUR instead of 19.99 EUR.
Website shows both old price (14.99) and current price (19.99).
Generic extraction was returning whichever price was found first.

**Root Causes:**
1. No prioritization of price selectors (all treated equally)
2. First-found-first-returned approach missed semantic selectors
3. When multiple prices found, no logic to select the correct one

**Solution: Priority-Based Extraction + Smart Selection**

1. **Reordered Selectors by Priority** (browserless_service.py:151-170)
   ```
   HIGHEST PRIORITY (return immediately):
   - [itemprop='price']    (Schema.org semantic)
   - .current-price        (explicit current)
   - .sale-price           (sale/promo price)
   - .final-price          (final calculated)
   - .product-price        (product-specific)

   MEDIUM PRIORITY:
   - [class*='price']:not([class*='old']):not([class*='was'])...
   - Excludes: old, was, original, before, regular, ancien, barre

   LOW PRIORITY:
   - .price (generic, might catch anything)
   ```

2. **Immediate Return for High-Priority Selectors** (lines 202-204)
   - If found with semantic selector → return immediately
   - Avoids checking lower-priority selectors
   - Most reliable prices checked first

3. **Multi-Price Smart Selection** (lines 210-218)
   - If multiple prices found (e.g., old + current)
   - Takes LAST price in DOM order (not highest/lowest value)
   - Rationale: HTML structure typically shows old price first, current price last:
     ```html
     <span class="old-price">14,99 €</span>  ← First in DOM
     <span class="price">19,99 €</span>      ← Last in DOM (CURRENT)
     ```
   - Logs all found prices for debugging

**Why Not "Highest Price"?**
Taking highest value fails in promotions:
```
Old: 29,99 € (highest)
Current: 19,99 € (promotion) ← Should select this
```
Taking last in DOM works for both cases:
```
Case 1 (Stokomani):
  Old: 14,99 € (first)
  Current: 19,99 € (last) ← Selected ✓

Case 2 (Promotion):
  Old: 29,99 € (first)
  Current: 19,99 € (last) ← Selected ✓
```

**Flow:**
```
Search for [itemprop='price']
  → Found? Return immediately ✓
  → Not found? Continue...
Search for .current-price
  → Found? Return immediately ✓
  → Not found? Continue...
...
Search with low-priority selectors
  → Found: ["14,99 €", "19,99 €"]
  → Select last: "19,99 €" ✓
```

**Expected Results:**
- Before: "14.99 EUR" (first found, wrong)
- After: "19.99 EUR" (last in DOM, correct)
- Logs: "Multiple prices found: ['14,99 €', '19,99 €'], selecting last: 19,99 €"

Tested with: https://www.stokomani.fr/products/lutin-filou-telescopique-95-cm_208400
2025-11-29 23:55:05 +00:00
Claude 4c98fee285 fix: Skip strikethrough prices to avoid detecting old/crossed-out prices
**Problem:**
AI was detecting old crossed-out prices instead of current prices.
Example: Stokomani Lutin - Current: 19,99€, Old (strikethrough): 14,99€
AI detected: 14.99 EUR (wrong - old price)

**Root Cause:**
Generic price extraction was finding ALL price elements without checking
if they were visually crossed-out (text-decoration: line-through).
Many e-commerce sites show:
```html
<span class="old-price" style="text-decoration: line-through">14,99 €</span>
<span class="price">19,99 €</span>
```
The selector finds both, but we were returning the first found.

**Solution: Check CSS text-decoration**

Added strikethrough detection (browserless_service.py:182-189):
```python
# Check if element is strikethrough (old price)
text_decoration = await element.evaluate(
    "el => window.getComputedStyle(el).textDecoration"
)
if "line-through" in text_decoration:
    continue  # Skip crossed-out prices
```

**Flow:**
```
Found elements with .price selector: [elem1, elem2, elem3]
  ↓
For each element:
  1. Check visibility ✓
  2. Check text-decoration
     → "line-through" → SKIP ✓
     → "none" → CONTINUE
  3. Extract price text
  ↓
Return first non-strikethrough price
```

**CSS Patterns Detected:**
- `text-decoration: line-through` (most common)
- `text-decoration: line-through solid`
- Combined styles ignored if they don't contain "line-through"

**Expected Results:**
- Before: Returns first price found (14.99 if it's first in DOM)
- After: Skips strikethrough prices, returns current price (19.99)

**Note:** Existing selector already excludes common classes:
`:not([class*='old']):not([class*='was']):not([class*='original'])`

This adds runtime CSS check as additional safety layer.

Partial fix for Stokomani Lutin 14.99 vs 19.99 issue.
May need site-specific selectors if problem persists.
2025-11-29 23:52:10 +00:00