Commit Graph
562 Commits
Author SHA1 Message Date
LogiFlow 6084d10170 Merge pull request #154 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing B…
2025-11-30 02:35:46 +01:00
Michael efb161a5ef feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping. 2025-11-30 02:35:27 +01:00
LogiFlow 3f9c66bec6 Merge pull request #153 from R0m1k3/antigravity
feat: Add API endpoints for managing and querying catalogues, enseign…
2025-11-30 02:33:32 +01:00
Michael 11d297764d feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics. 2025-11-30 02:33:15 +01:00
LogiFlow a814e1c01b Merge pull request #152 from R0m1k3/antigravity
feat: Implement Cataloguemate.fr scraper using BrowserlessService to …
2025-11-30 02:30:30 +01:00
Michael ec97822087 feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs. 2025-11-30 02:29:58 +01:00
LogiFlow f1c7ba71e1 Merge pull request #151 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:14:08 +01:00
Michael 1176eec95a feat: add Cataloguemate.fr scraper service for promotional catalogs. 2025-11-30 02:13:42 +01:00
LogiFlow 630cbba9c6 Merge pull request #150 from R0m1k3/antigravity
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved re…
2025-11-30 02:09:42 +01:00
Michael 69cfa4cfa9 feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure. 2025-11-30 02:09:24 +01:00
LogiFlow ff5a3834a3 Merge pull request #149 from R0m1k3/antigravity
feat: Add Cataloguemate.fr scraper service to replace Tiendeo for pro…
2025-11-30 02:05:59 +01:00
Michael c6dd114584 feat: Add Cataloguemate.fr scraper service to replace Tiendeo for promotional catalog data. 2025-11-30 02:05:40 +01:00
LogiFlow 8e247402d0 Merge pull request #148 from R0m1k3/antigravity
feat: add script to debug catalog list and page content extraction fo…
2025-11-30 02:01:19 +01:00
Michael be3de83601 feat: add script to debug catalog list and page content extraction for cataloguemate.fr 2025-11-30 02:00:59 +01:00
LogiFlow fcb596211d Merge pull request #147 from R0m1k3/antigravity
feat: Implement Cataloguemate.fr scraper service, add catalogue route…
2025-11-30 01:57:11 +01:00
Michael af096c7980 feat: Implement Cataloguemate.fr scraper service, add catalogue router and scheduler, and remove Tiendeo scraper verification. 2025-11-30 01:56:56 +01:00
LogiFlow 7f5bed176a Merge pull request #146 from R0m1k3/claude/amazon-france-search-page-01KmPqbdPCqxWXxEoFw9jgFo
Claude/amazon france search page 01 km pqbd p cqx w xx eo fw9jg fo
2025-11-30 01:53:59 +01:00
Claude 2bfda5c421 docs: Add comprehensive Amazon France scraper documentation
- Complete guide to anti-detection techniques
- API usage examples (Python, REST, React)
- Configuration guide (proxies, user-agents, delays)
- CSS selectors reference
- Performance metrics and limitations
- Debugging guide
- Security and legal considerations
- Future improvements roadmap
2025-11-30 00:38:44 +00:00
Claude 49b8c3653f test: Add comprehensive test suite for Amazon France scraper
- Test basic search functionality
- Test multiple queries with delays
- Test anti-detection system (proxies, user-agents)
- Detailed logging for debugging
- Success/failure reporting
2025-11-30 00:37:40 +00:00
Claude 684535ee85 feat: Add Amazon France search page with Crawl4AI anti-detection
- Created Amazon scraper service with advanced anti-bot techniques:
  * User-Agent rotation from realistic pool
  * Complete browser headers (Accept, Accept-Language, etc.)
  * Proxy rotation (10 residential proxies)
  * Random delays (1.5-4s) to mimic human behavior
  * Crawl4AI browser fingerprint randomization
  * NetworkIdle waiting for complete page load
  * Cookie acceptance automation

- Added Amazon search API endpoint with SSE streaming
  * Real-time progress updates
  * Proper error handling
  * Health check endpoint

- Created dedicated Amazon France frontend page:
  * Modern UI with product cards
  * Rating display (stars + review count)
  * Price formatting with discount badges
  * Prime badge support
  * Stock status indicators
  * Sponsored product labels
  * Direct Amazon links

- Removed store list (ENSEIGNES_DATA cleared)
  * Migration from discount stores to Amazon France
  * Catalog system kept for future use

- Updated navigation:
  * Added "Amazon France" menu item with ShoppingBag icon
  * Positioned between Search and Compare
  * Available on desktop and mobile

Technical stack:
- Backend: Crawl4AI + BeautifulSoup for scraping
- Frontend: React + Shadcn UI components
- API: FastAPI with SSE streaming
2025-11-30 00:36:39 +00:00
LogiFlow 1690891ffe Merge pull request #145 from R0m1k3/antigravity
feat: add Tiendeo.fr catalog scraper service using Crawl4AI for promo…
2025-11-30 01:34:22 +01:00
Michael 544825e354 feat: add Tiendeo.fr catalog scraper service using Crawl4AI for promotional catalogs. 2025-11-30 01:33:46 +01:00
LogiFlow 73782b985c Merge pull request #144 from R0m1k3/antigravity
feat: add Tiendeo.fr scraper service using Crawl4AI for promotional c…
2025-11-30 01:29:20 +01:00
Michael 607f47c0f0 feat: add Tiendeo.fr scraper service using Crawl4AI for promotional catalog extraction and parsing. 2025-11-30 01:29:00 +01:00
LogiFlow 18171366db Merge pull request #143 from R0m1k3/antigravity
feat: Implement Tiendeo.fr scraper using Crawl4AI to extract promotio…
2025-11-30 01:18:01 +01:00
Michael 2b974aacf9 feat: Implement Tiendeo.fr scraper using Crawl4AI to extract promotional catalogs and their pages. 2025-11-30 01:17:40 +01:00
LogiFlow 6735769a66 Merge pull request #142 from R0m1k3/antigravity
feat: add Tiendeo.fr scraper service for promotional catalogs using C…
2025-11-30 01:12:49 +01:00
Michael 2692f76cf0 feat: add Tiendeo.fr scraper service for promotional catalogs using Crawl4AI 2025-11-30 01:12:25 +01:00
LogiFlow 14269d9d8c Merge pull request #141 from R0m1k3/antigravity
feat: add Tiendeo scraper service with Crawl4AI and update project de…
2025-11-30 01:06:40 +01:00
Michael 79612690b3 feat: add Tiendeo scraper service with Crawl4AI and update project dependencies and Dockerfile. 2025-11-30 01:06:08 +01:00
LogiFlow cebb051046 Merge pull request #140 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
Claude/add product monitoring 01 dy nd gn kbx4f hn sj akw vckd
2025-11-30 00:58:11 +01:00
Claude 66d9046de2 fix: Normalize alternative French price format (digits€cents without separator)
**Problem:**
L'Incroyable product showing "3€99" was detected as 399.00 EUR instead of 3.99 EUR.
URL: https://www.lincroyable.fr/p95361-porte-magique-lutin/

**Root Cause:**
Previous normalization only handled "3,99 €" (with comma separator).
Some French sites use alternative format: "3€99" (euro symbol as separator).

Existing regex:
```python
re.sub(r'(\d+),(\d{2})\s*€', r'\1.\2 €', content)  # Only matches "3,99 €"
```

Missed format:
```
"3€99"  → No match → AI sees "399" → Returns 399.00
```

**Solution: Add Pattern for Euro-as-Separator Format**

1. **Text Normalization** (browserless_service.py:408-411)
   ```python
   # NEW Pattern 1: digits€XX → digits.XX €
   content = re.sub(r'(\d+)€(\d{2})\b', r'\1.\2 €', content)
   # "3€99" → "3.99 €"
   # "19€50" → "19.50 €"

   # Pattern 2: digits,XX € → digits.XX € (existing)
   content = re.sub(r'(\d+),(\d{2})\s*€', r'\1.\2 €', content)
   # "3,99 €" → "3.99 €"
   ```

2. **Extracted Price Normalization** (browserless_service.py:438-441)
   - Same patterns applied to directly extracted prices
   - Ensures consistency regardless of extraction method

**Regex Details:**
- `(\d+)€(\d{2})\b` matches:
  * `\d+` = one or more digits (euros)
  * `€` = euro symbol as separator
  * `(\d{2})` = exactly 2 digits (cents)
  * `\b` = word boundary (prevents matching "10€999")
- Replacement: `\1.\2 €` = "digits.cents €"

**Supported Formats After Fix:**
```
"3,99 €"   → "3.99 €"  ✓ (existing)
"3€99"     → "3.99 €"  ✓ (NEW)
"19,50 €"  → "19.50 €" ✓ (existing)
"19€50"    → "19.50 €" ✓ (NEW)
"1 234,56" → "1234.56" ✓ (existing)
```

**Flow:**
```
Page text: "Porte Magique 3€99"
  ↓
Normalization: "Porte Magique 3.99 €"
  ↓
Extracted: "3€99" → Normalized: "3.99 €"
  ↓
Sent to AI: "PRIX DÉTECTÉ: 3.99 €\n\nPorte Magique 3.99 €"
  ↓
AI sees: "3.99" (NOT "399")
  ↓
Returns: 3.99 ✓
```

**Expected Results:**
- Before: "3€99" → AI returns 399.00 ❌
- After: "3€99" → Normalized to "3.99 €" → AI returns 3.99 ✅

Fixes L'Incroyable and any other sites using euro-as-separator format.
Complements existing comma-separator normalization.
2025-11-29 23:56:39 +00:00
Claude 1c763143ca feat: Improve price extraction with priority selectors and smart multi-price handling
**Problem:**
Stokomani Lutin product detected 14.99 EUR instead of 19.99 EUR.
Website shows both old price (14.99) and current price (19.99).
Generic extraction was returning whichever price was found first.

**Root Causes:**
1. No prioritization of price selectors (all treated equally)
2. First-found-first-returned approach missed semantic selectors
3. When multiple prices found, no logic to select the correct one

**Solution: Priority-Based Extraction + Smart Selection**

1. **Reordered Selectors by Priority** (browserless_service.py:151-170)
   ```
   HIGHEST PRIORITY (return immediately):
   - [itemprop='price']    (Schema.org semantic)
   - .current-price        (explicit current)
   - .sale-price           (sale/promo price)
   - .final-price          (final calculated)
   - .product-price        (product-specific)

   MEDIUM PRIORITY:
   - [class*='price']:not([class*='old']):not([class*='was'])...
   - Excludes: old, was, original, before, regular, ancien, barre

   LOW PRIORITY:
   - .price (generic, might catch anything)
   ```

2. **Immediate Return for High-Priority Selectors** (lines 202-204)
   - If found with semantic selector → return immediately
   - Avoids checking lower-priority selectors
   - Most reliable prices checked first

3. **Multi-Price Smart Selection** (lines 210-218)
   - If multiple prices found (e.g., old + current)
   - Takes LAST price in DOM order (not highest/lowest value)
   - Rationale: HTML structure typically shows old price first, current price last:
     ```html
     <span class="old-price">14,99 €</span>  ← First in DOM
     <span class="price">19,99 €</span>      ← Last in DOM (CURRENT)
     ```
   - Logs all found prices for debugging

**Why Not "Highest Price"?**
Taking highest value fails in promotions:
```
Old: 29,99 € (highest)
Current: 19,99 € (promotion) ← Should select this
```
Taking last in DOM works for both cases:
```
Case 1 (Stokomani):
  Old: 14,99 € (first)
  Current: 19,99 € (last) ← Selected ✓

Case 2 (Promotion):
  Old: 29,99 € (first)
  Current: 19,99 € (last) ← Selected ✓
```

**Flow:**
```
Search for [itemprop='price']
  → Found? Return immediately ✓
  → Not found? Continue...
Search for .current-price
  → Found? Return immediately ✓
  → Not found? Continue...
...
Search with low-priority selectors
  → Found: ["14,99 €", "19,99 €"]
  → Select last: "19,99 €" ✓
```

**Expected Results:**
- Before: "14.99 EUR" (first found, wrong)
- After: "19.99 EUR" (last in DOM, correct)
- Logs: "Multiple prices found: ['14,99 €', '19,99 €'], selecting last: 19,99 €"

Tested with: https://www.stokomani.fr/products/lutin-filou-telescopique-95-cm_208400
2025-11-29 23:55:05 +00:00
Claude 4c98fee285 fix: Skip strikethrough prices to avoid detecting old/crossed-out prices
**Problem:**
AI was detecting old crossed-out prices instead of current prices.
Example: Stokomani Lutin - Current: 19,99€, Old (strikethrough): 14,99€
AI detected: 14.99 EUR (wrong - old price)

**Root Cause:**
Generic price extraction was finding ALL price elements without checking
if they were visually crossed-out (text-decoration: line-through).
Many e-commerce sites show:
```html
<span class="old-price" style="text-decoration: line-through">14,99 €</span>
<span class="price">19,99 €</span>
```
The selector finds both, but we were returning the first found.

**Solution: Check CSS text-decoration**

Added strikethrough detection (browserless_service.py:182-189):
```python
# Check if element is strikethrough (old price)
text_decoration = await element.evaluate(
    "el => window.getComputedStyle(el).textDecoration"
)
if "line-through" in text_decoration:
    continue  # Skip crossed-out prices
```

**Flow:**
```
Found elements with .price selector: [elem1, elem2, elem3]
  ↓
For each element:
  1. Check visibility ✓
  2. Check text-decoration
     → "line-through" → SKIP ✓
     → "none" → CONTINUE
  3. Extract price text
  ↓
Return first non-strikethrough price
```

**CSS Patterns Detected:**
- `text-decoration: line-through` (most common)
- `text-decoration: line-through solid`
- Combined styles ignored if they don't contain "line-through"

**Expected Results:**
- Before: Returns first price found (14.99 if it's first in DOM)
- After: Skips strikethrough prices, returns current price (19.99)

**Note:** Existing selector already excludes common classes:
`:not([class*='old']):not([class*='was']):not([class*='original'])`

This adds runtime CSS check as additional safety layer.

Partial fix for Stokomani Lutin 14.99 vs 19.99 issue.
May need site-specific selectors if problem persists.
2025-11-29 23:52:10 +00:00
Claude 531ed1bc15 fix: Convert French price format to English before sending to AI
**Problem:**
AI was misinterpreting French decimal comma as thousands separator:
- "3,99 €" detected as 399.00 EUR (instead of 3.99 EUR)
- "1,99 €" detected as 199.00 EUR (instead of 1.99 EUR)
- Caused by AI reading comma as grouping separator, not decimal point

**Root Cause:**
French format uses comma for decimals: "3,99 €" = three euros ninety-nine cents
English format uses dot for decimals: "3.99 €" = three euros ninety-nine cents

AI models (trained mostly on English) interpret:
- "3,99" → "3" and "99" separate → 399 (three hundred ninety-nine)
- Should be: "3.99" → 3.99 (three point ninety-nine)

**Solution: Pre-Convert Format Before AI Processing**

1. **Normalize ALL Text** (browserless_service.py:380-386)
   ```python
   # Convert French to English in all page text
   content = re.sub(r'(\d+),(\d{2})\s*€', r'\1.\2 €', content)
   # "3,99 €" → "3.99 €"
   # "12,50 €" → "12.50 €"
   # "1 234,56 €" → "1234.56 €" (also removes space thousands separator)
   ```

2. **Normalize Extracted Price** (browserless_service.py:400-410)
   ```python
   # Convert French format in directly extracted price
   normalized_price = re.sub(r'(\d+),(\d{2})', r'\1.\2', extracted_price)
   # "3,99 €" → "3.99 €"
   # Prepends: "PRIX DÉTECTÉ: 3.99 €" (not "3,99 €")
   ```

3. **Updated AI Prompt** (ai_schema.py:135-158)
   - Removed instructions to convert comma→dot (now done automatically)
   - Added: "The price has already been converted to English format for you"
   - Added explicit warnings about common mistakes:
     * "3.99 €" means 3.99 (NOT 399.00, NOT 3990.00)
     * "1.99 €" means 1.99 (NOT 199.00, NOT 1990.00)
   - Clear examples with both CORRECT and INCORRECT values

**Flow:**
```
Page text: "Lutin 3,99 € Ajouter au panier"
  ↓
Normalization: "Lutin 3.99 € Ajouter au panier"
  ↓
Extracted price: "3,99 €" → Normalized: "3.99 €"
  ↓
Sent to AI: "PRIX DÉTECTÉ: 3.99 €\n\nLutin 3.99 € Ajouter au panier"
  ↓
AI sees only English format: "3.99"
  ↓
AI Response: {"price": 3.99}  ✅ CORRECT
```

**Expected Results:**
- Before: "3,99 €" → AI returns 399.00 ❌
- After: "3,99 €" → Normalized to "3.99 €" → AI returns 3.99 ✅

**Regex Details:**
- `(\d+),(\d{2})\s*€` matches French prices: captures digits before comma, 2 digits after
- `\1.\2 €` replaces with: same digits, dot instead of comma, space + euro symbol
- Handles: 0,99 / 3,99 / 12,99 / 89,99 / 1234,99

Fixes issue where AI misread French decimal comma causing 100x price errors.
2025-11-29 23:51:00 +00:00
LogiFlow 8c35ace6a9 Merge pull request #139 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
feat: Add generic price extraction for all e-commerce sites
2025-11-30 00:45:37 +01:00
Claude b4dd4b4365 feat: Add generic price extraction for all e-commerce sites
**Problem:**
Price extraction was only working for Amazon. Other sites (B&M, Stokomani, etc.)
had prices missing from extracted text, causing AI to rely solely on screenshots.

Example error:
```
WARNING - No prices found in text with € symbol
AI Response: {"price": 9.99}  // Wrong! Real price was 1.99€
```

**Root Cause:**
- `page.inner_text('body')` doesn't capture all price elements
- Some sites use hidden elements, iframes, or CSS pseudo-elements
- Only Amazon had dedicated price extraction (_extract_amazon_price)
- Other sites got: `WARNING - Could not extract price from DOM`

**Solution: Universal Price Extraction**

1. **New Method: _extract_generic_price()** (browserless_service.py:145-208)
   - Tries common CSS selectors across all e-commerce sites:
     * Class-based: `.price`, `.product-price`, `.current-price`
     * Attribute-based: `[data-price]`, `[itemprop='price']`
     * French-specific: `[class*='prix']`, `[class*='tarif']`
     * Excludes old prices: `:not([class*='old'])`
   - Validates price format (contains € or comma with digits)
   - Falls back to regex: `\d+[,\.]\d{2}\s*€` on all text
   - Returns first valid price found

2. **Updated Extraction Logic** (browserless_service.py:319-338)
   - Amazon → uses `_extract_amazon_price()` (specific)
   - All other sites → uses `_extract_generic_price()` (universal)
   - Prepends "PRIX DÉTECTÉ: X,XX €" to text context
   - Logs success/failure for debugging

**Flow:**
```
Page load (bmstores.fr)
  ↓
Try selectors: .price, .product-price, [itemprop='price']...
  ↓
Found: "1,99 €" via selector .price
  ↓
Prepend: "PRIX DÉTECTÉ: 1,99 €\n\n[page text]"
  ↓
Send to AI
  ↓
AI sees explicit price → {"price": 1.99, "confidence": 0.95}
```

**Expected Results:**
- Before: `WARNING - No prices found in text with € symbol`
- After: `INFO - Found price via selector .price: 1,99 €`
- AI confidence improves (text + image vs image only)
- Fewer digit recognition errors (1 vs 9)

**Tested with:**
- https://bmstores.fr/produits/accessoire-noel-enfant/125651-lutin-farceur-35cm
- Price: 1,99 €

Fixes issue where non-Amazon sites had missing prices in text context,
causing AI to misread prices from screenshots only.
2025-11-29 23:43:42 +00:00
LogiFlow f8a9e45921 Merge pull request #138 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
fix: Improve AI price digit recognition to avoid 1.99/9.99 confusion
2025-11-30 00:39:18 +01:00
Claude 0529c885d3 fix: Improve AI price digit recognition to avoid 1.99/9.99 confusion
**Problem:**
AI was misreading prices: "1,99 €" detected as "9,99 €" with 90% confidence.
This is a critical OCR error where similar-looking digits (1, 7, 8, 9) are confused.
The model google/gemini-2.5-flash-lite made basic digit recognition mistakes.

**Root Cause:**
- AI vision models can confuse visually similar digits
- No specific instructions in prompt to verify digit accuracy
- No validation that detected price matches text context
- Model may be too lightweight for reliable OCR

**Solution: Three-Part Fix**

1. **Enhanced Prompt with Digit Verification** (ai_schema.py:145-149)
   - Added "CRITICAL - Read digits carefully" section
   - Explicit warning: "1,99 €" is NOT "9,99 €"
   - Instructions to double-check digits 1, 7, 8, 9 (visually similar)
   - Context check: Small prices (< 5€) are common
   - Verify price makes sense for product type

2. **More Examples in Prompt** (ai_schema.py:152-153)
   - Added: "1,99 €" -> 1.99 (NOT 9.99)
   - Added: "0,99 €" -> 0.99 (NOT 9.99)
   - Reinforces correct parsing of small prices

3. **Price Extraction Logging** (ai_service.py:349-355)
   - Extract ALL prices from text with regex: \d+[,\.]\d{2}\s*€
   - Log first 10 prices found: ["1,99 €", "2,99 €", ...]
   - Allows debugging: Is correct price in text?
   - Logs warning if NO prices found in text

4. **Debug Logs to INFO** (ai_service.py:347, 361)
   - Changed logger.debug → logger.info for text/prompt preview
   - Now visible in production logs for diagnosis

**Expected Improvement:**
- AI should now pay closer attention to digit shapes
- Lower confidence (0.5-0.7) if digits are unclear
- Better accuracy for small prices (0,99 - 4,99)
- Logs will show if problem is in text extraction or AI vision

**Testing:**
After restart, logs will show:
```
INFO - Cleaned text preview: 'Lutin de Noël 1,99 € Ajouter...'
INFO - Prices found in text: ['1,99 €', '2,99 €']
INFO - AI Response: {"price": 1.99, "price_confidence": 0.85}
```

If AI still returns wrong price but text has correct price,
then model is insufficient and should be upgraded to:
- google/gemini-2.0-flash-exp
- anthropic/claude-3-haiku
- openai/gpt-4o-mini
2025-11-29 23:36:47 +00:00
LogiFlow b4f957436e Merge pull request #137 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
fix: Add concurrency limits to prevent Browserless timeout saturation
2025-11-30 00:32:29 +01:00
Claude 10cdfcef5a fix: Add concurrency limits to prevent Browserless timeout saturation
**Problem:**
Search was launching ALL sites in parallel without any limit, causing Browserless
to be overwhelmed with simultaneous connections and timeout:
- 7 active sites = 7 parallel Browserless connections
- Each site scraping multiple products in parallel
- Result: "Timeout 60000ms exceeded" errors
- Most searches failed with 0 results

**Root Cause:**
```python
# search_service.py:489 (before fix)
for gen in generators:
    asyncio.create_task(producer(gen))  # ❌ ALL sites at once
```

With 7 sites × 3 products each = up to 21 concurrent Browserless connections.

**Solution: Two-Level Concurrency Control**

1. **Global Site Limit** (search_service.py:476)
   - Added `site_semaphore = asyncio.Semaphore(2)`
   - Max 2 sites can search in parallel
   - Others wait for a slot to free up
   - Example: 7 sites → only 2 active at a time

2. **Per-Site Product Limit** (search_service.py:306)
   - Reduced from `Semaphore(3)` to `Semaphore(2)`
   - Max 2 products scraped in parallel per site
   - Less pressure on Browserless

**Impact:**
- Before: 7 sites × 3 products = 21 concurrent connections ❌
- After: 2 sites × 2 products = 4 concurrent connections ✅
- 80% reduction in concurrent load on Browserless

**Flow:**
```
7 sites requested
  ↓
2 start searching (sites 1, 2)
5 wait in queue (sites 3-7)
  ↓
Site 1 completes → Site 3 starts
Site 2 completes → Site 4 starts
  ↓
Process continues until all sites done
```

**Expected Results:**
✅ No more Browserless timeout errors
✅ All sites complete successfully
✅ Sequential processing prevents saturation
✅ Search returns results from all sites

Fixes issue where search returned "Found 0 results" for most sites
due to Browserless connection timeouts.
2025-11-29 23:31:43 +00:00
LogiFlow 6e48e7ae30 Merge pull request #136 from R0m1k3/antigravity
feat: add Bonial and Tiendeo catalog scrapers
2025-11-30 00:24:06 +01:00
LogiFlow 130b187632 Merge pull request #135 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
fix: Add extract_text parameter to fix broken product search
2025-11-30 00:23:48 +01:00
LogiFlow 6ecf315380 Merge pull request #134 from R0m1k3/claude/add-product-monitoring-01DyNDGnKBX4fHNSjAkwVckd
feat: Add direct Amazon price extraction with CSS selectors
2025-11-30 00:23:28 +01:00
Claude 03793cea0a fix: Add extract_text parameter to fix broken product search
**Problem:**
Previous commit changed get_page_content() to always return visible text (inner_text)
instead of HTML source. This broke product search which uses BeautifulSoup to parse
HTML with CSS selectors. Search was returning 0 results for all sites.

**Root Cause:**
- Search service uses: `soup.select(config["product_selector"])` on HTML
- But was receiving plain text instead of HTML structure
- BeautifulSoup couldn't find any product links → 0 results

**Solution: Add extract_text parameter**

Modified `get_page_content()` signature:
```python
async def get_page_content(
    url: str,
    use_proxy: bool = False,
    wait_selector: str = None,
    extract_text: bool = False  # NEW parameter
) -> tuple[str, str]:
```

**Behavior:**
- `extract_text=False` (DEFAULT): Returns HTML source via `page.content()`
  → Used by search service for BeautifulSoup parsing

- `extract_text=True`: Returns visible text via `page.inner_text('body')`
  → Used by AI monitoring for price extraction
  → Includes Amazon price prepending ("PRIX DÉTECTÉ: ...")

**Changes:**

1. **browserless_service.py** (Lines 197-329)
   - Added `extract_text` parameter with default `False`
   - Conditional logic: if extract_text, use inner_text + Amazon extraction
   - Else: use page.content() (HTML source)
   - Preserves backward compatibility (default = HTML)

2. **scheduler_service.py** (Line 113)
   - Added `extract_text=True` for AI monitoring
   - Ensures AI gets visible text with Amazon prices

3. **search_service.py** (Line 104)
   - Added `extract_text=True` for scrape_item() AI analysis
   - Search itself uses default (HTML) for BeautifulSoup parsing

**Impact:**
✅ Search now works again (gets HTML for BeautifulSoup)
✅ AI monitoring gets visible text (better extraction)
✅ Amazon price extraction only runs when extract_text=True
✅ Backward compatible (default behavior = HTML)

Fixes issue where all searches returned: "Found 0 results for [site]"
2025-11-29 23:23:24 +00:00
Michael 34ac9b343a feat: add Bonial and Tiendeo catalog scrapers 2025-11-30 00:22:36 +01:00
Claude 76f6522528 feat: Add direct Amazon price extraction with CSS selectors
Critical fix for Amazon price detection - AI was returning price=null with 0.0 confidence
even though stock detection worked (1.0 confidence).

**Root Cause:**
- Amazon injects prices via JavaScript into hidden .a-offscreen elements (for screen readers)
- These elements contain the full price but are not visible
- AI vision models struggled to locate tiny price text in large screenshots
- Text filtering didn't always capture the exact price location

**Solution: Direct DOM Extraction + AI Fallback**

1. **New Method: _extract_amazon_price()** (browserless_service.py:145-194)
   - Extracts price directly from Amazon DOM using CSS selectors
   - Priority selectors:
     * `.a-price .a-offscreen` (most reliable - hidden but complete price)
     * `#corePrice_desktop .a-price .a-offscreen`
     * `#corePriceDisplay_desktop_feature_div .a-price .a-offscreen`
     * `#priceblock_ourprice` (older layouts)
     * Fallback: `span.a-price-whole` + `span.a-price-fraction`
   - Handles hidden elements (.a-offscreen) without visibility checks
   - Returns formatted price text: "89,99 €"

2. **Price Injection into Context** (browserless_service.py:304-309)
   - Prepends "PRIX DÉTECTÉ: 89,99 €" to page text
   - Gives AI explicit price information at the start of context
   - Only for Amazon product pages (detected via "/dp/" in URL)

3. **Improved AI Prompt** (ai_schema.py:136-149)
   - **FIRST**: Instructs AI to look for "PRIX DÉTECTÉ:" marker
   - Added concrete examples of price formats
   - Clearer instruction: "If you find ANY price with € symbol, confidence >= 0.8"
   - Reduces false negatives from overly cautious AI

**Flow:**
```
Amazon page load
  ↓
Extract price via .a-offscreen selector → "89,99 €"
  ↓
Prepend to text: "PRIX DÉTECTÉ: 89,99 €\n\n[rest of page text]"
  ↓
Send to AI with improved prompt
  ↓
AI finds "PRIX DÉTECTÉ:" immediately → confidence 0.8-1.0
```

**Why This Works:**
- Direct extraction is 100% reliable for Amazon's consistent DOM structure
- AI gets explicit price hint at start of text (most important info first)
- Prompt tells AI exactly where to look
- Even if direct extraction fails, AI can still find price in screenshot/text

**Expected Improvement:**
- Before: `{"price": null, "price_confidence": 0.0}`
- After: `{"price": 89.99, "price_confidence": 0.95}`

Tested with Amazon.fr MSI Mag product page.
2025-11-29 23:10:32 +00:00
LogiFlow 1b3f186eed Merge pull request #133 from R0m1k3/antigravity
feat: add Bonial.fr catalog scraper service using Playwright.
2025-11-30 00:00:10 +01:00
Michael a7aefb3e56 feat: add Bonial.fr catalog scraper service using Playwright. 2025-11-29 23:59:50 +01:00
LogiFlow 6935acc135 Merge pull request #132 from R0m1k3/antigravity
feat: Implement Bonial.fr scraper service to extract promotional cata…
2025-11-29 23:56:08 +01:00