Merge pull request #237 from R0m1k3/antigravity

feat: define canonical AI extraction schema with vision-first prompt …
This commit is contained in:
LogiFlow authored and GitHub committed 2025-12-22 10:38:36 +01:00
commit a42decf315
4 files changed
+128 -75

No files matched your search

+28 -39
View File
@@ -130,50 +130,39 @@ class AIExtractionMetadata(BaseModel):
# Prompt template for schema-first extraction
EXTRACTION_PROMPT_TEMPLATE = """Extract product price and stock status from this French e-commerce page.
# Prompt template for Vision-First extraction
EXTRACTION_PROMPT_TEMPLATE = """You are a Vision-First Price Extraction Agent.
Your Goal: Extract the main product price exactly as a human sees it on the screen.
**PRICE (IMPORTANT - French format):**
- **FIRST**: Look for "PRIX DÉTECTÉ:" at the start of text - this is the extracted price
- **The price has already been converted to English format for you**
- French original: "3,99 €" → Already shown to you as: "3.99 €"
- Just extract the number you see (e.g., "3.99" from "PRIX DÉTECTÉ: 3.99 €")
- Currency symbol is usually € at the end
- Look for: "PRIX DÉTECTÉ:", price tags, "Prix:", "€", numbers near "Ajouter au panier"
- Extract as DECIMAL NUMBER: If you see "3.99", return 3.99
- Ignore crossed-out/barré prices (old prices)
- **CRITICAL:** Ignore "Prix au litre", "Prix au kg", "P.U.", or unit prices usually shown in smaller text/parentheses (e.g., "4.60 € / L").
- **CRITICAL:** If you see both HT (Hors Taxe) and TTC (Toutes Taxes Comprises) prices, **ALWAYS select the TTC price**.
- Look for labels like "Taxe incluse", "TTC", "Prix payé". Ignore "HT" or "Hors Taxe".
- Example: If text has "1.15 € HT" and "1.38 € TTC", return 1.38 (NOT 1.15).
- If multiple prices, take the current/main price (not the original, not the unit price, not the HT price)
- **B&M STORES Specific:** The main price is often large and bold (e.g. "1.38€"), while unit price is small. ALWAYS take the main price (TTC). Ignore hidden HT prices (e.g. 1.15€).
**SOURCE OF TRUTH = IMAGE**
- The image provided is the **Absolute Truth**.
- The text provided below is scraped HTML content which may contain hidden/old prices.
- **IF IMAGE AND TEXT CONFLICT, TRUST THE IMAGE.**
- Only use the text if the image is completely unreadable or missing the price.
**CRITICAL - Common mistakes to avoid:**
- "3.99 €" means 3.99 (NOT 399.00, NOT 3990.00)
- "1.99 €" means 1.99 (NOT 199.00, NOT 1990.00)
- "0.99 €" means 0.99 (NOT 99.00, NOT 990.00)
- The decimal point separates euros from cents
- Small prices (< 10€) are very common for everyday items
**PRICE EXTRACTION RULES (French Format):**
1. **Visual Focus**: Look for the largest, boldest price on the screen. This is usually the main product price.
2. **Ignore Small Text**: Ignore "Prix au litre", "Prix au kg", or small unit prices (e.g., "(4.60 € / L)").
3. **Ignore Strikethrough**: Do not extract crossed-out prices (old prices).
4. **Ignore "HT"**: Always find the "TTC" (Tax Included) price. If you see "1.15 € HT" and "1.38 €", the visual price is 1.38.
5. **Ignore "Suggestions"**: Do not extract prices from "Other customers bought" or "Recommended products" sections.
- Examples of valid prices:
* "PRIX DÉTECTÉ: 1.99 €" -> 1.99 (NOT 199 or 1990)
* "PRIX DÉTECTÉ: 3.99 €" -> 3.99 (NOT 399 or 3990)
* "0.99 €" -> 0.99 (NOT 99)
* "89.99 €" -> 89.99 (NOT 8999)
* "1234.56 €" -> 1234.56
- If you find ANY price with € symbol, extract it with confidence >= 0.8
- If digits are unclear or blurry, reduce confidence to 0.5-0.7
- If unclear: set null and confidence < 0.5
**Output Format Cleaning:**
- "3,99 €" -> 3.99
- "1 234,56 €" -> 1234.56
- "0.99 €" -> 0.99
**STOCK:**
- TRUE if: "Ajouter au panier", "Acheter", "En stock", "Disponible", "Add to Cart", "Retrait 2h", "Click & Collect"
- FALSE if: "Rupture", "Indisponible", "Épuisé", "Out of Stock", "Notify Me"
- NULL if unclear or not shown
**STOCK STATUS RULES:**
- Check the button color and text.
- Green/Blue "Ajouter au panier" -> true
- Grey/Red "Rupture", "Indisponible" -> false
- If in doubt, look for "En stock" text.
**CONFIDENCE (0.0 to 1.0):**
- 0.9-1.0: Price clearly visible
- 0.5-0.8: Price found but partially obscured or uncertain
- Below 0.5: Cannot find price reliably
**CONFIDENCE SCORE:**
- 1.0: Price is clearly visible in the image and matches text.
- 0.9: Price is clearly visible in the image, even if text is missing.
- 0.5: Price found in Text ONLY (Image unclear).
- 0.0: No price found.
Respond ONLY with valid JSON:
{{
+12 -36
View File
@@ -1,40 +1,16 @@
# Task: Fix Screenshot Update Issue
# Task Board: Vision-Based Price Extraction
## Context
## 🚀 Current Focus
The user reports that in the "Suivi prix" (Tracking) section, triggering an update ("Mise à jour") does not update the screenshot. This has been identified as a caching issue due to static filenames.
- [x] Analysing current hybrid (Text + Image) extraction logic
- [x] Proposing "Vision First" strategy to owner
- [ ] **Implementation Phase**:
- [ ] Update `ai_schema.py` STRICT Vision Definition.
- [ ] Refine Prompt with "IMAGE IS TRUTH" directive.
- [ ] **Verification Phase**:
- [ ] Verify prompt generation.
## Current Focus
## 📝 Progress Log
Implementing logic to generate unique timestamped filenames for screenshots to bypass browser caching and ensure the latest image is displayed.
## Master Plan
- [x] Modify `tracking_scraper_service.py` to use timestamped filenames for tracking screenshots <!-- id: 0 -->
- [x] Verify that `item_service.py` correctly picks up the new files <!-- id: 1 -->
- [x] Create verification script `verify_screenshot_update.py` <!-- id: 2 -->
- [x] Run verification and confirm fix <!-- id: 3 -->
- [x] Modify `ItemService.get_items` to scan filesystem for latest screenshot <!-- id: 4 -->
- [x] Update `ItemService.delete_item` to clean up all related screenshots <!-- id: 6 -->
- [x] Create `verify_fix_item_service.py` to test the new logic <!-- id: 5 -->
## Current Focus
Improving price extraction reliability for B&M Stores and others.
- [x] Analyze `ai_schema.py` to check the extraction prompt <!-- id: 7 -->
- [x] Improve `ScraperService._extract_text` to be more targeted (e.g. main content only) <!-- id: 8 -->
- [x] Update AI prompt to better handle multiple prices (unit vs package) <!-- id: 9 -->
- [x] Verify extraction logic with simulation script <!-- id: 10 -->
## Progress Log
- Identified the issue: `ScraperService` overwrites `item_{id}.png`.
- Created implementation plan.
- User approved plan.
- Implemented filesystem scanning in `ItemService`.
- Verified fix with `verify_fix_item_service.py` successfully.
- Improved AI prompt to ignore unit prices (like "Prix au litre").
- Enhanced text extraction to remove menu/footer noise.
- Verified logic with `verify_extraction_logic.py` (simulated).
- **Fixed HT vs TTC issue**: AI was extracting 1.15€ (HT) instead of 1.38€ (TTC). Updated prompt to prioritize "Taxe incluse"/TTC prices.
- **2025-12-22**: Initialized task. Confirmed current system is *already* sending images, but likely confusing the AI with conflicting text data.
- **2025-12-22**: User selected **Option 2 (Vision Priority)**. Proceeding to rewrite AI System Prompt.
+50
View File
@@ -0,0 +1,50 @@
import sys
import unittest
# Mock modules to avoid ImportError for app dependencies we don't need for this specific test
from unittest.mock import MagicMock
sys.modules["sqlalchemy"] = MagicMock()
sys.modules["sqlalchemy.orm"] = MagicMock()
sys.modules["app.database"] = MagicMock()
sys.modules["app.utils.image"] = MagicMock()
sys.modules["app.utils.text"] = MagicMock()
sys.modules["app.utils.text"].filter_relevant_text = lambda text, max_length: text
# Mock pydantic
mock_pydantic = MagicMock()
class MockBaseModel:
pass
mock_pydantic.BaseModel = MockBaseModel
mock_pydantic.Field = MagicMock(return_value=None)
mock_pydantic.field_validator = MagicMock(return_value=lambda x: x)
sys.modules["pydantic"] = mock_pydantic
# Import the schema module
from app.ai_schema import get_extraction_prompt
class TestVisionPriorityPrompt(unittest.TestCase):
def test_vision_first_directives(self):
"""Verify that the prompt contains the Vision-First directives."""
# Scenario: Some random text context
page_text = "Some random text content from the page."
prompt = get_extraction_prompt(page_text)
print("\nGenerated Prompt Snippet:\n", prompt[:500], "...\n")
# Check for Critical Directives
self.assertIn("Vision-First Price Extraction Agent", prompt)
self.assertIn("**SOURCE OF TRUTH = IMAGE**", prompt)
self.assertIn("IF IMAGE AND TEXT CONFLICT, TRUST THE IMAGE", prompt)
# Check for stock rules
self.assertIn("STOCK STATUS RULES", prompt)
def test_prompt_without_text(self):
"""Verify prompt structure when no text is provided."""
prompt = get_extraction_prompt(None)
self.assertIn("**SOURCE OF TRUTH = IMAGE**", prompt)
self.assertNotIn("**Relevant text from page:**", prompt)
if __name__ == "__main__":
unittest.main()
+38
View File
@@ -0,0 +1,38 @@
# Walkthrough: Vision-First Price Extraction
In response to issues where the AI was being misled by hidden text (like unit prices or old prices in HTML), we have implemented a **Vision-Priority Strategy**.
## Changes Implemented
### 1. Updated AI System Prompt (`app/ai_schema.py`)
We completely rewrote the `EXTRACTION_PROMPT_TEMPLATE` to enforce the following rules:
- **Source of Truth = Image**: explicit instruction that the screenshot takes precedence over any text.
- **Conflict Resolution**: "IF IMAGE AND TEXT CONFLICT, TRUST THE IMAGE."
- **Visual Focus Rules**:
- Look for the largest/boldest price.
- Ignore small, styling-less text (often unit prices).
- Ignore crossed-out text.
### Verification
We verified the new prompt generation using `verify_vision_priority.py`.
**Generated Prompt Preview:**
```text
You are a Vision-First Price Extraction Agent.
Your Goal: Extract the main product price exactly as a human sees it on the screen.
**SOURCE OF TRUTH = IMAGE**
- The image provided is the **Absolute Truth**.
- The text provided below is scraped HTML content which may contain hidden/old prices.
- **IF IMAGE AND TEXT CONFLICT, TRUST THE IMAGE.**
```
## How to Test
1. Go to "Suivis Prix".
2. Force refresh an item that was previously incorrect (e.g., B&M item showing unit price).
3. The AI should now ignore the "hidden" unit price text and read the main price tag from the image.