Michael
2fae5350b3
feat: define canonical AI extraction schema with vision-first prompt and add vision priority verification script
2025-12-22 10:38:08 +01:00
Michael
83fa8ed734
feat: Add AI extraction schemas and prompt templates, and refine price extraction logic to prioritize TTC over HT.
2025-12-18 16:57:23 +01:00
Michael
384baa3575
feat: Introduce AI extraction schema with validation, prompt generation, and a verification script to improve price extraction reliability.
2025-12-18 16:32:32 +01:00
Michael
ef7cd7c1cf
feat: Add BMStores product parser and a new Playwright-based scraping service with popup handling.
2025-12-18 15:20:19 +01:00
Michael
51ede950ff
feat: Add scheduled item tracking service with dedicated web scraping and data extraction capabilities.
2025-12-18 13:23:53 +01:00
Michael
d50b18690e
feat: Add scheduler and scraper services for automated item tracking and price extraction.
2025-12-18 12:38:32 +01:00
Michael
4e9f2662bc
feat: add Playwright-based scraping service for tracking items.
2025-12-18 11:11:30 +01:00
Michael
5428459f36
feat: Add web scraping service for tracking and AI extraction verification script.
2025-12-18 11:05:09 +01:00
Michael
6fab64b229
feat: Add ScraperService with Playwright for robust web scraping, including browser management, configurable scraping options, and automated popup handling.
2025-12-18 10:56:06 +01:00
Michael
0d68859a5b
feat: add Playwright script to scrape gifi.fr product page and extract price information.
2025-12-18 10:42:31 +01:00
Michael
97664a4a16
fix: Improve ItemService screenshot management by prioritizing latest timestamped files and ensuring complete deletion.
2025-12-18 08:50:45 +01:00
Michael
4590273801
docs: revise task.md to address screenshot update caching instead of popup obscuration.
2025-12-18 08:34:02 +01:00
Michael
ed7f23d657
chore: Update task plan to focus on removing obscuring popups from product screenshots.
2025-12-18 08:14:56 +01:00
Michael
e930f4bbfa
feat: Implement main FastAPI application entry point, including database migrations, data seeding, and background task scheduling.
2025-12-02 18:52:41 +01:00
Michael
947152e132
feat: Implement scheduler_service for automated item price and stock tracking, including AI analysis and database updates.
2025-12-02 18:03:32 +01:00
Michael
fb89af5aba
Resolve merge conflict in scheduler_service.py
2025-12-02 17:52:48 +01:00
Michael
2c0a573a66
feat: Add web scraping service using Playwright and Browserless, and a new scheduler service.
2025-12-02 17:15:28 +01:00
Michael
d8b3014021
Resolve merge conflict in browserless_service.py: Keep navigation logic and fix screenshot path
2025-12-02 17:04:49 +01:00
Michael
241970cf70
feat: implement BrowserlessService for persistent browser management with auto-reconnection and enhanced price extraction.
2025-12-02 16:54:40 +01:00
Michael
8a8ef02168
feat: Introduce BrowserlessService for persistent web scraping with enhanced Amazon price extraction, alongside initial scheduler service and task tracking.
2025-12-02 13:59:39 +01:00
Michael
3bbab058c6
feat: implement persistent browserless service for robust web scraping and price extraction with auto-reconnection.
2025-12-02 11:46:32 +01:00
Michael
e7875f35a7
feat: Add French localization file with comprehensive translations for the application UI.
2025-12-02 09:51:56 +01:00
Michael
86aae8f628
Merge branch 'origin/antigravity' into antigravity
...
Resolved conflict in browserless_service.py by merging Amazon interstitial selectors.
Combined local robust selectors with remote additions for maximum coverage.
2025-12-02 09:37:31 +01:00
Michael
2b21cd176c
feat: add TrackingScraperService for Playwright-based web scraping with popup handling and smart scrolling capabilities.
2025-12-02 09:28:08 +01:00
Michael
5516562cb6
fix: Handle Amazon interstitials and dynamic screenshot updates
...
- Added Amazon 'Continue' interstitial selectors to BrowserlessService
- Ported robust popup handling from TrackingScraperService to BrowserlessService
- Updated ItemService to fetch dynamic screenshot URLs from PriceHistory
(fixes issue where forced updates didn't show new screenshots)
2025-12-02 09:27:12 +01:00
Michael
9111aa02b9
feat: Add Playwright script to detect and analyze Amazon "Continuer les achats" popup elements.
2025-12-01 17:26:47 +01:00
Michael
9178512465
feat: implement BrowserlessService for unified, persistent Playwright browser management with auto-reconnection and enhanced price extraction.
2025-12-01 17:14:12 +01:00
Michael
6eb821c40d
feat: Add automated item checking scheduler service and new dashboard item display components.
2025-12-01 17:07:50 +01:00
Michael
f0e1c6668d
feat: Implement new search service for multi-site e-commerce product search and item detail scraping.
2025-12-01 16:56:32 +01:00
Michael
f223dc5a29
feat: Increase Amazon limit to 50 and fix screenshot framing
...
- Increased default Amazon search results limit from 20 to 50 in backend and frontend
- Improved TrackingScraperService to target main product area on Amazon (avoiding reviews)
- Added specific selectors for Amazon product page (#imgTagWrapperId, #productTitle, etc.)
2025-12-01 16:18:35 +01:00
Michael
acfc5c6397
feat: Enhance Amazon stealth with comprehensive anti-detection
...
Added comprehensive stealth techniques to bypass Amazon bot detection:
- Complete HTTP headers (Accept, Accept-Language, Sec-Fetch-*, etc.)
- Extended chrome object with loadTimes, csi, app
- Permissions API override
- Plugins, languages, platform spoofing
- Battery API mocking
This significantly improves bot detection evasion compared to the basic
2-line stealth that was insufficient for Amazon's detection systems.
2025-12-01 13:55:57 +01:00
Michael
2429fa2018
Revert "fix: Remove homepage redirect to avoid Amazon bot detection"
...
This reverts commit 5d3d181f5e .
2025-12-01 13:26:05 +01:00
Michael
5d3d181f5e
fix: Remove homepage redirect to avoid Amazon bot detection
...
Amazon detects the homepage-then-search pattern as bot behavior and blocks
with 2065 byte pages. Now navigating directly to search URL to appear more
natural.
This reverts the homepage loading strategy from commit 85b19e7 as Amazon's
bot detection has evolved and now flags this pattern.
2025-12-01 13:25:51 +01:00
Michael
95846ec606
fix: Disable ImprovedSearchService auto-init to restore Amazon functionality
...
ImprovedSearchService and AmazonScraperService were both connecting to
Browserless at startup, causing connection conflicts and Amazon blocking
(2065 byte pages).
Now only AmazonScraperService initializes at startup, keeping Browserless
in continuous connection for Amazon. ImprovedSearchService will initialize
on-demand when needed.
This restores Amazon search to working state as it was at commit 85b19e7 .
2025-12-01 13:21:32 +01:00
Michael
8dd3082a9a
Merge branch 'antigravity' into main - Resolve conflicts
...
Resolved conflicts by keeping antigravity improvements:
- tracking_scraper_service.py: 47 popup selectors with RGPD/Amazon support
- main.py: No auto-init of TrackingScraperService (on-demand only)
This brings all popup handling improvements to main branch.
2025-12-01 13:13:58 +01:00
Michael
53a05f45e4
fix: Remove auto-initialization of TrackingScraperService to fix Amazon scraper
...
TrackingScraperService now initializes on-demand instead of at startup.
This prevents connection conflicts with AmazonScraperService and ImprovedSearchService
since all services connect to the same Browserless instance.
Fixes: Amazon scraper blocking issue (page too small - 2065 bytes)
2025-12-01 13:05:49 +01:00
Michael
e2a874bf6c
feat: implement TrackingScraperService for robust web scraping with Playwright, including shared browser management and popup handling.
2025-12-01 12:42:14 +01:00
Michael
a792060f08
feat: Add TrackingScraperService with enhanced popup handling for screenshots
...
- Add 47 popup selectors (vs 12 generic before)
- Support French RGPD platforms (Axeptio, Didomi, OneTrust, TarteAuCitron)
- Amazon-specific detection and handling (6 selectors)
- Retry logic with verification pass
- Double popup cleanup before screenshot
- Increased timeouts for slow animations
- New _verify_no_popups() helper function
This resolves popup visibility issues in the Suivi/Dashboard screenshots.
2025-12-01 12:27:51 +01:00
Michael
22e15515f6
Merge branch 'main' of https://github.com/R0m1k3/Priceflow
2025-12-01 11:32:14 +01:00
Michael
f27a75d9d6
feat: Implement initial application structure with routing, authentication flow, and responsive layout.
2025-12-01 11:32:12 +01:00
Michael
7592c2781e
feat: Implement the core FastAPI application, including database migrations, service orchestration, and a new tracking scraper service.
2025-12-01 09:32:37 +01:00
Michael
ffc2f0d632
feat: introduce centralized search configuration for scraping, including site selectors, proxies, and user agents.
2025-12-01 07:40:50 +01:00
Michael
36ae466583
feat: Introduce ImprovedSearchService utilizing Playwright for persistent browser-based e-commerce scraping and price extraction, along with new search configuration.
2025-12-01 01:30:06 +01:00
Michael
b35401a87b
feat: Add improved search and AI price extraction services, update Docker Compose, and introduce benchmark verification script.
2025-12-01 01:20:39 +01:00
Michael
8303e767f7
feat: implement a new persistent browser-based search service for e-commerce scraping.
2025-12-01 01:11:42 +01:00
Michael
8bad965e29
feat: Add ImprovedSearchService for persistent browser-based e-commerce search scraping.
2025-12-01 01:07:22 +01:00
Michael
759a28b2d4
feat: Add ImprovedSearchService for persistent browser-based e-commerce search and scraping.
2025-12-01 01:05:40 +01:00
Michael
4dbb9f6424
feat: add improved search service with persistent browser connection
2025-12-01 01:03:51 +01:00
Michael
93840804a4
feat: implement product search page with site selection, real-time results, and monitoring integration
2025-12-01 01:01:53 +01:00
Michael
39893e8c1e
feat: Add improved search service with persistent browser connection for e-commerce scraping.
2025-12-01 01:00:25 +01:00
Michael
b5c5961618
feat: Implement AI-powered price extraction using Gemma 2 and an improved search service with persistent browser connections for web scraping.
2025-12-01 00:57:05 +01:00
Michael
c159036583
feat: Implement comprehensive search configurations for discount stores and add Gifi-specific search and analysis tools.
2025-12-01 00:36:54 +01:00
Michael
63073d03b5
feat: introduce new search configuration, improved search service, and Carrefour search integration test.
2025-11-30 23:54:48 +01:00
Michael
bf97cd22e7
feat: introduce improved search service with persistent browser and modular parsers
2025-11-30 23:49:34 +01:00
Michael
7cb0e2f94b
feat: add improved search service with persistent browser connection and modular site-specific parsers
2025-11-30 23:45:51 +01:00
Michael
8bf9ddb2d7
feat: implement improved search service using persistent browser connection and modular parsers, alongside a BeautifulSoup check utility.
2025-11-30 23:43:11 +01:00
Michael
87568866e7
feat: introduce improved search service with persistent browser connection and modular site-specific parsers.
2025-11-30 23:40:22 +01:00
Michael
4dd77bd889
feat: add improved search service with persistent browser, modular parsers, and popup handling.
2025-11-30 23:35:40 +01:00
Michael
9855040f3e
feat: Implement improved search service with supporting models, schemas, and API routers, and refactor site verification script.
2025-11-30 23:33:05 +01:00
Michael
4720ed6b93
feat: Implement improved search service with persistent browser and modular site-specific parsers, including Gifi.
2025-11-30 22:21:08 +01:00
Michael
64067347d1
feat: implement new search service with multiple site parsers, search configuration, and supporting inspection/verification scripts.
2025-11-30 22:04:14 +01:00
Michael
95e38287cb
feat: Add new parsers for Auchan, Carrefour, L'Incroyable, and Stokomani, including search configuration and deployment updates.
2025-11-30 22:03:40 +01:00
Michael
b44082aeb4
feat: Implement new product parsers for Stokomani, Auchan, Carrefour, and Gifi, and add a centralized search configuration module.
2025-11-30 20:26:19 +01:00
Michael
e0d73f4d0e
feat: add LaFoirFouille parser to extract product search results from lafoirfouille.fr
2025-11-30 19:09:09 +01:00
Michael
fbd06cfb17
feat: Add initial parser for Auchan.fr to extract product search results.
2025-11-30 19:05:38 +01:00
Michael
078060f155
feat: Add abstract base parser with product result dataclass and implement Stokomani parser for search results.
2025-11-30 19:03:26 +01:00
Michael
ebe685fdf9
feat: add Stokomani parser to extract product search results from stokomani.fr
2025-11-30 18:17:07 +01:00
Michael
b3d5c56707
feat: Add Stokomani search results parser
2025-11-30 18:13:11 +01:00
Michael
c2f042beee
feat: Implement improved search service with centralized configuration and site-specific parsers for various retailers.
2025-11-30 18:05:34 +01:00
Michael
a8bffaec08
feat: Implement initial e-commerce site parsers with a base class and factory for extracting product search results.
2025-11-30 17:42:32 +01:00
Michael
ee13a0bccd
fix: resolve merge conflict in improved_search_service - unified domain matching logic
2025-11-30 17:37:16 +01:00
Michael
3db7db48fa
feat: Add improved search service with persistent browser connection for e-commerce scraping.
2025-11-30 17:32:58 +01:00
Michael
33d9042b10
feat: add search_config module for browserless, proxy, user agent, and site-specific search configurations, and create list_missing_sites script.
2025-11-30 16:57:44 +01:00
Michael
7ba8850838
feat: Implement new search service with comprehensive configuration for various e-commerce sites, proxies, and user agents.
2025-11-30 16:36:51 +01:00
Michael
f20bae52bd
feat: introduce core search configuration for sites, proxies, and user agents, alongside selector diagnosis and HTML dumping utilities.
2025-11-30 16:27:02 +01:00
Michael
79dc4bfafc
feat: Implement centralized search configuration with site definitions, proxy settings, and user agents, and add a debugging script.
2025-11-30 15:58:27 +01:00
Michael
7477d505e5
feat: Add ImprovedSearchService for persistent browser-based web scraping using Playwright.
2025-11-30 15:03:03 +01:00
Michael
510904d899
feat: Add Amazon product scraping service using Playwright and Browserless for persistent browser connections.
2025-11-30 14:56:01 +01:00
Michael
66d6f14d13
Merge branch 'antigravity' of https://github.com/R0m1k3/Priceflow into antigravity
2025-11-30 14:46:22 +01:00
Michael
e9f6c6e0be
feat: Add BrowserlessService for persistent, auto-reconnecting browser operations with enhanced price and content extraction.
2025-11-30 14:46:21 +01:00
Michael
3813173e01
feat: add search_config module centralizing site definitions, proxy settings, user agents, and cookie banner selectors.
2025-11-30 14:14:31 +01:00
Michael
b8a56ca85d
Merge branch 'antigravity' of https://github.com/R0m1k3/Priceflow into antigravity
2025-11-30 13:25:15 +01:00
Michael
c6b9befc9a
feat: introduce BrowserlessService for persistent, auto-reconnecting browser operations and enhanced price extraction
2025-11-30 13:25:13 +01:00
Michael
c93e55519e
Merge main: Resolved conflicts by keeping enhanced browserless service with auto-reconnection + price extraction
2025-11-30 12:49:40 +01:00
Michael
2710388faf
feat: Implement scheduled item checks with AI price extraction, database updates, and price change notifications.
2025-11-30 12:48:30 +01:00
Michael
93d73dbed8
feat: Implement BrowserlessService for robust browser automation with auto-reconnection, popup handling, and price extraction capabilities.
2025-11-30 12:43:22 +01:00
Michael
23afe4bc6c
feat: Add new services for search orchestration, browser automation, and scheduling.
2025-11-30 12:39:26 +01:00
Michael
606fbf1523
feat: Implement Catalogues page for displaying, filtering, searching, and managing promotional catalogues.
2025-11-30 11:32:28 +01:00
Michael
99650a992b
feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce new Catalogues frontend page.
2025-11-30 11:17:28 +01:00
Michael
03d79e45fb
feat: Introduce Cataloguemate.fr scraper service and associated catalogue administration UI.
2025-11-30 10:49:47 +01:00
Michael
d772e85ae4
feat: Add catalogue and enseigne API endpoints with scraping and scheduling services.
2025-11-30 10:39:16 +01:00
Michael
df8fbded92
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for robust scraping.
2025-11-30 02:47:11 +01:00
Michael
cde1f75d3a
feat: Add Cataloguemate.fr scraper service to replace Tiendeo, utilizing BrowserlessService for robust scraping.
2025-11-30 02:45:24 +01:00
Michael
101e5f0656
feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for catalog and page scraping
2025-11-30 02:42:53 +01:00
Michael
efb161a5ef
feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping.
2025-11-30 02:35:27 +01:00
Michael
11d297764d
feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics.
2025-11-30 02:33:15 +01:00
Michael
ec97822087
feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs.
2025-11-30 02:29:58 +01:00
Michael
1176eec95a
feat: add Cataloguemate.fr scraper service for promotional catalogs.
2025-11-30 02:13:42 +01:00
Michael
69cfa4cfa9
feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure.
2025-11-30 02:09:24 +01:00
Michael
c6dd114584
feat: Add Cataloguemate.fr scraper service to replace Tiendeo for promotional catalog data.
2025-11-30 02:05:40 +01:00