Commit Graph
100 Commits
Author SHA1 Message Date
Michael fa4d454063 feat: Add Amazon scraper service with persistent browser, anti-detection, and Pydantic schemas for product data. 2025-12-24 23:10:20 +01:00
Michael f8e4286bdf feat: add Amazon scraper service with persistent browser, stealth, and parsing capabilities for Amazon product data. 2025-12-24 23:01:10 +01:00
Michael 3a7dcd16ed feat: implement Amazon scraper service using Playwright with persistent browser connection and anti-detection techniques. 2025-12-24 22:49:02 +01:00
Michael cd18b47bc0 feat: Add core search configurations for various e-commerce sites, including proxy and user agent management, and introduce initial Amazon scraper service and proxy utilities. 2025-12-24 21:13:22 +01:00
Michael 43de609773 feat: add Amazon scraper service using Playwright with persistent browser and anti-detection features 2025-12-24 14:38:16 +01:00
Michael 0ab7ac5a38 feat: add Amazon scraper service using Playwright, Browserless, and stealth for persistent sessions. 2025-12-24 12:04:58 +01:00
Michael fd34bc66b2 feat: implement Amazon scraper service using Playwright with stealth features for product search and parsing. 2025-12-24 09:19:28 +01:00
Michael 38340cd59c feat: implement web scraping service using Playwright, including popup handling, smart scrolling, and specialized logic for Amazon and B&M stores. 2025-12-24 09:12:12 +01:00
Michael 931c5d42a0 chore: add script to reproduce Amazon tracking issues using a local browser. 2025-12-24 08:40:54 +01:00
Michael a0da02ab8d amazon tracking 2025-12-23 19:56:08 +01:00
Michael fcb23d090d feat: add scheduler service for automated item price and stock checks with scraping, AI extraction, and specific site parsers. 2025-12-23 10:47:15 +01:00
Michael 48f149d38f feat: implement improved search service with persistent browser connection and update search verification script. 2025-12-22 16:05:07 +01:00
Michael 39d4107e6d test: Update search verification query and target site. 2025-12-22 15:42:33 +01:00
Michael 0ff229de13 feat: add price verification script and update task.md to investigate comparator price extraction issues. 2025-12-22 15:24:09 +01:00
Michael 954672cfc0 feat: Add catalogue management API with enseigne, catalogue, and scraping endpoints, and integrate Cataloguemate scraper. 2025-12-22 14:46:50 +01:00
Michael 81b7cce947 feat: implement Cataloguemate.fr scraper using BrowserlessService to replace the Tiendeo scraper. 2025-12-22 14:12:16 +01:00
Michael 7842eead4a feat: add debug_catalog_images.py to analyze catalog image selectors and update task.md to reflect the new focus on catalog image fixes. 2025-12-22 13:49:27 +01:00
Michael d2fe4bf22b feat: Add Cataloguemate.fr scraper using BrowserlessService to replace the Tiendeo scraper. 2025-12-22 12:39:01 +01:00
Michael d528818c29 feat: add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for robust scraping. 2025-12-22 12:27:34 +01:00
Michael 1376b3f44f chore: remove obsolete debug and verification scripts and streamline walkthrough documentation. 2025-12-22 12:22:07 +01:00
Michael 4bf950c5eb feat: Implement HTTPX fallback in Cataloguemate scraper for enhanced reliability and add associated implementation plan and debug script. 2025-12-22 12:16:33 +01:00
Michael 17b3009b1d feat: Implement scheduler service for automated item price and stock tracking with hybrid extraction and bot detection. 2025-12-22 11:18:39 +01:00
Michael 2fae5350b3 feat: define canonical AI extraction schema with vision-first prompt and add vision priority verification script 2025-12-22 10:38:08 +01:00
Michael 83fa8ed734 feat: Add AI extraction schemas and prompt templates, and refine price extraction logic to prioritize TTC over HT. 2025-12-18 16:57:23 +01:00
Michael 384baa3575 feat: Introduce AI extraction schema with validation, prompt generation, and a verification script to improve price extraction reliability. 2025-12-18 16:32:32 +01:00
Michael ef7cd7c1cf feat: Add BMStores product parser and a new Playwright-based scraping service with popup handling. 2025-12-18 15:20:19 +01:00
Michael 51ede950ff feat: Add scheduled item tracking service with dedicated web scraping and data extraction capabilities. 2025-12-18 13:23:53 +01:00
Michael d50b18690e feat: Add scheduler and scraper services for automated item tracking and price extraction. 2025-12-18 12:38:32 +01:00
Michael 4e9f2662bc feat: add Playwright-based scraping service for tracking items. 2025-12-18 11:11:30 +01:00
Michael 5428459f36 feat: Add web scraping service for tracking and AI extraction verification script. 2025-12-18 11:05:09 +01:00
Michael 6fab64b229 feat: Add ScraperService with Playwright for robust web scraping, including browser management, configurable scraping options, and automated popup handling. 2025-12-18 10:56:06 +01:00
Michael 0d68859a5b feat: add Playwright script to scrape gifi.fr product page and extract price information. 2025-12-18 10:42:31 +01:00
Michael 97664a4a16 fix: Improve ItemService screenshot management by prioritizing latest timestamped files and ensuring complete deletion. 2025-12-18 08:50:45 +01:00
Michael 4590273801 docs: revise task.md to address screenshot update caching instead of popup obscuration. 2025-12-18 08:34:02 +01:00
Michael ed7f23d657 chore: Update task plan to focus on removing obscuring popups from product screenshots. 2025-12-18 08:14:56 +01:00
Michael e930f4bbfa feat: Implement main FastAPI application entry point, including database migrations, data seeding, and background task scheduling. 2025-12-02 18:52:41 +01:00
Michael 947152e132 feat: Implement scheduler_service for automated item price and stock tracking, including AI analysis and database updates. 2025-12-02 18:03:32 +01:00
Michael fb89af5aba Resolve merge conflict in scheduler_service.py 2025-12-02 17:52:48 +01:00
Michael 2c0a573a66 feat: Add web scraping service using Playwright and Browserless, and a new scheduler service. 2025-12-02 17:15:28 +01:00
Michael d8b3014021 Resolve merge conflict in browserless_service.py: Keep navigation logic and fix screenshot path 2025-12-02 17:04:49 +01:00
Michael 241970cf70 feat: implement BrowserlessService for persistent browser management with auto-reconnection and enhanced price extraction. 2025-12-02 16:54:40 +01:00
Michael 8a8ef02168 feat: Introduce BrowserlessService for persistent web scraping with enhanced Amazon price extraction, alongside initial scheduler service and task tracking. 2025-12-02 13:59:39 +01:00
Michael 3bbab058c6 feat: implement persistent browserless service for robust web scraping and price extraction with auto-reconnection. 2025-12-02 11:46:32 +01:00
Michael e7875f35a7 feat: Add French localization file with comprehensive translations for the application UI. 2025-12-02 09:51:56 +01:00
Michael 86aae8f628 Merge branch 'origin/antigravity' into antigravity
Resolved conflict in browserless_service.py by merging Amazon interstitial selectors.
Combined local robust selectors with remote additions for maximum coverage.
2025-12-02 09:37:31 +01:00
Michael 2b21cd176c feat: add TrackingScraperService for Playwright-based web scraping with popup handling and smart scrolling capabilities. 2025-12-02 09:28:08 +01:00
Michael 5516562cb6 fix: Handle Amazon interstitials and dynamic screenshot updates
- Added Amazon 'Continue' interstitial selectors to BrowserlessService
- Ported robust popup handling from TrackingScraperService to BrowserlessService
- Updated ItemService to fetch dynamic screenshot URLs from PriceHistory
  (fixes issue where forced updates didn't show new screenshots)
2025-12-02 09:27:12 +01:00
Michael 9111aa02b9 feat: Add Playwright script to detect and analyze Amazon "Continuer les achats" popup elements. 2025-12-01 17:26:47 +01:00
Michael 9178512465 feat: implement BrowserlessService for unified, persistent Playwright browser management with auto-reconnection and enhanced price extraction. 2025-12-01 17:14:12 +01:00
Michael 6eb821c40d feat: Add automated item checking scheduler service and new dashboard item display components. 2025-12-01 17:07:50 +01:00
Michael f0e1c6668d feat: Implement new search service for multi-site e-commerce product search and item detail scraping. 2025-12-01 16:56:32 +01:00
Michael f223dc5a29 feat: Increase Amazon limit to 50 and fix screenshot framing
- Increased default Amazon search results limit from 20 to 50 in backend and frontend
- Improved TrackingScraperService to target main product area on Amazon (avoiding reviews)
- Added specific selectors for Amazon product page (#imgTagWrapperId, #productTitle, etc.)
2025-12-01 16:18:35 +01:00
Michael acfc5c6397 feat: Enhance Amazon stealth with comprehensive anti-detection
Added comprehensive stealth techniques to bypass Amazon bot detection:
- Complete HTTP headers (Accept, Accept-Language, Sec-Fetch-*, etc.)
- Extended chrome object with loadTimes, csi, app
- Permissions API override
- Plugins, languages, platform spoofing
- Battery API mocking

This significantly improves bot detection evasion compared to the basic
2-line stealth that was insufficient for Amazon's detection systems.
2025-12-01 13:55:57 +01:00
Michael 2429fa2018 Revert "fix: Remove homepage redirect to avoid Amazon bot detection"
This reverts commit 5d3d181f5e.
2025-12-01 13:26:05 +01:00
Michael 5d3d181f5e fix: Remove homepage redirect to avoid Amazon bot detection
Amazon detects the homepage-then-search pattern as bot behavior and blocks
with 2065 byte pages. Now navigating directly to search URL to appear more
natural.

This reverts the homepage loading strategy from commit 85b19e7 as Amazon's
bot detection has evolved and now flags this pattern.
2025-12-01 13:25:51 +01:00
Michael 95846ec606 fix: Disable ImprovedSearchService auto-init to restore Amazon functionality
ImprovedSearchService and AmazonScraperService were both connecting to
Browserless at startup, causing connection conflicts and Amazon blocking
(2065 byte pages).

Now only AmazonScraperService initializes at startup, keeping Browserless
in continuous connection for Amazon. ImprovedSearchService will initialize
on-demand when needed.

This restores Amazon search to working state as it was at commit 85b19e7.
2025-12-01 13:21:32 +01:00
Michael 8dd3082a9a Merge branch 'antigravity' into main - Resolve conflicts
Resolved conflicts by keeping antigravity improvements:
- tracking_scraper_service.py: 47 popup selectors with RGPD/Amazon support
- main.py: No auto-init of TrackingScraperService (on-demand only)

This brings all popup handling improvements to main branch.
2025-12-01 13:13:58 +01:00
Michael 53a05f45e4 fix: Remove auto-initialization of TrackingScraperService to fix Amazon scraper
TrackingScraperService now initializes on-demand instead of at startup.
This prevents connection conflicts with AmazonScraperService and ImprovedSearchService
since all services connect to the same Browserless instance.

Fixes: Amazon scraper blocking issue (page too small - 2065 bytes)
2025-12-01 13:05:49 +01:00
Michael e2a874bf6c feat: implement TrackingScraperService for robust web scraping with Playwright, including shared browser management and popup handling. 2025-12-01 12:42:14 +01:00
Michael a792060f08 feat: Add TrackingScraperService with enhanced popup handling for screenshots
- Add 47 popup selectors (vs 12 generic before)
- Support French RGPD platforms (Axeptio, Didomi, OneTrust, TarteAuCitron)
- Amazon-specific detection and handling (6 selectors)
- Retry logic with verification pass
- Double popup cleanup before screenshot
- Increased timeouts for slow animations
- New _verify_no_popups() helper function

This resolves popup visibility issues in the Suivi/Dashboard screenshots.
2025-12-01 12:27:51 +01:00
Michael 22e15515f6 Merge branch 'main' of https://github.com/R0m1k3/Priceflow 2025-12-01 11:32:14 +01:00
Michael f27a75d9d6 feat: Implement initial application structure with routing, authentication flow, and responsive layout. 2025-12-01 11:32:12 +01:00
Michael 7592c2781e feat: Implement the core FastAPI application, including database migrations, service orchestration, and a new tracking scraper service. 2025-12-01 09:32:37 +01:00
Michael ffc2f0d632 feat: introduce centralized search configuration for scraping, including site selectors, proxies, and user agents. 2025-12-01 07:40:50 +01:00
Michael 36ae466583 feat: Introduce ImprovedSearchService utilizing Playwright for persistent browser-based e-commerce scraping and price extraction, along with new search configuration. 2025-12-01 01:30:06 +01:00
Michael b35401a87b feat: Add improved search and AI price extraction services, update Docker Compose, and introduce benchmark verification script. 2025-12-01 01:20:39 +01:00
Michael 8303e767f7 feat: implement a new persistent browser-based search service for e-commerce scraping. 2025-12-01 01:11:42 +01:00
Michael 8bad965e29 feat: Add ImprovedSearchService for persistent browser-based e-commerce search scraping. 2025-12-01 01:07:22 +01:00
Michael 759a28b2d4 feat: Add ImprovedSearchService for persistent browser-based e-commerce search and scraping. 2025-12-01 01:05:40 +01:00
Michael 4dbb9f6424 feat: add improved search service with persistent browser connection 2025-12-01 01:03:51 +01:00
Michael 93840804a4 feat: implement product search page with site selection, real-time results, and monitoring integration 2025-12-01 01:01:53 +01:00
Michael 39893e8c1e feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-12-01 01:00:25 +01:00
Michael b5c5961618 feat: Implement AI-powered price extraction using Gemma 2 and an improved search service with persistent browser connections for web scraping. 2025-12-01 00:57:05 +01:00
Michael c159036583 feat: Implement comprehensive search configurations for discount stores and add Gifi-specific search and analysis tools. 2025-12-01 00:36:54 +01:00
Michael 63073d03b5 feat: introduce new search configuration, improved search service, and Carrefour search integration test. 2025-11-30 23:54:48 +01:00
Michael bf97cd22e7 feat: introduce improved search service with persistent browser and modular parsers 2025-11-30 23:49:34 +01:00
Michael 7cb0e2f94b feat: add improved search service with persistent browser connection and modular site-specific parsers 2025-11-30 23:45:51 +01:00
Michael 8bf9ddb2d7 feat: implement improved search service using persistent browser connection and modular parsers, alongside a BeautifulSoup check utility. 2025-11-30 23:43:11 +01:00
Michael 87568866e7 feat: introduce improved search service with persistent browser connection and modular site-specific parsers. 2025-11-30 23:40:22 +01:00
Michael 4dd77bd889 feat: add improved search service with persistent browser, modular parsers, and popup handling. 2025-11-30 23:35:40 +01:00
Michael 9855040f3e feat: Implement improved search service with supporting models, schemas, and API routers, and refactor site verification script. 2025-11-30 23:33:05 +01:00
Michael 4720ed6b93 feat: Implement improved search service with persistent browser and modular site-specific parsers, including Gifi. 2025-11-30 22:21:08 +01:00
Michael 64067347d1 feat: implement new search service with multiple site parsers, search configuration, and supporting inspection/verification scripts. 2025-11-30 22:04:14 +01:00
Michael 95e38287cb feat: Add new parsers for Auchan, Carrefour, L'Incroyable, and Stokomani, including search configuration and deployment updates. 2025-11-30 22:03:40 +01:00
Michael b44082aeb4 feat: Implement new product parsers for Stokomani, Auchan, Carrefour, and Gifi, and add a centralized search configuration module. 2025-11-30 20:26:19 +01:00
Michael e0d73f4d0e feat: add LaFoirFouille parser to extract product search results from lafoirfouille.fr 2025-11-30 19:09:09 +01:00
Michael fbd06cfb17 feat: Add initial parser for Auchan.fr to extract product search results. 2025-11-30 19:05:38 +01:00
Michael 078060f155 feat: Add abstract base parser with product result dataclass and implement Stokomani parser for search results. 2025-11-30 19:03:26 +01:00
Michael ebe685fdf9 feat: add Stokomani parser to extract product search results from stokomani.fr 2025-11-30 18:17:07 +01:00
Michael b3d5c56707 feat: Add Stokomani search results parser 2025-11-30 18:13:11 +01:00
Michael c2f042beee feat: Implement improved search service with centralized configuration and site-specific parsers for various retailers. 2025-11-30 18:05:34 +01:00
Michael a8bffaec08 feat: Implement initial e-commerce site parsers with a base class and factory for extracting product search results. 2025-11-30 17:42:32 +01:00
Michael ee13a0bccd fix: resolve merge conflict in improved_search_service - unified domain matching logic 2025-11-30 17:37:16 +01:00
Michael 3db7db48fa feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-11-30 17:32:58 +01:00
Michael 33d9042b10 feat: add search_config module for browserless, proxy, user agent, and site-specific search configurations, and create list_missing_sites script. 2025-11-30 16:57:44 +01:00
Michael 7ba8850838 feat: Implement new search service with comprehensive configuration for various e-commerce sites, proxies, and user agents. 2025-11-30 16:36:51 +01:00
Michael f20bae52bd feat: introduce core search configuration for sites, proxies, and user agents, alongside selector diagnosis and HTML dumping utilities. 2025-11-30 16:27:02 +01:00
Michael 79dc4bfafc feat: Implement centralized search configuration with site definitions, proxy settings, and user agents, and add a debugging script. 2025-11-30 15:58:27 +01:00
Michael 7477d505e5 feat: Add ImprovedSearchService for persistent browser-based web scraping using Playwright. 2025-11-30 15:03:03 +01:00
Michael 510904d899 feat: Add Amazon product scraping service using Playwright and Browserless for persistent browser connections. 2025-11-30 14:56:01 +01:00