Commit Graph
100 Commits
Author SHA1 Message Date
Michael 4e9f2662bc feat: add Playwright-based scraping service for tracking items. 2025-12-18 11:11:30 +01:00
Michael 5428459f36 feat: Add web scraping service for tracking and AI extraction verification script. 2025-12-18 11:05:09 +01:00
Michael 6fab64b229 feat: Add ScraperService with Playwright for robust web scraping, including browser management, configurable scraping options, and automated popup handling. 2025-12-18 10:56:06 +01:00
Michael 0d68859a5b feat: add Playwright script to scrape gifi.fr product page and extract price information. 2025-12-18 10:42:31 +01:00
Michael 97664a4a16 fix: Improve ItemService screenshot management by prioritizing latest timestamped files and ensuring complete deletion. 2025-12-18 08:50:45 +01:00
Michael 4590273801 docs: revise task.md to address screenshot update caching instead of popup obscuration. 2025-12-18 08:34:02 +01:00
Michael ed7f23d657 chore: Update task plan to focus on removing obscuring popups from product screenshots. 2025-12-18 08:14:56 +01:00
Michael e930f4bbfa feat: Implement main FastAPI application entry point, including database migrations, data seeding, and background task scheduling. 2025-12-02 18:52:41 +01:00
Michael 947152e132 feat: Implement scheduler_service for automated item price and stock tracking, including AI analysis and database updates. 2025-12-02 18:03:32 +01:00
Michael fb89af5aba Resolve merge conflict in scheduler_service.py 2025-12-02 17:52:48 +01:00
Michael 2c0a573a66 feat: Add web scraping service using Playwright and Browserless, and a new scheduler service. 2025-12-02 17:15:28 +01:00
Michael d8b3014021 Resolve merge conflict in browserless_service.py: Keep navigation logic and fix screenshot path 2025-12-02 17:04:49 +01:00
Michael 241970cf70 feat: implement BrowserlessService for persistent browser management with auto-reconnection and enhanced price extraction. 2025-12-02 16:54:40 +01:00
Michael 8a8ef02168 feat: Introduce BrowserlessService for persistent web scraping with enhanced Amazon price extraction, alongside initial scheduler service and task tracking. 2025-12-02 13:59:39 +01:00
Michael 3bbab058c6 feat: implement persistent browserless service for robust web scraping and price extraction with auto-reconnection. 2025-12-02 11:46:32 +01:00
Michael e7875f35a7 feat: Add French localization file with comprehensive translations for the application UI. 2025-12-02 09:51:56 +01:00
Michael 86aae8f628 Merge branch 'origin/antigravity' into antigravity
Resolved conflict in browserless_service.py by merging Amazon interstitial selectors.
Combined local robust selectors with remote additions for maximum coverage.
2025-12-02 09:37:31 +01:00
Michael 2b21cd176c feat: add TrackingScraperService for Playwright-based web scraping with popup handling and smart scrolling capabilities. 2025-12-02 09:28:08 +01:00
Michael 5516562cb6 fix: Handle Amazon interstitials and dynamic screenshot updates
- Added Amazon 'Continue' interstitial selectors to BrowserlessService
- Ported robust popup handling from TrackingScraperService to BrowserlessService
- Updated ItemService to fetch dynamic screenshot URLs from PriceHistory
  (fixes issue where forced updates didn't show new screenshots)
2025-12-02 09:27:12 +01:00
Michael 9111aa02b9 feat: Add Playwright script to detect and analyze Amazon "Continuer les achats" popup elements. 2025-12-01 17:26:47 +01:00
Michael 9178512465 feat: implement BrowserlessService for unified, persistent Playwright browser management with auto-reconnection and enhanced price extraction. 2025-12-01 17:14:12 +01:00
Michael 6eb821c40d feat: Add automated item checking scheduler service and new dashboard item display components. 2025-12-01 17:07:50 +01:00
Michael f0e1c6668d feat: Implement new search service for multi-site e-commerce product search and item detail scraping. 2025-12-01 16:56:32 +01:00
Michael f223dc5a29 feat: Increase Amazon limit to 50 and fix screenshot framing
- Increased default Amazon search results limit from 20 to 50 in backend and frontend
- Improved TrackingScraperService to target main product area on Amazon (avoiding reviews)
- Added specific selectors for Amazon product page (#imgTagWrapperId, #productTitle, etc.)
2025-12-01 16:18:35 +01:00
Michael acfc5c6397 feat: Enhance Amazon stealth with comprehensive anti-detection
Added comprehensive stealth techniques to bypass Amazon bot detection:
- Complete HTTP headers (Accept, Accept-Language, Sec-Fetch-*, etc.)
- Extended chrome object with loadTimes, csi, app
- Permissions API override
- Plugins, languages, platform spoofing
- Battery API mocking

This significantly improves bot detection evasion compared to the basic
2-line stealth that was insufficient for Amazon's detection systems.
2025-12-01 13:55:57 +01:00
Michael 2429fa2018 Revert "fix: Remove homepage redirect to avoid Amazon bot detection"
This reverts commit 5d3d181f5e.
2025-12-01 13:26:05 +01:00
Michael 5d3d181f5e fix: Remove homepage redirect to avoid Amazon bot detection
Amazon detects the homepage-then-search pattern as bot behavior and blocks
with 2065 byte pages. Now navigating directly to search URL to appear more
natural.

This reverts the homepage loading strategy from commit 85b19e7 as Amazon's
bot detection has evolved and now flags this pattern.
2025-12-01 13:25:51 +01:00
Michael 95846ec606 fix: Disable ImprovedSearchService auto-init to restore Amazon functionality
ImprovedSearchService and AmazonScraperService were both connecting to
Browserless at startup, causing connection conflicts and Amazon blocking
(2065 byte pages).

Now only AmazonScraperService initializes at startup, keeping Browserless
in continuous connection for Amazon. ImprovedSearchService will initialize
on-demand when needed.

This restores Amazon search to working state as it was at commit 85b19e7.
2025-12-01 13:21:32 +01:00
Michael 8dd3082a9a Merge branch 'antigravity' into main - Resolve conflicts
Resolved conflicts by keeping antigravity improvements:
- tracking_scraper_service.py: 47 popup selectors with RGPD/Amazon support
- main.py: No auto-init of TrackingScraperService (on-demand only)

This brings all popup handling improvements to main branch.
2025-12-01 13:13:58 +01:00
Michael 53a05f45e4 fix: Remove auto-initialization of TrackingScraperService to fix Amazon scraper
TrackingScraperService now initializes on-demand instead of at startup.
This prevents connection conflicts with AmazonScraperService and ImprovedSearchService
since all services connect to the same Browserless instance.

Fixes: Amazon scraper blocking issue (page too small - 2065 bytes)
2025-12-01 13:05:49 +01:00
Michael e2a874bf6c feat: implement TrackingScraperService for robust web scraping with Playwright, including shared browser management and popup handling. 2025-12-01 12:42:14 +01:00
Michael a792060f08 feat: Add TrackingScraperService with enhanced popup handling for screenshots
- Add 47 popup selectors (vs 12 generic before)
- Support French RGPD platforms (Axeptio, Didomi, OneTrust, TarteAuCitron)
- Amazon-specific detection and handling (6 selectors)
- Retry logic with verification pass
- Double popup cleanup before screenshot
- Increased timeouts for slow animations
- New _verify_no_popups() helper function

This resolves popup visibility issues in the Suivi/Dashboard screenshots.
2025-12-01 12:27:51 +01:00
Michael 22e15515f6 Merge branch 'main' of https://github.com/R0m1k3/Priceflow 2025-12-01 11:32:14 +01:00
Michael f27a75d9d6 feat: Implement initial application structure with routing, authentication flow, and responsive layout. 2025-12-01 11:32:12 +01:00
Michael 7592c2781e feat: Implement the core FastAPI application, including database migrations, service orchestration, and a new tracking scraper service. 2025-12-01 09:32:37 +01:00
Michael ffc2f0d632 feat: introduce centralized search configuration for scraping, including site selectors, proxies, and user agents. 2025-12-01 07:40:50 +01:00
Michael 36ae466583 feat: Introduce ImprovedSearchService utilizing Playwright for persistent browser-based e-commerce scraping and price extraction, along with new search configuration. 2025-12-01 01:30:06 +01:00
Michael b35401a87b feat: Add improved search and AI price extraction services, update Docker Compose, and introduce benchmark verification script. 2025-12-01 01:20:39 +01:00
Michael 8303e767f7 feat: implement a new persistent browser-based search service for e-commerce scraping. 2025-12-01 01:11:42 +01:00
Michael 8bad965e29 feat: Add ImprovedSearchService for persistent browser-based e-commerce search scraping. 2025-12-01 01:07:22 +01:00
Michael 759a28b2d4 feat: Add ImprovedSearchService for persistent browser-based e-commerce search and scraping. 2025-12-01 01:05:40 +01:00
Michael 4dbb9f6424 feat: add improved search service with persistent browser connection 2025-12-01 01:03:51 +01:00
Michael 93840804a4 feat: implement product search page with site selection, real-time results, and monitoring integration 2025-12-01 01:01:53 +01:00
Michael 39893e8c1e feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-12-01 01:00:25 +01:00
Michael b5c5961618 feat: Implement AI-powered price extraction using Gemma 2 and an improved search service with persistent browser connections for web scraping. 2025-12-01 00:57:05 +01:00
Michael c159036583 feat: Implement comprehensive search configurations for discount stores and add Gifi-specific search and analysis tools. 2025-12-01 00:36:54 +01:00
Michael 63073d03b5 feat: introduce new search configuration, improved search service, and Carrefour search integration test. 2025-11-30 23:54:48 +01:00
Michael bf97cd22e7 feat: introduce improved search service with persistent browser and modular parsers 2025-11-30 23:49:34 +01:00
Michael 7cb0e2f94b feat: add improved search service with persistent browser connection and modular site-specific parsers 2025-11-30 23:45:51 +01:00
Michael 8bf9ddb2d7 feat: implement improved search service using persistent browser connection and modular parsers, alongside a BeautifulSoup check utility. 2025-11-30 23:43:11 +01:00
Michael 87568866e7 feat: introduce improved search service with persistent browser connection and modular site-specific parsers. 2025-11-30 23:40:22 +01:00
Michael 4dd77bd889 feat: add improved search service with persistent browser, modular parsers, and popup handling. 2025-11-30 23:35:40 +01:00
Michael 9855040f3e feat: Implement improved search service with supporting models, schemas, and API routers, and refactor site verification script. 2025-11-30 23:33:05 +01:00
Michael 4720ed6b93 feat: Implement improved search service with persistent browser and modular site-specific parsers, including Gifi. 2025-11-30 22:21:08 +01:00
Michael 64067347d1 feat: implement new search service with multiple site parsers, search configuration, and supporting inspection/verification scripts. 2025-11-30 22:04:14 +01:00
Michael 95e38287cb feat: Add new parsers for Auchan, Carrefour, L'Incroyable, and Stokomani, including search configuration and deployment updates. 2025-11-30 22:03:40 +01:00
Michael b44082aeb4 feat: Implement new product parsers for Stokomani, Auchan, Carrefour, and Gifi, and add a centralized search configuration module. 2025-11-30 20:26:19 +01:00
Michael e0d73f4d0e feat: add LaFoirFouille parser to extract product search results from lafoirfouille.fr 2025-11-30 19:09:09 +01:00
Michael fbd06cfb17 feat: Add initial parser for Auchan.fr to extract product search results. 2025-11-30 19:05:38 +01:00
Michael 078060f155 feat: Add abstract base parser with product result dataclass and implement Stokomani parser for search results. 2025-11-30 19:03:26 +01:00
Michael ebe685fdf9 feat: add Stokomani parser to extract product search results from stokomani.fr 2025-11-30 18:17:07 +01:00
Michael b3d5c56707 feat: Add Stokomani search results parser 2025-11-30 18:13:11 +01:00
Michael c2f042beee feat: Implement improved search service with centralized configuration and site-specific parsers for various retailers. 2025-11-30 18:05:34 +01:00
Michael a8bffaec08 feat: Implement initial e-commerce site parsers with a base class and factory for extracting product search results. 2025-11-30 17:42:32 +01:00
Michael ee13a0bccd fix: resolve merge conflict in improved_search_service - unified domain matching logic 2025-11-30 17:37:16 +01:00
Michael 3db7db48fa feat: Add improved search service with persistent browser connection for e-commerce scraping. 2025-11-30 17:32:58 +01:00
Michael 33d9042b10 feat: add search_config module for browserless, proxy, user agent, and site-specific search configurations, and create list_missing_sites script. 2025-11-30 16:57:44 +01:00
Michael 7ba8850838 feat: Implement new search service with comprehensive configuration for various e-commerce sites, proxies, and user agents. 2025-11-30 16:36:51 +01:00
Michael f20bae52bd feat: introduce core search configuration for sites, proxies, and user agents, alongside selector diagnosis and HTML dumping utilities. 2025-11-30 16:27:02 +01:00
Michael 79dc4bfafc feat: Implement centralized search configuration with site definitions, proxy settings, and user agents, and add a debugging script. 2025-11-30 15:58:27 +01:00
Michael 7477d505e5 feat: Add ImprovedSearchService for persistent browser-based web scraping using Playwright. 2025-11-30 15:03:03 +01:00
Michael 510904d899 feat: Add Amazon product scraping service using Playwright and Browserless for persistent browser connections. 2025-11-30 14:56:01 +01:00
Michael 66d6f14d13 Merge branch 'antigravity' of https://github.com/R0m1k3/Priceflow into antigravity 2025-11-30 14:46:22 +01:00
Michael e9f6c6e0be feat: Add BrowserlessService for persistent, auto-reconnecting browser operations with enhanced price and content extraction. 2025-11-30 14:46:21 +01:00
Michael 3813173e01 feat: add search_config module centralizing site definitions, proxy settings, user agents, and cookie banner selectors. 2025-11-30 14:14:31 +01:00
Michael b8a56ca85d Merge branch 'antigravity' of https://github.com/R0m1k3/Priceflow into antigravity 2025-11-30 13:25:15 +01:00
Michael c6b9befc9a feat: introduce BrowserlessService for persistent, auto-reconnecting browser operations and enhanced price extraction 2025-11-30 13:25:13 +01:00
Michael c93e55519e Merge main: Resolved conflicts by keeping enhanced browserless service with auto-reconnection + price extraction 2025-11-30 12:49:40 +01:00
Michael 2710388faf feat: Implement scheduled item checks with AI price extraction, database updates, and price change notifications. 2025-11-30 12:48:30 +01:00
Michael 93d73dbed8 feat: Implement BrowserlessService for robust browser automation with auto-reconnection, popup handling, and price extraction capabilities. 2025-11-30 12:43:22 +01:00
Michael 23afe4bc6c feat: Add new services for search orchestration, browser automation, and scheduling. 2025-11-30 12:39:26 +01:00
Michael 606fbf1523 feat: Implement Catalogues page for displaying, filtering, searching, and managing promotional catalogues. 2025-11-30 11:32:28 +01:00
Michael 99650a992b feat: Add Cataloguemate.fr scraper to replace Tiendeo and introduce new Catalogues frontend page. 2025-11-30 11:17:28 +01:00
Michael 03d79e45fb feat: Introduce Cataloguemate.fr scraper service and associated catalogue administration UI. 2025-11-30 10:49:47 +01:00
Michael d772e85ae4 feat: Add catalogue and enseigne API endpoints with scraping and scheduling services. 2025-11-30 10:39:16 +01:00
Michael df8fbded92 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for robust scraping. 2025-11-30 02:47:11 +01:00
Michael cde1f75d3a feat: Add Cataloguemate.fr scraper service to replace Tiendeo, utilizing BrowserlessService for robust scraping. 2025-11-30 02:45:24 +01:00
Michael 101e5f0656 feat: add Cataloguemate.fr scraper to replace Tiendeo and utilize BrowserlessService for catalog and page scraping 2025-11-30 02:42:53 +01:00
Michael efb161a5ef feat: Add Cataloguemate.fr scraper, replacing Tiendeo and utilizing BrowserlessService for catalog and page scraping. 2025-11-30 02:35:27 +01:00
Michael 11d297764d feat: Add API endpoints for managing and querying catalogues, enseignes, and scraping statistics. 2025-11-30 02:33:15 +01:00
Michael ec97822087 feat: Implement Cataloguemate.fr scraper using BrowserlessService to replace Tiendeo and fetch promotional catalogs. 2025-11-30 02:29:58 +01:00
Michael 1176eec95a feat: add Cataloguemate.fr scraper service for promotional catalogs. 2025-11-30 02:13:42 +01:00
Michael 69cfa4cfa9 feat: add Cataloguemate.fr scraper to replace Tiendeo for improved reliability and simpler structure. 2025-11-30 02:09:24 +01:00
Michael c6dd114584 feat: Add Cataloguemate.fr scraper service to replace Tiendeo for promotional catalog data. 2025-11-30 02:05:40 +01:00
Michael be3de83601 feat: add script to debug catalog list and page content extraction for cataloguemate.fr 2025-11-30 02:00:59 +01:00
Michael af096c7980 feat: Implement Cataloguemate.fr scraper service, add catalogue router and scheduler, and remove Tiendeo scraper verification. 2025-11-30 01:56:56 +01:00
Michael 544825e354 feat: add Tiendeo.fr catalog scraper service using Crawl4AI for promotional catalogs. 2025-11-30 01:33:46 +01:00
Michael 607f47c0f0 feat: add Tiendeo.fr scraper service using Crawl4AI for promotional catalog extraction and parsing. 2025-11-30 01:29:00 +01:00
Michael 2b974aacf9 feat: Implement Tiendeo.fr scraper using Crawl4AI to extract promotional catalogs and their pages. 2025-11-30 01:17:40 +01:00
Michael 2692f76cf0 feat: add Tiendeo.fr scraper service for promotional catalogs using Crawl4AI 2025-11-30 01:12:25 +01:00