mirror of
https://github.com/R0m1k3/Priceflow.git
synced 2026-10-12 01:39:25 +02:00
1.4 KiB
1.4 KiB
Implementation Plan - Catalog Scraper Fix
Problem
The catalog scraper (cataloguemate_scraper.py) fails to retrieve catalogs because the browserless service is likely being blocked or failing to connect (based on logs). However, the target site (cataloguemate.fr) is accessible via simple HTTP requests.
Proposed Changes
1. Modify app/services/cataloguemate_scraper.py
We will implement a fallback mechanism in scrape_catalog_list:
- Primary Method: Continue using
browserless_service.get_page_content. - Fallback Method: If
browserlessreturns no content or fails, usehttpx.AsyncClientto fetch the HTML directly. - Parsing: Use
BeautifulSoupto parse the HTML from either source.
[MODIFY] cataloguemate_scraper.py
- Import
httpxandrandom(for user-agents). - In
scrape_catalog_list, add atry/exceptblock or check for empty content. - If content is missing, call a new helper or inline
httpxrequest with standard headers.
Verification Plan
Automated Verification
Since we cannot run the full app locally (docker/dependencies issues), we will verify by code review and logic.
Manual Verification (User)
- User triggers the catalog update (via Admin or Scheduler).
- User checks logs to see if "Fallback to HTTP" message appears.
- User verifies that catalogs (e.g. B&M, Gifi) appear in the dashboard.