- Add User model with authentication
- Create JWT-based auth service
- Add login page with admin/admin default credentials
- Create admin panel with user management, settings, and site configuration
- Move settings to admin section (admin only)
- Change search page from multi-site checkboxes to single site selector
- Add logout functionality to layout
- Protect routes with authentication
- E.Leclerc: Fix URL from /cat/recherche~q= to /recherche?q=
(the old URL was showing promotions instead of search results)
- Auchan: Keep correct URL /recherche?text=
- Carrefour: Keep correct URL /s?q=
- Amazon FR/US: Simplify configuration, remove complex anti-bot
headers that were causing more issues than helping
- Simplify browserless search function (remove special Amazon handling)
- Add special handling for Amazon (FR and US):
- Custom User-Agent and HTTP headers mimicking real browser
- Don't block images (Amazon uses them for bot detection)
- Extra wait time and scroll for lazy-loaded content
- Timezone set to Europe/Paris
- Additional Sec-Fetch-* headers for authenticity
- Update Amazon product selectors with more robust patterns:
- a.a-link-normal for direct product links
- [data-asin]:not([data-asin='']) for filtering empty ASINs
- .s-main-slot targeting main search results
- Add &ref=nb_sb_noss parameter to avoid redirects
- Change default currency from USD to EUR for all French sites
- Update AI extraction prompt to use EUR by default
- Fix E.Leclerc search URL pattern
- Improve product selectors for Auchan, Carrefour, Amazon France
- Add Amazon France language parameter for better results
- Enhance cookie acceptance with site-specific selectors:
- Amazon (#sp-cc-accept)
- E.Leclerc (#onetrust-accept-btn-handler)
- Auchan (#popin_tc_privacy_button_2, #didomi-notice-agree-button)
- Carrefour (#onetrust-accept-btn-handler)
- Cdiscount (#footer_tc_privacy_button_2)
- And more CMP platforms (Didomi, OneTrust, Tarteaucitron, etc.)
- Improve cookie handling with iframe support and retry logic
- Remove ability to add/edit/delete search sites from settings UI
- Configure 16 French e-commerce sites with search URLs:
- Discount: Gifi, Stokomani, B&M, Centrakor, L'Incroyable, Action, La Foir'Fouille
- Grandes Surfaces: E.Leclerc, Auchan, Carrefour
- E-commerce: Amazon France, Cdiscount
- Électronique: Darty, Boulanger, Fnac
- Add automatic database seeding on application startup
- Add /api/search-sites/seed and /api/search-sites/reset endpoints
- Add cookie consent banner handling for sites that require it
- Simplify Settings page to only show site list with toggle for active state
- Add fallback to generic detection when configured selector fails
- Add many more CSS selectors for French e-commerce platforms (PrestaShop, Magento, WooCommerce)
- Add /produits/ pattern for bmstores.fr
- Improve generic container-based link detection
- Add navigation link filter to exclude non-product links
- Improve domain matching in direct_search_service to handle various formats
- Add /discover-all endpoint to discover search URLs for all sites missing one
- Add constants for magic numbers and fix linting issues
- Better domain cleaning in comparison functions
- Move imports to top-level in search_sites router
- Fix line too long issue with input selector
- Define constants for magic numbers (HTTP_OK, MIN_PRODUCTS_TO_MATCH)
- Use 'raise from' pattern for proper exception chaining
- Add search_url_discovery.py service that visits sites and finds search forms
- Automatically discover search_url when a new site is created without one
- Add POST /{site_id}/discover endpoint to force rediscovery
- Discovery runs in background to not block the API response
- Tries multiple methods: form analysis, common URL patterns
- Add search URLs for French discount stores: stokomani, gifi, bmstores,
centrakor, lafoirfouille, tedi, leclerc, carrefour, auchan
- Fix screenshot URL construction in search results
- Include screenshot_url in AI extraction results and error cases
- Use AIService.analyze_image() class method instead of module function
- Handle AIExtractionResponse tuple return type properly
- Add _clean_domain() to strip protocol/www from stored domains
- Fixes "https://https://" URL duplication issue
- Replace httpx with Playwright/Browserless for search pages
- Handles JavaScript-rendered content and bot detection
- Add wait selectors for each site to ensure content is loaded
- Block images and tracking scripts for faster loading
- Sequential site scraping to avoid overloading Browserless
- Remove duckduckgo-search dependency (unreliable results)
- Add direct_search_service.py that scrapes sites' search pages directly
- Add search_url and product_link_selector fields to SearchSite model
- Add beautifulsoup4 and lxml for HTML parsing
- Update schemas with new fields
- Add database migration for new columns
Direct search is more reliable as it queries each site's own search
functionality instead of relying on external search engines.
- Remove SearXNG container and configuration files
- Add duckduckgo-search library for web search
- Create duckduckgo_service.py with async search support
- Update search_service.py to use DuckDuckGo instead of SearXNG
- Fix BROWSERLESS_URL port to 3012
This eliminates the persistent limiter.toml directory issue and
provides a simpler, more reliable search solution.
- Mount ./searxng as /etc/searxng:rw instead of individual files
- Remove custom entrypoint.sh (no longer needed)
- Simplify settings.yml configuration
- Remove entrypoint script approach
- Mount limiter.toml directly as a file
- User must DELETE the SearXNG container in Portainer before redeploying
to clear the cached directory state