Web enricher rearchitect harden
We want to have a robust, performance optimized and reliable system with low cost and high success rate.
The purpose of the system is to be able to scrape data for SERP, Google AI mode and websites at scale. Both cheaply with standard queue and prioritized queue in dataforseo. Standard as default (batched at scale), priority when we do ad hoc requests for real time data fetch.
We save the data from any step in the process and for any type of data fetch (batch, ad hoc). The purpose is to extract as much valuable information as possible and to be able to revist and extract it later, if we missed something. Eg.
For the company data enrichment task; we run many SERP - the queries and all the dataforseo data should be saved. The SERP contains meta data for rank to identify a best match website in our website identification subtask. The SERP meta data in itself also contains information which can answer questions. The website raw data is also saved to later reextract information we might have missed. And then we save a sanitized / structured data extract of the data we have currently identied to be valuable. Eg. employee names, phone, email etc. or company address, official phone, email, taget market, value proposition, line of business, customer segments, commercial posture etc.
DataForSEO should use Denmark geo target.
The model to analyze the data / content being collected should by default be glm 5.3 flash, but easy to configure. This goes for rerank, data extract etc.
My impression is that the success rate of the current solution is not that good (website scrape). It's quite ineffective, not streamlined at all.
We should try to scrape website simply first and use advanced tricks immidiately after if that fails. Tiered Scraping Engine │ │ ├─ Tier 1: Fast HTTP │ (90% of sites: httpx/Trafilatura — zero browser overhead) │ └─ Tier 2: Crawl4AI │ (10% JS-heavy sites: Playwright worker pool)
The query formulation / creation is one step before the reuseable workflow pipeline described. Query formulation is task and situation dependent. Currently I see these tasks and scenarios (these are not conclusive): - Task: Find company contact information - Task: Find company business profile information - Task: Find employee contact information - Task: Find company competitors - Task: Find company customers
- Scenario: Batched 100-1000s on a schedule
- Scenario: adhoc request from Agent or in the frontends chat session