Skip to content

Web enrichment produces versioned company signals and discovered relationships

Status: accepted Date: 2026-08-14 Deciders: Chris (solo founder) Depends on: ADR-0011, ADR-0012, ADR-0013, ADR-0006, and ADR-0008

Context

The current web pipeline resolves company-owned web presence and extracts company contact points. It has good acquisition evidence, budgets, raw retention, and current read views. It does not yet publish the company profile or discovered relationships that the product needs. Its candidate ledger is an execution record, not a product signal model.

Registry contact points and web contacts also need different meanings. A CVR email or phone is a registry fact. A value read from a website is an observed signal. Consumers must see both sources without treating one as confirmation of the other.

Decision

  1. Web enrichment stays collection-owned and shortlist-scoped. It acquires evidence and publishes data. Analysis uses that data to compute scores and recommendations.
  2. The data path is raw observation, versioned signal, then current read model. New extractor versions append new signal observations. They do not overwrite the history needed for comparison and audit.
  3. The existing web-presence and company-contact signals remain. They are included in current_detail, with their source URL, collection time, extractor version, confidence, and evidence.
  4. Add company-profile signals. The first profile contract covers line of business, products or services, value proposition, target market, B2B/B2C posture, pricing evidence, customer segments, geography, and other explicit product facts. A missing signal means not observed or not extracted, not false.
  5. Add entity mentions and discovered relationships. Named customers, suppliers, partners, competitors, and people are first stored as evidenced mentions. A resolved mention publishes an observed relationship through ADR-0013. An unresolved mention stays a mention.
  6. Evidence is mandatory. Each interpreted signal and relationship carries the source URL, retrieval time, raw artifact reference, extractor and version, confidence, and a bounded evidence excerpt or locator.
  7. Registry and web values stay source-qualified. Current detail can show both. It does not silently choose one email, phone, address, or website as universal truth.
  8. Operational ledgers stay private. Provider task IDs, rate-limit state, budgets, caches, candidate decisions, and workflow traces remain internal. They support replay and audit but are not product read contracts.

Options considered

Option Benefit Cost Decision
Stop at website and contact discovery Already implemented No product profile or relationship data Rejected
Publish extractor output only as one current JSON object Simple read Loses version history, field evidence, and comparison Rejected
Publish versioned, evidenced signals and resolve mentions into graph edges Auditable detail; reusable by analysis; supports relationship history More explicit schemas and resolution work Chosen

Consequences

  • The enrichment schema needs profile-signal and entity-mention storage.
  • Current views select the best supported signal. History views expose all extractor versions and observations.
  • Re-extraction can run from retained raw pages without a new scrape.
  • A discovered edge cannot exist without a resolved target and cited evidence.
  • Personal-data minimization and the legal preconditions in ADR-0006 and ADR-0008 apply before a people mention is retained or resolved.

Open questions

  • Define the first profile field set and extraction schemas during the build. The categories above are fixed; exact value vocabularies can evolve with measured corpus evidence.