Skip to content

Company pass reference

How one company run finds and verifies the website, what it fetches, what it reads, and what it publishes (#842, ADR-0029). The architecture gives the overview; the operations page tells how to run it.

Leads

A lead is a place to look, not evidence. Leads are tried in this order:

  1. The website the company gave the registry, as given (scheme, host, path).
  2. The registry email domain, unless it is a mail provider (gmail.com, mail.tele.dk, teliamail.dk, and the others in FREEMAIL_DOMAINS).
  3. The site the AI Mode answer names in a sentence about the website (hjemmeside, webside, website, officiel).
  4. Up to three organic results. A result moves ahead when its domain spells the company name or its snippet carries the registry phone.
  5. The first other site the AI Mode answer cites.

Never a lead: a directory, social network, marketplace, news, or government domain (DIRECTORY_DOMAINS, vores-*.dk, *avis*.dk, erhvervsliv*.dk); a result whose title names a CVR; a URL whose path carries the CVR or a profile word (firmaer, virksomhed, erhvervsbasen, company, personer, medlemmer, medlemsvirksomheder, find-en-forhandler).

A registry website is visited as given. Every other lead is visited at its site root only (for a free Wix site, <account>.wixsite.com/<site>). When the root does not answer, the lead fails: a profile page on a portal (boligsiden.dk/ejendomsmaegler/…) prints the company's CVR, but it is not the company's website.

The AI Mode question asks for the website, the business, its customers, and its competitors in one Danish sentence. The organic keyword is the registered name and the town, without quotes.

Fetching

Step When
HTTP GET through the proxy, redirects followed to any domain, normal browser headers Always first
The same host and path over plain HTTP A home page that failed in transport over HTTPS
Real headless Chromium through the proxy The server answered but blocked, returned too little text (under 80 characters), or showed a challenge; or it did not answer and plain HTTP did not help
The same without the proxy Only when the proxy itself failed

A supporting page that does not answer is not retried in other ways, and a site that stops answering gets no more supporting requests in the pass. Certificates are always verified; an invalid certificate never gives content. Failures are named not-found, blocked, certificate, or unreachable (a gateway error from the proxy or a CDN counts as unreachable).

Ownership

Registry facts found on the site by exact matching:

Signal Match
cvr The eight CVR digits, with optional separators
name The trading name (legal form removed, Danish letters folded) in the text or title, all distinctive name words in the title, or the domain label spelling the name
phone The eight registry phone digits, with optional separators
email The registry email address, or its domain equal to the site's domain
address The registry street and house number
brand The domain label is one distinctive name word or all of them (gea.com, windelev.dk, nykredit.dk); a domain that spells the name gives both name and brand

A page is parked when it says the domain is for sale or parked, or shows a server default page. A "coming soon" or "under construction" page is parked only when it shows no cvr, phone, or address: a company that rebuilds its site often keeps its street and phone on the placeholder. A page is a registry listing when it carries three or more registry-extract words (branchekode, virksomhedsform, p-nummer, reklamebeskyttet, and others).

Evidence Decision Reason
Parked refuse parked
Registry listing refuse directory
The exact CVR accept, confidence 0.98 cvr
A labelled CVR of another company, and not ours model judge-*
Declared website (the registry website, or a site on its registrable domain) with name, phone, email, or address accept, 0.95 registry-website
name and one of phone, email, address, not a holding company accept, 0.9 identity
Not declared, and no signal at all (a brand counts as a signal) refuse no-identity
Anything else model judge-*

The model sees the registry facts, how the site was found, the signals, and up to 5,000 characters of the site's pages. It answers own, related, other, or unclear. own accepts (confidence at most 0.9). A declared website is also accepted for related and unclear (confidence at most 0.6); only other refuses it. A related site with a brand signal is accepted too; a related site's people are not published. Social profiles are published only for a site the company owns (not for judge-related).

Pages

From the links of the home page and the identity pages (and the sitemap when the home page has fewer than twelve links), at most eight pages in total, by topic quota: contact 1, about 1, people 2, offering 2, customers 2. A topic matches words in the link's path or text (kontakt, om os, firmaprofil, hvem vi er, medarbejdere, ydelser, referencer, and others in TOPICS). A short link text that starts with "Om" (Om PFP) is the about page. Shallow paths come first. Files, login, cart, search, tag, and feed paths are skipped. Pages must stay on the verified site's registrable domain.

Reading

Deterministic. Company contacts from every page (role mailboxes, phone numbers that a tel:, a label, or a country code marks, and postal addresses; enrichment/contacts.py), social profiles from the pages' links, and technology signatures from the retained HTML (enrichment/technology.py, the home page first within 1,000,000 characters).

Website model call. The pages are sent with their topic, at most 60,000 characters in total; short pages give their spare share to long ones. The answer has four lists: profile fields, classifications (sales_areas, target_markets), company mentions, and people with work contacts. Each item names its page and quotes it. The quote is found in the page text, allowing only whitespace, letter case, and dash or quote-mark differences; its exact range becomes the evidence locator. A sales area must quote its geography. People then pass enrichment/people.py: privacy, attribution, and objections. Output tokens: up to 16,384.

AI Mode model call. The answer's markdown is read into company_profile, competitors, and customers signals with quotes, JSON pointers, and the first citation after the quote. Competitors and customers also become company mentions (ai-mode-1.0).

Publication

One transaction writes the bundle and its pages, presence, contacts, extractor records, profile signals, filter evaluations, company mentions, people and their contacts, search retrievals, and source signals. Mentions then resolve: a normalized name that equals exactly one active company's publishes an observed edge. People resolve to this company's current registry participants by first given name and surname. The profile embedding follows after the commit.

Bounds

Bound Value Where
Sites checked per pass 4 MAX_SITE_VISITS
Organic leads 3 MAX_SEARCH_LEADS
Pages per site 9 including home MAX_PAGES
HTTP request 12 s HTTP_SECONDS
Browser page 25 s + 10 s settle BROWSER_SECONDS
Response body 5 MB MAX_BODY_BYTES
Model input for the site 60,000 characters MAX_INPUT_CHARS
AI Mode answer read 20,000 characters MAX_ANSWER_CHARS
Pass time limit 600 s PASS_SECONDS
Attempts per run 3 MAX_ATTEMPTS
Stale claim 20 minutes STALE_CLAIM
Companies per request 500 (configurable) WEB_MAX_REQUEST_COMPANIES

Outcomes

Outcome Reason Meaning
resolved cvr, registry-website, identity, judge-own, judge-related, judge-unclear The verified website and how it was verified (judge-related only for a declared website or a brand domain; judge-unclear only for a declared website)
resolved re-extract Retained evidence was read again
unresolved no-website No checked lead belongs to the company; search signals can still exist
blocked site-unreachable The company has a website we could not read: the registry website or the site AI Mode calls the website did not answer, or the registry email domain's server refused us (403 or a challenge); and no other stated site was refused. A 404 at the root, or a mail domain without a web server, is no-website
failed search-failed Neither search answered
failed no-registry-row, no-retained-evidence Nothing to work on
failed attempts-exhausted, error:<type> The run failed three times

The run log (company_enrichment_run.log) holds steps (every search, fetch with its attempts, lead list, ownership decision with its signals, and reading result with refused-quote examples), counts of published rows, and costs.

Tests

Behavior Test
A registry website is read into every published surface tests/test_web_pipeline.py::test_a_registry_website_is_read_into_every_published_surface
A quote not on the page drops only that claim tests/test_web_pipeline.py::test_a_quote_that_is_not_on_the_page_drops_only_that_claim
A redirect to the company's new domain is its website tests/test_web_pipeline.py::test_a_redirect_to_the_companys_new_domain_is_its_website, tests/test_web_fetch.py::test_a_redirect_to_the_companys_new_domain_is_followed
Search finds the site; a directory is never fetched tests/test_web_pipeline.py::test_search_finds_the_site_when_the_registry_names_none
No website still gives AI Mode signals tests/test_web_pipeline.py::test_a_company_without_a_website_still_gets_ai_mode_signals
An unreachable registry site is blocked tests/test_web_pipeline.py::test_an_unreachable_registry_site_is_blocked_not_unresolved
A mail domain without a web server, or a 404 root, is no-website; an email domain that refuses us is blocked tests/test_web_pipeline.py::test_an_email_domain_without_a_web_server_is_no_website, ::test_an_email_domain_whose_site_refuses_us_is_blocked, ::test_a_registry_website_without_a_home_page_is_no_website
An objected work contact is never published tests/test_web_pipeline.py::test_a_suppressed_work_contact_is_never_published
Duplicate delivery, retry, and the last attempt tests/test_web_pipeline.py::test_a_duplicate_delivery_is_a_no_op, ::test_an_error_releases_the_run_and_the_last_attempt_ends_it
Admission idempotency and one open run per company tests/test_web_pipeline.py::test_admission_is_idempotent_and_keeps_one_open_run_per_company
The sweep re-enqueues a lost run tests/test_web_pipeline.py::test_the_sweep_hands_lost_runs_back_to_the_queue
Re-extraction reads retained bytes only tests/test_web_pipeline.py::test_reextraction_reads_retained_bytes_without_search_or_fetch
A mention resolves only to one active company tests/test_web_pipeline.py::test_a_mention_resolves_only_to_one_active_company
Ownership rules tests/test_web_site_ownership.py
Leads and directory exclusion tests/test_web_leads.py
Fetch escalation and failure names tests/test_web_fetch.py
Real Chromium TLS, proxy tunnel, and hanging scripts tests/test_browser_tls.py (RUN_BROWSER_TESTS=1)
Classification and technology filters tests/test_web_classification_filters.py
Contact availability filters tests/test_company_contact_availability.py
The HTTP service and its callers tests/test_web_service.py