Evidenced web classification filters¶
Status: accepted Date: 2026-09-10 Issue: #654, decisions W1–W3
Decision¶
Publish sales_areas, target_markets, technologies, has_webshop,
has_google_tag, and has_facebook_pixel. These are observations in retained
company website evidence. They are not unqualified facts about every company
activity or every page on its website.
Commercial classification¶
The existing profile extractor gains a separate closed classification output. Existing free-text geography, target-market, and commercial-posture values are not converted into hard filters. Every classification requires an exact quote from a permitted input reference, a source URL, and the model's reported confidence. Confidence is a self-assessment, not a probability threshold.
| Dimension | Value | Required source statement |
|---|---|---|
| Sales area | local |
Service or delivery in a named town or local area. |
| Sales area | regional |
Service or delivery across a named region. |
| Sales area | national |
Explicit countrywide service or delivery, with the country named. |
| Sales area | international |
Explicit service or delivery across countries, with the geography named. |
| Target market | b2b |
The company explicitly serves business customers. |
| Target market | b2c |
The company explicitly serves private consumers. |
| Target market | b2g |
The company explicitly serves public-sector customers. |
An office address does not establish sales coverage. Languages do not establish international coverage. A procurement notice without company attribution does not establish B2G. Claims must concern the target company, not another group entity. Several supported values may coexist. Unresolved contradictory claims make that dimension unknown and keep the conflicting quotes in history.
Technology and webshop observation¶
A separate deterministic extractor reads HTML already present in the sealed source artifacts. It does not fetch pages, run JavaScript, fire tags, send events, or submit forms. It uses Beautiful Soup for DOM inspection and Tree-sitter's maintained JavaScript grammar for calls; comments, string examples, and escaped code samples cannot become installation calls.
The initial technology catalog is deliberately bounded:
| ID | Installation evidence |
|---|---|
wordpress |
A WordPress generator metadata element. |
woocommerce |
A WooCommerce asset path or a WooCommerce block element. |
shopify |
A Shopify CDN asset on the page. |
google_tag_manager |
A Google Tag Manager script URL with a GTM container ID. |
google_tag |
A Google tag loader and a matching literal gtag('config', ID) call. |
meta_pixel |
A Meta Pixel loader and a literal fbq('init', ID) call. |
Google Tag Manager alone is not a configured Google tag. Deferred consent-controlled installation code is installation evidence, not evidence that an event fired. Dynamically computed identifiers and server-only tagging are outside the detector coverage.
Webshop evidence requires an identifiable Shopify product-add form or a WooCommerce product-add control with a product identifier and submission control. Product identifiers must be positive decimal integers. A basket word, payment-provider mention, or technology asset alone does not qualify. Other shop implementations are outside this detector catalog.
These rules follow the official installation and storefront patterns:
- Google tag installation and tag identifiers.
- Tag Manager installation.
- Meta's maintained Pixel integration.
- WordPress generator markup.
- Shopify product-add form.
- WooCommerce product and cart blocks.
- Tree-sitter Python bindings and JavaScript grammar.
Coverage, publication, and freshness¶
Commercial classification uses the profile extractor's recorded bounded text coverage. HTML detection evaluates all retained successful page selections, with a total input limit of 1,000,000 characters. Missing or incompatible HTML and exceeded input limits are failed evaluations, never negative observations. Markdown-only evidence cannot support a negative technology or shop result. The registry-identity route limits text to the company identity block. Its HTML has no supported equivalent boundary and is unavailable to this detector. An inline Meta loader must match the official bootstrap syntax, including its script creation and insertion; an unused loader URL is insufficient. Re-extraction requires compatible retained artifacts and never acquires more pages silently.
A boolean false means no approved signature was detected in that declared HTML coverage. It does not mean site-wide absence. Both boolean selections exclude unknown values. A set filter matches any supported value. Other filter keys combine with AND; exact counts count each company once.
Each dimension keeps append-only evaluations with source-qualified evidence,
coverage, rule and extractor versions, source fetch times, and the frozen run
policy. The effective freshness field is policy.page_fresh_seconds, currently
2,592,000 seconds (30 days), frozen at admission. Expiry starts at the earliest
source fetch time in the evaluation, not at re-extraction or cache application.
The newest successful evaluation supersedes earlier values, including a negative or conflicting result. Failure preserves earlier successful evidence and records a separate latest-attempt status. Expired evidence remains readable but cannot match a hard filter. Search versions include the latest reached expiry boundary, so expiry changes the version before filter eligibility changes. Evaluation writes and search publication changes are one transaction.
Company detail publishes current eligibility, latest attempt, history, and provenance. Consumers use the public filter contract and do not interpret HTML or free text themselves.