How we measure business websites
Each state of the web report covers one city in one month. The figures change from city to city, the rules do not: they are written here, once. Each report then says what applies only to its own scan, such as the dates and the kind of scan.
Where the businesses come from
The businesses come from Overture Maps, an open dataset of places and businesses that several maps also use. We count the ones whose point falls within the city's administrative limits, the municipality rather than the province, after leaving out the categories that are not relevant to an online presence.
Each business's website is the one listed on its entry. This is not the official business register: these are the businesses a customer finds on a map, with the details the map carries.
What Overture Maps is, and why we use itWhat counts as a working website
A website works when at least one of its pages answers without an error. If our tool stops halfway, because of a timeout or a Lighthouse error, the site still counts as working: that is our limit, not a problem with the site.
A social profile, an app, or a page on a booking, delivery or marketplace platform does not count as a website, even when the business lists it as one. Every figure about the web covers working websites only.
Businesses and sites: two different units
Businesses are counted one by one. Everything about the web is counted by site: a shopping center whose 40 stores list the same website counts as one site, not 40. Basic structured data is the exception and is counted by business, because it describes each business.
Every figure says how many sites or businesses it is based on. A measurement that failed on a site does not count as a failure: that site is left out of that figure. That is why figures in the same report can cover slightly different numbers of sites.
Quick scans and full scans
In a quick scan the crawler reads each site's home page and the first 4 pages of its sitemap. Lighthouse and technology detection (Wappalyzer, after the JavaScript has run) run on the main page only.
In a full scan the crawler reads more pages per site; Lighthouse stays on the main page, so scores from different cities remain comparable.
Lighthouse and Core Web Vitals
Lighthouse is the Google tool behind PageSpeed Insights. We run it with its default settings: it simulates a mid-range phone on a slow mobile connection. These are lab measurements, the same for every site, not the times of real visitors: they are for comparing sites with each other, and they are longer than on a computer on a fiber connection.
Scores run from 0 to 100, in Lighthouse's bands: 0 to 49 poor, 50 to 89 needs improvement, 90 to 100 good. For times we use the thresholds Google publishes for the Core Web Vitals. Interactivity (INP) cannot be measured in the lab, so it does not appear.
| Measure | Good up to | Poor above |
|---|---|---|
| First content on screen (FCP) | 1.8 s | 3 s |
| Main content on screen (LCP) | 2.5 s | 4 s |
| Server response (TTFB) | 0.8 s | 1.8 s |
| Layout shift (CLS) | 0.1 | 0.25 |
Averages, medians and percentiles
The median is the value that splits the sites in half: half above, half below. The 25th and 75th percentiles hold the middle half of the sites, the 10th and 90th leave out the extremes. Percentiles are interpolated between the two closest values, as SQL's percentile_cont does.
Averages and shares are rounded to one decimal. When a report shows both the average and the median, the gap between them tells how much the extremes weigh.
Basic structured data
A business has complete basic structured data when the pages of its site carry two JSON-LD blocks: a LocalBusiness, or one of its subtypes (Restaurant, Dentist, Store...), with name, an address structured as a PostalAddress, telephone, URL and opening hours; and a WebSite with name and URL.
Each block can appear on any of the analyzed pages. A block missing even one of these counts as incomplete. An address written as free text does not count: Google would have to interpret it.
Sitemap, canonical and Search Console
A sitemap is valid when we find it, at the Sitemap line of robots.txt or at a common location, and it is a list of URLs, a urlset or a sitemap index. The canonical is the address the home page declares as its own main version.
Search Console counts as verified when we can see it from outside: a meta tag on the home page or a DNS record. Owners who verified it another way (a file uploaded to the server, Google Analytics or Tag Manager) are invisible to us, so the real share can be higher than the one we report.
The AI score
The AI score says how well an assistant such as ChatGPT, Claude or Gemini can read a site and tell what the business is. It runs from 0 to 100 and is the weighted average of ten checks, explained one by one on its own page. A check that could not be verified does not count as failed: the site is judged on the others.
How the AI score is calculatedSpam blocklists
We check the IP address of each working site's server against three blocklists: Spamhaus ZEN, Barracuda and SpamCop. A site is blocklisted when at least one of the three lists its server. Spamhaus's policy lists, which only flag addresses that should not send email directly (PBL), do not count. When a list does not answer, the site stays without an answer.
Being blocklisted is about email, not the website: email sent from that server is likely to land in customers' spam folders and never be read, while the site usually keeps working and showing up on Google as before. The server is often shared with hundreds of other sites, and the flag comes from one of them.
The national average
The national average is not an official statistic: it is the same report computed over every city in a country we have scanned up to that month, using each one's latest finished scan. A business found by several scans counts once.
We only show it once it covers at least two cities: with one, it would compare the city with itself. A value is in line with the average when the difference is within 5% of the average; beyond that, we say higher or lower.
How often we scan a city
A city can be scanned at most once a month. Each scan keeps its own report, which does not change when the city is scanned again: two months are two reports, and they can be compared.
While a scan is running, its report refreshes at most once a minute, and the pages of this site once an hour.
What we do not publish
Aggregate figures only. No single business's or single site's result appears in the reports or in the public data. Anyone who wants to see their own can ask for it on a free call.
The aggregate figures of every scan are also available as JSON from the dashboard's public API (GET /v1/scans/{scanId}), for anyone who wants to cite or reuse them.