Skip to main content

Website Intelligence

How conclusions are reached, and what they are worth

Any assessment is only as good as the evidence beneath it and the honesty with which its limits are stated. This page sets out both.

There is a particular failure mode in website analysis where a tool reports twelve hundred issues, ninety per cent of which are irrelevant, and the recipient — reasonably — concludes that none of it matters. The volume destroys the signal.

The opposite failure is a report that asserts a small number of confident conclusions with no indication of how they were reached, so the recipient cannot tell which to trust.

The method described here is designed to avoid both. Every finding is traceable to specific evidence, carries an explicit confidence level, and is ordered by the difference addressing it would make rather than by how easy it was to detect.

Evidence

What counts as evidence

Everything IXSEO examines is information a website publishes to anybody who asks for it: the HTML your server returns, the files at conventional locations, the records in public DNS, the headers on the response, the entries in published blocklist zones.

Nothing is inferred from private data, purchased datasets or modelled traffic estimates. This constrains what can be said, and that constraint is deliberate. A finding based on something you can verify yourself is worth more than one based on a figure you have to take on trust.

Where a check cannot be completed — a service was unavailable, a server refused the request, a record did not resolve — that is recorded as an unknown rather than treated as an absence. There is a meaningful difference between “no DMARC record exists” and “we could not determine whether a DMARC record exists”, and conflating them produces bad advice.

Evidence sources

  • HTML returned by your server, parsed without execution
  • robots.txt, XML sitemaps and llms.txt where present
  • HTTP response status, headers and redirect behaviour
  • Public DNS records for the domain and its mail routing
  • TLS certificate summary from the connection
  • Google PageSpeed Insights, where the service responds
  • Published blocklist and reputation zones
  • Home pages of compared competitors, one request each

Confidence

Three levels, stated on every finding

Confidence describes how firmly the evidence supports the conclusion. It is not a measure of how important the finding is, and the two are frequently unrelated.

How confidence levels are assigned
 BasisTypical example
HighDirectly observed, not open to reasonable disputeA page returns a server error, or no canonical declaration is present
ModerateStrong evidence with a stated assumption behind the readingContent appears thin for its commercial purpose, judged against comparable pages
LowIndicative signal, or a check that could not be completedA blocklist lookup did not resolve, or a pattern suggests an issue we cannot confirm passively

A score is a way of directing attention. It is not a measurement of performance, and it does not correspond to any position in any search result.

Scoring

How scores are constructed

Each module starts from a baseline and is reduced by the findings it produces, weighted by severity and by how reliably that signal predicts a real problem. Positive findings do not add points; they prevent deductions that would otherwise apply.

The maximum for a passive assessment is 96 rather than 100. This is not a stylistic choice. A passive review can establish that nothing observable is wrong; it cannot establish that nothing is wrong. Reserving the last few points marks that distinction on every report.

The overall score is a weighted combination of the modules that ran, not an average. Technical and SEO carry more weight because their failures tend to be more consequential and more immediately actionable. Where a module was not selected, it contributes nothing rather than being assumed adequate.

Every score is displayed with an explanation of what it reflects. A number without context invites misreading, and the most common misreading — that it corresponds to a ranking — is the one we work hardest to prevent.

Reading a band

85 and above
Sound. Remaining findings are refinements rather than corrections.
70 to 84
Functional with identifiable gaps. Worth addressing, not urgent.
55 to 69
Something material is limiting performance. Warrants attention this quarter.
Below 55
Significant problems, usually several interacting. Sequencing matters more than volume of effort.

Boundaries

What the method does not attempt

Being specific about the limits is part of the method. Each of these is a decision rather than an omission.

Prediction

  • Forecasting rankings for any search term
  • Estimating traffic gains from a change
  • Predicting citation in any generative system
  • Attributing past movement to a specific update with certainty

Private data

  • Anything requiring authentication to view
  • Internal analytics unless you supply access
  • Third-party modelled traffic figures as evidence
  • Competitor spend, revenue or internal performance

Active testing

  • Port scanning or service enumeration
  • Vulnerability probing of any kind
  • Submitting data to forms or endpoints
  • Any interaction beyond retrieving published pages

Certainty we do not have

  • Declaring a penalty without direct evidence
  • Declaring a vulnerability exploitable
  • Treating breach association as active compromise
  • Presenting inferred competitors as established ones

Standing statements

  • IXSEO performs passive public analysis unless a separate authorised engagement has been agreed.
  • Domain verification confirms control of an approved domain resource. It does not automatically authorise active testing of company infrastructure.
  • Domain verification does not automatically authorise active infrastructure testing.
  • Security Exposure findings identify publicly observable indicators and do not replace an authorised penetration test.
  • GEO assessments evaluate observable discoverability, structure, clarity and authority signals. They do not guarantee visibility or citations within third-party AI systems.
  • Scores are directional assessments based on the available evidence and should not be interpreted as search-engine rankings.

Frequently asked questions

Why does a site with no findings not score 100?
Because a passive assessment cannot prove the absence of problems, only the absence of observable ones. A score of 96 with no findings means we looked and found nothing wrong, which is a different statement from nothing being wrong. Reserving the top of the range is an honesty measure.
Are the scores comparable between websites?
Broadly, within the same module and for sites of similar type. They are constructed from the same signals with the same weights, so a technical score of 80 means roughly the same thing on two sites. They are less comparable across modules, because the underlying evidence differs in quality.
How are the weights decided?
By how reliably a signal predicts a real problem and how much difference addressing it tends to make. A missing title has a strong, well-established effect and is weighted accordingly. Structured data has a more contingent effect and is weighted more lightly. The weights are a matter of professional judgement and we describe them rather than presenting them as objective.
Do you use third-party ranking or traffic estimates?
Not as evidence for findings. Those estimates are modelled from panel data and clickstream sampling, with error margins wide enough to mislead when used for decisions. Where such data is referenced in a paid engagement it is labelled as a third-party estimate and never treated as measurement.
What happens when the evidence is ambiguous?
The finding is issued with low confidence and the ambiguity is stated in the text. It is not upgraded to a certainty because a definite recommendation reads better. Some findings exist principally to tell you that something warrants a look by someone with access we do not have.
How often does the method change?
Continuously and incrementally, as search engines and generative systems change what they do. Substantive changes to weighting or to what is checked are reflected here. We do not retrospectively restate old reports; a report is a record of what was observable when it was produced.

Apply the method to your website

The free snapshot uses the same evidence rules, the same confidence levels and the same scoring construction described on this page.