# How agent-ready are Finnish B2B sites? I scanned sixteen

2026-07-07

A scan of sixteen selected Finnish B2B sites records gaps in agent-readiness. This historical sample leads into the later study of 567 sites.

Over the past weeks I ran an independent agent-readiness scanner over sixteen Finnish company websites, mostly industrial and B2B, a few in healthcare. The scanner was isitagentready.com, which grades on a Level 0 to 5 scale. This is a small, non-random sample. The sites came from my own prospecting, not a statistical draw, so read it as a snapshot, not a census. The pattern was consistent enough to be worth writing down.

## Key figures

- Sixteen Finnish B2B sites scanned with an independent scanner, isitagentready.com.
- Almost all sites landed at isitagentready Level 1 of 5, a couple at Level 0, one at Level 2, none higher.
- The three most common gaps: HTML-only pages with heavy token overhead, missing structured data, and no action or capability layer.
- Largest token saving reported: about 16500 tokens of HTML against an estimated 1400 tokens of markdown on one site, a 91 percent saving.
- The two sites that published a real llms.txt sat at the top of the range.

Note added July 17: one reading of the llms.txt point is circular, since
the scanner scores llms.txt directly, so publishing one raises the score
by construction. The observation stands as a description of the measured
range, not as a causal claim about readiness.

## The numbers

On the isitagentready Level scale almost all of the sixteen landed at Level 1 of 5, the floor an ordinary CMS site reaches, a couple sat at Level 0, and only one reached Level 2. None reached Level 3 or above.

That does not mean these are broken websites. They load, and a person can use them without trouble. The scanner measures something else, whether an AI agent can read the site and act on it.

## The three gaps that showed up almost everywhere

Discoverability was usually fine, legibility was not. Most sites had robots.txt, a sitemap, sometimes explicit AI-bot rules, so an agent can find them. But the same sites served HTML only, often with heavy token overhead. One consumer-facing corporate site returned about 16500 tokens of HTML where a markdown version was estimated at 1400 tokens, a 91 percent saving. An agent can fetch the page, but reading it is slow and lossy.

The second gap was structured data, or the lack of it. Missing JSON-LD and product data was common, so an agent that reaches the site finds no structured statement of what the company makes or sells and has to infer it from the markup.

The third and most consistent gap was the action and capability layer. No markdown negotiation, no MCP server, no API discovery, no agent-auth metadata. One site that belongs to an AI company itself passed zero of eight checks in that discovery group. This is the layer that lets an agent move from finding a site to operating it, and it was absent almost everywhere.

## Why this matters now

AI agents are becoming a discovery and transaction channel. When an agent reads a site and cannot parse or act on it, the business risks being left out of the answer, which is a different problem from ranking lower. This scan read the sites. It did not record any assistant's answer or any search ranking. What it shows is that the sites in this sample are behind on being legible and actionable to the agents that increasingly read on a person's behalf.

The encouraging part is that the fixes are mostly known and mechanical. Serve markdown alongside HTML, add structured data, publish an llms.txt, expose the discovery manifests. Two of the sixteen had already started, they published a real llms.txt, and that is exactly why they sat at the top of the range.

For the larger sample, a later post ran the same scanner across [567 company sites](/blog/website-agent-readiness-567-sites).

To check where a site stands, the free llms.txt validator is at [turva.dev/llms-txt-validator](https://turva.dev/llms-txt-validator), and the agent-readiness audit and advisory work is at turva.dev.

Corrected 2026-09-25. Three passages said more than the scan measured. The token figures were an estimate from a scanner this site no longer uses, and they are now labelled as reported and estimated. Two sentences drew what an agent can answer and whether a business appears in an answer from configuration checks, and a third said most of these sites rank fine in search, which was never measured. They now say what the scan read. The counts did not change.

Corrected 2026-09-27. One more sentence still said these sites rank in search, which was not measured either. It now says only that they load and that a person can use them.

## Related

- [Common agent-readiness gaps in a measured sample](/guides/agent-readiness-gaps)
- [Measure agent-readiness with evidence](/guides/measurement-led-agent-readiness)
- [What an agent pays to read your site](/blog/cheaper-pages-for-agents)
- [Website agent readiness, measured on 567 company sites](/blog/website-agent-readiness-567-sites)
