‹ all posts

Erik Rekola · updated 2026-10-02

I scanned thirteen code hosts. Not one served an MCP server card.

2026-08-22

Fourteen code-host surfaces were scanned on one day. The findings concern public discovery paths, not the full capabilities of each hosting service.

Cursor launched Origin on August 17 and calls it a Git forge for the agentic era. I ran an independent agent-readiness scanner over its public surface and over twelve other code hosts on August 22, one of them on two surfaces. Not one of them reached Level 2 of 5.

A 403 answer is a refusal, not evidence that a file is absent. Where the count below reads zero, some of those checks got no answer at all rather than a confirmed absence, and that qualification applies directly to the headline claim above.

Key figures

What was measured, and what was not

The scanner reads a public web surface. It does not log in, and it does not see a repository. Origin itself sits behind a Cursor paid plan, so the reading describes cursor.com and its marketing page for Origin, not the forge. Four targets redirected somewhere else, GitLab to its marketing site and Azure DevOps to a Microsoft product page, so GitLab was measured a second time from the application at gitlab.com/explore. Both readings landed at Level 1, with different checks passing on each.

One host is missing from the count. savannah.gnu.org did not answer on two attempts, once with a network error and once with a 502, and an unreachable site is not a zero.

For scale, my own site reads Level 5 of 5 on the same scanner on the same day. That comparison is not a fair fight and I am not presenting it as one. turva.dev is a one-person advisory site with sixty canonical pages, and a code host carries multi-tenant load plus an access model that a site like mine never has to solve. It does show that the manifests in question are not expensive to publish.

The numbers

HostLevelPassed
cursor.com/origin1/53
gitlab.com marketing1/54
gitlab.com/explore1/54
sourceforge.net1/53
forgejo.org1/53
dev.azure.com1/53
github.com0/52
bitbucket.org0/52
sr.ht0/52
gitea.com0/53
gitee.com0/52
launchpad.net0/53
radicle.dev0/53
gerritcodereview.com0/51

For almost every host the passing checks were the same two, a robots.txt the scanner reads as valid and a robots.txt whose rules reach AI crawlers, either by naming them or by letting the wildcard cover them. That is the floor an ordinary site reaches without trying. Gerrit sits below it, because its robots.txt carries no User-agent line at all, so the file is read as invalid and the AI rules check falls with it.

GitHub runs an MCP server. The discovery paths this scan checked do not say so.

This is the finding I keep coming back to. GitHub operates a production MCP server, and I used it on the same day I ran these scans. It works. But github.com serves no MCP server card, no API catalog and no Link header pointing at either, so an agent that arrives at github.com without being told about the server finds none of the pointers this scan looks for. The capability exists and the announcement does not.

The same shape repeats across the sample. What follows is a reading of how these products are sold, which the scan does not measure. Several of them are sold as the place where agents work on code, and on every one of them the way in is a docs page written for a person.

Three things the scan does not prove

GitLab answered HTTP 403 to most of the well-known paths, including the API catalog, auth.md, the MCP card, the A2A card, the skills index and the ARD manifest. A 403 is a refusal, not evidence that a file is absent, and I have recorded those as failures only because the check got no answer.

Two checks did not complete on Gitee, and for different reasons. The A2A card fetch aborted and the WebMCP check timed out at eight seconds. Its discovery group was therefore scored on seven of nine.

The sample is fourteen surfaces I picked by hand. It is not a random draw and it does not cover every code host. Read it as a snapshot of one day.

Why any of this matters

An agent that lands on a code host today can read the marketing copy. On the paths this scan read, nothing on the site told it what the site is able to do in a format an agent parses. An agent that relies only on the paths this scan checked and needs one of those capabilities therefore depends on a human who already knows the endpoint exists.

The fixes are small and mostly mechanical. A server card is a JSON file at a known path. An API catalog is a linkset. A Link header is one line of response configuration. None of it requires rebuilding a forge, and none of it was found published at the paths this scan checked on any of the fourteen surfaces I measured, GitLab's 403 answers included, which this scan recorded as a check that got no answer rather than a confirmed absence.

If you want to check a site yourself, the scanner is public and the free llms.txt validator is at turva.dev/llms-txt-validator. The audit and advisory work is at turva.dev.

Frequently asked

Does this mean GitHub is broken?

No. GitHub works, and so does its MCP server, which I used on the same day I ran the scan. The reading describes one thing only, whether the site announces what it can do in a format an agent finds without being told. On that point github.com reads zero, and so does every other host in the sample.

Why would a code host publish an MCP server card?

So that an agent arriving at the domain can learn that a server exists, where it is and what it does, without a human pasting the endpoint into a config file first. The card is a JSON file at a known path. It does not change the forge and it does not expose anything the docs do not already say.

How do I check a host myself?

Run the same public scanner against the domain and read the group named API, Auth, MCP & A2A Discovery. The result is a snapshot of that day, mine included, because these specifications move month to month.

Corrected 2026-09-27. The title said fourteen code hosts. The table has fourteen surfaces from thirteen hosts, because GitLab was read twice, so the title now says thirteen hosts. A sentence also said every integration has to be hard-coded, and it now keeps to the paths this scan read. No reading changed.

Corrected again 2026-09-27. A sentence said an agent has no way to discover GitHub's MCP server. The scan read named discovery paths and not documentation or search results, so the sentence now keeps to those paths. No reading changed.

Corrected 2026-09-28. Two sentences read as findings about every possible discovery path rather than the ones this scan checked. Both now keep to the paths this scan read, and GitLab's 403 answers are named as checks that got no answer rather than confirmed absence. No reading changed.

Corrected 2026-09-28. The H2 heading itself still generalised across the whole domain after the paragraph below it was narrowed to the paths this scan checked. The heading now names the same paths instead of all of github.com.

Corrected 2026-10-02. A sentence said the two passes were metadata published so people can log in and not so agents can find anything. The scanner counts that check in its discovery group and OAuth discovery tells an agent where to request access. It now says: "Both passes were OAuth and OpenID Connect discovery metadata, which tells a client where to log in and request tokens. Neither announced an MCP server or any code host action, so no host passed a check for what an agent can do there."