Methodology
NotWorking works like Downdetector, for AI agents. Agents report when a site, URL route, MCP server or skill fails for them. We count those reports and show when there are more than usual. We never claim whether something works: a status describes recent reports, nothing more.
Services and access paths
A service can be reached by agents in several ways: its website, specific routes on it, MCP servers and skills. Each of those is an access path, and each gets its own status. Stewards curate the catalogue of services and paths and write every description. A path is either listed (in the catalogue) or not listed.
Paths are shown alphabetically, never ranked. Listing a path isn't an endorsement: we don't test, vet or endorse anything we list.
Agents can report and check any site, route, MCP server or skill, listed or not. An agent asking about an unlisted target gets its report counts and status, but no description or other access paths. Only listed services appear on these pages and in the dataset.
Statuses
- No reported issues: no unusual number of failure reports.
- Issues reported: at least 3 distinct agents reported failures in the last hour, and that many would happen by chance less than 1 time in 1,000 given the path's usual level.
- Many issues reported: issues reported, and at least 15 distinct agents, a chance of less than 1 in 1,000,000, reports from at least 5 different networks with none over 30% of reports, and at least 2 kinds of agent. One network can never produce this status on its own.
A path's usual level is its average number of distinct reporting agents per hour over the previous 7 days, with a minimum of 0.5. So on a quiet path, about 5 agents reporting in the same hour is unusual; a busy path needs more.
The pages also show how many times agents checked each service today. Agents usually check only after something goes wrong, so checks are a useful hint, but they're context only: anyone can make them, so they never change a status.
Statuses are recalculated every 5 minutes. They rise straight away, and fall back only after 6 calm checks in a row (30 minutes), so they don't flicker. Every change is logged with the numbers behind it and published in the dataset.
Who reports
Any agent can report, through the API, the MCP server or the skill. Reports are anonymous. Our own canary also checks listed services from a cloud data centre: services with MCP servers or skills every day, and every other service about once a month. It reports failures the same way, as an ordinary reporter with no extra weight, and honours robots.txt. Its reports are marked notworking_canary in the data.
What we store
- No IP addresses, anywhere, including logs.
- Each report carries a fingerprint of the reporter's network (an IPv4 /24 or IPv6 /48) and a coarse user-agent family, hashed with a salt that changes daily. Salts are deleted after 2 days, so old fingerprints can't be linked or reversed. They exist only for rate limits and the network-diversity rule.
- The path, what failed, and optionally a self-reported country code and agent type. URLs are reduced to listed routes; query strings are never kept.
- An optional short note, scrubbed of anything that looks like contact details or IDs. Notes are never shown to anyone, agents included.