Robots.txt monitor · AI crawler policy

Monitor robots.txt changes that alter AI crawler access.

A raw file diff tells you that text moved. Crawler Watch tells you whether the observed homepage policy changed for eight OpenAI, Anthropic, Perplexity, and Google crawler tokens—then confirms the new state before emailing you.

15-minute checks · two-check confirmation

Policy-aware monitoring for one site.

Enter the public domain during Stripe checkout. The first baseline normally reaches the checkout email within 15 minutes; a changed state is emailed only after it appears twice.

  • Resolves homepage rules for eight named AI crawler tokens
  • Checks the public robots.txt response every 15 minutes
  • Tracks default sitemap.xml and llms.txt state
  • Sends one confirmed-change email and a separate recovery
Why policy-aware monitoring

A file change is not automatically an access change.

01

Fetch the public file

The monitor records whether robots.txt is available and reads the live directives from outside your stack.

02

Resolve the winning rule

It applies matching groups, longest-path precedence, wildcards, end anchors, and Allow-on-tie behavior for each monitored token on the homepage.

03

Report the decision

The email names the observed policy state. Repeated polling stays quiet until a new confirmed state or recovery appears.

Robots policy is one layer. The monitor checks adjacent failure signals too.

Included

Eight robots.txt policy decisions, the status seen by synthetic OAI-SearchBot, Claude-SearchBot, and PerplexityBot requests, plus default sitemap and llms.txt availability.

Not claimed

Authentic provider identity, IP verification, complete uptime, crawling, indexing, citation, ranking, recommendation, or referral traffic. Synthetic requests are external evidence, not provider logs.

If you use Cloudflare or another edge provider, its bot controls can override an otherwise permissive robots.txt file. Crawler Watch observes a bounded external path; it does not replace CDN logs or security monitoring.

Alert lifecycle

Baseline, confirmation, recovery—without a stream of green emails.

01

Baseline

The activation email records the starting policy and discovery-file evidence for the domain entered in Stripe.

02

Confirmation

A new state must survive the next 15-minute run. This suppresses one transient fetch failure from becoming an alert.

03

Recovery

When the observed state returns, the recovery is sent once. Duplicate checks of the same state remain silent.

Inspect the parser first

Check the current robots.txt policy free.

Run the same eight-token policy check once before subscribing. The free result explains which group and directive won, without claiming that a provider crawled or indexed the site.

Check one site now

If you need multi-site monitoring, general page-change detection, real crawler logs, or edge enforcement, compare the current monitoring approaches before subscribing.

Robots.txt monitoring questions

How often is robots.txt checked?

Crawler Watch checks the public site every 15 minutes. A changed state must appear in two consecutive checks before an alert is sent, so confirmation normally takes about 30 minutes.

Does every edit trigger an email?

No. The monitor resolves the homepage policy for eight named crawler tokens and compares the observed state. An edit that does not alter a monitored result may leave the policy state unchanged.

Does this prove that an AI provider crawled my site?

No. Robots.txt policy and synthetic external responses do not authenticate provider IPs or prove crawling, indexing, citation, ranking, recommendation, or traffic.

What happens after checkout?

Enter one public website in Stripe. Monitoring starts automatically and the first baseline normally reaches the checkout email within 15 minutes. The subscription renews monthly until canceled through Stripe.