Fields
Every claim field documented across crawlers, one page per field.
Each field below groups every current claim of that field across all documented crawlers on one page, so a question like which crawlers obey robots.txt or which crawlers publish IP ranges can be answered by field instead of by crawler.
Each bar is one behaviour, such as obeying robots.txt, and how many of the crawlers have a documented answer for it.
| Field | Crawlers documented | Description |
|---|---|---|
| anti_circumvention_stance | 1 | Whether and how each crawler documents anti circumvention stance. |
| cites_sources | 9 | Whether and how each crawler documents cites sources. |
| fallback_to_googlebot_rules | 1 | Whether and how each crawler documents fallback to googlebot rules. |
| follows_noarchive | 1 | Whether and how each crawler documents follows noarchive. |
| honours_crawl_delay | 15 | Whether and how each crawler documents honours crawl delay. |
| observed_arrival_before_sitemap | 1 | Whether and how each crawler documents observed arrival before sitemap. |
| observed_fetches_json | 1 | Whether and how each crawler documents observed fetches json. |
| observed_fetches_markdown | 1 | Whether and how each crawler documents observed fetches markdown. |
| observed_fetches_robots_txt | 1 | Whether and how each crawler documents observed fetches robots txt. |
| observed_first_request_path | 1 | Whether and how each crawler documents observed first request path. |
| observed_format_pass_order | 1 | Whether and how each crawler documents observed format pass order. |
| observed_paths_fetched | 1 | Whether and how each crawler documents observed paths fetched. |
| observed_requested_mcp_endpoint | 1 | Whether and how each crawler documents observed requested mcp endpoint. |
| observed_verified_request_ratio_census | 1 | Whether and how each crawler documents observed verified request ratio census. |
| opt_out_mechanism | 1 | Whether and how each crawler documents opt out mechanism. |
| publishes_ip_list | 15 | Whether and how each crawler documents publishes ip list. |
| purpose_documented | 15 | Whether and how each crawler documents purpose documented. |
| reads_llms_txt | 9 | Whether and how each crawler documents reads llms txt. |
| respects_robots_txt | 15 | Whether and how each crawler documents respects robots txt. |
| robots_token_differs_from_ua | 2 | Whether and how each crawler documents robots token differs from ua. |
| robots_txt_fetch_marker | 1 | Whether and how each crawler documents robots txt fetch marker. |
| robots_txt_token | 1 | Whether and how each crawler documents robots txt token. |
| scope_of_crawling | 1 | Whether and how each crawler documents scope of crawling. |
| separate_search_and_training_tokens | 15 | Whether and how each crawler documents separate search and training tokens. |
| user_agent_string_full | 12 | Whether and how each crawler documents user agent string full. |
| user_fetch_exception_documented | 1 | Whether and how each crawler documents user fetch exception documented. |
| verifiable_by_rdns | 15 | Whether and how each crawler documents verifiable by rdns. |
| what_it_controls | 2 | Whether and how each crawler documents what it controls. |