A maker may want search engines to find a launch, AI search products to cite its documentation, and model-training crawlers to stay away, while a visitor may still need an assistant to fetch one page on request. Those choices are related, but they are not the same. One vague answer about whether bots are allowed hides the choices that matter.
IndieCrawl exists to make those choices visible. Enter a domain or an exact page URL. The checker follows the page to its final public destination, reads that site’s robots.txt, and applies the matching rules from RFC 9309 to the exact path for each named crawler. It also reads robots meta tags and X-Robots-Tag headers for page-level indexing instructions.
Every crawler result keeps the checked robots.txt file and the provider’s own documentation beside the answer. The indexing result links the page response it came from. IndieCrawl separates ordinary search, AI search, training and datasets, and user-triggered retrieval instead of blending them into an invented score.
The map describes published policy, not real network access. An Allow rule does not prove that a crawler can pass a firewall, CDN rule, login, or bot challenge. A noindex directive concerns indexing, not crawler access. Neither result promises a ranking, an AI citation, or use in training.
The core checker stays free. It asks for no account, card, or premium plan. It stores no lookup history or publishes a list of checked sites. Cookieless analytics measures visits.