# GeoAxis robots.txt # Default: allow crawling so the site appears in search + LLM answer engines # (Google, Bing, ChatGPT, Claude, Perplexity, Gemini, Meta AI, etc). # Block only dodgy scrapers: dataset hoarders and abusive bots. # robots.txt is voluntary — real enforcement happens at the WAF / auth layer. Sitemap: https://geoaxis.ai/sitemap.xml # ---- Blocked: dataset hoarders / abusive scrapers ---- # Listed individually because most crawlers honor only their own UA block. # SEO tools (Semrush, Ahrefs, DataForSEO, Serpstat) are intentionally NOT # blocked: blocking them blinds us from rank-tracking our own site in # those tools, and the data they expose is the data we need to compete. User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Omgilibot Disallow: / User-agent: Diffbot Disallow: / User-agent: MJ12bot Disallow: / User-agent: dotbot Disallow: / User-agent: PetalBot Disallow: / User-agent: TurnitinBot Disallow: / User-agent: Timpibot Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: Scrapy Disallow: / User-agent: BLEXBot Disallow: / User-agent: ZoominfoBot Disallow: / User-agent: Barkrowler Disallow: / User-agent: SeekportBot Disallow: / # ---- Default: everyone else allowed ---- # Authenticated surfaces (dashboard, login, claim) and the JSON API are protected # at the auth + WAF layer. We don't enumerate those paths here on purpose: # robots.txt is voluntary and listing them just gives recon to scanners. User-agent: * Allow: /