# Explore Sarajevo — crawl policy (updated August 2026). # # Strategy: allow search engines and live-retrieval agents (which # drive citation traffic). Block training scrapers (which consume # bandwidth without returning traffic). See CLAUDE.md Rule 13a. # # The PDFs (assets/booklets/, assets/guides/, assets/pdfs/) are indexable # content. They are listed in the sitemap and are not blocked here — Google # should index them (see CLAUDE.md Rule 6). User-agent: * Allow: / # ─── Classical search engines ─────────────────────────────── User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / # ─── Live-retrieval agents (drive citation traffic) ───────── # These agents fetch pages in real time to answer user queries. # Being cited in an AI Overview delivers +35% more organic clicks. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-Web Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # Google-Extended: Gemini training. Currently allowed — revisit # if bandwidth becomes a concern. User-agent: Google-Extended Allow: / # ─── Training scrapers (block) ────────────────────────────── # These crawl at high volume to build training datasets. # GPTBot: 1,255:1 crawl-to-refer ratio. # ClaudeBot: 20,583:1 ratio. # CCBot: Common Crawl training corpus. # Meta agents: Meta AI training. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Meta-ExternalFetcher Disallow: / User-agent: Bytespider Disallow: / User-agent: cohere-ai Disallow: / User-agent: Amazonbot Disallow: / User-agent: DiffBot Disallow: / Sitemap: https://exploresarajevo.com/sitemap.xml