# ============================================================================= # robots.txt for qoris.ai # Last updated: May 24, 2026 # ============================================================================= # Qoris is the operating layer for governed AI work. We welcome legitimate # crawlers — search engines, AI training crawlers from major model providers, # and AI retrieval/citation crawlers that bring our content into LLM responses # with attribution. We block scrapers, aggregators, and SEO crawlers that # extract value without contributing to discoverability. # # AI-readable references: # - LLM site map: https://qoris.ai/llms.txt # - Full content: https://qoris.ai/llms-full.txt # - Sitemap: https://qoris.ai/sitemap.xml # # For questions about our crawler policy: hello@qoris.ai # ============================================================================= # ----------------------------------------------------------------------------- # DEFAULT POLICY — All crawlers, all paths except internal # ----------------------------------------------------------------------------- # Sets the baseline. Most legitimate crawlers fall under this rule. User-agent: * Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ Disallow: /private/ Disallow: /*.json$ Disallow: /*?preview= Disallow: /*?draft= # ----------------------------------------------------------------------------- # SEARCH ENGINES — Explicitly allowed # ----------------------------------------------------------------------------- # Major search engines that drive traditional organic discovery. User-agent: Googlebot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ User-agent: Googlebot-Image Allow: / Disallow: /_next/ Disallow: /api/ User-agent: Bingbot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ User-agent: DuckDuckBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ User-agent: Slurp Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # ----------------------------------------------------------------------------- # AI TRAINING CRAWLERS — Major model providers (explicitly allowed) # ----------------------------------------------------------------------------- # These crawlers ingest content into foundation model training data. # Qoris allows them to maximize long-term brand presence inside LLM responses. # This includes content used by Claude, ChatGPT, Gemini, and other major LLMs. # Anthropic (Claude foundation model training) # We are a Claude Partner Network member. User-agent: anthropic-ai Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # OpenAI (ChatGPT foundation model training) User-agent: GPTBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Google AI training (Gemini, Vertex AI) # Distinct from Googlebot — Google-Extended controls whether content trains AI models User-agent: Google-Extended Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Apple AI training User-agent: Applebot-Extended Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Meta AI training User-agent: Meta-ExternalAgent Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Cohere training User-agent: cohere-ai Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Mistral training User-agent: MistralAI-User Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # ----------------------------------------------------------------------------- # AI RETRIEVAL & CITATION CRAWLERS — All allowed # ----------------------------------------------------------------------------- # These crawlers fetch content live to answer specific user queries with # citations back to qoris.ai. Almost always net-positive — they drive # inbound traffic and surface Qoris in LLM responses with attribution. # ChatGPT browsing / search mode (cites sources) User-agent: ChatGPT-User Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # OpenAI search index crawler (ChatGPT Search) User-agent: OAI-SearchBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Claude with web access (Anthropic retrieval, citations) User-agent: ClaudeBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Claude direct user-initiated fetches User-agent: Claude-Web Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Claude in Chrome browser extension User-agent: Claude-User Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Perplexity retrieval (cites sources by default) User-agent: PerplexityBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Perplexity user-initiated fetches User-agent: Perplexity-User Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # You.com retrieval User-agent: YouBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Phind (developer-focused AI search) User-agent: PhindBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Kagi search User-agent: KagiBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Google's AI-extended retrieval (separate from training) User-agent: GoogleOther Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # Brave search User-agent: BraveBot Allow: / Disallow: /_next/ Disallow: /api/ Disallow: /admin/ # ----------------------------------------------------------------------------- # SOCIAL & LINK PREVIEW CRAWLERS — All allowed # ----------------------------------------------------------------------------- # These fetch OG metadata for link previews on social platforms and chat tools. # Critical for outbound campaigns that share qoris.ai links. User-agent: facebookexternalhit Allow: / User-agent: Twitterbot Allow: / User-agent: LinkedInBot Allow: / User-agent: Slackbot Allow: / User-agent: Discordbot Allow: / User-agent: TelegramBot Allow: / User-agent: WhatsApp Allow: / User-agent: SkypeUriPreview Allow: / # ----------------------------------------------------------------------------- # DEVELOPER & DOCUMENTATION CRAWLERS — Allowed # ----------------------------------------------------------------------------- # GitHub link previews User-agent: GitHub-Camo Allow: / # Notion link previews User-agent: Notion-OG-Bot Allow: / # Figma file links User-agent: FigmaBot Allow: / # ============================================================================= # BLOCKED — Scrapers, aggregators, and bottom-feeder bots # ============================================================================= # These crawlers extract content for commercial scraping, SEO competitive # analysis, or AI training datasets we choose not to feed. Blocking them does # not affect legitimate discoverability. # ----------------------------------------------------------------------------- # AI training datasets we choose not to feed # ----------------------------------------------------------------------------- # Common Crawl — feeds dozens of LLM training datasets indiscriminately. # Many foundation models bypass Common Crawl in favor of direct crawlers. User-agent: CCBot Disallow: / # ByteDance (TikTok) AI training crawler User-agent: Bytespider Disallow: / # Amazon AI training crawler User-agent: Amazonbot Disallow: / # Diffbot — content extraction for resale to training datasets User-agent: Diffbot Disallow: / # Omgili / Webz.io — content scraping for resale User-agent: Omgilibot Disallow: / User-agent: Omgili Disallow: / # FacebookBot (separate from social preview) — training crawler User-agent: FacebookBot Disallow: / # Image scrapers for AI training User-agent: ImagesiftBot Disallow: / # ----------------------------------------------------------------------------- # SEO competitive analysis tools — extract content for competitor research # ----------------------------------------------------------------------------- # These tools provide value to subscribers; we are not subscribers. # Blocking prevents competitors from auto-monitoring our content changes. User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: SemrushBot-SA Disallow: / User-agent: SemrushBot-BA Disallow: / User-agent: MJ12bot Disallow: / User-agent: DotBot Disallow: / User-agent: rogerbot Disallow: / User-agent: BLEXBot Disallow: / User-agent: SerpstatBot Disallow: / User-agent: SeznamBot Disallow: / User-agent: SiteAuditBot Disallow: / # ----------------------------------------------------------------------------- # Aggressive crawlers and known bad actors # ----------------------------------------------------------------------------- # Petalsearch / Aspiegel (low-quality crawl) User-agent: PetalBot Disallow: / User-agent: AspiegelBot Disallow: / # Datanyze / Clearbit-style commercial enrichment scrapers User-agent: DataForSeoBot Disallow: / # Outdated / abandoned crawlers that still hit sites User-agent: linkdexbot Disallow: / User-agent: TurnitinBot Disallow: / # Generic aggressive crawlers User-agent: ZoominfoBot Disallow: / User-agent: VelenPublicWebCrawler Disallow: / User-agent: Sogou web spider Disallow: / User-agent: yacybot Disallow: / # ----------------------------------------------------------------------------- # AI summarization / answer-engine scrapers (no citation guarantee) # ----------------------------------------------------------------------------- # These bots summarize content without reliable attribution back to source. # Blocking until they implement consistent citation practices. User-agent: Bytedance Disallow: / User-agent: Andibot Disallow: / User-agent: ChatGPT Disallow: / # ============================================================================= # CRAWL HINTS — Where to find structured AI-readable content # ============================================================================= # These references help LLM-aware crawlers find Qoris's canonical content # without crawling every page individually. Sitemap: https://qoris.ai/sitemap.xml # Crawl rate hint for aggressive crawlers (seconds between requests) Crawl-delay: 1 # ============================================================================= # Notes for site maintainers # ============================================================================= # 1. This file is the canonical crawler policy. Update when: # - Adding new pages that should be crawled (no change usually needed) # - New AI training/retrieval crawlers emerge with significant traffic # - Strategic partnerships change AI training crawler policy # - New scraper/aggregator bots are observed in server logs # # 2. Review server logs quarterly to identify: # - New crawlers not yet categorized (allow or block?) # - Crawlers exceeding reasonable rates (consider blocking or Crawl-delay) # - Crawlers ignoring this file entirely (consider firewall-level blocks) # # 3. This file is read at https://qoris.ai/robots.txt. It must remain # accessible without authentication or redirects. # # 4. Robots.txt is a policy file, not a security mechanism. Confidential # or sensitive content should never live on URLs that depend on robots.txt # for protection. Use authentication, not Disallow rules. # # 5. Companion files: # - /llms.txt Short structured site map for LLMs # - /llms-full.txt Comprehensive content payload for LLM reasoning # - /sitemap.xml XML sitemap for search engines