# Ibex Insights: https://www.ibexinsights.co # AI-Powered Research, Traditional & Generative Search Optimization, # Marketing, and AI Agents for Education Institutions. # # We welcome AI search and answer engines. Indexing our content is how # higher-ed leaders find our work. User-agent: * Allow: / Disallow: /demo-engine.js # Build-time ETL source/notes are not content, also 404'd at the edge. Disallow: /data/etl/ # ---------------------------------------------------------------------- # Traditional search engines # ---------------------------------------------------------------------- User-agent: Googlebot Allow: / User-agent: Googlebot-News Allow: / User-agent: Googlebot-Image Allow: / User-agent: Bingbot Allow: / User-agent: Slurp Allow: / User-agent: DuckDuckBot Allow: / User-agent: Baiduspider Allow: / User-agent: YandexBot Allow: / # ---------------------------------------------------------------------- # AI / answer-engine crawlers (allow-listed for citation) # ---------------------------------------------------------------------- # OpenAI: training, browsing, search index User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: OAI-SearchBot Allow: / # Google: Gemini / AI Overview retrieval-augmented generation User-agent: Google-Extended Allow: / User-agent: GoogleOther Allow: / # Anthropic: Claude search & answer User-agent: anthropic-ai Allow: / User-agent: ClaudeBot Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-User Allow: / User-agent: Claude-SearchBot Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Apple Intelligence User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # Cohere User-agent: cohere-ai Allow: / User-agent: Cohere-AI Allow: / # Meta AI / Facebook share-card fetcher User-agent: Meta-ExternalAgent Allow: / User-agent: Meta-ExternalFetcher Allow: / User-agent: FacebookBot Allow: / User-agent: facebookexternalhit Allow: / # Amazon, Mistral, You, Phind, Brave, Diffbot User-agent: Amazonbot Allow: / User-agent: MistralAI-User Allow: / User-agent: YouBot Allow: / User-agent: PhindBot Allow: / User-agent: Bravebot Allow: / User-agent: Diffbot Allow: / # ByteDance / TikTok / Doubao User-agent: Bytespider Allow: / User-agent: TikTokSpider Allow: / # Common Crawl (training corpus) User-agent: CCBot Allow: / # Timpi (independent search) User-agent: Timpibot Allow: / # ---------------------------------------------------------------------- # Sitemaps # ---------------------------------------------------------------------- # /sitemap.xml is now a sitemap INDEX that recurses into all children below. # The children are also listed explicitly for crawlers that don't recurse indexes. Sitemap: https://www.ibexinsights.co/sitemap.xml Sitemap: https://www.ibexinsights.co/sitemap-pages.xml Sitemap: https://www.ibexinsights.co/data/sitemap-data.xml Sitemap: https://www.ibexinsights.co/data/sitemap-rankings.xml Sitemap: https://www.ibexinsights.co/data/sitemap-states.xml Sitemap: https://www.ibexinsights.co/data/sitemap-fields.xml Sitemap: https://www.ibexinsights.co/data/sitemap-bucketb.xml # ---------------------------------------------------------------------- # LLM-readable site descriptions (the "robots.txt for the AI era") # https://www.ibexinsights.co/llms.txt – short index # https://www.ibexinsights.co/llms-full.txt – full content map # ----------------------------------------------------------------------