# robots.txt for jlarkstories.com # # The text and images on this site are original creative work, or artwork # created by other artists and published here with their permission. They are # NOT free training data. # # Using any content from this site to train, fine-tune, or evaluate machine # learning or generative AI models, or to build a dataset assembled for that # purpose, is prohibited. Rights are expressly reserved, including for the # purposes of Article 4(3) of EU Directive 2019/790. # # Full terms: https://jlarkstories.com/terms/ # # Note: Ordinary search engines (Googlebot, Bingbot, # DuckDuckBot) and AI *answer* bots that fetch a page because a human asked # about it (PerplexityBot, ChatGPT-User, Claude-User, OAI-SearchBot) are # allowed ON PURPOSE, so this site can still be found and cited. Only crawlers # that collect content for model training are blocked below. # --------------------------------------------------------------------------- # Foundation-model training crawlers # --------------------------------------------------------------------------- # Note: Google-Extended and Applebot-Extended are training-only controls. # Blocking them does NOT affect Google Search results or Siri. User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-Web User-agent: Google-Extended User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: Applebot-Extended User-agent: Meta-ExternalAgent User-agent: meta-externalagent User-agent: FacebookBot User-agent: Amazonbot User-agent: Bytespider User-agent: PanguBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: AI2Bot User-agent: AI2Bot-Dolma User-agent: MistralAI-Train Disallow: / # --------------------------------------------------------------------------- # Dataset and corpus builders # --------------------------------------------------------------------------- # These matter most for artwork. They are how images end up in image-text # training sets. User-agent: CCBot User-agent: img2dataset User-agent: ImagesiftBot User-agent: Diffbot User-agent: omgili User-agent: omgilibot User-agent: Webzio-Extended User-agent: Timpibot User-agent: Kangaroo Bot User-agent: YouBot User-agent: VelenPublicWebCrawler User-agent: Scrapy Disallow: / # --------------------------------------------------------------------------- # Commercial AI / SEO content scrapers # --------------------------------------------------------------------------- User-agent: SemrushBot-OCOB User-agent: SemrushBot-SWA Disallow: / # --------------------------------------------------------------------------- # Everyone else: welcome # --------------------------------------------------------------------------- User-agent: * Allow: / # Emerging IETF AIPREF syntax (draft). Reserves AI training while leaving # search and human-initiated retrieval permitted. Parsers that don't # understand this line ignore it. Content-Usage: train-ai=n Sitemap: https://jlarkstories.com/sitemap.xml