# Updated 2026-Aug-21 @ 12.00 # NOTE: Blindly allowing AI crawlers complete access to site will likely impact performance both of Epicor Commerce and your ERP. # The format below has been arrived at over time collating suggestions from other customers SEO specialists. # # Please discuss with support before making changes to verify likely impact and efficacy # ADD A COMMENT AT TOP TO INDICATE IF CHANGES HAVE BEEN MADE # If customer uses SemRush then need to change to Allow: / # https://www.semrush.com/bot/ lists the following bots and uses # AI-powered digital marketing platform [USA] # CS0005422882 SemRush crawl strategy differs to Google et al # If you want to enable change to Allow: / # DO NOT DELETE THIS SECTION to use wildcard [*] rules User-Agent: SemrushBot User-Agent: SiteAuditBot # SemrushBot for different SEO and technical issues User-Agent: SemrushBot-BA # Backlink Audit tool User-Agent: SemrushBot-SI # On Page SEO Checker tool and similar tools User-Agent: SemrushBot-SWA # Checking URLs on your site for SWA tool User-Agent: SemrushBot-OCOB # Content Toolkit User-Agent: SemrushBot-FT # Plagiarism Checker and similar tools User-Agent: SemrushBot-ESI # Enterprise Site Intelligence service User-Agent: SemrushBot-SA # Older / auxiliary Site Audit and SEO analysis traffic User-Agent: SemrushBot-EO # Experimental / internal User-Agent: SplitSignalBot # SemRush for SEO A/B testing Allow: / Disallow: /catalog/product_compare/ # CS0005590925 prevent duplicate content issues Disallow: /customer/ Disallow: /rest/ Crawl-Delay: 5 # Google User-Agent: Googlebot Allow: / User-Agent: Googlebot-Image Allow: / # Default configuration User-Agent: * # Custom for SKC Allow: /sampling-guides/*/detail/?id # NOTE: Google never honours the crawl-delay directive, but other bots do # NOTE: Set to maximum of 30, some crawlers may ignore very high values Crawl-Delay: 3 # NOTE: Relatively, new, proposal by CloudFlare, therefore unlikely to be honoured # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. Content-signal: search=yes,ai-train=no # Put customer specific allows here # NOTE: Sitemap directives should be very bottom of file [if not being auto-generated by Magento] # DELETE OR CHANGE AS REQUIRED: Some sites may require these, function not totally clear # Customer specifically requested product_compare be disallowed - CS0005274748 Disallow: /catalog/product_compare/ Allow: /catalog/category/view/ Allow: /catalog/product/view/ # Disable customer specific items # Disable system related directories # Disallow: /404/ Disallow: /admin/ Disallow: /ajax/ Disallow: /ajaxcartpro/ Disallow: /ajaxpost/ Disallow: /api/ Disallow: /app/ Disallow: /bin/ Disallow: /cgi-bin/ Disallow: /dev/ Disallow: /downloader/ Disallow: /generated/ Disallow: /includes/ Disallow: /index.php Disallow: /lib/ Disallow: /log/ Disallow: /magento/ Disallow: /phpserver/ Disallow: /pkginfo/ Disallow: /report/ Disallow: /setup/ Disallow: /staging/ Disallow: /stats/ Disallow: /tmp/ Disallow: /update/ Disallow: /var/ Disallow: /vendor/ Disallow: /SiteAdmin/ # Speculative blocks, benefit unproven # Disallow: /*.aspx # Block internal APIs Disallow: */rest/ Disallow: */graphql # Disable standard search pages # (single-store + multi-store) e.g. /us/catalogsearch/ # Disallow: */catalogsearch/ Disallow: */instantsearch/ Disallow: */search/ # Disable 3rd-party search modules # Disallow: /eccsearch/ Disallow: /amasty_xsearch/ Disallow: /mageworx_searchsuiteautocomplete/ Disallow: /mfproductsearch/ Disallow: /rapidhawksearch/ Disallow: /searchwidget/ # Yoma related?? Disallow: /searchwidget/ # Disable checkout & customer account # Disallow: /checkout/ Disallow: /customcheckout/ Disallow: /onestepcheckout/ Disallow: /punchout/ Disallow: /b2b/ Disallow: /customerconnect/ Disallow: /sales/ Disallow: */customer/ Disallow: */quickorder/ Disallow: */quickview/ Disallow: /review/ Disallow: /sendfriend/ Disallow: /tag/ Disallow: */wishlist/ Disallow: /media/captcha/ Disallow: /media/customer/ Disallow: /media/downloadable/ # Multi-store related URL's # Disallow: */switch/ Disallow: */redirect/ # Disable 3rd party URL's # # seen to cause problems - CS0005288887 Disallow: /amlocator/ # EC 2025.2.6, ECC-13923, Codazon_GoogleAMPManager module causes issues Disallow: /amp/ Disallow: /amfile/ Disallow: /amga4/ Disallow: /amasty_customform/ Disallow: /amasty_banners/ Disallow: /ammostviewed/ Disallow: /silkdistributor/ Disallow: /yoma_checkout/ Disallow: /autorelated/ # Ensure only main, clean URLs are indexed by search engines # This reduces wasted crawl budget and excessive requests to server Disallow: /*? # However we need to allow PLP pages /apparel.html?p=2 # But block things like /apparel.html?p=2&this=that # /catalogsearch/result/index/?ajax_nav=1&brand=5512&p=2&q=air+compressor Allow: /*?p= Disallow: /*?p=*& #Some customer may want to allow categories for Google #Allow: /*cat= # Need to allow specific Google and Bing et al related marketing parameters # NOTE: AI suggests that most correct version is /*?utm* Allow: /*?utm Allow: /*?gad Allow: /*?gclid Allow: /*?dclid Allow: /*?srsltid Allow: /*?msclkid Allow: /*?_gl # Need to allow following for Google Merchants # CS0005543044, but also check http response for 406 etc Allow: /amfeed/feed/download* # Bots you may or may not want, by default blocked User-Agent: BrightEdgeOnCrawl User-Agent: BrightEdge Crawler User-Agent: Google-Extended # Train generative AI models i.e. Gemini User-Agent: GoogleOther # Google Internal R&D User-Agent: GoogleOther-Image # Not tied to: Google Search or ranking. User-Agent: GoogleOther-Video # User-Agent: GPTBot # OpenAI training crawler User-Agent: OAI-SearchBot # OpenAI discovery/search crawler User-Agent: ChatGPT-User # OpenAI user-initiated fetch User-Agent: TerraCotta # Ceramic AI - https://ceramic.ai User-Agent: ClaudeBot # Anthropic/Claude crawler User-Agent: Claude-SearchBot User-Agent: Claude-Web # Anthropic web access crawler User-Agent: anthropic-ai # Anthropic crawler identifier User-Agent: PerplexityBot # Perplexity crawler/indexer User-Agent: Perplexity-User # Perplexity user-initiated fetch User-Agent: DuckDuckBot # DuckDuckGo crawler User-Agent: DuckAssistBot # DuckDuckGo assistant/AI answers fetcher User-Agent: Applebot # Apple crawler User-Agent: Applebot-Extended # Apple AI usage/training control User-Agent: Amazonbot # Amazon crawler User-Agent: Amzn-SearchBot # Amazon search crawler User-Agent: AmazonProductDiscovery User-Agent: AmazonSellerInitiatedListing User-Agent: facebookexternalhit User-Agent: meta-externalads User-Agent: meta-externalagent # Meta AI crawler User-Agent: meta-externalfetcher # Meta fetcher User-Agent: meta-webindexer # Meta web index crawler User-Agent: AhrefsBot # Allowed (potential future SEO tooling) User-Agent: AI2Bot # Allen Institute crawler (datasets) User-Agent: Ai2Bot-Dolma # AI2 Dolma dataset crawler User-Agent: Baiduspider # Baidu crawler (China) User-Agent: Baiduspider-image # https://help.baidu.com/question?prod_id=99&class=0&id=3001 User-Agent: Baiduspider-video User-Agent: Baiduspider-news User-Agent: Baiduspider-favo User-Agent: Baiduspider-cpro User-Agent: Baiduspider-ads User-Agent: Baiduspider-render User-Agent: CCBot # Common Crawl (training corpora) User-Agent: MojeekBot # Mojeek crawler (independent index) User-Agent: PetalBot # Huawei Petal Search crawler User-Agent: Qwantbot # Qwant crawler User-Agent: SeznamBot # Seznam crawler (Czech) User-Agent: Slurp # Yahoo crawler (search largely Bing-powered) User-Agent: YandexBot # Yandex crawler (Russian search engine) http://yandex.com/bots User-Agent: YandexUserProxy # Yandex related proxy/crawler User-Agent: YaDirectFetcher # Yandex related User-Agent: YandexRenderResourcesBot # User-Agent: Yeti # Naver crawler (South Korea) User-Agent: YouBot # You.com AI search crawler # Nuisance bots you almost certainly do not want User-Agent: AdIdxBot User-Agent: AIBeaverProspectSearch # UK - https://www.docbeaver.co/ User-Agent: aiHitBot # (aihitdata.com) — business intelligence crawler User-Agent: AliyunSecBot # Alibaba Cloud (Aliyun) — security/monitoring crawler User-Agent: Amazing-SearchBot User-Agent: AspiegelBot User-Agent: AZBlogTextHarvester User-Agent: BacklinksExtendedBot User-Agent: Barkrowler # scraper UA seen in blocklists User-Agent: BitSightBot # security ratings / risk intelligence crawler User-Agent: Brightbot # likely Bright Data–associated crawler name (unverified) User-Agent: BrokenLinksBot User-Agent: BufferLinkPreviewBot User-Agent: Bytespider User-Agent: CCBot # NOTE: robots compliance questionable User-Agent: CheckMarkNetwork # User-Agent: Cliqzbot User-Agent: ClueWeb-Crawler # Academic research crawler by Prof Chenyan Xiong group @ Carnegie Mellon User-Agent: cohere-ai User-Agent: cohere-training-data-crawler User-Agent: CoreAutoPDFPretrainingBot # NOTE: robots compliance is not confirmed User-Agent: Crawlspace User-Agent: coccocbot-image User-Agent: coccocbot-web User-Agent: DataProvider User-Agent: DataForSeoBot User-Agent: Diffbot User-Agent: Discordbot User-Agent: DomainStatsBot User-Agent: Dormouse User-Agent: DotBot # link intelligence / Moz Link Index crawler User-Agent: ev-crawler User-Agent: Exabot User-Agent: ExaSearchBot User-Agent: FriendlyCrawler User-Agent: GeedoProductSearch User-Agent: GeedoShopProductFinder User-Agent: HaloBot User-Agent: iaskspider User-Agent: IbouBot User-Agent: ICC-Crawler User-Agent: ImagesiftBot User-Agent: ImageMetaCrawler User-Agent: imageSpider User-Agent: img2dataset User-Agent: ISSCyberRiskCrawler User-Agent: Kangaroo Bot # Both variants seen in lists User-Agent: KangarooBot User-Agent: KareBot User-Agent: LanaiBotmarch User-Agent: LinkupBot User-Agent: MatchorySearch User-Agent: MauiBot User-Agent: MJ12bot # Majestic-12 User-Agent: oBot # IBM X-Force (IBM Germany R&D) — content security crawler User-Agent: Odin User-Agent: omgili User-Agent: omgilibot User-Agent: OpenindexSpider User-Agent: Orbbot User-Agent: Owler # competitive intelligence/company info User-Agent: PanguBot # PanGu LLM training/data collection User-Agent: panscient.com # NOTE: robots compliance is not confirmed User-Agent: PavingBot # NOTE: robots compliance is not confirmed User-Agent: Pomothy-Bot User-Agent: proximic User-Agent: RyteBot # RyteBot from crawling your site for tools provided by Ryte.com GMBH [Germany]: User-Agent: SaluteBot User-Agent: Scrapy User-Agent: SeekportBot User-Agent: SeobilityBot User-Agent: SERankingBacklinksBot User-Agent: SerpMinerBot User-Agent: serpstatbot User-Agent: ShapBot User-Agent: ShapBot-Extended # Validity not confirmed User-Agent: Shap-User User-Agent: Sidetrade indexer bot User-Agent: SiteAnalyzerbot User-Agent: SleepBot User-Agent: Sogou # Validity not confirmed, AI suggests very old ~2008 User-Agent: SpiderLing User-Agent: stepstoneCrawlBot User-Agent: Thinkbot User-Agent: TikTokSpider User-Agent: Timpibot User-Agent: TinEye User-Agent: TinEye-Web User-Agent: Turnitin # plagiarism database crawler User-Agent: TurnitinBot User-Agent: TwitterBot User-Agent: VelenPublicWebCrawler # business dataset crawler (robots-aware) User-Agent: Webzio-Extended User-Agent: WellKnownBot User-Agent: YelpBot User-Agent: YouBot User-Agent: ZoominfoBot # ZoomInfo — data broker/company info [distinct from Zoom messaging] # NOTE: Google never honours the crawl-delay directive, but other bots do # NOTE: Set to maximum of 30, some crawlers may ignore very high values Crawl-Delay: 10 Disallow: / # Disable 3rd party URL's seen to cause problems - CS0005288887 # [if you decide to change to allow bots access] # Disallow: /amlocator/ Disallow: /amfile/ Disallow: /amp/ # See also # https://github.com/ai-robots-txt/ai.robots.txt/blob/main/robots.txt Sitemap: https://trafficsafetyproducts.net/media/sitemap/trafficsafety_sitemap.xml