User-agent: * Disallow: /_service/*/display_message/* Disallow: /_service/ Allow: /_service/*.jpg$ Allow: /_service/*.png$ Allow: /_service/*/feed/* Allow: /_service/*/podcast/* # Same rules again for websites served under a path prefix (website.url, e.g. /no or /sweden). # robots.txt paths are matched from the host root, so "/_service/" above does not cover # "/no/_service/" and prefixed sites would be uncovered. The Allow lines are mirrored too, otherwise # the broader Disallow would stop podcast and image indexing on exactly those sites. # # These MUST start with "/" - RFC 9309 defines a rule value as path-pattern = "/" *UTF8-char-noctl, # and a crawler is required to ignore a rule it cannot parse. "/*/_service/" is the correct form; # a bare "*/_service/" is not guaranteed to be read at all. This matters because the nginx deny in # limit_request_rate.conf ($is_crawler, which does now include Googlebot, bingbot, Applebot, # YandexBot and Bytespider) only covers the /_service//download/ routes. The rest of the # /_service/ space - and the user-directed fetchers $is_crawler deliberately excludes - is governed # by these lines and nothing else. # "*" matches "/" as well, so one rule covers prefixes of any depth. Disallow: /*/_service/*/display_message/* Disallow: /*/_service/ Allow: /*/_service/*.jpg$ Allow: /*/_service/*.png$ Allow: /*/_service/*/feed/* Allow: /*/_service/*/podcast/* # Do NOT re-add "Allow: /_service/*/download/*". A more specific Allow beats the Disallow above in # every major crawler, so that one line invited crawlers into the module download routes # (Payment/Shop/Document DownloadWindow - member documents and purchased digital goods, no SEO # value). Meta's crawler then re-fetched the same ~850 files ~78x per 12h: ~41 GB/day of egress and # the single largest source of load on the estate. Podcast/feed indexing is covered by the two # Allow lines above and is unaffected. Disallow: /_data/ Disallow: /_import/ Disallow: /_ext/ Disallow: /_maintenance/ Disallow: /_module/ Disallow: /_system/ Disallow: /_static/ Disallow: /_temp/ Disallow: /fckeditor/ Disallow: /imageedit/ Disallow: /magpierss/ Disallow: /maps/ Disallow: /mediaplayer/ Disallow: /messagesender/ Disallow: /uploads/ Disallow: /kottedzhnye-poselki/poselki-v-moskovskoj-oblasti/tag-query/* Disallow: /podbor-uchastka Disallow: */tag-query/*_* Disallow: */tt/*/tt/* Disallow: */tt-or/*/tt-or/* Disallow: */tag-query/(* Disallow: */tag-query/%28* Disallow: */start=*&end=* Disallow: /a/search/ Disallow: /fundraising_contract/ Disallow: /eat/kebabhouse Disallow: /actions/akcia/doubling-bonuses-glam Disallow: /actions/akcia/only-modest-prices Disallow: /enc/enc/small-povaryata Disallow: /actions/akcia/kooza Disallow: /shops/chester Disallow: /eat/kebabhouse Disallow: /actions/akcia/doubling-bonuses-glam Disallow: /actions Disallow: /actions/akcia/only-modest-prices Disallow: /m/ Disallow: /mobile/ Disallow: /payment_invoice/ Disallow: /fundraising_donation/ User-agent: AhrefsBot Disallow: / User-agent: AspiegelBot Disallow: / # Brand-monitoring / backlink-SEO scrapers: no traffic benefit to any room, and among the most # expensive per request we serve (~0.25 s of page generation each). User-agent: AwarioBot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: DotBot Disallow: / User-agent: MJ12bot Disallow: / User-agent: SemrushBot Disallow: / User-agent: ZoominfoBot Disallow: /