Rattlesnakes By Mail

Bingbot vs Google-Extended

Bingbot and Google-Extended compared on 13 documented fields, from 19 claims, last verified 2026-09-14.

FieldBingbotGoogle-ExtendedComparison
cites_sourcesyesnot documenteddifferent
follows_noarchiveyesnot documenteddifferent
honours_crawl_delayyesn/adifferent
opt_out_mechanismyesnot documenteddifferent
publishes_ip_listyesnodifferent
purpose_documentedyespolicy_onlydifferent
reads_llms_txtnot_documentednot documenteddifferent
respects_robots_txtyesn/adifferent
robots_token_differs_from_uanot documentedyesdifferent
separate_search_and_training_tokensnoyesdifferent
user_agent_string_fullMozilla/5.0 (compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm)not documenteddifferent
verifiable_by_rdnsyesn/adifferent
what_it_controlsnot documentedGemini model training and grounding, not Search inclusiondifferent

Differences

FieldBingbot statementGoogle-Extended statement
cites_sourcesCopilot Search in Bing delivers clearly cited sources alongside generative answers to support publishers and content owners.not documented
follows_noarchiveBingbot excludes noarchive-tagged content from Copilot and Chat answers and from generative AI training data.not documented
honours_crawl_delayMicrosoft states that a crawl-delay directive in robots.txt always takes precedence over Bing Webmaster Tools crawl control settings for Bingbot.Google does not publish any Crawl-delay behaviour for the Google-Extended token.
opt_out_mechanismMicrosoft lets publishers block Bingbot content from Copilot answers and AI training using noarchive and nocache tags.not documented
publishes_ip_listMicrosoft publishes Bingbot's IP address ranges as a JSON file at bing.com/toolbox/bingbot.json for verification purposes.Google publishes no IP list for Google-Extended because Google-Extended has no separate HTTP request user agent string.
purpose_documentedBingbot serves as Microsoft's standard crawler and handles most of Bing's daily web crawling.Google-Extended is a standalone product token that controls whether Google uses crawled content to train Gemini models.
reads_llms_txtMicrosoft does not publish whether Bingbot reads llms.txt files.not documented
respects_robots_txtBingbot honors robots.txt directives and other Microsoft-supported control mechanisms that content owners set for crawling.Google-Extended is a robots.txt token that publishers set to control use of crawled content for Gemini training.
robots_token_differs_from_uanot documentedGoogle-Extended is a robots.txt control token and does not correspond to a crawler user agent string.
separate_search_and_training_tokensMicrosoft controls Bingbot's use in AI training through meta tags rather than publishing a separate training-specific crawler token.Google-Extended controls Gemini training use only and does not affect inclusion in Google Search.
user_agent_string_fullBingbot sends a user agent string containing bingbot/2.0 and a link to Microsoft's bingbot.htm page.not documented
verifiable_by_rdnsMicrosoft verifies Bingbot traffic through reverse DNS lookups that resolve to hostnames ending in search.msn.com.Google does not publish a reverse-DNS verification method for Google-Extended.
what_it_controlsnot documentedGoogle-Extended controls AI training and grounding in Google systems and does not control Google Search inclusion.

Sources