Skip to content

The Data Scientist

AI Crawlers Now Match Googlebot in Traffic: What Your Website Analytics Are Really Telling You 

AI Crawlers Now Match Googlebot in Traffic: What Your Website Analytics Are Really Telling You 

Introduction

Websites receive traffic in different ways, and many owners are unfamiliar with these methods. For a long time Google and other search engines had dominant bots for web crawling, but the development of AI and with it AI web crawling, has created a new generation of web crawlers that are already rivaling search engines in the traffic they receive.

An analysis of over 240,000 bot visits during a 30-day period reveals an interesting trend. AI crawlers now match Googlebot in traffic, highlighting how AI-driven bots are becoming a consistent presence in website crawl activity rather than an occasional one. This indicates that AI crawlers represent as much bot traffic as classic search engines.

The increase in AI crawlers may affect content creators, web app developers, marketers, and any digital professional who needs to understand and work around web traffic.

AI Crawlers and Search Engines Are in a Dead Heat

For this data collection, 240,060 visits from 24 known bots were recorded. The visits spanned over 5,000 unique URLs. The bot visits were grouped into 5 categories:

  • AI crawlers
  • Search engines
  • SEO tools
  • Social Media Preview Bots
  • Archive Bots

One among the many interesting findings of this study is the fact that AI crawlers and search engine crawlers both accounted for exactly 35% of total bot traffic.

The remainder of the traffic was captured by:

  • SEO Tools: 21.2%
  • Social Preview Bots: 8.6%
  • Archive Services: 0.2%

Which Bots Crawl Websites the Most?

Bots behave differently. Some bots will only crawl newly indexed pages while others will crawl the same page over and over.

During the 30 day study the following crawlers were the most active:

Googlebot – 36,840 hits

Bingbot – 35,610 hits

AmazonBot – 33,040 hits

MajesticBot – 31,860 hits

ChatGPT-User – 24,350 hits

ClaudeBot – 17,430 hits

AhrefsBot – 14,550 hits

LinkedInBot – 14,340 hits

Over the 30 days, Googlebot was the most active bot with an average of 1,200 hits per day. AmazonBot and ChatGPT-User also represent how web data collection is aggressively pursued by AI companies.

The Most Efficient Crawlers

Heavy traffic can also mean inefficient crawling.

A unique metric was hits per unique path. The bots that ranked the best for this metric were the bots that visited a page once rather than requesting the same content over and over.

The best crawlers included:

AhrefsBot

Applebot

MajesticBot

SemrushBot

Bingbot

While bots like ClaudeBot and Googlebot make frequent requests, they do so to find updates in the documents they keep and update their lists as frequent as possible.

The Bots that Never Leave

Some crawlers behave quite differently.

Rather than exploring thousands of URLs, some crawlers will make repeated requests for the same limited set of pages.

Notable examples include:

LinkedInBot

FacebookBot

ChatGPT-User

LinkedInBot, for instance, made almost 1,000 hits for each URL. This shows it crawls to refresh the metadata it uses to generate content previews. Likewise, ChatGPT-User made repeated requests to the same limited set of URLs. This behavior shows how AI crawlers update or verify the information that is used for content generation.

Which Bot Goes Crawling the Most?

Undoubtedly, as far as the total number of unique URLs crawled, MajesticBot was the most prolific crawler. It crawled more than 2,000 unique pages during the assessment. Crawlers in order of coverage:FFF

  • Bingbot
  • Googlebot
  • AmazonBot
  • AhrefsBot

These crawlers prefer broader coverage of the crawled sites over repeated coverage of certain pages.

For those in the SEO industry, this emphasizes the idea that backlink analysis tools can help discover new pages quickly since those tools crawl for new pages in the internet at a high frequency.

Rare But Valuable Visitors

There were a number of bots that were only infrequent visitors to the crawled sites during the month. They included, but were not limited to:

  • Internet Archive
  • Baiduspider
  • DuckDuckBot
  • Twitterbot
  • Pinterestbot
  • Google-Extended (Gemini)
  • Screaming Frog

While these bots were not frequent visitors, each of these bots has a specific purpose, from aiding in the web history preservation process, to social media previews and inclusion to specialized indexing.

For those who require visibility in more niche or international markets, these bots can be more important.

Importance for Website Traffic

For website owners, the equal distribution of search engine bots and AI crawlers is more than just an interesting statistic. It reflects a fundamental shift in how information is shared and consumed on the internet.

In the past, websites existed primarily to publish information for people and search engines. Today, many publishers are recognising that Web Content is Now Just Free Training Data for AI companies, which continuously collect publicly available information to improve their models and AI-powered tools.

Publishings that are freely available online have provided a really rich resource for AI companies. We can see this trend in traffic data. As certain AI systems known as ‘crawlers’ continue to access more data from websites, more and more people who post online are beginning to worry how their post data is being utilized and whether they should be entitled to more control.

What Publishers Need to Think About

AI crawling will lead publishers away from analyzing only activities from Googlebot, and more toward server logs and analytics on bots crawling your site.

Analyzing bot activity can facilitate answers for questions including:

Which companies’ bots are accessing our site?

How often do they access our pages?

Is there a restriction in place for how often they may crawl our site, and are they adhering to it?

Are there bots tha should be blocked via our robots.txt file?

How quickly does bandwidth become exhausted?

Recognizing these factors will help publishers make contact decissions to what extent they will need to scale infrastructure and what will be the methods to secure accessibility to their content and data policies.

The Coming Age of Web Crawling

Content AI is changing the web beyond search engines being the primary consumers of web content.

AI is crawling the web at the same scale as traditional search engines, changing the web and how we think about web traffic.

Even though Googlebot is still the most widely used bot, collectively AI bots are at the same level of activity, and the AI traffic will continue to increase due to its popularity.

The age of web crawling by AI is rapidly approaching. For business owners, web developers, and web marketers, understanding crawlers and the bots accessing your content will become a critical practice in managing your site.