AI Bots Are Changing. And We Have to Change With Them.

[gtranslate]

AI services today do not send only one type of bot to websites. The same provider can run a crawler for collecting data, a bot for AI search, and a client that opens a page directly because a user asked for it. This is where the internet is starting to change.

In the past, it was quite easy to tell the difference between a human visitor, a traditional search engine, and a harmful bot. Today, a new and fast-growing group of traffic is appearing between them – AI bots.

And as quickly as they change, WEDOS Global Protection has to change too.

That is why we have recently made another major step. We have started not only to identify known AI bots, but also to divide them by what they actually do.

GPTBot is not ChatGPT-User

OpenAI is a good example.

One company can generate several very different types of traffic.

GPTBot can systematically crawl large amounts of content.

OAI-SearchBot collects data for AI search.

ChatGPT-User can appear when a specific user asks ChatGPT to open a certain page.

From the server’s point of view, these are automated HTTP requests in all cases.

But their purpose is very different.

That is why we now divide AI traffic into groups such as these:

TypeWhat it does
trainingsystematically collects content for work with AI models
search_indexcrawls websites for AI search and creating answers
on_demandthe request is created directly by an action of an AI assistant user

This division is not only for statistics. It allows us to treat different types of traffic in different ways.

Data is valuable. And everyone wants it as fast as possible

Operators of AI systems need data. A lot of data. And ideally, they want it as fast as possible.

A crawler does not have to behave like a normal visitor who opens a page, reads it for a while, and then clicks somewhere else. It can visit hundreds or thousands of URLs in a short time.

Of course, this does not automatically mean an attack. But from the server’s point of view, the result can look very similar. Even a legitimate bot can create enough load to affect normal visitors. We have already seen cases like this in the past.

That is why we do not want to simply mark known AI bots as trusted and allow them without limits.

We need to know who they are, what they do, and when necessary, limit their traffic before it starts causing problems for the customer.

One user request does not always mean one HTTP request

Requests created directly by users of AI services are even more interesting. For example, a user asks a question and expects an answer within a few seconds. To do this, the AI service may need to get information from several pages or download several resources from one page at the same time.

One user request therefore does not always mean one HTTP request. It can create several requests, sometimes even dozens of them almost at the same time.

It also depends on how the cache of a specific AI service is designed. Not every AI system stores content it has already downloaded and uses it again for the next request. Some traffic can therefore be repeated.

And of course, the user on the other side does not want to wait. This creates pressure to collect data in parallel as quickly as possible. From the AI service’s point of view, this is logical. From the web server’s point of view, however, it can look like very aggressive automated traffic.

A known AI bot does not mean unlimited access

That is why this information:

this IP belongs to a known AI provider

does not mean:

bypass all protections

It only means that we know more about the source of the traffic.

The whole principle can be simplified into three steps:

verification → classification → action

First, we want to know who is sending the request. Then we want to know what the bot is used for. Only after that do we decide how to handle it. The result can be normal access, a rate limit, verification using Proof of Work, or a complete block.

The important point is that the identity of a bot and its permissions are not the same thing.

Why dividing AI bots is important

A systematic crawler that collects large amounts of content is different from a request created by a specific user of an AI assistant. If a crawler starts downloading too aggressively, we can limit its traffic.

With on_demand traffic, however, we know that a real person may be waiting behind the request, using an AI service and expecting an answer. This is why we do not want to put all AI bots into one list and apply the same rule to all of them.

We need to be able to say:

we know you
we know what you do
but if you overdo it, we will slow you down

This is an important difference compared to a traditional whitelist.

AI traffic will continue to grow

Today, we deal with crawlers, AI search, and requests created by users.

The next step will increasingly be AI agents that browse websites for users, compare offers, collect information, or communicate directly with web services.

The simple division into:

human / bot

will therefore not be enough in the long term. And soon, this will not be enough either:

good bot / bad bot

We need to know who is coming, why they are coming, how they behave, and how much load they create. This is the direction in which we are now developing WEDOS Global Protection. AI bots are not automatically a problem. But they are also not traffic that should get unlimited access without any checks.

And as AI systems crawl more and more of the internet, it will become even more important to understand this difference.