Advanced bot filtering

Find out whether that traffic spike was real.

A crawler can make a quiet site look busy, and most analytics tools either hide what they removed or never noticed it. StepMetrics shows you the difference between a visit proven human, a visit proven automated, and a visit nobody can vouch for.

The crawler we found

Over fourteen days in August 2026 a crawler walked one of our sites by following its own links, ten to twenty seconds apart, from a rotating residential proxy pool, sending an ordinary Chrome on Windows user agent. Because the exit address changed between every hop, and the visitor identity takes the address as an input, each hop became a separate visitor in a separate country. Every one was counted as a person.

On that site, 77 of 104 visits counted as human were this one crawler.

Nothing about it looked unusual in a report. That is the point: the traffic you most need filtered is the traffic that does not announce itself.

Engagement signals

The tracker used to send one beat per pageview and nothing after. So a one-page visit produced an identical record whether somebody read it for four minutes or a script fetched it and left. No rule stacked on top of that silence could tell them apart, and a rule that guesses on silence is how real readers get deleted.

So we collected the missing observation instead. The browser now reports when a visit is actually interacted with, which a fetch-and-move-on crawler does not produce. One of the tempting rules we tested first, treating a visit that arrives from our own pages as a rotating proxy, would have flagged a site's own operator returning from a verification email four days running. It was thrown out.

Bot, engaged or unverified

Most tools report traffic as human or bot, which forces every unknown visit into the human column and quietly inflates the number you care about. StepMetrics keeps the unknown visible.

Unverified is an honest answer, not a hedge. It means we do not have evidence either way, and you can see the traffic and judge it yourself.

Auditable exclusions

A visit is only marked as a bot on something it did or declared: it named itself a crawler, the browser reported that it was being driven by automation, it navigated faster than a person operates a browser, or it moved as part of a group that behaved like a machine pool. The reason is stored with the visit, and a bot without a stated reason cannot exist in the database because the schema refuses it.

Filtered traffic is kept for the same retention period as everything else, so you can look at what was removed and disagree with it.

Avoiding false positives

The rules are deliberately precision-first. A crawler left in unverified is on your screen, where you can see it and judge it. A real reader wrongly filed as a bot is gone from your numbers and invisible, and you will never know to look.

That is why there is no rule keyed on a single pageview, a missing referrer, a zero duration, or a burst of countries. The first three describe an ordinary bounce exactly as well as a crawler. The fourth depends entirely on your traffic: three visits from three countries in two minutes is a crawl on a site with seven visits a day and a normal Tuesday on a site with seven thousand.

Slow crawlers

A scraper in a hurry is easy. A crawler that goes slowly, a handful of visits a day from a different address each time, walks through any time-based rule. What survives however slowly it moves is not timing, it is ownership: those addresses belong to networks that rent compute. Measured across four days on one site:

NetworkVisitsAddressesBrowser fingerprintsEngagedMost pages read
A compute provider4847101
A second compute provider66101
A consumer VPN65229

Forty-seven addresses used once each is not forty-seven people. The third row is the one that matters: those are real readers behind a consumer VPN, and they are on a datacenter network too. They are not filtered, because more than one fingerprint appeared, somebody engaged, and somebody read a second page. The network only decides which groups are eligible to be judged. The behaviour still has to convict them.

Crawlers that come back

Once a pool of machines has been caught, it stays caught. The same crawler that scraped one of our sites came back to two others about once a week each, too rarely for any pattern to show on either. Each visit carried the exact browser string of the pool already caught, from the same rental network, and did nothing a person does. A visit that matches like that is filtered on every site, while any scroll, tap or second page still spares it.

Configurable network list

Which hosting networks count is configuration on your account, not something frozen into a release. If your real audience sits behind one particular cloud, you can remove it. The default list is deliberately cautious and leaves out the big consumer edge networks entirely, because those carry ordinary people through privacy relays and office connections. An incomplete list misses crawlers; a list with a consumer network in it deletes readers. Only one of those mistakes can be undone.

Limits

No filter catches everything, and we would rather say so than imply a clean number. Unverified means we do not have enough evidence either way, so some crawlers sit there. This is not ad-fraud protection and it does not block anything; it decides what counts in your reports.

How a visit gets classified

Try it on your site.

Add your first site for free. See where visitors come from, what they look at, and how many take the next step.

No credit card required.