Problems with web crawler/bot swarms (hosted sites)

We are hosted by III and we’ve experienced a significant uptick in our site being swarmed by web crawlers to the point that Leap (we are a Polaris library) becomes unusable. The reliable indicator is an “unable to connect to ERMS” error message when logging in to Leap. We’ve had three incidents in the last month. Innovative suggested setting up a geofence, but when I asked for the problem IP addresses, only one of the three incidents involved IP addresses from outside the U.S.–the others were Google and OpenAI/GPTBot. Are other hosted libraries experiencing this problem lately? If so, has anyone gotten Polaris to set up anything for their server that has helped?

What is the domain/name URL of the hosted sites? I don’t need the full thing; I’m just wondering if it is in your library’s domain or not. If it is in your domain, you might have some additional options available to you.

Regardless, it seems like we’re hearing this enough that it should probably be something that the @steering-committee should bring up (they probably already have but this gives them more fuel) with Clarivate to address to make sure they are addressing it in a more systematic way.

Hi Emma, We here at GMILCS are having similar issues, though ours seem to happen almost daily. Luckily, things seem to resolve themselves, requiring only a refresh or reboot at the moment.

Stark Library (a Sierra library also hosted by Clarivate) has had at least two severe week+ long bot incidents, the first occurring in early Feb 2025 and another a couple of months ago.  At IUG 2026, I learned that Clarivate was implementing Cloudflare on its hosted systems this summer.  Does anyone know the status of that rollout?

1 Like

All these replies are interesting, thank you!

Bob, I’m sorry to hear Stark’s incidents have lasted so long. Has III not been able to identify and block the originating IPs?

I’m glad to hear Clarivate is planning to implement Cloudflare. I too hope for a status update on that. I put in a general “what are you doing about bots” ticket and got a very generic reply just today that didn’t tell me much:

“Bot and crawlers are being addressed incrementally to avoid blocking patron or otherwise ‘good’ traffic. If traffic is being blocked wholesale/in bulk, it’s much more likely library or patron IP’s will be blocked in error so we are hyper-aware and sensitive to what we’re blocking basically. We are adding additional entries to the robots file that hopefully should prevent or limit this activity as well.”

Hi Wes,

The domain will be something like libraryname.polarislibrary.com, ex. ours is irving.polarislibrary.com. Sadly not in our domain.

Yeah, then it will rely on Clarivate putting it behind a reverse proxy to protect like Cloudflare or bunny.net.

The fact they think the bots are going to respect robots.txt is WILD to me.

They’ve been playing a game of Whack-a-Mole, because none of the origin IPs are used for very long before the bad actor(s) switch to a different IP.  We use BiblioCommons for our discovery layer, website CMS, mobile app and events calendar, and at one point Clarivate got a little overzealous with the IP blocking and blocked requests coming from BiblioCommons, disabling the public catalog in the process.

If these miscreants/ne’er-do-wells/scofflaws have a profit motive, I’d like to understand what that is.  Is some incredibly well-heeled and possibly evil media retailer, most likely with a supervillain-style clean shaven head (none come to mind at the moment) attempting to drive consumers away from the free lending model of public libraries and toward retail?  Nah, that’s a crazy idea.

1 Like

We are a hosted Sierra system.

The last ticket w/support that I had open for this - 10634918 - was not due to any performance issues, but just me noticing when I go to pull the search stats from the Web Management Reports. Some months I see spikes in the number of searches, and then report to III. Usually I get a response back that they will look into it and block suspicious Ips (those doing high number of searches).

For this latest ticket, I worked with someone named Filip, and here is a recap of our discussion, after I asked for more details after getting this first response:

“I reviewed the search requests in the access logs, and it looks like most of the recent increase is coming from traffic outside the US.
Please let me know how you’d like us to proceed from here.”

After I asked for more details (they host us, I’m not expert!)…

"I understand that your patrons are using the Vega catalog, and not the webpac catalog based on our previous conversations. There are a couple ways we can address this to alleviate your concerns with the WebPac traffic.

First lets cover what is actually happening, and why the hosting center is not directly blocking the traffic. Normal crawler activity generally speaking is a single IP, or small series of IP addresses, sending large volume of requests to the webpac catalog. The issue you are reporting is due to a large volume of IP addresses sending a small amount of requests. This is a new traffic pattern being used to get around normal crawler mitigation tools.
The two ways I can think of to handle this would be to either block based on geo location, or block based on the actual URL used. The down side to blocking based on geo location is if the user is abroad they will not be able to log into their account. They would still have normal access to vega though, as this block only affects the application server. The other option would be to block the search url for example: Fauquier County Public Library / All Locations (which I pulled from the log as an actual example). Blocking based on search will not impact a patron logging into their account, it would only affect someone accessing a search page on the webpac.

My suggestion would depend on your expected workflow for your patrons. If the expected workflow is strictly to use Vega, then blocking the searches will have the least impact on your patrons. Please review the provided information and let me know how you would like to proceed. Let me know if any of the provided information need clarification to determine next steps."

I then asked what the blocking experience would be for the end-user…

"If we implement geo-blocking, anyone trying to access Sierra/Search from outside the allowed countries would be blocked automatically. They would not be redirected to Vega. Access would only be allowed if the request is coming from one of the approved countries, or if their IP address has been specifically whitelisted.

What other customers choose to do really depends on their needs and the workflow they want to support for their patrons. Some may choose completely blocking the WebPAC searches, while others chose to keep searches allowed, and decide to go with geo blocking.

Based on the server resource utilization logs, they do not show any spikes or signs that the app server is being slowed down. At the moment, the main issue appears to be the increased number of searches in WebPAC, which may be affecting the accuracy of your search statistics."

I didn’t end up putting in any blocks at this point, since our system doesn’t seem to be impacted, just the search results are not reliable (and honestly, not sure if they ever truly were). However, being that we pay ~17K for the hosting part for Sierra, I’d hope that they are proactively monitoring. Anyway, Filip provided a lot more information than I normally get from support about crawling, which you may find of use. I just had to ask a few times to get more details.

Alison Pruntel
Manager, Technology & Materials
Fauquier County Public Library
11 Winchester Street
Warrenton, VA 20186
540-422-8515

https://fauquierlibrary.orghttps://fauquierlibrary.org/


From: Emma Olmstead-Rumsey webguru@innovativeusers.org
Sent: Thursday, July 16, 2026 12:43 PM
To: Pruntel,Alison Alison.Pruntel@fauquiercounty.gov
Subject: IUG Forums: Problems with web crawler/bot swarms (hosted sites)

CAUTION: This email originated from outside of the organization. Do not follow instructions, click links, or open attachments unless you know the content is safe.

Someone replied to a topic you are Watching.
[https://forum.innovativeusers.org/letter_avatar_proxy/v4/letter/e/b5a626/45.png]
Emma Olmstead-Rumseyhttps://forum.innovativeusers.org/u/eolmstead eolmstead
July 16

We are hosted by III and we’ve experienced a significant uptick in our site being swarmed by web crawlers to the point that Leap (we are a Polaris library) becomes unusable. The reliable indicator is an “unable to connect to ERMS” error message when logging in to Leap. We’ve had three incidents in the last month. Innovative suggested setting up a geofence, but when I asked for the problem IP addresses, only one of the three incidents involved IP addresses from outside the U.S.-the others were Google and OpenAI/GPTBot. Are other hosted libraries experiencing this problem lately? If so, has anyone gotten Polaris to set up anything for their server that has helped?


Visit Topichttps://forum.innovativeusers.org/t/problems-with-web-crawler-bot-swarms-hosted-sites/2992/1 or reply to this email to respond.

You are receiving this because you enabled mailing list mode.

Unsubscribe from these emails<>.

1 Like