Attack of the Bots! AI Harvesters and Libraries Presentation by: Kate Dohe Director, Digital Programs & Initiatives University of Maryland Libraries 2025.10.28 The economic incentives for AI are leading to more crawlers on the web than ever before, harvesting vastly more content from websites ranging from independent blogs to major e-commerce platforms. These crawlers frequently ignore best practices for web harvesting, like following instructions in robots.txt or abiding by rate limits. IT Begins...With The Crawl Generative AI requires massive amounts of training data, often sourced from the open web with crawlers. This technology is decades old, and web crawlers for search engines are an established part of web traffic. Bad Robots... SWARM! AI bot operators often send large networks of bots to a single web property to download as much content as quickly as they can, which puts substantial stress on a server’s resource usage. SNEAK! Because website administrators can block or filter self-identified bots, many AI bots deliberately impersonate human traffic. This makes it difficult to ban them or control their traffic without affecting regular site visitors. Swamp! The large volume of bots and these aggressive tactics are effectively the same as a distributed denial-of-service (DDoS) attack, and take websites offline by overwhelming the server with content requests. In the Library... Libraries produce large amounts of high-quality open digital content in our digital collections, institutional repositories, research guides, catalogs, and archive management systems. This makes us a tempting target. Library technology teams tend to be lean, even in large research universities. It doesn’t take much to put our systems and people underwater. With The Killer Bots Many of our systems are designed for searching, refining, and linking to large amounts of content. Bots exploit this by generating unusual search queries to send to our systems, and then follow infinite filter permutations and links to more content, consuming resources for hours. They Came in the Night! AN attack at UMD Libraries Continent Unique visitors Asia 3455 North America 302 South America 145 Europe 73 Africa 10 Central America 5 Oceania 4 At the Witching Hour... Our Digital Collections site was targeted by bots in the early hours of October 4 and continuing through the 6 , mostly originating from Asia, as indicated by our web analytics and server monitoring systems. th th Bad Requests During events like this one, we see a significant drop in “successful” HTTP requests. This happens for two reasons: 1. The Bots request large amounts of bad links, returning error statuses. 2. As the bots consume more of our bandwidth and system resources, legitimate users experience timeouts and slow response rates from the application. Invasion! https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:trombone006::6 trombone&f[1]=instrumentation:alto e-flat saxophone&f[2]=instrumentation:clarinet&f[3]=instrumentation:flute&f[4]=instrumentation:oboe&f[5]=instrumentation:percussion &f[6]=instrumentation:trombone https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:e-flat clarinet001::1 e-flat clarinet&f[1]=instrument_count:eng- hrn001::1 eng-hrn&f[2]=instrument_count:flute001::1 flute&f[3]=instrumentation:clarinet&f[4]=instrumentation:e-flat clarinet&f[5]=instrumentation:horn&f[6]=instrumentation:oboe&f[7]=instrumentation:trombone https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:clarinet002::2 clarinet&f[1]=instrument_count:violin002::2 violin&f[2]=instrumentation:b-flat saxophone&f[3]=instrumentation:bassoon&f[4]=instrumentation:clarinet&f[5]=instrumentation:cornet&f[6]=instrumentation:flute https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:baritone001::1 baritone&f[1]=instrument_count:clarinet003::3 clarinet&f[2]=instrument_count:flute001::1 flute&f[3]=instrumentation:b-flat clarinet&f[4]=instrumentation:cor- bb&f[5]=instrumentation:xylophone&f[6]=larger_categories:Band https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:clarinet012::12 clarinet&f[1]=instrument_count:enghrn001::1 enghrn&f[2]=instrument_count:flute008::8 flute&f[3]=instrument_count:sax-ten-bb002::2 sax-ten- bb&f[4]=instrument_count:trombone006::6 trombone&f[5]=instrumentation:flute&f[6]=instrumentation:sax-bari-eb https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:cl-alt-eb001::1 cl-alt-eb&f[1]=instrument_count:e-flat clarinet001::1 e-flat clarinet&f[2]=instrument_count:sax-ten-bb001::1 sax-ten-bb&f[3]=instrument_count:trombone003::3 trombone&f[4]=instrumentation:alto e-flat saxophone&f[5]=instrumentation:cl-alt-eb&f[6]=instrumentation:snare drum https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:b-flat clarinet001::1 b-flat clarinet&f[1]=instrument_count:baritone e-flat saxophone001::1 baritone e-flat saxophone&f[2]=instrument_count:clarinet003::3 clarinet&f[3]=instrument_count:oboe002::2 oboe&f[4]=instrumentation:clarinet&f[5]=instrumentation:trombone https://digital.lib.umd.edu/scores/search?f[0]=instrument_count:b-flat contrabass clarinet001::1 b-flat contrabass clarinet&f[1]=instrument_count:baritone001::1 baritone&f[2]=instrument_count:timpani001::1 timpani&f[3]=instrumentation:clarinet&f[4]=instrumentation:sax-ten- bb&f[5]=instrumentation:timpani&f[6]=instrumentation:trombone&f[7]=instrumentation:trumpet&f[8]=instrumentation:tuba https://digital.lib.umd.edu/scores/search?f[0]=collections:004::American Bandmasters Association (ABA) Score Collection&f[1]=instrument_count:eng-hrn001::1 eng-hrn&f[2]=instrumentation:cor- bb&f[3]=instrumentation:flute&f[4]=instrumentation:oboe&f[5]=instrumentation:percussion Meanwhile, we start to see thousands of requests like these in our web analytics: direct traffic to search queries of a specialized musical scores collection, with infinite permutations of the collection’s search filters, seconds apart from each other. Organic, human traffic does not look like this. https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:trombone006::6 https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:e-flat https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:clarinet002::2 https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:baritone001::1 https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:clarinet012::12 https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:cl-alt-eb001::1 https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:b-flat https://digital.lib.umd.edu/scores/search?f%5B0%5D=instrument_count:b-flat https://digital.lib.umd.edu/scores/search?f%5B0%5D=collections:004::American No Silver Bullet Responding to Bot Attacks >> Large-scale IP address blocking >> Humanity checkers and CAPTCHAs >> Enhanced web application firewalls >> Honeypots and bot traps >> Authentication and whitelisting users >> Withdrawing content Bleeding US Dry These events are expensive. Running the bot-specific services necessary to keep our systems online costs us thousands each year. IT staff members are diverted--often outside normal business hours--to work on these until service is restored. Many libraries may have no choice but to restrict open access to our content. We don’t want to do this, but until harvesting traffic dissipates or significant resources are allocated to our infrastructure, this may be our only option to stay online. What The Bots Cost Libraries Our patrons experience frustrating slowdowns and timeouts on our sites, meaning they cannot conduct their research or access digital materials reliably. This pushes more complaints to our front-line access services teams to troubleshoot. Surviving to the end! We want to share our content! At UMD Libraries, many of our services have APIs to programmatically query and retrieve our data. We have an Open Data resource site (https://opendata.lib.umd.edu/) for such uses. Other research libraries often do the same. If you must use bots to harvest web content, please respect the rules - follow the robots.txt instructions, limit your request rate, use published sitemaps instead of searches. Being a Harvesting hero If you want collections data and resources that may not be available via API, contact the library directly first. If it is within our power to share the data, most of us are happy to work with researchers on a solution. https://opendata.lib.umd.edu/ https://opendata.lib.umd.edu/ More Scary Stories Weinberg, Michael. “Are AI Bots Knocking Cultural Heritage Offline?” GLAM-E Lab, June 2025. https://glamelab.org/products/are-ai-bots-knocking-cultural-heritage-offline/. Casden, Jason, David Romani, Tim Shearer, and Jeff Campbell. “Mitigating Aggressive Crawler Traffic in the Age of Generative AI: A Collaborative Approach from the University of North Carolina at Chapel Hill Libraries.” The Code4Lib Journal, no. 61 (October 2025). https://journal.code4lib.org/articles/18489. Wrobel, Tom. “Ethical Harvesting: The Impact of AI Model Training on Cultural Heritage Institutions.” 2025. https://www.canva.com/design/DAGykGqgByE/e10RPoQXSlwiMiF9A0zpkw/view. Dohe, Kate. “Guest Post - ‘Have You Proved You’re Human Today?’ Open Content and Web Harvesting in the AI Era.” The Scholarly Kitchen, October 7, 2025. https://scholarlykitchen.sspnet.org/2025/10/07/guest-post-have-you-proved-youre- human-today-open-content-and-web-harvesting-in-the-ai-era/. === Special thanks to Dan Bowling, Peter Eichman, and Jeremy Gottwig from UMD Libraries’ Software Systems, Development, and Research for their input and resource sharing for this presentation. Recommended Reading and Credits https://glamelab.org/products/are-ai-bots-knocking-cultural-heritage-offline/ https://journal.code4lib.org/articles/18489 https://www.canva.com/design/DAGykGqgByE/e10RPoQXSlwiMiF9A0zpkw/view https://scholarlykitchen.sspnet.org/2025/10/07/guest-post-have-you-proved-youre-human-today-open-content-and-web-harvesting-in-the-ai-era/ https://scholarlykitchen.sspnet.org/2025/10/07/guest-post-have-you-proved-youre-human-today-open-content-and-web-harvesting-in-the-ai-era/