Consumed at Scale: AI-Driven Extraction and the Political Economy of OpenStreetMap
Large-scale scraping attacks on OpenStreetMap's infrastructure have exposed a fundamental tension between open geodata as a community-maintained commons and its role in current commercial AI development dynamics. This presentation examines the implications of AI-driven data extraction for OSM’s infrastructure and licence compliance guidelines as well as its consequences for OSM’s position within the geospatial data ecosystem.
Since its founding in 2004, OSM has navigated recurring tensions between voluntary contribution and corporate participation [1], however, the recent surge in AI-driven large-scale scraping of OSM infrastructure represents a qualitative escalation of this dynamic. In late January 2026, an OSM-affiliated account publicly appealed to journalists to investigate coordinated scraping activity; according to the responsible engineer Grant Slater, a single week saw 100,000 IP addresses querying OSM servers simultaneously, each making only a few requests [2, 3]. The OSM case is not isolated: in February 2026, a volunteer-maintained mapping project called Vaguely Rude Places saw daily requests jump from the low thousands to the hundreds of thousands, exhausting its monthly tile allocation in a single day and forcing the map offline. This was not due to technical failure, but because AI crawler demand had crossed a cost threshold the system was never designed to accommodate [4, 5]. These incidents illustrate how AI-/automated scraping translates directly into financial and operational costs for projects that were built around human-scale usage patterns.
OSM's appeal for AI is substantial and growing. The dataset combines near-global coverage, fine-grained semantic tagging (place types, road classifications, building attributes, land use categories), a rich versioned edit history, and free availability under an open licence. For general-purpose language and multimodal models, OSM data can ground spatial analysis by linking named places to structured geographic attributes. Recent research has demonstrated a 70% improvement in geospatial prediction performance on certain tasks when LLMs are augmented with OSM auxiliary data [6]. For specialised geospatial AI (GeoAI), OSM geometries can serve directly as labelled training data for object detection, routing, and location intelligence systems, with documented applications in urban planning, hazard prevention, climate modelling, and traffic monitoring. OSM, in short, functions as a uniquely accessible and semantically rich proxy for human spatial knowledge and is increasingly being treated as such by commercial AI developers, without the community's consent and without meaningful reciprocity.
The precise mechanisms behind these access patterns remain unclear. OSM's infrastructure was built to meet the behavioural patterns of human users; AI-driven consumption accelerates and changes these patterns in ways the infrastructure was not designed to absorb [5]. Slater has described OSM as being “at constant war” with scrapers, noting that traffic frequently arrives through residential proxy networks that make it “hard to distinguish normal users from automated collection” [7]. While the origins and intentions behind these access patterns remain largely opaque, they may in part be attributed to user queries directed at AI chatbots and assistants that generate maps or location-based outputs by drawing on openly accessible geodata in real time, a question that is increasingly discussed within online open-source communities and academia.
This paper examines the implications of AI-driven data extraction for OSM as an open-source project, focusing on two interrelated dimensions: infrastructural strain and licence compliance. It situates these dynamics within the broader political economy of open data and AI development, and discusses emerging governance concerns at the community, institutional, and regulatory level. We draw on a combination of document and content analysis, and reviews of relevant technical and legal literature. Primary sources include OSM Foundation communications, community forum discussions, and reports from the 2026 scraping events. These are contextualised through secondary literature on geodata and its intersection with digital commons [8, 9, 10, 1], and digital capitalism [11, 12]. The analysis is qualitative and interpretive, situated within a critical geography and science and technology studies (STS) framework [13], and deliberately foregrounds the political-economic and governance dimensions of these dynamics [14].
Two central dimensions structure our contribution. First, on the infrastructural effects: coordinated scraping exploits API infrastructure built and maintained largely by volunteers, generating financial and labour costs that fall disproportionately on the community. When volunteer capacity is consumed by bot-detection and infrastructure triage, the risk of contributor burnout increases. Simultaneously, the origins of these requests are structurally difficult to determine, as the technical mechanisms applied render automated and human traffic effectively indistinguishable.
Second, on licence compliance: the Open Database License (ODbL) governing OSM data imposes conditions on derivative databases, including requirements for open publication of training datasets constituting substantial extractions and for attribution in model documentation [15]. Whether these obligations are systematically met by commercial AI developers remains largely unverified and unenforced at this point. Notably, the OSMF's own attribution guidelines acknowledge that sophisticated models can be “overtrained” to the point of recreating their training set, in which case the model output itself may constitute a derivative database subject to ODbL terms.
Taken together, these dynamics point to a qualitative shift in OSM’s position within the geospatial data ecosystem: from a collaborative infrastructure for mapping the world to an involuntary resource for AI systems that seek to model it. This shift reflects a deeper tension between the logic of open data, which presupposes a community of reciprocal contributors, and the logic of commercial AI development, which tends to treat open data as an input to be optimised rather than a commons to be maintained. The resulting risk is not enclosure in the classical sense of a commons being directly fenced off or privatised by external actors. Rather, it lies in a more paradoxical dynamic: large-scale, largely unaccountable extraction by commercial AI systems may push open mapping communities towards defensive forms of closure, such as more restrictive access controls, authentication requirements, or gated data infrastructures. Such protective enclosure would be a response to appropriation rather than its initial form, but it could nevertheless alter OSM’s character as a digital commons and open-source project. We therefore identify these dynamics as a new and qualitatively distinct form of commercial pressure on OSM — one that operates not through direct participation in the project, but through large-scale consumption of its resources — and raise questions about how community-maintained infrastructures can adapt to AI-driven consumption patterns, how enforceable attribution mechanisms for OSM-derived data might be developed, and what meaningful reciprocity between commercial AI companies and the open mapping community could look like in practice.
Hannah Boettcher is a doctoral researcher and research associate at the Institute of Geography at Friedrich-Alexander-University Erlangen-Nürnberg. Their research is situated in the field of digital geography, combining political- and social-geographical perspectives on digital spatial data with computational and qualitative methods. Hannah Boettcher is a founding and editorial member of the journal GRID (https://gridisnotajournal.de/), and their work includes social- and political-geographical research on OpenStreetMap as well as the social-scientific and discourse-analytical use of GIS-based methods. Methodologically, they combine qualitative approaches with computational text and data analysis, including natural language processing (NLP) and topic modelling, in order to investigate how digital spatial data and technologies reshape knowledge production and spatial governance.