From Text to Terrain: Automated Geolocation Analysis of Dark Web Data
How to turn noisy dark web text into actionable geospatial intelligence - fast, lightweight and at scale?
This talk introduces a lightweight NLP pipeline that can be combined with CIRCL's AIL Framework to uncover hidden geographic patterns in unstructured dark web data. Combining fastopic filtering techniques, created at this year's hackathon.lu, with RueBERTA, a RoBERTa model trained on weakly labeled data, the system extracts and parses location-related entities ranging from cities and POIs to full addresses. It can even catch your secret PO-Box!
The result: an interactive map that transforms chaotic underground text into intuitive, explorable geographic insights.
This project presents a NLP-based geolocation pipeline, whose development started at Hackathon.lu. It is designed to work on dark web data from the AIL Framework and other noisy text data formats. The goal of the project is to transform unstructured textual data into actionable geospatial intelligence by automatically extracting, normalizing and visualizing real-world locations hidden in underground datasets.
The pipeline combines several processing stages. First, a lightweight keyword-based filtering method called fastopic performs coarse-grained detection of location- and address-related content. This significantly reduces the search space and provides first location metadata.
The filtered text samples are then processed using RueBERTA, a RoBERTa-based model trained on weakly labeled data, to extract fine-grained geographic entities such as cities, points of interest, full and partial addresses, postal boxes etc.. Extracted entities are normalized through address parsing and enrichment steps before being linked against external geolocation databases such as Nominatim.
To improve reliability and reduce ambiguity, we compute confidence scores by combining contextual information from the fastopic filtering stage with the responses returned by the geolocation services. The final results are displayed through an interactive map interface, allowing analysts to visually explore geographic patterns within dark web datasets.
Coming from a bachelors in linguistics and driven by a long-standing enthusiasm technology and IT that started at kindergarden age, Laura Bernardy combined her passion for language and computer science through a Master’s degree in Computational Linguistics.
Her early research focused on NLP for Luxembourgish and other low-resource languages, where limited data availability and noisy texts pose significant technical challenges. Over time, this interest expanded towards another underexplored low-resource domain: dark web data. Similar to low-resource linguistic settings, underground data is highly noisy, fragmented, multilingual, and difficult to process – while additionally introducing challenges related to cybersecurity and threat intelligence.
Today, Laura is pursuing a PhD at the SnT Luxembourg, focusing on the intersection of NLP, dark web analysis, and CTI research. Her work explores how modern language technologies can be applied to extract actionable intelligence from unstructured underground data while remaining practical.