BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.com//euroscipy-2026//speaker//3U8HLG
BEGIN:VTIMEZONE
TZID:Poland
BEGIN:DAYLIGHT
DTSTART:20250720T000000
TZNAME:CEST
TZOFFSETFROM:+0200
TZOFFSETTO:+0200
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251026T030000
RDATE:20261025T030000
TZNAME:CET
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260329T030000
RDATE:20270328T030000
TZNAME:CEST
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:Building a Scientific Taxonomy at Scale with Graph Clustering\, Em
 beddings\, and LLMs - Daniele Raimondi
DTSTART;TZID=Poland:20260720T121000
DTEND;TZID=Poland:20260720T123000
DTSTAMP:20260907T090220Z
UID:pretalx-euroscipy-2026-PRCPUX@pretalx.com
DESCRIPTION:Scientific organizations struggle to extract actionable insigh
 ts from publication data tagged with inconsistent and noisy keywords. Tran
 sforming hundreds of thousands of such keywords into a **110\,000+ concept
 **\, semantically consistent taxonomy\, and attaching them hierarchically 
 at scale\, requires more than ad-hoc normalization: it demands careful sys
 tem design.\n\nThis talk presents a production-grade pipeline that extends
  **OpenAlex's 4-level framework** (Domain → Field → Subfield → Topic
 ) with a granular **Concept layer**\, resulting in a **5-level scientific 
 taxonomy**. The system combines **SPECTER2 embeddings** to model semantic 
 similarity\, **Leiden graph clustering** to group 100K+ concepts\, and **Q
 drant** for efficient vector-based hierarchical attachment.\n\nA central c
 ontribution is a strategic\, multi-stage integration of **LLMs**. Rather t
 han using LLMs end-to-end\, we deploy them at **5 targeted points** where 
 semantic judgment matters most: concept granularity filtering\, field clas
 sification across **26 domains**\, cluster renaming\, explanation generati
 on and validation of topic assignments using multi-embedding comparisons. 
 Deterministic methods ensure scalability and reproducibility\, while LLMs 
 provide semantic precision where embeddings alone fall.\n\nThe resulting t
 axonomy is used in production to automatically tag **millions of publicati
 ons**\, enabling real-time trend detection and portfolio-level analytics t
 hat support strategic decision-making.
LOCATION:Room 1.19 (Ground Floor\, Shannon)
URL:https://pretalx.com/euroscipy-2026/talk/PRCPUX/
END:VEVENT
BEGIN:VEVENT
SUMMARY:Automating Scientific Paper Classification at Scale with Retrieval
 –Reranking and LLMs - Daniele Raimondi
DTSTART;TZID=Poland:20260721T152000
DTEND;TZID=Poland:20260721T154000
DTSTAMP:20260907T090220Z
UID:pretalx-euroscipy-2026-G7BJFJ@pretalx.com
DESCRIPTION:Scientific organizations struggle to extract actionable insigh
 ts from publication data tagged with inconsistent and noisy author keyword
 s. Automatically assigning papers to consistent\, semantically grounded co
 ncepts is essential for reliable trend detection\, search\, and analytics.
 \n\nThis talk presents a production-grade\, **two-stage classification pip
 eline** that tags hundreds of thousands of scientific papers against a **1
 10K+ concept taxonomy**. Given a fixed hierarchical taxonomy extending Ope
 nAlex's **4-level structure** with a granular concept layer\, the system c
 ombines **vector-based retrieval**\, **cross-encoder reranking**\, and tar
 geted **LLM validation** to achieve scalable and accurate paper classifica
 tion.\n\nIn **Stage 1 (Candidate Retrieval)**\, paper metadata (title\, ab
 stract\, author keywords) is embedded using **SPECTER2** and queried again
 st **Qdrant** to retrieve a small\, high-recall candidate set from over **
 110\,000 concepts**. In **Stage 2 (Reranking and Filtering)**\, **cross-en
 coder models** perform fine-grained semantic matching\, while **LLMs (Azur
 e OpenAI)** are selectively applied to resolve ambiguous cases and produce
  confidence-scored assignments.\n\nDeployed on **millions of publications*
 *\, the system standardizes noisy keywords and enriches paper metadata wit
 h semantically consistent concept tags\, enabling downstream analytics at 
 scale.
LOCATION:Room 1.38 (Ground Floor\, Turing)
URL:https://pretalx.com/euroscipy-2026/talk/G7BJFJ/
END:VEVENT
END:VCALENDAR
