Open Source Conference Luxembourg

Data That Travels: Cross-Domain Insights from our Open Source Journey in Education

Higher education runs on sprawling edtech stacks built on proprietary, opaque data models — fragmentation that deepens vendor lock-in, drives up costs, and blocks institutions from building the trustworthy, AI-ready tools their students need.

At DataKind, we've spent years building the alternative: open source infrastructure and products now used by 100+ broad-access institutions serving over 1 million students. This talk shares how we use Apache Spark, Apache Airflow, and Apache ECharts to power two freely available, open core education products designed to support student graduation success.

Education is where this panel starts, but not where it ends. Our work in education has sharpened our view on an open source data infrastructure approach — one that turns fragmented institutional data into a reusable foundation that can unlock impact across student success support, institutional technology transformation, and sector AI innovation.

We will share our work on a deployable open data model and associated infrastructure for staging, mapping, transformation, and access, built around four pillars of trust (modular pipelines for institutional autonomy, public versioned code, standardized data dictionaries, and schema lineage tracking) and discuss practical challenges — including data quality, system interoperability, and explainability — to translate open infrastructure into tools that non-technical teams can actually use.

Attendees will leave with access to open source public-good infrastructure for student success; frameworks for evaluating shared data models and AI-ready infrastructure; practical approaches to institutional deployment; and opportunities to join a community building an interoperable foundation on which an ecosystem — not a single organization — can innovate.


What does open source infrastructure have to do with whether a first-generation college student makes it to graduation? Quite a lot — and what we've learned answering that question in education turns out to matter well beyond it.

At DataKind, we build data and AI products for social impact, including tools that help locally serving, broad-access higher education institutions make better use of student data. These institutions serve first-generation students, working adults, and learners returning to education, often while operating with limited technical capacity and without access to the expensive enterprise systems used by well-resourced universities. In that environment, affordability, interoperability, and long-term maintainability aren't nice-to-haves — they're core design requirements.

Two products anchor this work. Edvise uses predictive modeling, powered by Apache Spark, to identify students who may be at risk of not graduating — and just as importantly, to surface the factors behind those predictions so student support staff can act on them before a student has to ask. A companion data integration platform, orchestrated by Apache Airflow, ingests information from student information systems, learning management systems, and CRMs, harmonizing it into a shared open model centered on student success indicators, with Apache ECharts powering the analytics layer staff use to explore it. Together, these products are freely available to over 100 collaborating institutions today, representing more than 1 million students — and early time-and-motion testing has shown the orchestrated stack cutting standard reporting turnaround from 40 days to 4 hours.

These products were codesigned with institutional actors who understood the value of open source components and had lived experience with the consequences of vendor lock-in — a pattern especially visible across an institutional landscape spread across the hyperscalers. That codesign process is also where this panel's broader argument comes from: the open data model underneath these tools — built on modular pipelines, public versioned code, standardized data dictionaries, and schema lineage tracking — isn't specific to education. By cleaning and harmonizing institutional data into a predictable model and access layer, this same kind of infrastructure can dramatically lower the barrier to building advising agents, financial-aid tools, pathway applications, analytics, and other AI-enabled supports — in education or elsewhere. Public models, mappings, transformation code, APIs, and context layers can act as a "Rosetta Stone" across a learner's — or any user's — journey, letting institutions, researchers, credential providers, and entrepreneurs build on a shared foundation rather than recreate it.

This panel brings together the people building that foundation to discuss what it actually takes to make it work in practice — data quality, system interoperability, explainability, and the unglamorous effort of translating open infrastructure into tools non-technical teams can use — and what education's experience with these problems can teach other domains navigating the same terrain.

Caitlin Augustin

Caitlin is the Vice President of Product and Programs at DataKind, a global nonprofit advancing the use of data, technology, and AI for social impact. At DataKind, Caitlin leads the development of reusable, interoperable data and AI tools codesigned for real-world implementation across sectors such as education, community services, and humanitarian response.

Dave Parry

Dave is Director of Strategic Partnerships at DataKind, where he develops mutually beneficial partnerships that enable the organization to build scalable, data-driven solutions for social impact partners around the world. Based in London, he works at the intersection of data science and social good, helping mission-driven organizations unlock the potential of their data.

Larry Kilroy