DINAcon 2026

Sufficiency over Size: Making AI More Sustainable with Open-Source RAG

When AI gives wrong answers, the reflex is to reach for a bigger model. This reflex is expensive and energy-hungry, yet it rarely fixes the problem. Most organizational knowledge sits in documents such as guidelines, contracts, and reports, and what is missing is not intelligence but access. This talk presents an open-source architecture for Retrieval-Augmented Generation (RAG) that supports small open models, grounded in documents prepared by Docling, an open-source project developed in Switzerland, and retrieved through hybrid search with OpenSearch. The principle throughout is sufficiency over size: by providing the right knowledge, we do not need ever-bigger models.


The scaling reflex

When an AI system answers incorrectly, the standard response is to increase its size. However, adding billions of parameters makes such models ever more expensive and energy-hungry to run. In organizational settings, this often misses the point, because the failure was rarely a lack of intelligence. For most questions people ask inside an organization, the answer already exists in a PDF, a policy, a report, or a spreadsheet. The model simply never saw it.

Sufficiency over size

Retrieval-Augmented Generation (RAG) provides organizational context to a model. The system first retrieves the relevant passages from the organization's own documents and lets the model answer from them, with sources attached. A model that is handed the right three paragraphs does not need hundreds of billions of parameters. Small open-weight models, served locally on a standard laptop, deliver grounded, citable answers at a fraction of the energy and cost of frontier-scale systems, with data that never leaves the machine. This is the sufficiency principle applied to AI: meeting the actual need precisely, rather than increasing size and hoping quality follows.
Quality through open document understanding
Yet many RAG projects stall after a promising demo, and the reason is rarely the model. Default RAG treats documents as flat text: tables collapse into noise, headings lose their hierarchy, scanned pages disappear, and naive text chunks return fragments without context. The answers look fluent and are quietly wrong. The decisive levers sit earlier in the pipeline:

  • Document understanding instead of text extraction. Docling, developed in Switzerland and open-sourced in 2024, parses PDFs, Office files, and scans into structured, AI-ready data that preserves tables, hierarchy, and provenance. It has become one of the most widely adopted open-source projects in this space and the basis for an emerging open document standard.
  • Hybrid retrieval instead of vector search alone. OpenSearch combines exact keyword matching with semantic vector search, so that legal terms, product codes, and paraphrased questions all find the right passage.
  • Transparent orchestration instead of black boxes. Langflow makes the pipeline itself visible: every step from ingestion to answer is inspectable and adjustable, a precondition for trust and auditability.

Investing here is what allows the model to stay small.

Open from end to end

The components are integrated in OpenRAG, an openly licensed distribution that makes the complete pipeline reproducible from public repositories. With Langflow, every answer points back to its sources. Through Ollama support, even the language model runs locally: the entire system, from document parsing to the final answer, operates on a single machine with no cloud dependency at all. No component locks its users in, as each can be replaced, inspected, and operated independently.
What attendees take away

  1. A practical map of where default RAG breaks and which levers fix it, applicable to any stack, not only the one shown.
  2. The sufficiency principle as a decision heuristic: when retrieval and preparation beat model size, and what that means for energy and cost.
  3. An overview of the open-source RAG ecosystem and its governance: open licenses, active communities, and what that means for longevity.

Technical concepts (parsing, chunking, hybrid search, model right-sizing) are explained from first principles, so the session works for engineers and decision-makers alike.

Philipp Gordetzki

Philipp Gordetzki is an AI Engineer at IBM Switzerland and Associated Researcher at the University of St. Gallen. He recently defended his doctoral dissertation at the University of St. Gallen. His research on AI for resilient supply chains and AI-human creativity has been presented at international conferences and has been accepted for publication in Information Systems Research. In his work as an AI Engineer, he has designed and built RAG-based pilot systems for clients across industries.