Press Release

Europe’s Open-Source Data Stack Is Having a Moment

By Yann Ranchere, Co-Founder & CEO, Motley — August 2026

Why open source, data control and AI are pulling a European data ecosystem into view

Look at the recent funding announcements one by one and they do not obviously belong together. Kestra raised a $25 million Series A. Polars raised €18 million. dltHub announced an $8 million raise. Lightdash raised $11 million. Motley, my company, raised a $1.5 million pre-seed.

Put these rounds side by side, and a broader picture starts to emerge.

Over the past few weeks I have been mapping the European open-source data landscape, from databases and transformation engines to orchestration, BI, feature infrastructure and semantic layers. What stood out was not simply the number of companies. It was how much of the stack Europe now covers.

DuckDB and Polars have each attracted close to 40,000 GitHub stars. Kestra is approaching 27,000. QuestDB is around 17,000, Lightdash around 6,000 and dlt above 5,000. GitHub stars are an imperfect measure of adoption, but the impact is hard to miss.

Five years ago it was easy to point to a handful of exceptional European open-source data projects. Today it is possible to map a large part of the stack.

Independence first

Europe has long had stricter rules around privacy and data. But the broader issue behind many of these companies is control: who owns the data, where the software runs, who can access it and how dependent the customer becomes on a vendor.

That concern now extends well beyond regulation. European institutions are explicitly talking about technological sovereignty, open-source alternatives and reducing dependency on non-European providers. Enterprise buyers are asking the same questions in more practical terms.

Open source gives companies more choice over those questions.

DuckDB runs locally rather than requiring data to be sent to a separate database service. Nao can be self-hosted. SLayer can run entirely inside a customer’s own environment. Hopsworks supports on-premises and air-gapped deployments as well as managed cloud.

These are different products with different architectures, but they share something important: the customer has options.

For European companies, privacy matters. But control over data and infrastructure is the bigger theme.

Europe already builds the engines

The modern data stack is still often described as a Silicon Valley story. That is increasingly hard to reconcile with the projects themselves.

DuckDB has become a major analytical database project. Polars has built a high-performance DataFrame and query engine used far beyond Europe. Kestra has built a substantial orchestration community. QuestDB has established itself in time-series infrastructure. KNIME has been building open-source analytics software since 2006.

Further up the stack, dlt has become an important Python-native data-loading project. Bruin combines ingestion, transformation and quality in a single framework. Hopsworks spans feature infrastructure, data and machine-learning workloads.

They are not all direct competitors, and they do not fit neatly into a single architecture. They complement each other and overlap on some aspects. But more importantly, European open-source data companies are showing up across many of the layers that modern data teams rely on.

AI is changing the top of the stack

The part of the map changing fastest is the interface between enterprise data and AI agents.

Giving an agent access to a database is relatively easy. Giving it enough context to use that data correctly is much harder.

Nao approaches the problem through context engineering: bringing together schemas, metadata, modelling, rules and other information an analytics agent needs. Lightdash is building governed metrics and business context for use by both people and agents.

At Motley, SLayer is built as an agent-first semantic layer designed to provide developers and data teams with a highly expressive, reliable, and high-performance foundation for analytics-focused agentic workflows.

The problem is straightforward. A human analyst who sees an ambiguous column can ask someone what it means. An agent needs more of that meaning to be explicit: which metric definition is correct, which joins are allowed, which data a user can see and which business rules apply.

As more analytical work is delegated to agents, that context becomes infrastructure.

And the shift is not limited to companies building directly for analytics agents. dltHub is moving into agentic data engineering. Kestra is extending orchestration across data and AI workloads. Hopsworks is bringing agents into its data and AI platform.

AI is not creating a separate stack. It is changing the requirements of the existing one.

Open source is becoming the adoption layer

There is another pattern across many of these companies: open source is increasingly how the product enters an organisation rather than the thing being sold.

Polars keeps its core library open source while building commercial cloud products around it. dlt remains an open-source Python library while dltHub sells managed infrastructure and production tooling. Lightdash lets teams self-host its open-source product while selling managed cloud and enterprise capabilities. SLayer is MIT-licensed, while Motley Cloud provides the managed layer around it.

The details differ from company to company, but the underlying logic is familiar.

Developers can inspect the software, try it locally and often run it themselves. The commercial relationship begins when a company wants managed infrastructure, enterprise controls, collaboration, support or simply does not want to operate another production system.

That model fits particularly well with the control question. Customers do not have to choose between owning their infrastructure and paying a vendor. They can keep the former while paying for the convenience of the latter.

Distributed across Europe

The projects in the map are spread across Europe: London, Paris, Amsterdam, Zurich, Berlin, Stockholm and Brussels all appear.

There is no obvious European capital for open-source data infrastructure. Instead, the ecosystem is distributed across several technology hubs.

That may end up being one of its strengths. Open-source communities are already international by default, and European founders, contributors, employees and investors regularly move across borders.

The ecosystem does not need to exist in one city to exist.

An ecosystem, not a handful of projects

The growth of European open-source data companies is a trend.

Europe now has significant open-source projects across databases, transformation, ingestion, orchestration, analytics, machine-learning infrastructure and the emerging semantic layer for AI.

There is no single reason they have appeared.

Europe has strong technical talent. Open source has become a powerful way to win developer adoption. AI is creating new infrastructure requirements. And European companies are increasingly sensitive to where their data and software run, who controls them and how easily they can move away from a provider.

Those forces are starting to reinforce one another.

It looks like an ecosystem.

Yann Ranchere is Co-Founder and CEO of Motley, the company behind SLayer, an open-source semantic layer for governed querying by AI agents and applications

Author:

Related Articles

Back to top button