Case study

When a report is no longer enough to build on

Länsi-Uusimaa Wellbeing Services County (LUVN), a Finnish wellbeing services county, brought software engineering methods into its data production. Two years, and no locks.

LUVN was technically ahead of its peers before any of this began. That is what made its problem the harder one: the county had got far enough that the next bottleneck sat in the way the work was done, rather than in the platform underneath it.

Key facts
Client
Länsi-Uusimaa Wellbeing Services County (LUVN)
Sector
Public health and social care — a wellbeing services county of roughly 480,000 residents
Duration
2023 → ongoing
Invinite's role
Architecture and method for data production, built together with LUVN's data engineering team
Environment
Azure Data Lake, Databricks Lakehouse
Methods
PySpark/Python, unit-tested transformation functions, CI/CD, data contracts
Permission to publish
Given by the client. The quote is the client's own wording.

Where it started

In 2023, as soon as the wellbeing services county was formed, LUVN made a call that few of its peers made that early: it skipped conventional data warehousing altogether and moved straight to Azure Data Lake and Databricks Lakehouse. By early 2025 the first reports were running in production.

The county had already seen the wall it was heading towards. A lakehouse cannot be built one reporting application at a time. Every new report carried its own definition, its own query and its own version of the truth, and each one depended on the person who had written it. What it needed was a refined data foundation that several uses could rely on — and, for the data engineering team, automated testing and quality assurance built into the method.


Why this is a harder problem than it looks

Replacing a warehouse with a lakehouse is a procurement. What comes after it is not. When every data product defines its own truth, an error does not appear as an error. It appears as two different numbers for the same thing, in the meeting where that number was supposed to settle a decision. Fixing it is handwork, handwork does not scale, and the more data products there are, the less anyone dares to touch them.

No tool solves this. It is a question of method, and a method changes only when the people working inside it change how they work. That is why this was never a delivery project.

What was done

Four changes, each of them checkable in the code.

  1. 01

    From SQL to PySpark and Python

    The single largest change. It turns data engineering into software development instead of query writing, and with it every tool of software development becomes available as it is.

    You cannot meaningfully unit-test a SQL script. You can unit-test a function.

  2. 02

    From long pipelines to small transformation functions

    Long, hand-validated data pipelines were broken into transformation functions that a developer can test locally before anything runs.

    An error shows up on a developer's machine in minutes, instead of in production months later.

  3. 03

    DevOps and CI/CD across the data lifecycle

    Testing and deployment were automated. Quality control moved out of someone's memory and into a pipeline that runs every time.

    Quality that depends on someone remembering to check is just luck.

  4. 04

    Data contracts

    Contracts are where domain layer modelling ends up. They describe, technically and logically, the core data that stands for things in the real world. They carry the metadata that matters, state which stage of its lifecycle the data is in, set out quality rules and ownership — and they travel with the code from development into production.

    This is the point where quality validation becomes automatic in every layer of the architecture. It is also the structure that later makes AI dependable: a language model gets an unambiguous account of what each data structure means in business terms.

What changed

  • A weekly release rhythm

    After the first release, new functionality reached production every week, rather than once per release window.

  • A way of working that transfers

    The result was a documented way of working that does not rest on one team or one person. It has since been applied elsewhere.

  • Two years, no lock-in

    The engagement has run two years with no contractual or technical tie holding it in place. Every decision to continue has been the client's own.

Where this page gives a number, it also says how the number was measured. If we cannot say that, the number does not belong here.

The client's view

“Moving to a Data as Software model has fundamentally transformed how we approach data production — we want to understand the data and produce knowledge. The past two years of close collaboration have allowed us to treat data as a core product — improving both the speed and reliability of our data pipelines and data products. For a wellbeing services county, this means more accurate, data-driven insights that ultimately support better decision-making and better care for our citizens.”

Henna Degerlund, Chief Data Architect, Länsi-Uusimaa Wellbeing Services County

Why this was not a delivery project

The method did not arrive with Invinite as a finished model. It came out of two things: LUVN had a clear direction for its platform, and an unusual ability to judge very complex proposals about data and code architecture. We designed the technical architecture to solve LUVN's problems, and they took it apart critically. A method developed this far would not have emerged without that exchange.

The way of working was built along the way. Documentation, workflows and everything learned belong to both teams, because they were written together throughout — which is also why there is nothing left to hand over at the end.

“For two years the client has chosen to continue, month by month. It is the only measure of continuity we trust.”


What this makes possible next

A disciplined data foundation is what AI requires before it can work at all. Anyone leading data in a health organisation needs to recognise that the last step cannot be taken first. At LUVN the order runs in three stages.

  • 01

    AI-assisted development

    An agent reads the domain definitions and the project's tasks, and helps the data engineer write the code that produces the data.

  • 02

    Support for defining data

    When data is defined in code and its exact meaning is spelled out in data contracts, a language model has an unusually good basis for understanding both the data and the code behind it — and for checking that new definitions follow the agreed architecture.

  • 03

    Tools for the business

    Only once the first two have proved dependable for data specialists does AI reach business users. That requires the underlying data to stay well defined, good quality and controlled in terms of who may see what.

Data contracts decide this. They give a language model an unambiguous account of what each data structure means in business terms. That is the difference between an answer that sounds plausible and an answer you can check.

What this case does not prove

Part of why this worked is that the client was exceptionally capable. LUVN could assess architecture proposals critically, and did. In an organisation with no technical ability of its own to push back, the same method would either move more slowly or leave the supplier deciding alone, and the second of those is the shell of Data as Software rather than the thing itself. So the first question in any assessment we run is who, on the client's side, is able to say no.

Data as SoftwareDatabricks Lakehousedata contractsPySparkwellbeing services countydata production

GET IN TOUCH

Have you got a project that needs taking all the way into use? Let's talk.

If we can't help, we'll say so.

contact@invinite.fiLinkedIn / @inviniteHelsinki · Tampere