← All Use Cases

Data foundation and operations for research data

Most AI initiatives in research do not fail at the model. They fail because the data cannot be found, the analysis cannot be repeated, and nobody can operate the prototype. We build the foundation that makes models viable in the first place.

Data foundation and operations for research data

The situation

Research data is created in a decentralised way: in ELN and LIMS systems, in instrument exports, in spreadsheets on network drives, in individual researchers' analyses. Every source has its own nomenclature, and the same substance, the same assay and the same cell line carry three different names across three systems.

The consequence does not only affect AI. It affects traceability: an analysis that cannot be reconstructed two years later is equally worthless for a publication and for a regulatory submission. And it affects economics, because the same preparation effort recurs in every project.

Typical data basis

Exports from ELN and LIMS systems, instrument data, accumulated spreadsheet estates, plus the reference systems to normalise against – established ontologies for substances, assays and cell lines.

Why such initiatives fail

The complete data model first. Eighteen months of modelling, no usable interim result, and eventually somebody asks what it is for. Anyone who cannot show an answerable question within a few months loses the funding before the foundation stands.

Normalisation without ownership. Ontology matching is carried out once as a project and then maintained by nobody. Within a year the nomenclature diverges again and the effort recurs. Without named ownership and a maintenance process, FAIR is a short-lived condition.

Monitoring without a recipient. A pipeline runs in production but nobody in the business owns it. The first schema change upstream breaks it, the alert lands in a mailbox nobody reads – and three months later someone is still working from stale figures.

Approach

Structure from the use case, not from the data model. We start with the questions the data foundation is meant to answer. A complete data model nobody uses costs more than an incomplete structure that works.

Normalise nomenclature against standards. Only once substances, assays and cell lines are matched against ontologies do datasets become linkable – and only then does the F in FAIR mean more than an intention.

Reproducibility as infrastructure, not as discipline. Pipelines in established workflow systems, containerised tools, pinned versions, versioned datasets. An analysis then remains traceable years later without anyone having to remember.

Design for operation from the start. Monitoring, a retraining strategy, drift detection, documented model states. The transition from prototype to operation is where most initiatives stall – we treat it as part of the project, not as a follow-on topic.

Regulation as an architectural question. GxP requirements do not apply in early research; there we work independently and fast. Where an application touches regulated processes, we supply the documentation your validation requires; system validation itself stays with you.

What you get

A linked, queryable data foundation instead of separate sources, reproducible analysis pipelines, monitoring in operation, and documentation that withstands an audit – including the evidence that becomes relevant under the EU AI Act.

Where our experience comes from

We have been building data foundations for research organisations for years – and we operate them too:

Alongside our Data Infrastructure and Data Operations services and an information security management system certified to ISO/IEC 27001.

Last updated: 4 August 2026