A governed lakehouse for real-time analytics and data science.

A data platform ingesting operational databases, APIs, files, and event streams into a governed lakehouse for self-service analytics, reporting, and machine-learning workloads.

Client

Multi-product technology company

Industry

Data infrastructure

Duration

28 weeks

Team

Data architect · 2 data engineers · 2 platform engineers

Delivered

2026

Product, finance, operations, and data-science teams were pulling different answers from production databases, SaaS exports, event topics, and manually maintained files. Pipelines were owned as isolated scripts, dashboards disagreed on core metrics, backfills were risky, and data scientists spent more time assembling training datasets than testing models. The company needed shared infrastructure that could serve streaming operations, BI, and machine learning without creating another central bottleneck.

Constraints that shaped the build
  • Streaming events, database CDC, API pulls, and file deliveries had to land through one observable ingestion model.
  • Historical backfills and schema evolution could not interrupt fresh production data or silently rewrite trusted results.
  • Access controls had to follow organisation, domain, environment, and data-sensitivity boundaries.
  • Analytics and data-science teams needed self-service SQL and curated datasets without direct access to production systems.
01

Define data contracts

We grouped sources by business domain and established ownership, schemas, freshness expectations, quality rules, retention, and compatibility before migrating pipelines.

02

Build the lakehouse in layers

Immutable raw data, validated domain tables, and curated data products were stored as Apache Iceberg tables with catalogue, lineage, and reproducible transformations.

03

Unify batch and streaming

Kafka and Flink handled low-latency event flows while Airflow and Spark orchestrated larger transformations, backfills, and data-science preparation against the same lakehouse model.

04

Create a self-service control plane

A Next.js platform exposed datasets, lineage, freshness, quality, pipeline runs, ownership, access requests, and cost signals without requiring every user to understand the infrastructure.

Ingestion fabric

Java and Spring Boot services for source registration, CDC, event ingestion, API and file connectors, schema validation, replay, and dead-letter handling.

Governed lakehouse

S3 and Apache Iceberg storage with partitioning, schema evolution, snapshots, catalogue metadata, lineage, retention, and domain-level access controls.

Analytics and data science

Trino SQL, curated marts, notebooks, reproducible feature datasets, MLflow experiment tracking, and controlled publication of model outputs.

Data operations control plane

Pipeline health, freshness, quality tests, dependencies, backfills, incidents, ownership, access, and infrastructure cost in one React and Next.js interface.

CDC

continuous ingestion from operational databases

SQL + ML

analytics and data-science workloads on shared governed data

Iceberg

schema evolution, snapshots, reproducible backfills, and time travel

Lineage

source-to-dashboard ownership, freshness, and quality visibility

Core technology
  • Java
  • Spring Boot
  • React
  • Next.js
  • Kafka
  • Apache Flink
  • Apache Spark
  • Airflow
  • S3
  • Apache Iceberg
  • Trino
  • PostgreSQL
  • MLflow
  • Kubernetes
  • AWS
What mattered most

The platform succeeded when data ownership and operability were treated as product features. Storage and compute mattered, but contracts, lineage, quality, replay, and self-service made the infrastructure usable beyond the data team.

Does this resemble the problem in front of you?

Share the context, constraints, and where you are stuck. We’ll reply with useful questions and a clear next step.

Tell us about it