Work

Five projects. Same two people on every one.

SAP warehouses, edge streaming, marketing attribution and automated data quality — across industrial, finance, telco and education. Every one of them needed at least two of our three capabilities, which is the whole argument for having them in the same team.

Projects delivered
Five, across four sectors
Most common starting point
SAP data rebuilt by hand every month
Capabilities per project
Two of three, every time
People on your project
The two who scoped it

Selected work

Client names withheld where we do not yet have written permission.

01
Industrial

Automated recognized revenue and analytical P&L

Financial and hiring reports lived inside SAP and were rebuilt by hand every month — days of work before anyone trusted the numbers. We modelled the raw SAP tables into a Databricks warehouse feeding live Power BI dashboards the finance team reads directly.

  • SAP
  • Databricks
  • Power BI
Capabilities
  • Data architecture
  • Business intelligence
Read the case
02
Finance

SAP accounting insights in Power BI

Accounting data arrived as manual SAP dumps into Excel, uploaded to a local server and crossed against the CRM by hand. We automated the export into Databricks, modelled it against the CRM, and turned a monthly cost figure into a daily one.

  • SAP ECC
  • Databricks
  • Spark
  • Power BI
Capabilities
  • Data engineering
  • Business intelligence
Read the case
03
Telco

Exploiting edge data

Network health and OSS metrics sat scattered across edge devices with no unified view, so problems surfaced only after they had already hit the network. We built a real-time path from Node-RED at the edge through Kafka and PySpark into live Grafana dashboards, across a hybrid on-premise and cloud setup.

  • Node-RED
  • Kafka
  • Airflow
  • PySpark
  • Grafana
Capabilities
  • Data architecture
  • Data engineering
Read the case
04
Education

Modelling ROI and sales funnel from digital marketing

Marketing data sat in parquet files with no foreign keys to join it to the financial data, so nobody could say which campaigns paid for themselves. We built the ETL from landing to gold and the relationships between the two verticals, then surfaced the target KPIs in Power BI.

  • Databricks
  • Power BI
  • ETL
  • SEO data
Capabilities
  • Data engineering
  • Business intelligence
Read the case
05
Data platform

Data quality reporting powered by AI

Quality problems went undetected until corrupted data had already reached the BI layer. We centralised the signals from Azure Data Factory jobs, Databricks jobs and table schemas behind the DQX framework, and used an LLM to turn the raw rule output into reports the client’s team actually reads.

  • Azure Data Factory
  • Databricks
  • DQX
  • LLM
Capabilities
  • Data architecture
  • Data engineering
Read the case

Capabilities per project

No project used just one of the three.

This is the argument for keeping architecture, engineering and BI in the same team, shown rather than claimed. Every handover between these three is a place where a project normally loses two weeks.

Which of the three capabilities each project required
Project Data architecture Data engineering Business intelligence
Industrial Recognized revenue & P&L Yes No Yes
Finance SAP accounting in Power BI No Yes Yes
Telco Edge data in real time Yes Yes No
Education Marketing ROI modelling No Yes Yes
Data platform AI-powered quality reporting Yes Yes No

Read the three capabilities in detail: data architecture, data engineering, business intelligence.

The stack behind these five

What we actually ran, not what we could list.

Sources

  • SAP ECC
  • Local CRM
  • Node-RED edge devices
  • Parquet landing files
  • Azure Data Factory jobs

Processing

  • Databricks
  • Spark / PySpark
  • Kafka
  • Airflow
  • DQX quality rules

Platform

  • Azure
  • Medallion architecture
  • Delta tables
  • Hybrid on-prem / cloud
  • CI/CD dev · test · prod

Delivery

  • Power BI
  • Grafana
  • Semantic models
  • LLM-generated reporting

Questions

What people ask after reading these.

Will connecting to our SAP slow down the ERP?

No. We extract on a schedule from the tables we need, outside business hours where the system allows it, and every transformation happens afterwards in Databricks — never inside SAP. Three of these five projects run on SAP ECC data. In none of them does the reporting layer add load to the ERP, because the ERP is only ever read from.

Our data is a mess. Do we need to clean it before you start?

No — that is the work, not a prerequisite. Marketing files with no keys to join on, accounting arriving as manual Excel dumps, quality problems nobody caught until they reached the dashboard: those are the starting conditions of three of these cases, not obstacles we asked the client to remove first. If your data were already clean you would not need us.

How soon do we see something working?

The first useful artefact is the current-state map, and that lands in the first two to three weeks. A first pipeline delivering data end to end typically follows within six to ten weeks of getting access. The variable that moves those dates is almost never technical — it is how long it takes to get credentials for the source systems.

Do we have to move to Databricks?

No. Databricks appears in four of these five because it fitted what those clients already had — mostly Microsoft estates with teams fluent in SQL. The telco project runs on Node-RED, Kafka and Grafana across a hybrid on-premise setup, and no Databricks at all. We are certified on Google Cloud and Azure and we build on AWS. The architecture follows your estate, not our preference.

What happens to the platform when you leave?

You own it, and your team runs it. Every project ends with documented transformations, CI/CD across dev, test and production, and a technical handover session. We write the governance rules — naming, schema approval, alert routing — precisely so the platform survives us. A dependency on us would be a design failure, not a business model.

Can you work alongside our internal IT or data team?

That is the usual arrangement. In the industrial project the governance patterns existed specifically so the client’s own team could work in parallel without breaking production. We can take a defined scope end to end, or embed alongside your people on your roadmap and your tools. What we will not do is build something your team is then locked out of.

Start here

If one of these looked familiar, it probably was.

Manual closes, data that only agrees with itself by accident, dashboards nobody trusts. Thirty minutes, no slides. We will tell you what we would do, roughly what it costs, and whether you need us at all.

Start a conversation

Or see the three capabilities these projects came from.