Skip to content

Retail

Rebuilding the data pipeline behind a retailer’s daily reporting

Overnight batch jobs that regularly ran into the working day replaced by incremental Spark pipelines, so store and head-office reports are ready within 30 minutes of the data arriving.

Client Project7 months
Client
Grocery and home retailer, 900 stores, Saudi Arabia
Industry
Retail
Region
Saudi Arabia
Duration
7 months
Team
9 engineers
  • Data engineering
  • Cloud infrastructure
  • DevOps and CI/CD

Results

Transactions processed a day
2B+
Shorter pipeline runtime
40%
From source to report
<30 min
Data accuracy against source
99.99%

The challenge

Point-of-sale, e-commerce, loyalty and supplier data came from more than 50 sources into an on-premise warehouse through hand-written SQL jobs. The nightly run took nine hours and failed often enough that regional managers stopped trusting morning reports.

Volumes were growing past two billion rows a day, and the finance team needed the same figures the stores saw, not a separate reconciliation each month.

Before and after

How things ran when we started, and once the work shipped.

  • Before: Nine-hour nightly batch, often late
  • After: Incremental loads, reports ready in 30 minutes
  • Before: Hand-written SQL jobs with no tests
  • After: Tested PySpark jobs deployed through CI
  • Before: Monthly reconciliation between finance and stores
  • After: One curated layer behind 500+ daily reports

Approach

How the work was done

4 phases over 7 months, with a working demo at the end of every week.

  1. 01

    Weeks 1–5

    Source inventory and data contracts

    Catalogued every source, owner and schema, and agreed data contracts with the teams that produce them so upstream changes stop breaking the pipeline silently.

  2. 02

    Weeks 4–12

    Lakehouse foundation

    Built an S3 and Delta Lake lakehouse on AWS with Terraform, and moved ingestion to incremental loads using change data capture from the store systems.

  3. 03

    Weeks 10–24

    Pipeline rebuild

    Rewrote the transformations as tested PySpark jobs orchestrated by Airflow, running old and new side by side until outputs matched to the row.

  4. 04

    Weeks 24–30

    Reporting and cut-over

    Pointed 500+ scheduled reports at the new warehouse layer, added data quality checks with alerts to owners, and switched off the legacy jobs.

Architecture

Data lands incrementally, is validated before anyone sees it, and every report reads from the same curated layer.

  • Change data capture from store and e-commerce databases with Debezium
  • Delta Lake on S3 with bronze, silver and gold layers
  • PySpark on EMR, orchestrated by Airflow
  • Great Expectations data quality checks with owner alerts
  • Redshift serving layer for Power BI and scheduled reports

Stack

  • Apache Spark
  • Delta Lake
  • AWS EMR
  • Airflow
  • Debezium
  • Redshift
  • Great Expectations
  • Power BI
“Regional managers now open the dashboard at 8am and trust what they see. Finance and operations finally argue about the business, not about whose numbers are right.”
Director of Data and AnalyticsGrocery and home retailer, Saudi Arabia

Have a project like this?

Tell us where things stand today. On a 30-minute call an engineer will sketch how we would approach it, the timeline and a price range.