Warehouse wisdom.
Lakehouse speed.
We fuse classic modelling with cloud-native muscle, so your data moves at lakehouse speed. Proven patterns from Kimball, Inmon and Data Vault, built on a medallion lakehouse in Fabric or Databricks.
No single model fits every estate
A one-size-fits-all model is outdated. We blend these patterns to match your data sources, governance needs and growth plans.
| Pattern | What it is | Best for |
|---|---|---|
| Kimball dimensional | Denormalised facts and dimensions in a star schema, built for fast BI. | Interactive Power BI and Fabric dashboards, where query speed and user-friendly models matter. |
| Corporate Information Factory | A normalised enterprise warehouse feeding subject-area marts. | Regulated industries that need a single version of the truth before data is re-shaped for analytics. |
| Hybrid Kimball and Inmon | A CIF core for governance, with star-schema marts for speed. | Large enterprises balancing auditability with self-service BI. |
| Data Vault 2.0 and 2.1 | Hubs, links and satellites for agility and full history. | Cloud lakehouses where schema drift is common and automated CI/CD is key. |
| BEAM✲ analysis | Business-event-centred requirements gathering, modelled with the people who ask the questions. | Agile projects that iterate directly with business users to avoid a model that misses the point. |
Unify, transform and accelerate insight
A modern warehouse brings all your enterprise data, structured and unstructured, onto a scalable, secure lakehouse. You get self-service analytics, real-time insight and a foundation for AI, without touching the performance of your source systems.
Sources
Sales, operations, finance, customers
- ERP and CRM
- Databases and APIs
- Files and streams
Bronze
Raw, as it arrived
- Metadata-driven Data Factory loads
- Incremental, with watermarks
- Kept for replay and audit
Silver
Integrated and historised
- Cleansed and conformed
- 3NF or Data Vault
- Business keys reconciled
Gold
Shaped for analysis
- Star-schema marts
- Conformed dimensions
- Served to Power BI
Across every layer
- OneLake or Delta Lake storage
- Microsoft Purview lineage
- Decoupled compute for reporting
- CI/CD through DEV, TEST and PROD
The star schema, still the fastest way to an answer
Ralph Kimball's bottom-up approach models each business process as a fact table at a declared grain, surrounded by the dimensions people slice it by. Conformed dimensions are shared across facts, so 'customer' or 'product' means the same thing in every report.
- Built for the way people ask questions: by date, product, customer, place
- Fast in Power BI, where star schemas are what the engine is tuned for
- Planned with a bus matrix so marts join up instead of drifting apart
Fact
fact_sales
- FKdate_key
- FKproduct_key
- FKcustomer_key
- FKstore_key
- Σquantity_sold
- Σnet_amount
- Σdiscount_amount
Dimension
dim_date
- PKdate_key
- date
- month
- quarter
- financial_year
Dimension
dim_customer
- PKcustomer_key
- customer_name
- segment
- region
Dimension
dim_product
- PKproduct_key
- product_name
- brand
- category
Dimension
dim_store
- PKstore_key
- store_name
- channel
- country
The Corporate Information Factory
Often called the father of the data warehouse, Bill Inmon laid the foundation for enterprise data architecture.
“A subject-oriented, integrated, time-variant and non-volatile collection of data in support of management decision-making.”
Inmon's top-down approach starts with a centralised, normalised enterprise data warehouse as the single source of truth. Every source system feeds it through robust pipelines, and reporting tools and data marts consume cleansed, governed data downstream.
The key advantage is decoupling. If a source system changes, only the load needs to adapt, leaving analytics, dashboards and reports untouched. That separation between operational and analytical systems is especially valuable when integrating third-party or external data.
Best forIdeal for organisations that put data quality, governance and stability first across a complex enterprise estate.
All the data, all of the time
Created by Dan Linstedt as an alternative to both Kimball and Inmon, Data Vault 2.0 separates business keys (hubs), relationships (links) and context (satellites) into a structure that is adaptable, auditable and built for long-term history.
Where traditional models strive for a single version of the truth, often by cleansing away data that doesn't conform, Data Vault keeps a single version of the facts: raw and unfiltered, with business rules applied downstream. Highly parallel loading makes it a natural fit for big data and a mix of streaming, structured and unstructured sources.
- Built-in auditability for compliance-heavy industries
- End-to-end lineage and traceability
- Faster parallel loading from many sources
- Resilient to source system changes, without breaking what sits downstream
Hub
hub_customer
customer_id
- Satsat_customer_details
- Satsat_customer_address
Link
link_order_line
customer · order · product
- Satsat_order_line
Hub
hub_product
product_code
- Satsat_product_details
- Satsat_product_price
Hub
hub_order
order_number
- Satsat_order_status
- Hubunique business keys
- Linkrelationships between them
- Satdescriptive data, with full history
Store everything. Analyse anything.
A data lake holds all your enterprise data in its raw, native format, from database records and CSV files to IoT feeds, PDFs, images and video, with schema applied when it is read rather than when it is written.
Why a data lake
Raw files on their own lose metadata: types, business rules and validation. Modern platforms pair the lake with governance such as Microsoft Purview, and a lakehouse layer, to add structure, lineage and usability.
- Cost-effective scalability for big data
- Flexible storage of disparate sources with no upfront schema
- Batch and streaming ingestion
- A natural home for analytics, AI and machine learning
What Delta Lake adds
Delta Lake is an open-source storage layer on top of Azure Data Lake Storage that brings warehouse reliability to the lake. It is the table format under both Databricks and Fabric.
- ACID transactions for reliable data
- Time travel and versioning
- Schema enforcement and evolution
- Faster reads and writes through file statistics and compaction
- Native with Apache Spark, Databricks and Fabric