DataXcelerator · Lakehouse edition

Add a source.
Not a project.

DataXcelerator is our metadata-driven ingestion framework for the Azure lakehouse. Describe a source in a few rows of metadata and it lands in your data lake as open parquet: incrementally, privately and on a schedule, ready for Databricks, Microsoft Fabric or anything else that reads a lake.

REST and OData APIsSharePoint listsSQL databasesAny Data Factory sourceFree in dev and UAT
The short version

Organisations turn to DataXcelerator when

The first API integration was a project. So was the second. By the fifth, the same login, paging, retry and incremental logic has been written four different ways by four different people.

Every new source has become a project

DataXcelerator writes the hard parts once, as templates. Authentication, paging, batching, incremental loads and failure handling are solved problems. A new source is configuration, and a new table is a row.

Nightly loads are fragile and expensive

Full reloads hide problems until the bill or the batch window runs out. Every mapping is incremental by default, with a watermark, a maximum batch window and a rerun path for anything that fails.

Nobody can say what runs where

Connections, schedules, column mappings and failures live in one metadata database, not in a notebook or someone's head. Audit it with a query. Secrets never leave Key Vault.

How it works

One framework between your systems and your lakehouse

Three parts. A metadata database that describes what to load. Data Factory templates that do the loading. An open landing zone that any engine can read.

How DataXcelerator connects your systems to your lakehouseArchitecture diagram. Your systems (REST and OData APIs, SharePoint lists, SQL databases, files and any other Data Factory source) feed Azure Data Factory running the DataXcelerator templates, inside a managed virtual network with private endpoints and no secrets in pipelines. A metadata database on Azure SQL describes every load and Key Vault supplies secrets and certificates. Data Factory writes to a landing zone in ADLS Gen2 as parquet, CSV or JSON in year, month and day partitions, which Databricks, Microsoft Fabric and any other engine that reads parquet can use.REST and OData APIsSharePoint listsSQL databasesFiles and any otherData Factory sourceAzure Data FactoryDataXcelerator templates: oneexecutor, a template per source type,batching, retries and failure loggingMetadatadatabaseAzure SQLKey Vaultsecrets,certificatesLanding zoneADLS Gen2parquet, CSV or JSONyear / month / daypartitionsDatabricksMicrosoft FabricAny engine thatreads parquetYOUR SYSTEMSDATAXCELERATOR ON AZUREYOUR LAKEHOUSEMANAGED VIRTUAL NETWORK · PRIVATE ENDPOINTS · NO SECRETS IN PIPELINESdescribeseveryload

Control plane

A small Azure SQL database holds every source, connection, table mapping, column mapping, watermark and failure. It is the single place to look when someone asks what loads, from where, how often.

Data plane

Generic Data Factory pipelines read the metadata at run time and do the work: fetch a credential, split the load into batches, copy to parquet, record the outcome. No pipeline is written for a specific table.

Landing zone

Open files in your own storage account, partitioned by load date. Databricks and Fabric read them directly, as can any engine that reads parquet. There is no proprietary format and no lock-in to one engine, or to us.

After every load, Data Factory also starts the downstream lakehouse job and refreshes the Power BI semantic models, so one schedule carries the data from source to report.

Data Factory linked services list showing Key Vault, Azure SQL and storage over private endpointsData Factory linked services list showing Key Vault, Azure SQL and storage over private endpoints
The linked services of a live deployment: SQL, storage and Key Vault reached over managed private endpoints, REST, SharePoint and HTTP for the sources, Databricks for the lakehouse. Connection names redacted.
Every run, every table

What happens when the schedule fires

The same seven steps for a sales ledger with twenty million rows and a lookup table with twenty. The metadata decides the details.

Schedule fires

A Data Factory trigger starts the control master. Sources load in parallel. The lakehouse refresh and Power BI wait for all of them.

Look up the mapping

The template reads one row: source, destination, pattern, watermark and fan-out settings. Nothing is hard-coded in a pipeline.

Fetch the credential

The connection says how to authenticate: a Key Vault secret, an OAuth2 token, a certificate. Pipelines never see a password in clear text.

Fan out

If the source is split by site, branch or region, one mapping becomes one run per value, each with its own watermark.

Batch the window

The gap between the last watermark and now is cut into batches no larger than the mapping allows. A six-month backfill behaves like a nightly run.

Copy to parquet

Each batch is copied to the lake with the column mapping and data types generated from metadata, into year, month and day folders.

Record the outcome

Success advances the watermark. Failure writes the exact URL, iterator and date range to the failure log, where the reprocess template picks it up next run.

Reruns are not a special case.

Because every batch is addressed by mapping, iterator and date window, a failed one can be replayed in isolation without touching what already landed.

Data Factory monitor showing a nightly control master run with every template executor succeededData Factory monitor showing a nightly control master run with every template executor succeeded
One night's run in the Data Factory monitor: the control master at 07:00, the per-table template executors fanning out in parallel, every one Succeeded. The source systems shown are OpenTable and Zembra, both third-party platforms.
Configure, don't code

A new table is a row, not a pipeline

This is what adding a table looks like. One record in the metadata database, plus a one-activity wrapper pipeline that passes its id to the template executor.

tdcdx.table_mappingenabled
source_system
EPOS
connection
epos-api · REST · OAuth2 · paged
template
REST API, batched
source_url
/sales/headers?site={{iterator_value}} &from={{watermark_value}}&to={{watermark_value_to}}
destination
landing/epos/sales_header/
watermark
modified_on · from 2023-01-01 · max 7 days per batch
iterator
sites (one run per site)
columns
24 mapped · types enforced · nested JSON flattened

What the framework works out for you

  • The paging, the token refresh and the authorisation header the API expects
  • The date batches between the last watermark and now, within the limit you set
  • The column translator Data Factory needs, generated from the column rows at run time
  • The partition folder for every file, so downstream readers can prune by date
  • The failure record if a batch breaks, and the watermark update when it does not

Agent skills included

DataXcelerator ships with skills for coding agents such as Claude Code. Point one at your metadata database and it drafts new mappings, column lists and lakehouse transformations in the framework's own conventions, for your engineers to review rather than write.

Under the bonnet

The metadata model

Nine tables in one schema. Everything the templates do is a lookup against them, and everything they learn is written back to them.

The DataXcelerator metadata model: nine tables in one schemaMetadata model diagram. table_mapping, one row per table holding source, destination, watermark, batch window and options, sits at the centre. It takes its source and destination from table_connection and source_system, its pipeline pattern from table_template, and its columns from table_mapping_column. It fans out through table_iterator_configuration to table_iterator_value, with one table_iterator_watermark per mapping and iterator. Failures go to table_mapping_failure and are replayed by the reprocess template.table_connectionhost, protocol, auth method,paging ruletable_templatewhich pipeline pattern runstable_mapping_failurewhat broke, where, for whichwindowsource_systemthe systems you ingest fromtable_mappingone row per table: source,destination, watermark, batchwindow, optionstable_mapping_columnsource path to sink columnand typetable_iterator_configurationhow a mapping fans outtable_iterator_valueeach site, branch or regiontable_iterator_watermarklast loaded value per mappingand iteratorsource and destinationone watermark per mapping × iteratorpatternfan outreplayed by the reprocess template

The template hierarchy

Data Factory folders map onto the levels below. Everything under L99 is the product; everything under L01 is yours.

L00 · Control
One control master per schedule. It runs the source masters in parallel, then the lakehouse job, then the semantic-model refresh, and stops the clock on anything that overruns.
L01 · Sources
One wrapper per table: a single activity that calls the template executor with a mapping id. This is the only pipeline you add per table, and it is six lines of JSON.
L99 · Templates
The executor switches on the mapping's template. Level 1 handles credentials and fan-out (REST batched, REST non-batched, SharePoint list, reprocess failures). Level 2 cuts the date batches. Level 3 copies one batch to parquet and records the outcome.
Utilities
Key Vault secret, OAuth2 token, Databricks cluster and SQL warehouse start and stop, job lookup by name, Power BI dataset refresh. Reused by every template, replaceable by you.
Deployment
The factory is ARM-parameterised per environment and shipped by an Azure DevOps pipeline, with managed private endpoints to SQL, storage, Key Vault and Databricks. A naming-standard procedure keeps the metadata database tidy as it grows.
Data Factory resource tree: L00 control, one L01 folder per source system, L04 Power BI, L99 DataXcelerator templates and utilitiesData Factory resource tree: L00 control, one L01 folder per source system, L04 Power BI, L99 DataXcelerator templates and utilities
The resource tree of a live factory. L00 control at the top, one L01 folder per source (here OpenTable, SharePoint and Zembra), L04 for the Power BI refreshes, and the L99 product folders at the bottom.
The L00 control master pipeline: two source masters run in parallel, then a Databricks job, then a Power BI refreshThe L00 control master pipeline: two source masters run in parallel, then a Databricks job, then a Power BI refresh
L00 in practice. The source masters run in parallel; when they finish, the lakehouse job runs, then the semantic models refresh. Five activities for the whole estate.
A DataXcelerator utility pipeline that parses a paged JSON API responseA DataXcelerator utility pipeline that parses a paged JSON API response
One of the utilities: walking a paged JSON response. It resolves the collection path and the paging path, decides whether there is a next page, and returns success or failure to the caller. Written once, used by every API source.
Downstream

From landing zone to star schema

The lake is the contract. Our Databricks templates take it from there; Fabric reads the same files.

Bronze

Streaming tables that read each landing folder as files arrive. One declarative statement per table, no code, and nothing re-read that has already been processed.

Silver

Incremental change capture (SCD type 1 or 2) with the cleansing in plain SQL: trim and standardise, derive business keys, apply trading-day rules, hash surrogate keys.

Gold

Materialised views in a dimensional model: facts and conformed dimensions, clustered for query performance and ready for a semantic model in Power BI.

Lakeflow Declarative PipelinesServerless computeDatabricks Asset Bundles, dev and prdJob orchestration bronze → silver → goldUnity Catalog volumesFailure alerts to our support desk
In production

Built on real sources, not demo data

DataXcelerator grew out of client platforms we run today. The examples below are what it ingests for them.

Restaurant group

EPOS sales, reservations, gift cards, loyalty, online reviews and SharePoint budgets, per site, every night, into one sales-and-guest model that the whole business reports from.

Belting manufacturer

An expensive third-party data warehouse replaced with an owned Azure platform and a single version of the truth across global branches.

Dairy

[ONE LINE ON THE SOURCES AND OUTCOME, TO CONFIRM]

Specialist lender

[ONE LINE ON THE SOURCES AND OUTCOME, TO CONFIRM]

See the case studies
How you buy it

Yours to run. Yours to extend.

Free in development and UAT

Install it, build with it, prove it on your own sources. Nothing to pay until it goes live.

Licensed in production, with support

A production environment carries a licence and our support desk behind it. It usually arrives inside an engagement; it is also sold on its own.

Templates, not a black box

It is SQL, JSON and YAML in your tenant. Modify it, extend it, add your own templates. The licence excludes resale, derivative products and use outside the licensed environments.

What is in the box

  • Metadata database: schema, views, procedures and naming standards
  • Data Factory templates and utility pipelines
  • Azure DevOps deployment pipelines, parameterised per environment
  • Databricks asset bundle with bronze, silver and gold templates
  • Documentation for every layer, and the agent skills
  • Support from the people who built it
Next step

Bring us the source you have been putting off.

A free initial consultation. We will tell you how many rows of metadata it is, and how soon it could be in your lake.