Data
Warehouses, pipelines, and the screens someone acts on.
Most of this data already exists. It sits in systems that were never built to talk to each other, in shapes nobody can read quickly, and it arrives after the moment it was needed.
What we build
Eight shapes this work usually takes. Most projects are two or three of them at once.
Warehouses and semantic models
A single store fed on a schedule, and a model on top that names the relationships so people can query it without an analyst translating first.
Pipelines between systems
Scheduled ingestion, transformation and reconciliation, running where the data already lives.
Consolidated views
One place that reads from several operational systems at once, with each profile seeing its own version of it.
Monitoring dashboards
The screen a team keeps open, with the indicators that actually change what someone does next.
Volume moves and migrations
Moving large sets between systems without losing rows, history or the ability to reconcile what arrived.
Exploration interfaces
Maps, filters and drill-downs for material too large or too granular to read as a table.
Structure out of documents
Classification and extraction from PDFs, scans and handwritten forms that were never born as data.
Corpus for assistants
The retrieval layer an assistant answers from: chunking, embeddings and permissions, before any model is chosen.
Where it usually breaks
Every one of these projects started with data the organization already had. None of them started with a storage problem.
It’s in several systems
The Inter-American Development Bank monitors loans across operational systems that were never built to talk to each other. Smart Portfolio reads from all of them into a single set of indicators, with access rules that give each role its own view of the same data.
Two systems, two protocols, one warehouse
An operator ran its business across a core platform that only exposed SOAP and a CRM that spoke REST. We built a REST gateway over the SOAP interface so the rest of the estate could read it at all, then daily pipelines into a warehouse and a semantic model, with the dashboards built on the model rather than on the raw tables.
Delivered on Microsoft Fabric, with a data engineer on the team.
It’s not in a shape anyone can read
Harvard’s India Policy Insights holds health and population data at a granularity that defeats a spreadsheet. We built it as interactive maps, so a policymaker can compare districts and find what matters without opening a GIS tool or asking an analyst first.
Harvard University. High-granularity geospatial data on PostgreSQL, rendered in the browser.
It arrives after the moment it was needed
EmpowerHealth runs several care programs at once, and a weekly report is too late to change what happens in any of them. The analytics update as the programs run, so the team sees a program drifting while there is still something to do about it.
EmpowerHealth. Real-time analytics across parallel programs, under HIPAA.
What the dashboard is for
The hard part of a monitoring screen isn’t adding indicators. It’s deciding which ones earn the space, and what each one should make somebody do.
We start from the decision, not from the data
The first question isn’t what you can measure. It’s what someone is supposed to do differently when a number moves. An indicator that changes nobody’s behaviour is a maintenance cost with a chart on top.
Who sees what is settled before the schema
In most of this work different roles are entitled to different slices of the same data. That isn’t a permission layer added at the end. It shapes how the indicators are built, and getting it wrong later means rebuilding them.
The model comes before the screen
A dashboard built straight on the source tables breaks the first time a source changes. A semantic model in between names the entities and their relationships once, so the screens on top stay readable and the next question doesn’t need a new pipeline.
What we work in
Storage and modelling
SQL Server, Azure SQL and Cosmos DB for relational and document workloads, PostgreSQL where the rest of the stack already runs on it, and Azure Storage underneath. Semantic models on top, so the relationships are named once rather than rebuilt per screen.
Movement
Azure Data Factory for scheduled transfer and transformation, and direct integration where a system can be read in place instead of copied. REST gateways over legacy interfaces when the source system can’t be read any other way.
The unified route
Microsoft Fabric where ingestion, storage, transformation, analysis and visualization should sit in one place rather than five, which is most of the time when a team doesn’t already have a data platform.
What people look at
Power BI where a team already lives in it, and React where the screens belong inside the product instead of beside it. GraphQL when one interface reads across several services and access control has to live in the query.
Input that isn’t data yet
Azure Document Intelligence for PDFs, scans and forms, including handwritten records, and generated Word and PowerPoint where the output has to leave the system.
Questions we get before the first call
Do you build data warehouses?
Yes. A recent one consolidated two operational systems on a daily schedule into a warehouse with a semantic model and dashboards on top, built on Microsoft Fabric with a data engineer on the team. What we don’t do is a multi-year platform programme measured in terabytes. That is a different kind of team, and we’ll say so in the first conversation.Do we have to move our data for you to work with it?
Only where it earns it. A copy made to be readable is a second thing to keep in sync and a second place to get access control wrong, so where a system can be read in place on a schedule, that’s what we do. Where the queries are heavy enough to hurt the operational system, a warehouse is the answer and we’ll build one.What do you build the screens in?
Power BI or Microsoft Fabric where a team already works in them and wants to keep exploring on their own. React or Angular, inside the product, where the screen belongs next to the thing it’s about. That is what the published work shows: in all three cases the data was part of a product rather than a reporting exercise.Different teams can’t see the same things. Can you handle that?
It’s the normal case, not the exception. Smart Portfolio gives each role at the IDB its own view of the same loan data, and that was decided before the indicators were built rather than filtered afterwards. Reversing it later means rebuilding them.
Bring us the question nobody can answer quickly.
In a 45-minute working session we’ll map where that answer currently lives, what it would take to put it in one place, and whether it needs a screen at all. Bring the question; we’ll bring the systems it touches.
