Dmytro Onyshchenko
← All case studies

Multi-Source Data Access with Cube.js

A Cube.js semantic layer that gives backend services one clear interface for reading and joining data from multiple data sources at runtime.

Cube.jsSemantic layerRESTGraphQLSQLCaching

Customer

An industrial machinery business that engineers energy-efficient heating and cooling products and solutions for buildings, industry, infrastructure, and the food cold chain, with about 11,300 employees worldwide and €3.1bn in net sales.

Problem

The platform has multiple data sources, and services need to read data from them and join it at runtime. Doing this inside each service duplicates logic and couples services to how and where data is stored.

Architecture

The idea is to add a presentation layer on top of the data sources with Cube.js. Any backend service calls Cube.js, which reads from multiple data sources. Cube.js gives a clear interface for data reads and brings features that help here:

  • a data model of cubes and views, where a view can create a single facade over two separate data sources;
  • aggregations;
  • caching and pre-aggregations;
  • access control based on a security context from the validated token;
  • REST, GraphQL and SQL APIs.

My contribution

  • Designed how Cube.js fits into the existing infrastructure, so services could read from multiple data sources through one layer without reworking them.
  • Built, configured, and deployed Cube.js as part of that infrastructure.
  • Defined the cubes and views and the main configuration, which expose data from separate sources as a single, consistent model.

Outcomes

  • Improved data layer: services read through one semantic layer with a clear, consistent interface instead of each integrating with the data sources directly.
  • Removed duplication: cross-source joins and aggregations are defined once in the data model instead of being repeated in every service.
  • Improved performance through Cube.js caching and pre-aggregations.
  • Better encapsulation of data: services depend on cubes and views, not on how and where the data is stored.
  • Centralized access control, based on the security context from the validated token.