Keywords: applied category theory, data fabric
Summary: Original proposal about data fabrics for ACT 2025.
Matt Cuffaro, Sean L. Wu, and I wrote this proposal for ACT 2025 which gives a concise overview and motivation for data fabrics:
Robust data science frameworks are those which can adapt to complex, evolving threats in public health. While the OMOP Common Data Model (OMOP CDM) provides an international standard for recording patient electronic health records, such datasets, like patient health records, census survey microdata, health system indicators, and environmental data, requires coordination to federate this data so that it is consumed in a timely way. However, delays from integrating and analyzing these datasets can have severe implications within a public health crisis.
To address these challenges in data integration and analysis, novel mathematical approaches that can handle complex relationships while maintaining computational efficiency are increasingly necessary. We turn our attention to AlgebraicJulia, an open-source ecosystem of Julia packages operating under the philosophy that applied category theory provides a mathematical basis for reasoning about the design and construction of scientific software. The central data structure, the “attributed C-Set,” (ACSet), defined in [2], has diverse and ubiquitous application throughout AlgebraicJulia, from specifying multiphysics models to defining large meshes. However as an in-memory columnar database, it is restricted by the RAM specifications of the client machine, and, by an analogy between database tables and objects in a category, makes large schema (such as the OMOP CDM) difficult to specify. Large (up to terabyte-sized) datasets and complex schema makes working with existing workflows more attractive and in turn slows adoption of AlgebraicJulia tools by working scientists in more data-intensive fields, such as public health.
Our response to the general problem of ergonomically specifying large database schema, as well as the specific challenge of handling the schema and data of the OMOP CDM in the AlgebraicJulia framework takes inspiration from the data fabric, or federated database architectures which provide a unified access protocol for accessing, mutating, and collating data [1]. They were originally developed in order to manage growing data models that were increasingly becoming distributed and mediated by microservices, but have recently found application in academic science [3], [4] We will demonstrate how our implementation of a data fabric meets our requirements of (1) quickly specifying a large database schema which (2) may be decomposed into smaller schema populated by different data sources which (3) may be queried and manipulated through a common access protocol. We borrow an industry term, ”data fabric,” to call our contribution, which is a diagram of ACSets, whose nodes are data sources and edges are foreign key constraints between data sources, which also has database reflection and virtualization, or caching of data.
In this talk we will give an exposition of ACSets and the use-case from public health, specifying the OMOP CDM, to motivate our implementation of the “data fabric.” As a use case example, we will demonstrate how such a framework could assist in the analysis of climate-impacted diseases such as heart attack by integrating this patient data with environmental and census data. Our intent is to show that with relatively little developer time, we can replicate important features from enterprise and scientific uses in databases into an open-source ecosystem while remaining both consistent with the design philosophy of AlgebraicJulia and interoperable with the broader Julia scientific ecosystem. A category-theoretic perspective, by mathematically formalizing what it means to think about database specification and transformation, gives an additional layer of abstraction that helps us reason about the practical challenges in federating data. We argue that this abstraction helps us develop a robust framework for data management in open-source, mathematically-guided science.
T. Priebe, S. Neumaier, and S. Markus, “Finding your way through the jungle of big data architectures,” in 2021 IEEE International Conference on Big Data (Big Data), IEEE, 2021, pp. 5994–5996.
E. Patterson, O. Lynch, and J. Fairbanks, “Categorical data structures for technical computing,” Compositionality, vol. 4, 2022.
I. Buleje, V. S. Siu, K. Y. Hsieh, et al., “A versatile data fabric for advanced IoT-based remote health monitoring,” in 2023 IEEE International Conference on Digital Health (ICDH), IEEE, 2023, pp. 88–90.
A. Panta, X. Huang, N. McCurdy, et al., “Web-based visualization and analytics of petascale data: Equity as a tide that lifts all boats,” in 2024 Ieee 14th Symposium on Large Data Analysis and Visualization (Ldav), IEEE, 2024, pp. 1–11.
Zelko, Jacob S. Data Fabrics ACT 2025 Proposal. https://jacobzelko.com/notes/aaa-0292/. September 28, 2026.