Eight things you cannot do with unlinked data
Every integration vendor promises to cut reconciliation effort, and most of them can, so a page
that only counts saved hours proves nothing about
this architecture.
What follows is the harder argument: eight questions a life-science organization actually asks,
where the classical estate loses on structure, not efficiency.
Running each
regulated object as a governed holon changes what is possible, not just what it costs, and every
figure below is measured from a running synthetic holarchy rather than estimated.
The money is at the bottom of the page, where it belongs.
Where these numbers come from
Behind this page sits complete synthetic pharmaceutical use cases built as holons: the MRX-302 product, the AXION-3 Phase III study, a clinical-supply lot, a lab result, a quality event, a marketing-authorization submission, and the explorer that renders them. Six holons, all built, all validated on every build.
How: 11,204 walkable nodes, from a 5,783-triple supply lot to a 399,024-triple clinical study, every one of them projected from the same four-graph pattern.
How: a headless browser walks the real pages and asserts what they draw, alongside per-holon verifiers and a holarchy-wide join check. A claim that stops being true stops the build.
How: drift is a failure, not a warning. If a documented count, a file digest or a cross-holon link no longer matches the graph, verification fails rather than reporting a discrepancy nobody reads.
“Which hospital runs the site of the trial that evidences this asset?”
An ordinary question that crosses four business domains: portfolio, product, clinical, and operations. Classically it crosses four systems too, and there is no single place to ask it. Someone builds an integration, or a warehouse join, or emails three people.
The obstacle is not query performance, it is that the join does not exist as a fact. Each system holds its own key for the same thing, so the relationship lives in ETL code, in a mapping spreadsheet, or in a person's head. Every new question needs a new pipeline, and every pipeline is another copy to keep in step. The estate grows quadratically in connections while the questions stay linear.
In the cellular model a node another holon speaks for is a door, not a foreign key. Diving it hands you to that holon and the walk continues, so one path crosses four graphs without the walker ever being told they changed system. No integration is written, because nothing is being integrated: the identity was shared to begin with.
How: organization → R&D portfolio → MRX-302 → AXION-3 → a site → the institution that runs it. Asserted end to end by the test suite on every build.
How: 296 nodes across the holarchy are spoken for by another holon and resolve as doors. None of them is a pipeline; each is a shared identity declared once.
How: the walk above was never designed. It is a consequence of the links existing, which is why the next cross-domain question costs nothing to answer either.
“Is this object valid right now?”
Not “did it pass validation when it was loaded”, and not “will it pass when we next run the checks”. Right now, as you ask.
Classically, validation is an event at a boundary: a load job, a nightly rule run, a pre-submission sweep. Between two runs the estate has no opinion about its own correctness, and the gap is where the work hides. The pre-filing “consistency sweep” that costs weeks exists precisely because validity decayed silently since the last check.
A holon carries a membrane: SHACL shapes in its boundary graph, validated against its own interior on every build, reporting intact or compromised. Validity stops being a report and becomes a property of the object. During construction of this very corpus the membrane caught a capability that had been added without the ownership link its own rule requires, and refused the build until it was fixed — which is the behaviour, not an anecdote about it.
How: every build runs pyshacl over each holon's interior graphs against its own boundary shapes. All six currently report INTACT, 0 violations, 0 warnings.
How: the check is part of building the object, not a job scheduled after it. There is no window in which the object is changed but unverified.
How: a compromised membrane is not a dashboard row. It stops the artifact being produced, which is what makes the green state mean something.
“Prove to me that this data is governed”
The question an inspector actually asks. The classical answer is a governance policy in SharePoint, a stewardship RACI, and a sample of records pulled by hand to show the policy was followed.
Because the policy and the data are different artifacts in different tools, the only bridge between them is human testimony. Nobody can execute a PDF. So assurance is sampled, periodic, and expensive, and the honest answer to “is it governed today?” is “it was, in March, for the records we looked at”.
Here a governance question is a query with a row count. Each holon carries its own competency questions and answers them against its own graph; the answer is either rows or a failure, never a reassurance. The Case page walks 500 of them for a single study, and the FAIR governance ontology adds 29 more that every holon must satisfy.
How: the FAIR governance ontology's own 29 questions are executed against each holon's TriG. All six holons answer all 29 — 174 answered questions, re-run on every build.
How: all answered from the same governed data, with no purpose-built extract behind any of them. Walk them on the Case page.
How: an auditor can re-run every one of them, on the live graph, and get the same rows. Sampling is replaced by re-execution.
“Is this documentation still true?”
Every data estate has an architecture diagram that was accurate once. Drift — between what the documentation claims and what the system is — is not a failure of discipline. It is the expected behaviour of keeping the description in a different place from the thing.
A diagram in Visio, a catalogue in a data-governance tool, a runbook in Confluence: none of them is connected to the system's own state, so none of them can notice when it changes. Teams respond with review cycles, which are just scheduled attempts to re-synchronise two things that were never joined.
The explorer that renders this holarchy is itself a holon in it. It describes its own modules, widgets, rules, and history in its own graph, measures its own files at build time, and a verifier compares every documented count and digest against reality. There is no version of this repository in which the documentation is stale and the build is green.
How: modules, widgets, tiers, rules, events, documents, diagrams and a qualification dossier — 1,375 walkable nodes, all generated from measurement rather than typed by hand.
How: a documented count that no longer matches a fresh measurement is reported as drift and fails verification. A document cannot go stale while the build passes.
How: per-holon verifiers plus a holarchy-wide check that every declared cross-holon link resolves to a holon that actually exists. All report zero.
“What does this field mean, and may I use it?”
Asked constantly, answered badly. The definition is in a data dictionary, the licence is in a contract, the provenance is in a load log, and the vocabulary it should conform to is in a standard nobody linked to.
The four answers live in four places, and none of them travels with the value. Copy a column into a report and its meaning does not come along. This is why the same fact gets re-derived, re-approved and re-argued in every system it reaches, and why a data dictionary is out of date the week after it is written.
Every predicate a holon asserts carries its own governance: an identifier, the vocabulary it belongs to, its licence, its provenance, and where its definition is published — expressed in the FAIR Governance Ontology, queryable, and validated by shape. Meaning is attached to the fact, so it survives every projection the fact appears in.
How: from a real enterprise unification of 20+ years of clinical data — ~3,500 studies, ~900,000 subjects, 3M+ source variables onto ~4,000 governed targets across 50 domains. The mechanism on this page is what makes that ratio reachable.
How: identifier, knowledge representation, vocabulary, licence and provenance are asserted for every attribute a build actually uses — derived, not hand-written, so coverage cannot lag behind the data.
How: 11 of 17 are met. The 6 that are not — retrieval protocol, access control, metadata persistence, a deployed catalog, registration in a search index — are recorded as failed with the reason, rather than scored green. An assessment that never fails is not an assessment.
“Show the same record to a clinician and to a data steward”
They need different things from one record. The clinician needs the institution running the site. The steward needs the shapes governing the class, the named graphs, and the generic/synthetic split.
The classical answer is a second layer: a semantic model over the warehouse, a curated BI dataset, a business glossary mapped to physical columns. Every one of those is another copy with its own refresh, its own owner and its own drift. Two audiences, two artifacts, two things to keep true.
Here both audiences read the same build. A view is a filter over what was already made, not a second construction, so switching costs a re-slice rather than a rebuild — and the camera, the position and everything already revealed survive the switch. The business view names nothing after the data: the folders called Attributes and Rules simply are not there, and a record's own facts stand in the record's own space.
How: 49.1% business, 50.9% data, 0 untagged. Assigned at export time from each node's own class and attribute names — there is no third, unclassified default.
How: one build is kept whole and filtered per view. Nothing is duplicated, so nothing can disagree.
How: asserted by the test suite — the same space yields 7 business nodes and 3 data nodes, and a detour through one view and back restores every node the visitor had revealed.
“How did this object get to where it is?”
Not a change log of rows, but the object's own history: what happened to it, in order, attributable, and still attached to the thing it happened to.
Audit trails are bolted on per system and truncated per retention policy, so an object's story is scattered across as many histories as it has systems, each with a different notion of time and identity. Reconstructing one narrative means joining audit tables that were never designed to be joined — which is why regulatory reconstruction is a project, not a query.
The event graph is one of the object's four graphs, not an appendix to it. Every fact carries the event that produced it, so provenance is a property of the record. Because the history is structured rather than logged, it can be read as evidence and as shape — the same 800 dated events that prove what happened also draw the study's own arc over time.
How: read straight off the study holon — no audit-table join, no retention window, no separate history store to reconcile against.
How: each event is placed by whichever dated attribute it actually carries, so the timeline is derived from the facts rather than maintained as a second dataset.
How: because provenance is attached to the fact and never separated from it, “what happened to this object” is answered where the object lives.
“Will this still work two orders of magnitude up?”
The question that kills most architectures, because the honest answer is usually “not without re-platforming”.
Classical designs are sized for a volume. A model that works for one product line is rebuilt for the portfolio, then rebuilt again for the enterprise, because the structure and the scale were decided together. Each rebuild re-litigates the same modelling arguments and strands the analytics built on the previous shape.
The four-graph pattern is fractal: it is the same at every level, so the object that holds a single lab result and the object that holds a 399,024-triple trial are described, validated, projected and walked identically. Growth is more instances of the same shape, not a different shape. And no space ever draws more than its budget, so a holon with 32,826 recorded instances stays as readable as one with forty.
How: 5,783 triples for the clinical-supply lot, 399,024 for the AXION-3 study. Same four graphs, same builders, same verifiers, same explorer.
How: complexity is offered on demand rather than dumped: what does not fit waits behind a node of its own, and raising the budget never hides what was already revealed.
How: the holarchy went from five built holons to six, and from one taxonomy to a 317-concept capability library, without a single builder changing its pattern.
A governance scorecard the graph computes about itself
No survey, and no invented score either. Because every holon's state, rules, history and views are linked, the holarchy can be asked what it is actually like — and answer with counts. These three reports are produced by the DISYNPHiA holon's own data-governance capability, re-measured against all 11,204 nodes on every build, and they are shown here exactly as they came out.
| What was measured | Result | Read it as |
|---|---|---|
| Nodes classified business or data | 100% | 0 untagged — 49.1% business, 50.9% data |
| Nodes carrying exactly one data scope | 100% | 0 untagged — 36.9% generic, 63.1% synthetic |
| Nodes whose class carries a curated definition | 66.1% | 7,401 of 11,204 — 3,803 not yet curated |
| Holons with an intact membrane | 6 / 6 | 0 violations, 0 warnings |
| FAIR competency questions answered | 29 / 29 | on every one of the six holons |
| Drift between documentation and graph | 0 | enforced — drift fails the build |
The efficiency case, kept in its place
The eight use cases above are the argument, because they are the part a classical estate cannot answer at any budget. The savings below are real too, but they are the ordinary consequence of not doing the same work repeatedly — and unlike everything above them, they are planning estimates, not measurements. They are marked as such and belong at the bottom of the page.
| Holon | Effort | Time | Illustrative annual saving |
|---|---|---|---|
| Clinical Study | −45% data management | ~40% faster to lock | ≈ $1.1M / 25 studies |
| Dossier Submission | −50% assembly & QC | ~3 weeks earlier filing | ≈ $13k / submission |
| Batch / Lot | −35% release review | 5 → 2 days to disposition | ≈ $430k / 600 batches |
| Quality Event | −40% per investigation | 45 → 30 days to closure | ≈ $480k / 500 events |
| Safety Case | −30% per ICSR | weeks → days to signal | ≈ $2M / 10,000 cases |
| Medicinal Product | −60% product-data upkeep | days → seconds per query | ≈ $250k / portfolio |
| Lab Result | −50% review & transcription | hours → minutes to release | ≈ $1.6M / 200,000 results |
Each holon saves work on its own, but the return is shared: the governed product identity feeds the study, the batch, the safety case and the submission at once. Map a fact once and every domain that reuses it stops paying to re-create it — which is why per-holon figures understate the result when a whole organization runs this way.
Ask for a cross-domain question answered without a new pipeline. Ask whether the object knows if it is valid right now. Ask for governance you can execute rather than read. Ask what happens to the documentation when the system changes. Ask what breaks at ten times the volume. The efficiency claims will sound similar everywhere; these five will not.
The point: value here is not a slide, not a survey, and not a saving. It is a set of questions that become answerable — and every claim on this page is a count taken from six graphs that fail their own build when the claim stops being true. See it end to end on the Case page.