Digital twin

A causal model of your business, built from your data.

RootCause learns the causal structure over your ontology, at enterprise variable counts, in hours. Bring your own hypotheses and test them against the same data. Every edge is inspectable, attributable and versioned.

Read the method notes
One neighborhood of a 2,400 variable runtested 91 · retained 15 · undetermined 3tested 91 · retained 15 · undetermined 3
  • Promo spend

    • causesWeb traffic
  • List price

    • causesWeb traffic
    • is related to, direction not settled:Stock coverdirection not settled
  • Weatheroutside the business

    • causesWeb traffic
  • Lead time3 lags

    • causesStock cover
    • causesCarrier
  • Web traffic

    • causesUnits sold
  • Stock cover

    • causesUnits sold
    • causesOn-time rate
  • Unknown influencenot in your data

    • is an inferred cause ofSupport loadinferred
    • is an inferred cause ofChurn rateinferred
  • Carrier

    • causesOn-time rate
    • is related to, direction not settled:Support loaddirection not settled
  • Units sold2 lags

    • causesGross margin
    • causesRepeat rate
  • On-time rate

    • causesChurn rate
    • is related to, direction not settled:Gross margindirection not settled
  • Support load

    • causesChurn rate
  • Ends atRepeat rateGross marginChurn rate

Observed Unknown influenceOrientedUnderdetermined

Three relationships here carry an open circle at both ends. The data settled that they are related and could not settle which way the arrow runs, so the model keeps them marked. One inferred common cause sits outside the measured columns and is labelled as such.

Why this is hard

Causal inference works. It has not scaled.

The methods are decades old and well validated. They are also expensive. Conventional causal discovery scales badly in the number of variables, which is why the technique has historically been reserved for a handful of carefully framed questions per year, run by specialists.

Enterprise data does not arrive as a handful of carefully framed questions. It arrives as thousands of variables across dozens of systems.

RootCause runs discovery at sub-quadratic scaling, which changes what is practical. Thousands of models can run at once. Each one rebuilds when the data behind it moves, so the graph in front of a planning meeting matches the business the meeting is about.

Built for parallel scale, with thousands of models kept current as their inputs change.

Two ways in

Learn the structure, or bring your own.

Both routes carry equal weight. Automated discovery covers the whole ontology. Hypothesis testing suits a team that already has a view of how the business works and wants the data to arbitrate it.

Automated discovery

Point RootCause at the ontology and it proposes the causal structure. Conditional independence testing, orientation within equivalence classes where the data supports it, and explicit reporting of what it cannot resolve. Output is a graph with per-edge evidence.

Hypothesis testing

You already believe things about your business. Encode them as edges and RootCause tests each one against the data. Supported, contradicted or underdetermined, with the statistics behind the verdict. Domain knowledge that survives becomes part of the model. Domain knowledge that fails is often the more valuable finding.
Your edge, against the discovered structurePromo spend → Units sold
  • Promo spend

    • causesWeb traffic
    • causesUnits sold
  • List price

    • causesUnits sold
  • Web traffic

    • causesGross margin
  • Units sold2 lags

    • causesGross margin
  • Ends atGross margin

  • Also in the modelUnknown influenceSupport loadChurn rate

Observed Unknown influenceOrientedUnderdetermined
Test a proposed edge
Supportedverdict
Effect estimate as the conditioning set growsDashed line is zero effect.
Test statisticFisher z = 14.6, p < 0.001
Effect+0.28 [0.21, 0.35]
Conditioning setList price, Web traffic, Week of year
Sample1.84M rows, 148 weeks

The association survives every conditioning set tested. Reversing the edge fits the noise structure worse, so the direction is settled and the edge enters the model.

Inspectability

No unexplained arrows.

Click any edge and see what put it there. The conditional independence tests that survived, the conditioning sets, the estimated strength and its interval, the rows involved, and the alternative structures that were considered and rejected. Where orientation could not be determined from data alone, the edge says so. Where a relationship depends on an assumption, the assumption is named on the edge.

Per-edge evidence

Tests, conditioning sets, strength and interval, on every edge.

Named assumptions

Assumptions attach to the edges that depend on them, and propagate to any answer that uses them.

Versioned

Graphs are versioned. Every answer records which version produced it.
Method

What is actually running.

Discovery uses SPARC-fast, our causal discovery engine and a Screening, Pruning, Agreement and Reconciliation Cascade. It is built around a conditional independence test with O(n log n) complexity in place of the quadratic tests conventional approaches rely on. That single substitution is most of the scaling result.

Orientation within Markov equivalence classes uses noise structure that idealized formulations discard. Real-world data propagates noise through causal pathways asymmetrically, and that asymmetry carries orientation information. Where the classical treatment says a direction is unidentifiable, real data frequently identifies it. Where it does not, the edge stays undetermined.

The metadata layer feeds this directly. Known seasonality, monotonic constraints and inferred structural properties enter as priors, so discovery spends its budget on the structure it has to learn.

Runtime against variable countConventional discoverySPARC-fast
Conventional discoverybecomes impractical hereConventional discoverySPARC-fastvariables in the model →runtime →

The axes carry no values. Runtimes and skeleton accuracy are published with the benchmark configuration and the dataset list, because a chart without its configuration is worth nothing to anyone qualified to check it.

Failure modes

What this cannot tell you.

Three limits, stated here so you do not have to find them yourself. Each one is reported on the affected edge inside the product.

Discovery needs variation

A lever that has never been moved has no observable effect to estimate. The model says so and leaves the effect unestimated.

Unmeasured causes stay unmeasured

RootCause flags where a common cause is likely and marks the affected edges underdetermined. No direction is assigned to them.

Aggregation matters

A relationship visible at weekly grain can vanish at monthly. The ontology records grain, so the model reasons about it explicitly and reports when a result is grain-sensitive.

From model to decision

The answers run on top of the graph.

The graph itself is a research artifact. What a planning meeting needs is the interventional answer, and that answer is computed on the graph.

Causal Modeling

Bring a hypothesis you are confident about.

The interesting result is usually the one that fails.