The evidence engine

Replay that must earn the right to speculate.

Most simulation tools will run any counterfactual you type in, on any baseline, and hand you a confident number. Mine won't. Before a single what-if is permitted, the engine has to reproduce the recorded event — the actual RoCoF, the actual frequency nadir — with nothing added. Only then does it get to ask "and what if a device had been there?" The output is a dossier a regulator can take apart: versioned, provenance-stamped, and stamped with the exact horizon cutoff it was generated under.

0
phases, P0 → P6
0k+
verified events catalogued
0
event detectors
£0m
reconstructed in one period
The rule

No what-if without a reproduced baseline.

The rule is simple and non-negotiable: no counterfactual is permitted until the engine has first reproduced the recorded event — RoCoF and frequency nadir — with nothing added. Not "roughly matched". Not "calibrated to". Reproduced, scored, and recorded as such. A baseline that has not earned that score is not a baseline; it is a guess wearing a lab coat, and the engine treats it accordingly.

This is not a documentation convention I promise to follow. It is enforced in the verbs. The placeholder 2019 baseline that ships in the event catalogue scores Failed — because it should — and the scenario and evidence verbs refuse to speculate on top of it. They do not warn and proceed. They refuse. That refusal ships in the product, on purpose, where a prospective licensee can trip over it and understand exactly what they are buying.

Every consultancy deck in this industry contains a counterfactual nobody can check. This engine's counterfactuals arrive chained to a reproduction of reality that anyone can check — and if the chain is broken, nothing arrives at all.

The motif of this page
"Honest failure is a feature, and it is load-bearing."

A tool that cannot say Failed cannot be trusted when it says Reproduced. The scoring gate exists precisely so that the word "Reproduced", when it does appear in a dossier, means something a sceptic can lean on.

baseline: Failed → downstream verbs refuse baseline: Reproduced → counterfactuals unlock

There is no override flag. If you find yourself wanting one, you are asking the tool to lie to you more efficiently.

The seven phases

From guard rails to dossier, in the only defensible order.

Each phase is a CLI verb with a versioned artifact schema, and each is gated on the one before it. The order is the argument: the rules come first, the credibility gate comes second, and the speculation comes last — if it comes at all.

Guard & provenance gridsim-frequency-export/1

Phase zero builds the rules before the features, which is the only order in which rules ever survive. ForwardInferenceGuard hard-wires the no-forward-inference horizon into every downstream verb — not as a config option but as a constructor argument nothing can omit. ProvenanceTag makes every quantity carry its origin and trust level from the moment it enters. And the frequency export schema gridsim-frequency-export/1 fixes the wire format for every trace the engine will ever emit, versioned from day one so a dossier produced today is still parseable when a regulator opens it in five years.

The credibility gate replay-event

The gate everything else waits behind. A GridEvent from the EventCatalog — a real, recorded disturbance — is replayed with no device attached, and the ReplayValidator scores whether the engine reproduced the recorded RoCoF and frequency nadir with nothing added. The baseline must score Reproduced; anything less and phases P2 and P5 are locked shut. This is where the engine earns — or fails to earn — the right to say anything at all about alternative histories.

Scenario sweeps scenario

Only behind a reproduced baseline do the counterfactuals run. A ScenarioSpec hands the ScenarioEngine an inertia × device × corner grid — the same recorded event, re-run across the declining-inertia risk curve the GB grid is actually descending: 260 → 155 → 120 → 102 → 50 GVA·s. Every cell is a historical counterfactual, never a forecast: "what would this recorded event have done at that inertia, with that device" — a question about the past with one variable changed, which is the only kind of question the architecture permits.

Fleet dynamics fleet

A FleetSpec drives FleetDynamics through the MultiDeviceFrequencyIntegrator, and the physics is kept honest where the industry routinely fudges it: intrinsic ½Jω² rotational inertia is N-additive across devices, fast frequency response is not, and the engine never blends the two into one flattering headline number. Device rotational inertia is carried in its own labelled column of the output, so a reviewer can see exactly which megawatt-seconds are physics and which are control response. Add ten devices and the inertia column adds; the response column does whatever the coupled dynamics say it does.

Cross-validation crossval

The independent residual instrument — the same one that holds the solver to machine epsilon on the validation page — run here as a phase in its own right. Every claim the pipeline has produced so far is re-derived from the raw exports by code that shares nothing with the code that produced it. It is one thing to mark your own homework; it is another to hand the exam script to an invigilator you wrote to be hostile. Nothing graduates to a dossier without passing through this phase.

Evidence dossiers evidence

AvoidedCost and EvidenceReport assemble everything upstream into a versioned gridsim-evidence/1 dossier — JSON for machines, Markdown for humans, both stamped with the horizon cutoff they were generated under. The economics are computed the defensible way and no other: avoided cost equals the demand actually shed by recorded LFDD action, multiplied by the value of lost load. Reconstructed from what happened — never projected from what a vendor hopes will happen. If the recorded event shed nothing, the dossier says the avoided cost was nothing, however disappointing that is commercially.

The historical state estimator gda-replay

The destination the six phases before it exist to make trustworthy. GdaTimeSeries feeds the GdaReplayLoop: hand it any past timestamp and it returns the grid's full solved electrical state at that instant — per-bus voltages and angles, flows, the frequency picture — corrected against the public record and provenance-stamped column by column. Any past timestamp; no future ones, because the guard from P0 is still standing here at P6, exactly where it was installed before any of this was built.

Formula corner

Two equations. Neither of them negotiable.

The frequency integration and the money both reduce to short identities. The engineering is in refusing to decorate them.

df/dt = ΔP · f0 / (2 · H · S)

The swing equation — the basis of every frequency trace the integrator produces. ΔP is the active-power imbalance in MW (the infeed lost, minus whatever response has arrived), f0 the nominal frequency (50 Hz), H the system inertia constant in seconds, and S the rated apparent power base in MVA — so 2·H·S is the stored kinetic energy term, the GVA·s figure quoted everywhere on this site. It runs inside the MultiDeviceFrequencyIntegrator per device, per timestep, with the P3 additivity rules applied as stated: device ½Jω² adds into H·S, device response adds into ΔP, and the two are never laundered into each other.

AvoidedCost = MWshed(recorded) × VoLL

The avoided-cost identity, and the entire economic methodology of a dossier. MWshed(recorded) is the demand actually disconnected by low-frequency demand disconnection during the recorded event — a figure that exists in the public record, not in a model. VoLL is the value of lost load, a published regulatory number. Multiply them. That is the whole trick: the counterfactual device is only ever credited with cost that was demonstrably incurred without it. No willingness-to-pay surveys, no projected market growth, no discount-rate theatre. If the number looks small, that is because it is honest.

Post-event dossiers

Real events, reconstructed and filed.

These are not demo assets. Each is a dossier the engine has actually produced from the recorded event, and each ships as JSON plus Markdown, stamped with the no-forward-inference cutoff it was generated under.

9 August 2019 — the GB blackout day

The day the lights actually went out: the lightning strike, the double infeed loss, the LFDD action that disconnected over a million customers. Reconstructed as a full dossier from the recorded event — the frequency trace, the RoCoF excursion, the demand shed, and the avoided-cost arithmetic run against the disconnection that really happened. The one day everyone in this industry cites; here it is as a checkable artifact rather than an anecdote.

gridsim-evidence/1 LFDD recorded

513.6 MW infeed loss @ 57.9 GVA·s

A 513.6 MW infeed-loss reconstruction at 57.9 GVA·s system inertia — the inertia figure derived as a fuel-mix proxy and labelled as exactly that, with the full generation-mix breakdown attached so a reviewer can re-derive it fuel by fuel. Produced 3 July 2026. A thoroughly ordinary event, which is rather the point: the pipeline treats an unremarkable Tuesday with the same rigour as a national incident.

gridsim-evidence/1 mix breakdown attached

Interconnector trip — 26 June 2026

An interconnector-trip dossier from late June 2026 — a subsea link dropping its flow and the GB frequency response that followed, caught by the detector fleet, replayed through the credibility gate, and filed. Interconnector losses are the growth category of GB frequency events as the import share rises; the catalogue is accumulating them as they happen, each one another baseline future counterfactuals can be gated against.

gridsim-evidence/1 cutoff-stamped

Every dossier records the horizon cutoff in force when it was generated — so it can prove, years later, that it could not have peeked at anything it should not have seen.

The detector fleet

13 detectors. 92,000+ events, none of them anecdotes.

The event catalogue is not hand-curated. Thirteen detectors sweep the lake's historical record and file everything they find under the versioned gda-events/1 schema — each detector acceptance-tested against known history before its output counts. If a detector cannot find 9 August 2019 on its own, it does not ship.

Detector family Watching for Fed by
Physics — the grid's pulse
frequency-excursion Departures from the statutory band — depth, duration, recovery shape 1-second frequency, 392M samples
rocof-event Rate-of-change-of-frequency spikes — the signature of sudden imbalance 1-second frequency, differentiated
infeed-loss Generation units dropping off the system, sized in MW Per-unit metered volumes & physical notifications
interconnector-trip Subsea links losing their flow — the growing category Interconnector flow telemetry
Market — the grid's invoice
price-spike Imbalance price excursions and their physical antecedents Settlement price streams
constraint-bind Boundary corridors hitting their limits, with the cost of managing them Constraint-limit & balancing-cost data
gda-events/1 — versioned schema acceptance-tested vs known history 92k+ events catalogued every event = a candidate baseline

The families above are the shape of the fleet; the thirteen detectors divide the work between them — several variants per family, tuned to different magnitudes and timescales. What matters is the consequence: the credibility gate never runs short of recorded reality to be tested against.

The market-physics gap

The physics and the invoice, on the same axis.

For any half-hour since the records begin. This is the proof that the engine's economics are reconstruction, not estimation.

Balancing spend
£0m
one settlement period
Net energy moved
0 MW
that is the whole net effect
Separate actions
0
priced individually
Weekly gap
9 figures
and rising

One representative settlement period: £2.65 million of balancing spend to move 271 MW net. Not 271 MW of action — 271 MW of net effect, achieved through 993 separate balancing actions, each one priced through the bid-offer ladder, action by action, from the settlement record. The engine reconstructs the whole ladder: which unit, which bid, which offer, what it cost, and what the physics was doing at that exact half-hour. Scale the exercise to a week and the gap between what the electricity was worth and what balancing it cost runs to nine figures.

The point is not the outrage — the point is the method. Anyone can gesture at balancing costs; the engine puts the recorded frequency trace and the recorded invoice on the same time axis and lets you scrub through both. When a dossier claims a device would have displaced a given action, the action it names is a real one, with a real price, from a real half-hour.

The output

What a scenario answer actually looks like.

Three questions, three verdicts — including the one the architecture insists on being able to give.

scenario / feasibility results GATED · REPRODUCED BASELINE
// Q1 — unit-trip counterfactual at a recorded historical instant
{
  "op": "scenario", "kind": "unit_trip",
  "t": "2026-06-11T14:30:00Z",              // settled past — servable
  "baseline": { "score": "Reproduced" },     // the gate, already passed
  "verdict": "SECURE",
  "rocofHzPerS": -0.0004,                    // low ten-thousandths
  "nadirHz": 49.951,
  "bindingConstraint": null
}

// Q2 — 50 MW load addition, same instant
{
  "op": "feasibility", "kind": "load_add", "mw": 50,
  "t": "2026-06-11T14:30:00Z",
  "verdict": "FEASIBLE",
  "headroomMw": 812,                         // >800 MW to spare
  "bindingConstraint": "B6 thermal (not binding at +50 MW)"
}

// Q3 — anything newer than the guard allows
{
  "op": "scenario", "kind": "unit_trip",
  "t": "2026-07-10T09:00:00Z",              // inside the 25 h moat
  "verdict": "INFEASIBLE_HORIZON",
  "detail": "requested instant violates the no-forward-inference
             cutoff; no result was computed",
  "result": null                             // not a worse answer. no answer.
}

The third verdict is the honest one most tools cannot give. Ask about anything the guard has not yet released and the answer is not a degraded estimate with a caveat — it is a refusal, machine-readable and final. INFEASIBLE_HORIZON is not an error state; it is the product working.

Evidence is only as good as the instrument that keeps it honest.

The cross-validation phase you met at P4 has a page of its own — and the evidence engine ships at Professional tier and above.