Skip to content

Benchmarks

Every row of the table below is computed from SYNTHETIC data, either the validation backbone (seed=42, 4 loops, 48 h @ 60 s; four fault archetypes, one per loop, injected during day 2) or seeded noise where the dataset column says so. Each real-data section after the table is written by its study runner and carried through this generator untouched: everything from the first ## REAL heading onward is preserved verbatim. See docs/DATA.md.

A code cell that reads tagledger 0.1.0 names this package before its rename to tsdive; its commit is in the history before the rename.

Regenerated by make bench; CI compares this file byte-for-byte.

stage metric value dataset caveat
1 profile longest constant run, 400-sample freeze planted mid-window 400 samples over 33.25 h; equals the planted run random walk seed=42, 7 d @ 300 s a measurement with no flag; stall at the window end reads 0 s
2 feature extraction replay-stable yes backbone seed=42
2 feature columns per window 19 backbone seed=42
3 MAD screen detection @level_shift 0% backbone seed=42 loop0 provisional
3 MAD false-flag rate (clean day) 0.0% backbone seed=42 loop0 provisional
3 MAD baseline, 119 of 120 samples on one value ZeroSpreadBaseline 0.5 lattice
4 regime screen detection @level_shift 100% backbone seed=42 loop0 MODE-keyed
4 segment break offset from level_shift start (samples) 5 backbone seed=42 loop0 48h the PV lags its own SP step
4 segment break offset from level_shift end (samples) 235 backbone seed=42 loop0 48h no break there: the step is 0.3 of the window's MAD-sigma
4 segment breakpoints over 48 h 10 backbone seed=42 loop0 48h 11 true change points: the SP steps every 4 h
4 segment breakpoints over a clean day 5 backbone seed=42 loop0 day1 the 5 SP steps day 1 holds; none invented
4 segment breakpoints over 48 h (drift loop) 11 backbone seed=42 loop1 48h the 11 SP steps; a ramp is not a step and adds none
4 compare, tag ranked first sim-unit-2:FIC002.PV +0.21 sigma, spread x0.49 backbone seed=42 8 tags (4xPV, 4xSP) the stuck_sensor loop: the freeze holds the PV at the level it reached before its setpoint stepped down, which lifts the day's median and halves its MAD
4 compare, level shift of the level_shift loop's PV +0.0005 sigma (rank 4 of 8) backbone seed=42 8 tags (4xPV, 4xSP) the injected 8-unit step covers 4 h of a day whose PV cycles the three regimes 40/55/70, so the day's MAD is 14.5 and its median moves 0.01
4 compare, largest absolute pair delta -0.265 backbone seed=42 8 tags (4xPV, 4xSP) sim-unit-0:FIC000.SP against sim-unit-2:FIC002.PV, 0.149 -> -0.116: the stuck PV stops following a setpoint. All four loops run one setpoint schedule, so this delta ties across all four SP tags
4 compare, pairs carrying a delta and no bootstrap interval 22 of 28 backbone seed=42 8 tags (4xPV, 4xSP) every pair holding an SP: the differenced setpoint is 0 at all but 5 samples a day, so a block resample that draws none of them leaves the column flat
4 compare, pair interval coverage of a zero delta (MCSE) 0.925 (0.013) white noise seed=42, 2 tags, rho 0.6, 240 differenced rows per period 400 no-change replicates against a nominal 0.95; each interval resamples both periods
4 compare, SPE contribution of the top tag 69% (sim-unit-2:FIC002.PV) backbone seed=42 8 tags (4xPV, 4xSP) over 254 SPE breach rows of 1430; day 1's components keep 0.993 of day 1's variance and 0.905 of day 2's
4 switchback, claim rate at a zero shift (MCSE) 0.060 (0.012) AR(1) phi 0.8 seed=42, 400 records of 480 samples, 16 blocks of 30 one balanced schedule per record, p over 1000 drawn assignments, against a nominal 0.05
4 switchback, detection of a 0.5 sigma shift (MCSE) 0.455 (0.025) AR(1) phi 0.8 seed=42, 400 records of 480 samples, 16 blocks of 30 a step in every B block, sigma = 1.4826 MAD of the record; the 95% interval excludes 0
5 SPC rule hits inside drift interval 5 backbone seed=42 loop1
5 rule BEYOND_3SIGMA hits 0 backbone seed=42 loop1
5 rule RUN_9_SAMESIDE hits 5 backbone seed=42 loop1
5 rule TREND_6 hits 43 backbone seed=42 loop1
6 T2 breaches inside variance_burst interval 0 backbone seed=42 4xPV empirical 99% limit
6 SPE breaches inside variance_burst interval 238 backbone seed=42 4xPV variance faults surface in residuals
6 T2 breaches total (day 2) 9 backbone seed=42 4xPV includes drift/spike loops
6 T2 and SPE breaches with one PV in thousandths unchanged backbone seed=42 4xPV
7 IsolationForest ROC-AUC (fault vs normal) 0.920 backbone seed=42 48x4h windows SYNTHETIC only

REAL: 3W v2.0.0 (real WELL-* instances)

Descriptive profile statistics only. No detector metric appears in this section: when it was written, the splitter in data/splits.py sliced rows inside a stratum rather than holding whole groups out, so a well landed on both sides of the cut and any stage-3-and-above number over 3W would have measured memorisation of a well's own signature. Stage-3-and-above 3W numbers appear in the detector section below, computed under group_holdout.

This section is not regenerated by make bench; it is produced by examples/studies/3w_profile/run_study.py and make_report.py, and carried through the generator unchanged. Rows are over every real instance, with no subset. Method and caveats (in particular how low the flatline thresholds sit on a 1 Hz record) are in examples/studies/3w_profile/REPORT.md.

stage metric value dataset (manifest sha256) code split
1 real instances fetched and converted 1119 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 archives written (instance x present variable, plus 2 label tags) 14347 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 instances profiled 1119 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 tags profiled 14347 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 mean coverage 1.000 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 median coverage 1.000 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 mean valid fraction (rows with a usable value) 0.964 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 median valid fraction (rows with a usable value) 1.000 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 tags whose valid fraction is below 1.000 while coverage is 1.000 19.7% (2824 of 14347) 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 measurement tags with flatline FIRED (1 h window vs earlier windows) 31.3% (2294 of 7337) 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 tags NOT ASSESSED for flatline (all reasons) 49.3% 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 NOT ASSESSED: state-valued tag (role=MODE) 48.5% 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 NOT ASSESSED: archive shorter than two 1 h windows 0.8% 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 NOT ASSESSED: final window entirely NaN 0.1% 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 measurement tags holding one distinct value across the whole archive 22.4% (1643 of 7337) 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 tags refused (any typed refusal) 0.0% 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 absent-variable rate (all-NaN columns per instance) 59.9% (16.2 of 27) 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)
1 non-MODE tags whose unit the alias table cannot resolve 0.0% (0 of 7337) 3W v2.0.0 real 14bc1397d3ee tagledger 0.1.0 @ 46e8d3b1 all real instances; no train/test (descriptive profile statistics only, no detector numbers)

REAL: 3W v2.0.0, detectors (three study designs)

Detector numbers over real data, under three designs, one row per (tool, design). pooled-cross-well fits every model on training-fold windows pooled across wells and scores wells it never saw; GroupSplit.leakage_check() runs before any fit. own-history takes each tag's baseline from that instance's own first 3 windows, chosen by position and never by label, and scores the later ones; stages 6 and 7 fit across instances either way, so they keep the well holdout and z-score each column on the same per-instance baseline (per-instance-standardised). onset-aligned cuts every instance relative to its own fault onset instead of its first sample, so each contributes the same six windows and the clock control has nothing left to read.

The pooled rows for the per-tag tools are marked (artefact, see study). mad_baseline and the SPC limits describe one tag; fitting them on other wells' absolute levels measures how far this well sits from those wells, and that is what produced the below-chance folds (MAD 0.178 on fold 1). The pooled rows are kept and marked as an artefact of the design. The population screen is not marked: its question is cross-instance by construction, so pooling is the design it asks for.

Read the clock control row before any detector row. It reads no sensor value. Its score is how many windows into its own instance the window sits, and at AUC 0.900 it is above every detector under the own-history design. Under the onset-aligned design the same control reads 0.626, and exactly 0.500 once the fixed spacing that design puts between its pre and post windows is removed; that pair is what the aligned rows are measured against. Method, refusal reasons and what survives the clock are in examples/studies/3w_detectors/REPORT.md.

The onset-aligned rows carry four numbers because they answer four questions. AUC ranks every pre window against every post window. The paired hit rate asks only whether an instance's own post window outscores its own pre window, which any score that rises with time wins outright. The clock control takes it 100%/100%. FAR and the post1 detection rate use each instance's own alarm limit: the largest score the tool gave that instance's three baseline hours.

This section is not regenerated by make bench; it is produced by examples/studies/3w_detectors/run_detectors.py and make_report.py, and carried through the generator unchanged. Every row was rerun on tsdive 0.7.0. Only the per-instance-standardised MSPC rows moved, because fit_pca scales each column on its training standard deviation since 0.7.0.

Target: a 1 h window is positive when any row inside it carries fault class 1-9 or transient 101-109. 6,992 windows over 1,119 real instances and 40 wells, 51.4% positive, six common measurement variables.

stage metric value dataset (manifest sha256) code design split
3 MAD screen (regime-blind), max robust z AUC 0.580 (0.178-0.854), refused 0.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
3 MAD screen (regime-blind), max robust z AUC 0.867, refused 35.7% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 own-history own history, first 3 windows per instance, label-blind; no group holdout
4 regime screen keyed by LABEL_state AUC 0.493 (0.158-0.759), refused 4.2% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
4 population screen, frozen (well, variable) refused on 100% of windows: its baseline is the held-out well's own windows, which this split holds out (PopulationTooSparse) 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well 5-fold group holdout by well, seed 42, stratified on folder label
4 population screen under the split its question has AUC 0.508 (0.493-0.524), refused 12.3% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well 5-fold group holdout by instance, seed 42 (NOT a held-out-well number)
5 SPC individuals rule hits AUC 0.545 (0.319-0.797), refused 0.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
5 SPC individuals rule hits AUC 0.754, refused 35.7% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 own-history own history, first 3 windows per instance, label-blind; no group holdout
6 MSPC Hotelling T2 AUC 0.502 (0.337-0.735), refused 53.4% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
6 MSPC Hotelling T2 AUC 0.744 (0.542-1.000), refused 70.2% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 per-instance-standardised 5-fold group holdout by well, seed 42, stratified on folder label
6 MSPC SPE AUC 0.617 (0.434-0.826), refused 53.4% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
6 MSPC SPE refused on 100% of windows: after fit_pca scales each column, every fold keeps all six components, so SPE is 0 up to rounding 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 per-instance-standardised 5-fold group holdout by well, seed 42, stratified on folder label
7 IsolationForest over window features AUC 0.532 (0.442-0.777), refused 0.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well (artefact, see study) 5-fold group holdout by well, seed 42, stratified on folder label
7 IsolationForest over window features AUC 0.706 (0.663-0.813), refused 36.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 per-instance-standardised 5-fold group holdout by well, seed 42, stratified on folder label
none clock control (window position in its own instance, no sensor read) AUC 0.900, refused 35.7% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 own-history own history, first 3 windows per instance, label-blind; no group holdout
3 MAD screen (regime-blind), max robust z AUC 0.799, paired hit 0.812/0.917, FAR 62.5%, detect post1 91.7% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned own history, 3 baseline windows before onset, label-blind; no group holdout; 48 evaluable instances
5 SPC individuals rule hits AUC 0.713, paired hit 0.812/0.812, FAR 43.8%, detect post1 85.4% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned own history, 3 baseline windows before onset, label-blind; no group holdout; 48 evaluable instances
6 MSPC Hotelling T2 AUC 0.636 (0.562-0.900), paired hit 0.893/0.778, FAR 96.4%, detect post1 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned, per-instance-standardised 5-fold group holdout by well over the 21 evaluable wells, seed 42; fitted on training instances' baseline and pre windows only; 28 of 48 instances scored
6 MSPC SPE refused on 48 of 48 instances: after fit_pca scales each column, every fold keeps all six components, so SPE is 0 up to rounding 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned, per-instance-standardised 5-fold group holdout by well over the 21 evaluable wells, seed 42; fitted on training instances' baseline and pre windows only; 28 of 48 instances scored
7 IsolationForest over window features AUC 0.793 (0.745-0.949), paired hit 0.875/0.875, FAR 100.0%, detect post1 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned, per-instance-standardised 5-fold group holdout by well over the 21 evaluable wells, seed 42; fitted on training instances' baseline and pre windows only; 48 evaluable instances
none clock control (window position in its own instance, no sensor read) AUC 0.626, paired hit 1.000/1.000, FAR 100.0%, detect post1 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned own history, 3 baseline windows before onset, label-blind; no group holdout; 48 evaluable instances
none clock control, design offset removed (the instance's onset position, constant across its windows) AUC 0.500, paired hit 0.000/0.000, FAR 0.0%, detect post1 0.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned own history, 3 baseline windows before onset, label-blind; no group holdout; 48 evaluable instances
4 frozen windows flagged by the population screen 4.6% (242 of 5,315) 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well 5-fold group holdout by instance, seed 42 (NOT a held-out-well number)
4 whole-life-frozen tags indicted 9 of 748 reachable (1,643 exist; the rest are outside the common variable set) 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well 5-fold group holdout by instance, seed 42 (NOT a held-out-well number)
6 windows carrying all six common variables 46.8% (3,269 of 6,992) 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 pooled-cross-well 5-fold group holdout by well, seed 42, stratified on folder label
none windows an own-history baseline leaves scorable 64.3% (4,495 of 6,992); 1,935 are baseline windows and 562 sit in the 468 instances of 3 windows or fewer 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 own-history own history, first 3 windows per instance, label-blind; no group holdout
none instances carrying a fault row at all 43.1% (482 of 1,119) 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned all real instances; the other 637 are the negatives-only pool
none instances the onset-aligned design can score 4.3% (48 of 1,119); the rest have fewer than 5 whole windows before onset or fewer than 2 after it, or no fault row at all 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned 48 instances over 21 wells, 288 windows
none evaluable instances whose onset is a transient class (101-109) 100% (48 of 48); 0 have a steady onset 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned 48 instances over 21 wells, 288 windows
none aligned test windows whose rows disagree with the label the design gives them 0 of 144 3W v2.0.0 real 14bc1397d3ee tsdive 0.7.0 @ 11d27513 onset-aligned 48 instances over 21 wells, 288 windows

REAL: TEP (Rieth 2017 simulation, subset)

Descriptive profile statistics only. TEP is a simulation, so the profile stage finds nothing to refuse. Zero gaps, zero duplicate timestamps, coverage and valid fraction 1.000 everywhere, no typed refusal on any tag. The one place the simulator touches the real world is its units, and half of them do not resolve.

This section is not regenerated by make bench; it is produced by examples/studies/tep_profile/run_study.py and make_report.py, and carried through the generator unchanged. Method and caveats (in particular why the timestamps are synthetic) are in examples/studies/tep_profile/REPORT.md.

stage metric value dataset (manifest sha256) code split
1 simulation runs converted 420 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 archives written (run x variable, plus 1 label tag) 22,260 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 runs profiled 420 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 tags profiled 22,260 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 mean coverage 1.000 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 median coverage 1.000 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 mean valid fraction (rows with a usable value) 1.000 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 gaps of any class 0 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 duplicate or non-monotonic timestamps 0 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 samples carrying a BAD or UNCERTAIN severity 0 TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 tags refused (any typed refusal) 0.0% (0 of 22,260) TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 non-MODE tags whose unit the alias table cannot resolve 50.0% (10,920 of 21,840): kPa gauge, kW, kscmh, mol% TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 measurement tags with flatline FIRED (3 h window vs earlier windows) 7.7% (1,688 of 21,840) TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 tags NOT ASSESSED for flatline (all reasons) 1.9% TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 NOT ASSESSED: state-valued tag (role=MODE) 1.9% TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 measurement tags holding one distinct value across the whole run 0.0% (0 of 21,840) TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 flatline fires that are a whole-window flat (plant trip, faults 6 and 18) 51.3% (866 of 1,688, across 33 runs) TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)
1 distinct values of the flatline signal-1 reference (p99 change interval) 83 values, 180 s on 12,249 tags, 360 s on 6,933 tags, 900 s on 2,054 tags, 180 s being the sample period; see REPORT.md TEP Rieth 2017 95f369c2b2b8 tagledger 0.1.0 @ 46e8d3b1 20 runs x 21 faults, training splits only; no train/test (descriptive profile statistics only, no detector numbers)

REAL: 3W Chronos-Bolt zero-shot

A pretrained forecaster, amazon/chronos-bolt-small (revision 772f3d25, chronos-forecasting 2.3.1), scores the detector study's 288 onset-aligned windows with nothing fitted on 3W. The score is the mean absolute error of the one-hour median forecast in units of the 10-90% band, over six variables at one-minute means from a four-hour context. Method, refusals and the per-fold spread are in examples/studies/3w_chronos/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/3w_chronos/run_chronos.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none Chronos-Bolt small zero-shot: instances scored 33 of 48 (68.8%) 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 3c673377 onset-aligned no fit on 3W; own history, 3 baseline windows before onset, label-blind; no group holdout; 33 of 48 evaluable instances scored
none Chronos-Bolt small zero-shot: instances refused 15 of 48; a baseline window with under 60 context minutes or a constant variable 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 3c673377 onset-aligned no fit on 3W; own history, 3 baseline windows before onset, label-blind; no group holdout; 33 of 48 evaluable instances scored
none Chronos-Bolt small zero-shot forecast surprise (no fit on 3W) AUC 0.675 (0.469-0.819), paired hit 0.909/0.545, FAR 18.2% (-0.068 over the 25.0% floor), detect post1 27.3% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 3c673377 onset-aligned no fit on 3W; own history, 3 baseline windows before onset, label-blind; no group holdout; 33 of 48 evaluable instances scored
none clock control on the same 33 instances (window position in its own instance, no sensor read) AUC 0.621 (0.618-0.676), paired hit 1.000/1.000, FAR 100.0% (+0.750 over the 25.0% floor), detect post1 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 3c673377 onset-aligned no fit on 3W; own history, 3 baseline windows before onset, label-blind; no group holdout; 33 of 48 evaluable instances scored

REAL: 3W conformal test martingale

A conformal test martingale (conformal_p_values, mixture_martingale, martingale_alarm in tsdive.eval) over the largest robust z per minute, calibrated on each instance's third window and run over the rest of its record. Under exchangeability of the calibration and stream scores the alarm fires with probability at most delta over the whole record (Ville's inequality); the permutation check makes exchangeability true by shuffling the scores and measures the same alarm. Method, groups and refusals are in examples/studies/3w_conformal/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/3w_conformal/run_conformal.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none conformal martingale alarm, delta 0.05: FAR over the whole normal record FAR 76.3% (bound 5.0%); permutation check 0.04% over 10 seeds; reversed order (calibrate on the first stream window, stream the calibration window) 13.2% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e own-history own history: 2 fit windows, 1 calibration window, label-blind; no group holdout; 537 of 538 normal and 53 of 53 mixed instances scored; cache manifest bfff4f2e897b
none conformal martingale alarm, delta 0.01: FAR over the whole normal record FAR 73.9% (bound 1.0%); permutation check 0.00% over 10 seeds; reversed order (calibrate on the first stream window, stream the calibration window) 11.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e own-history own history: 2 fit windows, 1 calibration window, label-blind; no group holdout; 537 of 538 normal and 53 of 53 mixed instances scored; cache manifest bfff4f2e897b
none worst-baseline threshold on the same scores: FAR over the whole normal record FAR 67.2% (+0.422 over the 25.0% floor) 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e own-history own history: 2 fit windows, 1 calibration window, label-blind; no group holdout; 537 of 538 normal and 53 of 53 mixed instances scored; cache manifest bfff4f2e897b
none clock control martingale (position in the record, no sensor read), delta 0.05: FAR over the whole normal record FAR 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e own-history own history: 2 fit windows, 1 calibration window, label-blind; no group holdout; 537 of 538 normal and 53 of 53 mixed instances scored; cache manifest bfff4f2e897b
none conformal martingale alarm, delta 0.05, mixed records FAR 71.7%, detect 28.3%, median delay 0 windows; worst-baseline FAR 71.7%, detect 24.5% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e own-history own history: 2 fit windows, 1 calibration window, label-blind; no group holdout; 537 of 538 normal and 53 of 53 mixed instances scored; cache manifest bfff4f2e897b
none conformal martingale alarm, delta 0.05 FAR on pre 47.9% (+0.229 over the 25.0% floor), first alarm in post0 33.3%, by end of post1 37.5%, crossed by end of post1 85.4%; permutation check 0.00% over 10 seeds 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e onset-aligned own history: 2 fit windows, 1 calibration window before onset, label-blind; no group holdout; 48 evaluable instances; cache manifest 6f4f097fcd4a
none conformal martingale alarm, delta 0.01 FAR on pre 45.8% (+0.208 over the 25.0% floor), first alarm in post0 33.3%, by end of post1 35.4%, crossed by end of post1 81.2%; permutation check 0.00% over 10 seeds 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e onset-aligned own history: 2 fit windows, 1 calibration window before onset, label-blind; no group holdout; 48 evaluable instances; cache manifest 6f4f097fcd4a
none worst-baseline threshold on the same scores FAR on pre 52.1% (+0.271 over the 25.0% floor), first alarm in post0 33.3%, by end of post1 37.5% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e onset-aligned own history: 2 fit windows, 1 calibration window before onset, label-blind; no group holdout; 48 evaluable instances; cache manifest 6f4f097fcd4a
none clock control martingale (position in the record, no sensor read), delta 0.05 FAR on pre 100.0%, first alarm in post0 0.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.2.0 @ 5d40975e onset-aligned own history: 2 fit windows, 1 calibration window before onset, label-blind; no group holdout; 48 evaluable instances; cache manifest 6f4f097fcd4a

REAL: SKAB

The package's screen, spc and mspc calls and the clock control on SKAB, the Skoltech Anomaly Benchmark, where every labelled record returns to normal after its fault. Each record is scored against its own first three 60 s windows; the profile runs first and its findings (zero-MAD sensors, gapped records) become refusal rows. Method, tables and the conformal bed are in examples/studies/skab/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/skab/run_skab.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
1 sensors with a zero MAD over the whole record (profile, before any detector) Pressure 35 of 35 records; Volume Flow RateRMS 1 of 35 records SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
1 sensors left out of a record's score (zero MAD over the baseline windows) Pressure 30 of 35 records; Volume Flow RateRMS 13 of 35 records SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
3 MAD screen, largest robust z over the usable tags: ranking, labelled records AUC 0.677 (valve1 0.561, valve2 0.549, other 0.843), 577 windows scored, 4 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
3 MAD screen, largest robust z over the usable tags: worst-baseline alarm, labelled records FAR before onset 56.1% (+0.311 over the 25.0% floor), detect 93.9%, median delay 0 windows, FAR after recovery 74.7%, recovered within 2 windows 18.5%; 33 records with a threshold SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
3 MAD screen, largest robust z over the usable tags: false-alarm rate over the anomaly-free record FAR 99.4% (+0.744 over the 25.0% floor) over 164 windows, 0 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
5 SPC individuals rules, hit count over the usable tags: ranking, labelled records AUC 0.735 (valve1 0.676, valve2 0.710, other 0.807), 577 windows scored, 4 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
5 SPC individuals rules, hit count over the usable tags: worst-baseline alarm, labelled records FAR before onset 78.7% (+0.537 over the 25.0% floor), detect 97.0%, median delay 0 windows, FAR after recovery 82.8%, recovered within 2 windows 3.6%; 33 records with a threshold SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
5 SPC individuals rules, hit count over the usable tags: false-alarm rate over the anomaly-free record FAR 98.8% (+0.738 over the 25.0% floor) over 164 windows, 0 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC Hotelling T2, window mean: ranking, labelled records AUC 0.775 (valve1 0.656, valve2 refused, other 0.931), 402 windows scored, 179 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC Hotelling T2, window mean: worst-baseline alarm, labelled records FAR before onset 89.1% (+0.641 over the 25.0% floor), detect 100.0%, median delay 0 windows, FAR after recovery 100.0%, recovered within 2 windows 0.0%; 26 records with a threshold SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC Hotelling T2, window mean: false-alarm rate over the anomaly-free record refused, 164 of 164 windows refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC SPE, window mean: ranking, labelled records AUC 0.749 (valve1 0.587, valve2 refused, other 0.930), 213 windows scored, 368 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC SPE, window mean: worst-baseline alarm, labelled records FAR before onset 90.4% (+0.654 over the 25.0% floor), detect 100.0%, median delay 0 windows, FAR after recovery 95.0%, recovered within 2 windows 9.1%; 14 records with a threshold SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
6 MSPC SPE, window mean: false-alarm rate over the anomaly-free record refused, 164 of 164 windows refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none clock control (window position in its own record, no sensor read): ranking, labelled records AUC 0.716 (valve1 0.705, valve2 0.725, other 0.725), 581 windows scored, 0 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none clock control (window position in its own record, no sensor read): worst-baseline alarm, labelled records FAR before onset 100.0% (+0.750 over the 25.0% floor), detect 100.0%, median delay 0 windows, FAR after recovery 100.0%, recovered within 2 windows 0.0%; 34 records with a threshold SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none clock control (window position in its own record, no sensor read): false-alarm rate over the anomaly-free record FAR 100.0% (+0.750 over the 25.0% floor) over 164 windows, 0 refused SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none conformal martingale alarm, delta 0.05, anomaly_free records false alarm on 1 of 1 records (bound 5.0%); permutation check 0.0% over 10 seeds SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none conformal martingale alarm, delta 0.01, anomaly_free records false alarm on 1 of 1 records (bound 1.0%); permutation check 0.0% over 10 seeds SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none conformal martingale alarm, delta 0.05, labelled records FAR before onset 72.7% (bound 5.0%), alarm inside the span 27.3%, after recovery 0.0%, median delay 29 s; 33 scored, 1 refused; permutation check 0.3% over 10 seeds SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none conformal martingale alarm, delta 0.01, labelled records FAR before onset 72.7% (bound 1.0%), alarm inside the span 27.3%, after recovery 0.0%, median delay 34 s; 33 scored, 1 refused; permutation check 0.0% over 10 seeds SKAB c0d612939333 tsdive 0.7.0 @ 11d27513 own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free

REAL: baseline drift (3W and SKAB)

Three own-history baselines scored by one code path on the 3W window caches and the SKAB archives, with the clock control beside them. static fits the median and MAD of the first three windows' minute medians, differenced fits the same over the minute-to-minute differences, and rolling refits over the 3 windows before every scored window. Method, groups and refusals are in examples/studies/baseline_drift/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/baseline_drift/run_drift.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none static baseline: false-alarm rate on 3W records with no fault window FAR 71.2% per window (+0.462 over the 25.0% floor), 60.0% of records raise one 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none static baseline: 3W records whose fault arrives later AUC 0.855, FAR before the first fault window 70.9% (+0.459 over the 25.0% floor), detect 98.0%, median delay 0.0 windows, median 10.0 fault windows firing in a row 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none differenced baseline: false-alarm rate on 3W records with no fault window FAR 41.4% per window (+0.164 over the 25.0% floor), 68.0% of records raise one 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none differenced baseline: 3W records whose fault arrives later AUC 0.569, FAR before the first fault window 40.1% (+0.151 over the 25.0% floor), detect 81.6%, median delay 0.5 windows, median 2.0 fault windows firing in a row 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none rolling baseline: false-alarm rate on 3W records with no fault window FAR 29.3% per window (+0.043 over the 25.0% floor), 72.0% of records raise one 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none rolling baseline: 3W records whose fault arrives later AUC 0.575, FAR before the first fault window 27.9% (+0.029 over the 25.0% floor), detect 89.8%, median delay 2.0 windows, median 2.0 fault windows firing in a row 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none clock control clock control baseline: false-alarm rate on 3W records with no fault window FAR 100.0% per window (+0.750 over the 25.0% floor), 100.0% of records raise one 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none clock control clock control baseline: 3W records whose fault arrives later AUC 0.893, FAR before the first fault window 100.0% (+0.750 over the 25.0% floor), detect 100.0%, median delay 0.0 windows, median 38.0 fault windows firing in a row 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe own-history own history: 3 fit windows, threshold from window positions 3 to 5, label-blind; nothing fitted across records, so no group holdout applies; folds are wells; 25 records with no fault window and 49 whose fault arrives later; cache manifest bfff4f2e897b
none static baseline, onset-aligned instances FAR on pre 54.2% (+0.292 over the 25.0% floor), first alarm in post0 72.9%, detect post1 83.3% 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe onset-aligned own history: 3 baseline windows before the onset, label-blind; no group holdout; 48 instances; cache manifest 6f4f097fcd4a
none differenced baseline, onset-aligned instances FAR on pre 29.2% (+0.042 over the 25.0% floor), first alarm in post0 54.2%, detect post1 60.4% 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe onset-aligned own history: 3 baseline windows before the onset, label-blind; no group holdout; 48 instances; cache manifest 6f4f097fcd4a
none rolling baseline, onset-aligned instances refused: no threshold window scored: a rolling baseline needs 3 windows before the scored window, and the threshold windows of this design are the instance's first three, which have no predecessors in the cache 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe onset-aligned own history: 3 baseline windows before the onset, label-blind; no group holdout; 48 instances; cache manifest 6f4f097fcd4a
none clock control clock control baseline, onset-aligned instances FAR on pre 100.0% (+0.750 over the 25.0% floor), first alarm in post0 100.0%, detect post1 100.0% 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ bede72fe onset-aligned own history: 3 baseline windows before the onset, label-blind; no group holdout; 48 instances; cache manifest 6f4f097fcd4a
none static baseline: SKAB labelled records AUC 0.617, FAR before onset 36.6% (+0.116 over the 25.0% floor), detect 82.3%, median delay 0.0 windows, FAR after the span 67.4%, recovered within 2 windows 30.8% SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none static baseline: SKAB anomaly-free record FAR 95.7% (+0.707 over the 25.0% floor) over 161 windows SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none differenced baseline: SKAB labelled records AUC 0.547, FAR before onset 28.2% (+0.032 over the 25.0% floor), detect 75.8%, median delay 1.0 windows, FAR after the span 35.6%, recovered within 2 windows 86.4% SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none differenced baseline: SKAB anomaly-free record FAR 68.3% (+0.433 over the 25.0% floor) over 161 windows SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none rolling baseline: SKAB labelled records AUC 0.530, FAR before onset 29.0% (+0.040 over the 25.0% floor), detect 84.9%, median delay 1.0 windows, FAR after the span 36.8%, recovered within 2 windows 76.0% SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none rolling baseline: SKAB anomaly-free record FAR 47.2% (+0.222 over the 25.0% floor) over 161 windows SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none clock control clock control baseline: SKAB labelled records AUC 0.591, FAR before onset 100.0% (+0.750 over the 25.0% floor), detect 100.0%, median delay 0.0 windows, FAR after the span 100.0%, recovered within 2 windows 0.0% SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free
none clock control clock control baseline: SKAB anomaly-free record FAR 100.0% (+0.750 over the 25.0% floor) over 161 windows SKAB c0d612939333 tsdive 0.3.0 @ bede72fe own-history own history: first 3 windows of each record, label-blind; nothing fitted across records, so no group holdout applies; folds are the upstream folders; 34 labelled records, 1 anomaly-free

REAL: 3W audit effect (the checks the detectors run)

The own-history design run twice over the same cache and one code path. checked is the shipped behaviour, where own_history_baselines drops an (instance, variable) pair whose baseline raises InsufficientQuality or whose MAD scale is zero. unchecked keeps every pair and uses a zero scale the way naive code does: the MAD screen returns +inf for any window median off the centre, and the SPC limits collapse onto the centre. Method, groups and refusals are in examples/studies/3w_audit_effect/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/3w_audit_effect/run_audit_effect.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none (instance, variable) pairs the own-history checks drop zero scale 690 of 3870 (541 hold one distinct baseline value, 149 are rounded to a lattice); InsufficientQuality 0 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
none tags holding one distinct value over the whole archive, reachable by this cache 1643 of 7337 measurement tags; 748 on the six common variables, 504 in an instance this design scores 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
3 MAD screen, max robust z, checked baselines AUC 0.8665 over 4495 stream windows (0 scoring +inf); FAR on instances with no fault window 73.5% (+0.484 over the 25.0% floor), 0 of them with an infinite threshold; detect 100.0%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
3 MAD screen, max robust z, unchecked baselines AUC 0.8183 over 4495 stream windows (364 scoring +inf); FAR on instances with no fault window 69.0% (+0.440 over the 25.0% floor), 58 of them with an infinite threshold; detect 90.6%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
5 SPC individuals rule hits, checked baselines AUC 0.7540 over 4495 stream windows (0 scoring +inf); FAR on instances with no fault window 53.4% (+0.284 over the 25.0% floor), 0 of them with an infinite threshold; detect 100.0%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
5 SPC individuals rule hits, unchecked baselines AUC 0.7474 over 4495 stream windows (0 scoring +inf); FAR on instances with no fault window 55.2% (+0.302 over the 25.0% floor), 0 of them with an infinite threshold; detect 100.0%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
none clock control (window position in its own instance, no sensor read), checked baselines AUC 0.8997 over 4495 stream windows (0 scoring +inf); FAR on instances with no fault window 100.0% (+0.750 over the 25.0% floor), 0 of them with an infinite threshold; detect 100.0%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b
none clock control (window position in its own instance, no sensor read), unchecked baselines AUC 0.8997 over 4495 stream windows (0 scoring +inf); FAR on instances with no fault window 100.0% (+0.750 over the 25.0% floor), 0 of them with an infinite threshold; detect 100.0%, median delay 0.0 windows 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 54c04f3d own-history own history, first 3 windows per instance, label-blind; no group holdout; 538 instances with no fault window, 53 whose fault arrives later, 468 refused as too short; cache manifest bfff4f2e897b

REAL: shift intervals (3W, SKAB, TEP, Turbine Upgrade, synthetic)

Before/after level intervals (naive, hac, ewc, block_bootstrap), raw and adjusted on the other tags of the record, scored for coverage on synthetic AR(1) series, TEP fault onsets and injected turbine shifts, and for claim rates on placebo dates where nothing was changed. Rows marked post hoc use a refusal chosen after the 3W result. Method, refusal rules and the pre-registered criteria are in examples/studies/shift_intervals/REPORT.md.

This section is not regenerated by make bench; it is produced by examples/studies/shift_intervals/run_shift.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none naive level interval, raw: coverage at a zero shift, no drift, n 480 per period, pooled over rho (MCSE) phi 0: 0.963 (0.008), phi 0.5: 0.767 (0.017), phi 0.9: 0.315 (0.019), phi 0.98: 0.157 (0.015) SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none hac level interval, raw: coverage at a zero shift, no drift, n 480 per period, pooled over rho (MCSE) phi 0: 0.963 (0.008), phi 0.5: 0.932 (0.010), phi 0.9: 0.872 (0.014), phi 0.98: 0.763 (0.017) SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none ewc level interval, raw: coverage at a zero shift, no drift, n 480 per period, pooled over rho (MCSE) phi 0: 0.957 (0.008), phi 0.5: 0.957 (0.008), phi 0.9: 0.878 (0.013), phi 0.98: 0.618 (0.020) SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none block_bootstrap level interval, raw: coverage at a zero shift, no drift, n 480 per period, pooled over rho (MCSE) phi 0: 0.955 (0.008), phi 0.5: 0.918 (0.011), phi 0.9: 0.710 (0.019), phi 0.98: 0.452 (0.020) SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none hac level interval, raw, R1 refusing: coverage under a 1 SD drift, phi 0.5, n 480 0.010 (0.006) over the answered replicates, R1 refuses 52.2% SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none hac level interval, adjusted, covariate stepping by half the target's shift mean bias -0.445 delta; R3 flags 47.0% of replicates over the 24 cells SYNTHETIC tsdive 0.3.0 @ 12653ee5 synthetic AR(1) 200 replicates per cell, seed from the cell's parameters; nothing fitted across series
none naive level interval, raw, base: 3W placebo date claim rate pooled 83.1%, well-averaged 77.4%, record-level 99.4%; 1802 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none hac level interval, raw, base: 3W placebo date claim rate pooled 65.6%, well-averaged 56.5%, record-level 94.0%; 1802 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none hac level interval, raw, base+R1: 3W placebo date claim rate pooled 47.3%, well-averaged 44.1%, record-level 62.0%; 547 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none ewc level interval, raw, base: 3W placebo date claim rate pooled 63.8%, well-averaged 54.3%, record-level 93.5%; 1802 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none ewc level interval, raw, base+R1: 3W placebo date claim rate pooled 40.2%, well-averaged 40.7%, record-level 51.6%; 547 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none ewc level interval, raw, base+R4 (post hoc): 3W placebo date claim rate pooled 39.5%, well-averaged 32.7%, record-level 41.1%; 448 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none block_bootstrap level interval, raw, base: 3W placebo date claim rate pooled 70.5%, well-averaged 62.5%, record-level 96.7%; 1802 of 3228 (instance, tag) rows answered 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none hac level interval, adjusted, base: 3W placebo date claim rate pooled 74.1%, well-averaged 56.7%, record-level 90.3%; 371 of 3228 (instance, tag) rows answered; width ratio adjusted/raw 0.927 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none hac level interval, adjusted, base+R1+R3: 3W placebo date claim rate pooled 60.3%, well-averaged 50.8%, record-level 74.6%; 78 of 3228 (instance, tag) rows answered; width ratio adjusted/raw 0.996 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 placebo date nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 2 windows before and 2 after, label-blind; cache manifest bfff4f2e897b
none hac level interval, raw, base: 3W fault onset detection 75.9% of 199 answered (instance, tag) rows, beside the placebo claim rate 65.6% 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 known change date (fault onset) nothing fitted across records, so group holdout by well holds by construction; 48 instances, 21 wells, 3 windows before the onset and 2 after; cache manifest 6f4f097fcd4a
none ewc level interval, raw, base: 3W fault onset detection 76.4% of 199 answered (instance, tag) rows, beside the placebo claim rate 63.8% 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 12653ee5 known change date (fault onset) nothing fitted across records, so group holdout by well holds by construction; 48 instances, 21 wells, 3 windows before the onset and 2 after; cache manifest 6f4f097fcd4a
none hac level interval, raw, base: SKAB placebo dates and fault onset claim rate 62.5% on the anomaly-free pairs, 43.9% on the pre-onset halves; detection 54.7% at the onset SKAB c0d612939333 tsdive 0.3.0 @ 12653ee5 placebo date; known change date (fault onset) nothing fitted across records, so group holdout by folder holds by construction; 34 labelled records, 1 anomaly-free, 1 s rows as recorded
none ewc level interval, raw, base: SKAB placebo dates and fault onset claim rate 62.5% on the anomaly-free pairs, 38.7% on the pre-onset halves; detection 53.3% at the onset SKAB c0d612939333 tsdive 0.3.0 @ 12653ee5 placebo date; known change date (fault onset) nothing fitted across records, so group holdout by folder holds by construction; 34 labelled records, 1 anomaly-free, 1 s rows as recorded
none hac level interval, raw, base: turbine placebo weeks and injected shifts placebo claim rate 24.4% of 41 week pairs; injected coverage (MCSE) r 0.02: 0.756 (0.067), r 0.05: 0.805 (0.062), r 0.09: 0.805 (0.062) Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 12653ee5 placebo date; injected shift nothing fitted across records, so group holdout by pair holds by construction; 44 week pairs on 2 turbine pairs, before the upgrade, label-blind; adjacent weeks share weather
none ewc level interval, raw, base: turbine placebo weeks and injected shifts placebo claim rate 41.5% of 41 week pairs; injected coverage (MCSE) r 0.02: 0.585 (0.077), r 0.05: 0.585 (0.077), r 0.09: 0.585 (0.077) Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 12653ee5 placebo date; injected shift nothing fitted across records, so group holdout by pair holds by construction; 44 week pairs on 2 turbine pairs, before the upgrade, label-blind; adjacent weeks share weather
none hac level interval, adjusted, base: turbine placebo weeks and injected shifts placebo claim rate 32.4% of 37 week pairs; injected coverage (MCSE) r 0.02: 0.676 (0.077), r 0.05: 0.703 (0.075), r 0.09: 0.784 (0.068) Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 12653ee5 placebo date; injected shift nothing fitted across records, so group holdout by pair holds by construction; 44 week pairs on 2 turbine pairs, before the upgrade, label-blind; adjacent weeks share weather
none hac level interval, base: vortex generator retrofit, 5,000 rows each side raw -0.564 [-0.827, -0.301] SD, adjusted -0.0005 [-0.0246, +0.0235] SD; no truth known Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 12653ee5 known change date (retrofit) one record; nothing fitted across records
none hac level interval, raw, base: TEP fault-free coverage per variable group, detection over faults 1-20 coverage (MCSE) xmeas_1_22 0.911 (0.006), xmeas_23_41 0.898 (0.007), xmv_1_11 0.908 (0.009); answered 100.0%; detection 1.0% to 82.4% over (fault, group) TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 12653ee5 fault onset with an ensemble truth truth from runs 1-250, scored runs 251-350, samples 1-160 against 161-320; nothing fitted across runs; cache sha256 ee6e1d65cb2c
none ewc level interval, raw, base: TEP fault-free coverage per variable group, detection over faults 1-20 coverage (MCSE) xmeas_1_22 0.924 (0.006), xmeas_23_41 0.942 (0.005), xmv_1_11 0.938 (0.007); answered 100.0%; detection 0.3% to 81.7% over (fault, group) TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 12653ee5 fault onset with an ensemble truth truth from runs 1-250, scored runs 251-350, samples 1-160 against 161-320; nothing fitted across runs; cache sha256 ee6e1d65cb2c
none ewc level interval, raw, base+R4 (post hoc): TEP fault-free coverage per variable group, detection over faults 1-20 coverage (MCSE) xmeas_1_22 0.950 (0.005), xmeas_23_41 0.939 (0.006), xmv_1_11 0.955 (0.007); answered 83.7%; detection 0.5% to 67.8% over (fault, group) TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 12653ee5 fault onset with an ensemble truth truth from runs 1-250, scored runs 251-350, samples 1-160 against 161-320; nothing fitted across runs; cache sha256 ee6e1d65cb2c

REAL: switchback schedules (3W, TEP, Turbine Upgrade, SKAB)

Randomized switchback schedules on records where nothing was changed, with a known shift injected into the B blocks through a first-order lag. The randomization arm is scored beside block_t, hac, ewc and naive intervals on the same schedules and beside the first-half against second-half split of the shift interval study. The rows read the confirmatory run on fresh assignments; the first registration and its verdict are in examples/studies/switchback/REPORT.md, with the method and refusals.

This section is not regenerated by make bench; it is produced by examples/studies/switchback/run_switchback.py and make_report.py, and carried through the generator unchanged.

stage metric value dataset (manifest sha256) code design split
none randomization claim rate at a zero injected shift over every open (L, washout) (MCSE) raw pooled 0.041 (0.002), worst cell 0.054; adjusted pooled 0.039 (0.002), worst cell 0.050; first-half against second-half split of the same records: hac 65.2%, ewc 63.5% of 1770 (record, target) units 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 4 draws each, label-blind; cache manifest bfff4f2e897b; fresh draws, seed from (bed, record, L, draw + 100000)
none highest claim rate at a zero injected shift over tau and washout, raw, other arms on the same schedules block_t 0.051, hac 0.211, ewc 0.184, naive 0.666 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 4 draws each, label-blind; cache manifest bfff4f2e897b; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization raw: coverage of the exact truth pooled over delta > 0 at washout ceil(3 tau); detection at delta 0.25, tau 0 coverage 0.9588 (0.0019); L 15 min: 0.113, smallest delta at 0.8 none; L 30 min: 0.051, smallest delta at 0.8 none 3W v2.0.0 real 14bc1397d3ee tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by well holds by construction; 538 instances with no fault window, 16 wells, 4 draws each, label-blind; cache manifest bfff4f2e897b; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization claim rate at a zero injected shift over every open (L, washout) (MCSE) raw pooled 0.050 (0.001), worst cell 0.051; adjusted pooled 0.050 (0.001), worst cell 0.050; first-half against second-half split of the same records: hac 6.1%, ewc 3.9% of 5200 (record, target) units TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across runs; fault-free runs 251-350, 10 draws each; cache sha256 49ffa27703fe; fresh draws, seed from (bed, record, L, draw + 100000)
none highest claim rate at a zero injected shift over tau and washout, raw, other arms on the same schedules block_t 0.050, hac 0.130, ewc 0.084, naive 0.366 TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across runs; fault-free runs 251-350, 10 draws each; cache sha256 49ffa27703fe; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization raw: coverage of the exact truth pooled over delta > 0 at washout ceil(3 tau); detection at delta 0.25, tau 0 coverage 0.9505 (0.0011); L 2 h: 0.475, smallest delta at 0.8 none; L 4 h: 0.509, smallest delta at 0.8 0.5 TEP (Rieth 2017 simulation, testing split) f57ab3ee3443 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across runs; fault-free runs 251-350, 10 draws each; cache sha256 49ffa27703fe; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization claim rate at a zero injected shift over every open (L, washout) (MCSE) raw pooled 0.048 (0.005), worst cell 0.060; adjusted pooled 0.053 (0.005), worst cell 0.060; first-half against second-half split of the same records: hac 100.0%, ewc 100.0% of 2 (record, target) units Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by pair holds by construction; 2 pairs before the upgrade, 250 draws each; fresh draws, seed from (bed, record, L, draw + 100000)
none highest claim rate at a zero injected shift over tau and washout, raw, other arms on the same schedules block_t 0.056, hac 0.090, ewc 0.194, naive 0.876 Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by pair holds by construction; 2 pairs before the upgrade, 250 draws each; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization raw: coverage of the exact truth pooled over delta > 0 at washout ceil(3 tau); detection at delta 0.25, tau 0 coverage 0.9517 (0.0052); L 1 d: 0.922, smallest delta at 0.8 0.25; L 3 d: 0.732, smallest delta at 0.8 0.5 Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by pair holds by construction; 2 pairs before the upgrade, 250 draws each; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization claim rate at a zero injected shift over every open (L, washout) (MCSE) raw pooled 0.045 (0.004), worst cell 0.045; adjusted pooled 0.053 (0.003), worst cell 0.060; first-half against second-half split of the same records: hac 71.4%, ewc 71.4% of 7 (record, target) units SKAB c0d612939333 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift one record, nothing fitted across records; 500 draws; fresh draws, seed from (bed, record, L, draw + 100000)
none highest claim rate at a zero injected shift over tau and washout, raw, other arms on the same schedules block_t 0.046, hac 0.252, ewc 0.412, naive 0.687 SKAB c0d612939333 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift one record, nothing fitted across records; 500 draws; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization raw: coverage of the exact truth pooled over delta > 0 at washout ceil(3 tau); detection at delta 0.25, tau 0 coverage 0.9549 (0.0037); L 5 min: 0.352, smallest delta at 0.8 none; L 10 min: 0.334, smallest delta at 0.8 none SKAB c0d612939333 tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift one record, nothing fitted across records; 500 draws; fresh draws, seed from (bed, record, L, draw + 100000)
none randomization adjusted on the turbine covariates: median width ratio adjusted/raw, tau 0, washout 0 L 1 d: 0.163, L 3 d: 0.139 Turbine Upgrade 31b2b0c7e33a tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift nothing fitted across records, so group holdout by pair holds by construction; 2 pairs before the upgrade, 250 draws each; fresh draws, seed from (bed, record, L, draw + 100000)
none registered criteria first registration on the first draws: W1 fail, W2 fail, W3 fail, W4 pass (passes on Turbine Upgrade), W5 pass, parked; corrected criteria on fresh draws: W1' pass, W2' pass, W3' pass, W4 pass (passes on Turbine Upgrade), W5 pass, supported 3W, TEP, Turbine Upgrade, SKAB tsdive 0.3.0 @ 46c064d5 switchback schedule, injected shift see the rows above; the corrected criteria were registered after the first run