Open questions · active
Does intervention-effect separation improve early warning beyond passive historical risk?
The historical-only benchmark now shows that multivariate prefixes can predict future critical risk in some systems, especially stochastic network collapse, and more modestly in sepsis and r/place. The unresolved question is whether a learned intervention-effect score provides earlier lead time, better discrimination or calibration, lower false-alarm rates on non-critical histories, and more mechanism-relevant temporal localization than the historical-only alert head and passive baselines under independent-entity, leakage-safe rolling evaluation.
Recent evidence · validated
Earlier episodic causal validation exposed leakage and mask collapse
The episodic/meta-batch causal path did not yield a valid positive result. Query visibility was derived from the query critical label and could include the target or later windows, so apparent causal separation depended on oracle truncation. Under a fixed historical prefix, no evaluated simulation simultaneously beat the support baseline, maintained stable positive intervention-effect separation, and avoided fixed-position mask collapse. This negative result motivates the current historical-only baseline, explicit non-critical prefixes, deterministic intervention scoring, and a stricter rolling comparison of intervention versus passive risk.
Outputs in progress · active
Leakage-safe critical-event research platform and benchmark suite
The repository now contains learnable feature/time intervention masks, deterministic intervention-effect scoring, an alert head with explicit abstention, non-critical/no-alert loss support, historical-prefix data preparation, logistic baselines, multi-seed ensemble evaluation, and frozen-test artifacts. It includes synthetic critical systems and historical benchmarks for stochastic network collapse, PhysioNet sepsis, r/place, CHB-MIT, and NASA IMS, plus documented source exclusions where future leakage or insufficient independent events prevent valid testing. The current worktree passes 185 tests with PYTHONPATH set to the repository root.