A case cannot quietly change the rules.
Insights are deferred to later versions instead of retroactively improving the tested framework.
EVIDENCE & LIMITS
EDF is a working research framework, not a finished scientific instrument. This page separates what the repository demonstrates, what the current evidence suggests, and what remains unanswered.
Challenger, Apollo 11, Apollo 13, Boeing 737 MAX, Pixar, and Toyota can all be expressed through the same EDF fields. This establishes internal representational consistency, not external explanatory validity.
Current cases often distinguish the nearest failure from broader governance, authority, information, and design controls. No blinded matched-baseline study has yet established better control-point quality.
Historical High and Very High labels measured internal structural agreement in a shared research process. The repository also lacks enough raw execution provenance to fully reconstruct those R1 computations.
The current site demonstrates an intended compact workflow. Completion time, training burden, and independent usability remain unmeasured.
HOW VALIDATION IS PROTECTED
EDF freezes a specification during a validation cycle and records proposed improvements separately. Evidence packages define what a run may use, reducing hindsight contamination.
Insights are deferred to later versions instead of retroactively improving the tested framework.
Future confirmatory runs must hash and freeze the evidence actually available to the analysis.
Validation Protocol v1.2 requires exact inputs, raw outputs, failed runs, model and evaluator versions, and executable scoring.
KNOWN GAPS
Transparency about limitations is part of the framework’s credibility, not a footnote to it.
Current work does not yet show whether people unfamiliar with EDF produce comparable diagnoses.
Control-point quality, completeness, calibration, and inter-rater agreement need operational measures.
EDF–0 is intended to be rapid, but completion time and training burden have not been established.
A better-looking diagnosis is not enough. Future work should test whether EDF leads to better interventions.
READ THE PRIMARY MATERIAL
The machine-readable source of truth for current claim states.
Read source ↗FINDINGSWhy historical agreement labels do not establish independent reproducibility.
Read source ↗POLICYThe evidence language and promotion gates applied to current claims.
Read source ↗PROTOCOLThe controlled, provenance-preserving protocol for future runs.
Read source ↗