EVIDENCE & LIMITS

Trust should be earned in public.

EDF is a working research framework, not a finished scientific instrument. This page separates what the repository demonstrates, what the current evidence suggests, and what remains unanswered.

INTERNAL DEMONSTRATION

The same field structure can organize varied repository-authored cases.

Challenger, Apollo 11, Apollo 13, Boeing 737 MAX, Pixar, and Toyota can all be expressed through the same EDF fields. This establishes internal representational consistency, not external explanatory validity.

Confidence: medium–high
SUGGESTIVE

EDF may help surface system-level control points.

Current cases often distinguish the nearest failure from broader governance, authority, information, and design controls. No blinded matched-baseline study has yet established better control-point quality.

Confidence: low
UNSUPPORTED

Cross-executor reproducibility has not been demonstrated.

Historical High and Very High labels measured internal structural agreement in a shared research process. The repository also lacks enough raw execution provenance to fully reconstruct those R1 computations.

Confidence in positive claim: low
OPEN

EDF–0 may be fast and learnable enough for routine use.

The current site demonstrates an intended compact workflow. Completion time, training burden, and independent usability remain unmeasured.

Confidence: low

HOW VALIDATION IS PROTECTED

The framework does not silently rewrite itself to fit the cases.

EDF freezes a specification during a validation cycle and records proposed improvements separately. Evidence packages define what a run may use, reducing hindsight contamination.

VERSION INTEGRITY

A case cannot quietly change the rules.

Insights are deferred to later versions instead of retroactively improving the tested framework.

EVIDENCE ISOLATION

A diagnosis is limited to declared material.

Future confirmatory runs must hash and freeze the evidence actually available to the analysis.

EXECUTION PROVENANCE

A result must be reconstructable before it can support a performance claim.

Validation Protocol v1.2 requires exact inputs, raw outputs, failed runs, model and evaluator versions, and executable scoring.

KNOWN GAPS

The interesting questions are still ahead.

Transparency about limitations is part of the framework’s credibility, not a footnote to it.

01

Independent analysts

Current work does not yet show whether people unfamiliar with EDF produce comparable diagnoses.

02

Quantitative measures

Control-point quality, completeness, calibration, and inter-rater agreement need operational measures.

03

Time and learning curve

EDF–0 is intended to be rapid, but completion time and training burden have not been established.

04

Decision outcomes

A better-looking diagnosis is not enough. Future work should test whether EDF leads to better interventions.

READ THE PRIMARY MATERIAL

Do not take the website’s word for it.