ABL-014
No Trace Persistence
architecture-criticalscore 1.000 · rank 1removes trace_persistence · component trace
Hypothesis. Suppressing trace fields tests explainability and auditability contribution without changing decisions.
Mechanism. Trace persistence does not alter the final decision; it is the evidence surface that exposes r, w, z, a, m, θ, R, and Π. Its large score is an auditability result, not a prediction-quality one.
Takeaway. Changes no prediction metric yet carries the largest score — auditability treated as a first-class scientific result.
ABL-001
No Diagnostics
architecture-criticalscore 0.775 · rank 2removes diagnostics · component G
Hypothesis. Removing the diagnostic vector z will reduce governance response to corrupted or unsafe conditions.
Mechanism. Diagnostics sit at the start of the governance chain. Removing them starves severity and downstream threshold behaviour of a primary input, so degradation shows up broadly across safety, trace, and stability.
Takeaway. The entry point of the governance chain; its removal degrades safety, trace, and stability broadly.
ABL-015
Consensus-Only MAVS
architecture-criticalscore 0.775 · rank 3removes governance_stack · component G,A,P,Θ,Π
Hypothesis. Removing diagnostics, severity, organs, adaptive threshold, and hard veto tests whether governance explains MAVS-GC advantages beyond consensus.
Mechanism. Tests whether consensus alone reproduces MAVS-GC. Holding benchmark identity fixed, the governance stack contributes independently of the consensus computation.
Takeaway. Consensus alone does not explain MAVS-GC; the governance stack contributes independently.
ABL-003
Static Severity
architecture-criticalscore 0.701 · rank 4removes adaptive_severity · component A
Hypothesis. A fixed severity level tests whether adaptive per-example severity adds value beyond constant caution.
Mechanism. Severity aggregation converts diagnostic signals into scalar governance pressure. A static level weakens the adaptive bridge between detected stress and threshold policy.
Takeaway. Adaptive per-example severity beats constant caution — large governance-stability and trace effects.
ABL-002
No Severity
materialscore 0.650 · rank 5removes severity · component A
Hypothesis. Computing diagnostics but forcing a = 0 isolates the contribution of severity aggregation.
Mechanism. Absent severity severs the link between detected stress and governed threshold, so effects concentrate in trace and governance-stability terms.
Takeaway. Severity aggregation matters, but its effect concentrates in trace and governance-stability terms.
ABL-004
Random Severity
materialscore 0.588 · rank 6removes semantic_severity_alignment · component A
Hypothesis. Matched-distribution random severity tests whether semantic row alignment matters.
Mechanism. Misaligned severity keeps the distribution but breaks per-row semantics, measuring whether stress information stays attached to the correct example.
Takeaway. Semantic row-alignment matters: shuffling it still degrades trace and governance stability.
ABL-006
Static Rebalancer
materialscore 0.411 · rank 7removes contextual_weight_adaptation · component W
Hypothesis. Static weighted ensemble weights separate static weighting from contextual governance.
Mechanism. The rebalancer controls how specialist evidence enters consensus. Static weights recover some governance-stability loss but are not contextual governance.
Takeaway. Static weighting recovers some stability but is not a substitute for contextual governance.
ABL-005
No Rebalancer
materialscore 0.299 · rank 8removes contextual_rebalancer · component W
Hypothesis. Uniform or fixed weights test whether adaptive specialist weighting contributes to robustness.
Mechanism. A narrow effect indicates the benchmark responds more to detecting and responding to stress than to reweighting specialists alone.
Takeaway. Contextual rebalancing has a narrow effect here — stress detection dominates over reweighting.
ABL-008
No Severity Term
materialscore 0.256 · rank 9removes threshold_severity_term · component Θ
Hypothesis. Setting λ = 0 tests whether flags raise the acceptance threshold under stress.
Mechanism. Θ maps severity and mitigation into the acceptance threshold. Removing the severity term flattens the response to stress; mixed predictive effects are expected as stricter thresholds trade F1 for safety.
Takeaway. The severity term in Θ drives the strongest robustness and specialist-failure contributions.
ABL-010
Static Threshold
materialscore 0.256 · rank 10removes adaptive_threshold · component Θ
Hypothesis. A fixed θ tests whether adaptive threshold policy contributes to robust decisions.
Mechanism. A fixed θ removes row-adaptivity from acceptance, operating through the same severity-linked profile as the severity term.
Takeaway. Adaptive thresholding matters through the same severity-linked profile as the severity term.
ABL-013
Shuffled Threshold
materialscore 0.206 · rank 11removes row_aligned_threshold_semantics · component Θ
Hypothesis. Shuffling θ across examples tests whether row-level threshold semantics matter.
Mechanism. Shuffling keeps the threshold distribution but breaks per-row alignment, isolating semantic threshold placement from average strictness.
Takeaway. Row-level threshold semantics matter for trace stability even when accuracy is unchanged.
ABL-007
No Organs
neutralscore 0.000 · rank 12removes organs · component P
Hypothesis. Removing mitigation evidence tests whether organs prevent over-rejection while preserving safety.
Mechanism. Organs provide bounded mitigating evidence. A neutral aggregate effect means the mitigation pathway did not dominate the paired contributions in this benchmark state.
Takeaway. Inactive in this aggregate evidence — a boundary condition, not proof of irrelevance.
ABL-009
No Mitigation Term
neutralscore 0.000 · rank 13removes threshold_mitigation_term · component Θ
Hypothesis. Setting δ = 0 tests whether mitigation controls over-rejection.
Mechanism. The mitigation term relaxes the threshold under supporting evidence. Neutral here because over-rejection pressure was not activated in the prepared rows.
Takeaway. Neutral here; its role is latent until over-rejection pressure appears.
ABL-011
No Hard Veto
neutralscore 0.000 · rank 14removes hard_veto · component Π
Hypothesis. Disabling hard veto tests extreme-risk protection under severe stress.
Mechanism. Π is the decision policy. A neutral result means the observed rows did not expose extreme-risk veto conditions large enough to move the aggregate.
Takeaway. Neutral in aggregate — the benchmark rows did not expose extreme-risk veto conditions.
ABL-012
Soft Veto
neutralscore 0.000 · rank 15removes hard_veto_discreteness · component Π
Hypothesis. Replacing hard veto with a threshold penalty tests smoothness versus safety.
Mechanism. Softening the veto removes its discreteness. Neutral here; veto behaviour needs targeted stress regimes to activate measurably.
Takeaway. Neutral here; veto discreteness needs targeted stress to activate.