Back to paper

Evidence record · 2026

The MAVS-GC benchmark program

The complete verified record behind MAVS-GC: clean accuracy, robustness under corruption, stability, dynamic sequential validation, and the Diagnostic Sciences correlated-failure fix.

Saif Malik · MAVS Research Program · Chapters 10A–10D + Specials-1
10A

Clean accuracy

Chapter 10A prevents overclaiming. Under clean benchmark conditions, MAVS-GC is competitive but not dominant: it produced positive metric deltas in 79 of 288 comparisons, improved accuracy over Veto MAVS in 2 of 8 and over the Static Weighted ensemble in 0 of 8. Governance shifts the error profile — often raising precision and reducing false positives while lowering recall/F1 — rather than raising the accuracy ceiling.

10B

Robustness under corruption

Chapter 10B is the strongest classical signal: under stress, MAVS-GC fails more safely. Across four datasets and nine corruption families, governed consensus suppresses unsafe acceptance by up to ~200×.

Failure behaviour under corruption

Chapter 10B · accuracy vs. unsafe acceptance

Accuracy (higher is better) Unsafe acceptance (lower is better)
Pure MAVS-GCours
89.95%
1.35%
Mean / Veto
74.31%
27.29%
Single model
59.46%
45.42%

Under specialist-failure corruption, Pure MAVS-GC keeps accuracy high while unsafe acceptance stays near zero — roughly 20× lower than ensemble baselines and 34× lower than a single model.

Figure 1. Accuracy vs. unsafe acceptance under corruption (Chapter 10B). Toggle between specialist failure and high corruption.
10C

Reproducibility & stability

Clean-condition reproducibility gains are limited, but stability preservation strengthens as corruption increases.

MetricPure MAVS-GCBaseline
Prediction stability0.9716150.952713
Decision stability0.9757700.958762
Consensus stability0.9793320.963946
Trace stability0.9679760.959693
Table 2. Behavioural stability under corruption (Chapter 10C): Pure MAVS-GC vs. aggregation baseline.
10D

Dynamic validation

Chapter 10D moves from static rows to sequential episodes with corruption schedules, recovery periods, and hidden safety labels. The full minimum run produced 383,200 trace records across 13 experiments (E1–E5), 2,530 episodes, and 12,161 failure cards, with trace and audit-trace completeness at 1.0000.

The finding is precise: MAVS-GC achieved near-zero unsafe acceptance across the reported rows, but behaved conservatively — sometimes paying for safety with elevated false rejection. It is auditable and safe, not universally superior.

EnvironmentRewardUARFRRRank
Text Safety Stream0.99320.00000.00742 / 14
Tool-Use Security0.89740.00000.17034 / 14
Synthetic Ops0.85780.00000.18045 / 14
Table 3. E3 governance-baseline comparison (MAVS-GC rows). Unsafe acceptance stays at zero across all three environments.

Negative result — motivates DS-CF

On correlated representation collapse, MAVS-GC avoided unsafe acceptance (UAR 0.0000) but collapsed into rejection: FRR 1.0000, mean reward 0.0250, collapse sensitivity −1.0000. Chapter 10D explicitly does not claim MAVS-GC solves correlated failure — which is exactly the weakness the next result fixes.

See how Diagnostic Sciences (DS-CF) fixes this ↓
S-1

Specials-1 · Diagnostic Sciences (DS-CF)

The Diagnostic Sciences correlated-failure fix retargets governance from punishing correlation to punishing harmful correlation. It is a governance-only change (no training) that removes the false-rejection collapse from Chapter 10D while preserving zero unsafe acceptance.

Correlated-failure fix (DS-CF) — before vs. after

Phase 4 A/B · governance-only, no model training · traces 1.0000 complete

Unsafe acceptance (UAR)
before
0.000
after
0.000
preserved at zero
False rejection (FRR)
before
0.316
after
0.000
removed in core A/B
Mean reward
before
0.688
after
0.844
improved

DS-CF changes governance from punishing correlation to distinguishing harmful correlation from safe consistency. It removes the false-rejection failure mode while keeping unsafe acceptance at zero — validated across 1,794 governance decisions.

Figure 4. Correlated-failure A/B: original MAVS-GC vs. DS-CF.
False rejection across independent holdouts

Phase 5 · original MAVS-GC vs. DS-CF · all DS-CF unsafe acceptance = 0

original FRR DS-CF FRR
DS-CF diagnostic ablationreward 0.819
0.488
0.000
Synthetic correlated failurereward 0.838
0.319
0.000
Multi-agent triagereward 0.826
0.259
0.074
Tool-use securityreward 0.938
0.158
0.000
Cyber triagereward 0.896
0.125
0.000
Dynamic corruption · text safetyreward 0.917
0.105
0.000
External taxonomy projectionreward 1.000
0.000
0.000
Stress schedule sweepreward 0.250
0.000
0.000

Average false rejection drops from 0.182 to 0.009 across eight families with zero unsafe acceptance. The only residual is multi-agent triage, where safe-consistency evidence was masked — an evidence-availability limit, not a hard-veto failure.

Figure 5. False rejection across eight independent holdout families (Phase 5).

A trace audit over 1,794 governance decisions found zero raw-correlation-only vetoes, 94 valid conjunctive hard vetoes, zero hard-veto rule violations, and 40 / 40 ambiguous cases escalated correctly.

11B

Governance ablations

Chapter 11B decomposes the robustness signal by removing one governance mechanism at a time. Trace persistence, diagnostics, and the severity→threshold chain explain most of the degradation.