Validating artifact schemas and preparing the command center…
Validating artifact schemas and preparing the command center…
Adapt · Red-team arena
Under an equal full-campaign budget, LinUCB outperformed random and rule-mutation search across every fixed seed; the novelty bonus itself produced no measurable improvement.
Contextual bandit
0.807
95% CI 0.519 to 1.212
Random search
0.126
95% CI 0.020 to 0.235
Rule mutation
0.503
95% CI 0.167 to 1.047
Connected points aid seed tracking; they do not imply temporal continuity.
| Seed | Bandit | Random | Rule mutation |
|---|---|---|---|
| 20260812 | 1.579 | 0.303 | 1.562 |
| 20260813 | 0.859 | 0.242 | 0.286 |
| 20260814 | 0.463 | -0.011 | 0.157 |
| 20260815 | 0.535 | 0.074 | 0.394 |
| 20260816 | 0.599 | 0.024 | 0.117 |
A lightweight contextual bandit—not a long-horizon MDP.
Read the defender
Recent detection, fidelity and family coverage form the context.
Choose a campaign
The policy balances promising actions with uncertain ones.
Observe immediate reward
A valid campaign is scored after one defender response.
Update the policy
The next selection uses what the arena just learned.
Matched-pool hardening-stage campaigns—not the equal-budget strategy comparison above.
R pre-fidelity = value + evasion + detection cost + resource cost + novelty
R final = R pre-fidelity × fidelity
Mean fidelity multiplier
× 0.923
Mean final reward
-0.175
At V2, approved value, evasion and novelty are zero in every diagnostic row. Only detection and resource costs remain; fidelity scales their sum.
This receives the same visual weight as the positive baseline comparison.
Two-sided paired test
p = 0.625
Δ = -0.017
The evidence does not support keeping novelty as a performance claim. The term remains logged for auditability.