Validating artifact schemas and preparing the command center…
Validating artifact schemas and preparing the command center…
Evidence & governance
The application validates a seven-artifact evidence manifest before rendering. Calibration, negative results, scope boundaries and reproducibility remain inspectable alongside positive findings.
7
Approved evidence artifacts
1.0
V2 campaign diagnostic artifact
1.0.0
Representative campaign artifact
Operating thresholds are calibrated on a temporal validation split disjoint from held-out evaluation.
train
16,800
legitimate rows
validation
5,600
legitimate rows
test
5,600
legitimate rows
Open any API record to inspect the same payload consumed by the interface.
| Claim | Artifact | Field / derivation | Boundary | Open |
|---|---|---|---|---|
| Adaptive attacker reward comparison | artifacts/full_multiseed/aggregate.json | contextual_bandit_reward / random_reward / rule_mutation_reward / significance_tests | Full fixed-budget, five-seed synthetic protocol. | JSON |
| All seed-level strategy and generation points | artifacts/full_multiseed/seed_results.csv | *_reward, v*_held_out_*, v*_fresh_* | No single seed promoted. | JSON |
| Leakage-free threshold selection | artifacts/precomputed/calibration_audit.json | calibration_split, evaluation_split, overlaps, thresholds | Thresholds use validation legitimacy only; test rows are excluded. | JSON |
| Prevalence collapse | artifacts/precomputed/prevalence_metrics.csv | scenario, prevalence, precision_mean, recall_mean, fpr_mean | Seed 20260812 only. | JSON |
| Matched-pool V2 campaign diagnostic | artifacts/precomputed/v2_campaign_diagnostics.csv | valid, detection_rate, fidelity, reward terms, seed summaries | Fixed action pool held constant across generations. | JSON |
| Eight archetypes and scenario variants | artifacts/precomputed/attack_atlas.json | all attack-card fields | Equal-priority hypotheses; not likelihood rankings. | JSON |
| Stateful illustrative rollout | artifacts/precomputed/representative_campaign.json | generator, fidelity, summary, events | Synthetic simulator output; not cardholder data. | JSON |
The full archived protocol and the illustrative quick run remain explicitly separate.
Full archived evidence
python experiments/run_multiseed.py --protocol full
Use for submission-grade replication; results live in the full_multiseed artifacts.
Illustrative quick run
python experiments/run_benchmark.py --quick
Use only for a judge demonstration; never compare its output with the full protocol.
Full-protocol confidence interval method: nonparametric bootstrap 95% CI for the seed-level mean.