Validating artifact schemas and preparing the command center…
Validating artifact schemas and preparing the command center…
Reality check · Prevalence collapse
When fraud prevalence falls from the experimental mix to a realistic payment base rate, false positives vastly outnumber true positives—even though ROC-AUC, recall and FPR are unchanged by the mathematical reweighting.
V2 experiment precision · seed 20260812
90.17%
At 9.39% experiment prevalence.
V2 realistic precision · seed 20260812
1.31%
Exact reweighting to 0.015% prevalence.
V2 legitimate FPR · seed 20260812
1.125%
Held fixed during exact prevalence reweighting.
Only the artifact-supported endpoints are observations. Dashed guides are explicitly non-evidentiary.
Observed, downsampled stress and exact-reweighted scenarios remain separate.
| Defender | Scenario | Prevalence | Precision | Recall | FPR | ROC-AUC |
|---|---|---|---|---|---|---|
| V0 | experiment prevalence | 9.385% | 91.89% | 89.83% | 0.821% | 0.9886 |
| V0 | deployment stress downsampled 0.015pct | 0.018% | 1.91% | 89.75% | 0.821% | 0.9843 |
| V0 | deployment exact prevalence reweighted 0.015pct | 0.015% | 1.61% | 89.83% | 0.821% | 0.9886 |
| V1 | experiment prevalence | 9.385% | 90.40% | 97.41% | 1.071% | 0.9961 |
| V1 | deployment stress downsampled 0.015pct | 0.018% | 1.59% | 97.25% | 1.071% | 0.9931 |
| V1 | deployment exact prevalence reweighted 0.015pct | 0.015% | 1.35% | 97.41% | 1.071% | 0.9961 |
| V2 | experiment prevalence | 9.385% | 90.17% | 99.66% | 1.125% | 0.9998 |
| V2 | deployment stress downsampled 0.015pct | 0.018% | 1.55% | 99.50% | 1.125% | 0.9997 |
| V2 | deployment exact prevalence reweighted 0.015pct | 0.015% | 1.31% | 99.66% | 1.125% | 0.9998 |