MEASURED, WITH LIMITS
The tradeoffs, in the open.
Explore two recorded synthetic evaluations. These results describe these datasets; they do not establish real-world detection quality.
Recorded September 30, 2026. 6,000 synthetic examples split 3,600 / 1,200 / 1,200 by source and template groups; 600 positives and 600 negatives in the test set. Validation had no extracted negative candidates, limiting model selection.
| Method | Precision | Recall | F1 |
|---|---|---|---|
| Regex only | 100.00% | 50.00% | 66.67% |
| Rules + entropy (default) | 66.56% | 99.50% | 79.76% |
| Historical forest filter (hypothetical) | 100.00% | 65.00% | 78.79% |
| Actual / predicted | Negative | Positive |
|---|---|---|
| Negative | 300True negatives | 300False positives |
| Positive | 3False negatives | 597True positives |
Precision measures how many flagged rows carry the positive label. Recall measures how many positive rows are found. F1 balances both. Labels describe invented examples, not verified credentials.
Read this experiment’s aggregate evidence →Why the scanner keeps every candidate.
The historical model filter reduces recall from 99.50% to 65.00%, rejecting 207 additional positive examples beyond the rules’ three misses. Optional CLI scoring therefore annotates every finding without hiding any. On the fresh holdout, XGBoost and the retrained forest tie; each still retains 250 authored negative fixtures.
Different populations. Separate conclusions.
Do not compare changes across these two holdouts as model improvement: their groups and distributions differ. Neither dataset has independent human adjudication or unseen real-repository evaluation. Both holdouts are now observed and must not become tuning data. Sprint 4 reproduced the historical results; it was not a new evaluation.
Protocol, interpretation, and reproduction