MEASURED, WITH LIMITS

The tradeoffs, in the open.

Explore two recorded synthetic evaluations. These results describe these datasets; they do not establish real-world detection quality.

Recorded September 30, 2026. 6,000 synthetic examples split 3,600 / 1,200 / 1,200 by source and template groups; 600 positives and 600 negatives in the test set. Validation had no extracted negative candidates, limiting model selection.

Sprint 3 · historical holdout — synthetic holdout, 1,200 rows
MethodPrecisionRecallF1
Regex only100.00%50.00%66.67%
Rules + entropy (default)66.56%99.50%79.76%
Historical forest filter (hypothetical)100.00%65.00%78.79%

Rules + entropy (default)

Precision66.56%
Recall99.50%
F179.76%
Confusion matrix — actual rows, predicted columns
Actual / predictedNegativePositive
Negative300True negatives300False positives
Positive3False negatives597True positives

Precision measures how many flagged rows carry the positive label. Recall measures how many positive rows are found. F1 balances both. Labels describe invented examples, not verified credentials.

Read this experiment’s aggregate evidence →

Why the scanner keeps every candidate.

The historical model filter reduces recall from 99.50% to 65.00%, rejecting 207 additional positive examples beyond the rules’ three misses. Optional CLI scoring therefore annotates every finding without hiding any. On the fresh holdout, XGBoost and the retrained forest tie; each still retains 250 authored negative fixtures.

Different populations. Separate conclusions.

Do not compare changes across these two holdouts as model improvement: their groups and distributions differ. Neither dataset has independent human adjudication or unseen real-repository evaluation. Both holdouts are now observed and must not become tuning data. Sprint 4 reproduced the historical results; it was not a new evaluation.

Protocol, interpretation, and reproduction