EXPERIMENTAL LOCAL ML

A score is another signal.

The integrated random forest is an optional annotation tool. Rules plus entropy remain the default, and no finding is removed by model score.

13 features

Lengths, entropy, character composition, format indicators, and assignment context. Labels, IDs, provenance, and split membership are excluded.

100 trees

The historical forest uses maximum depth 6 and minimum leaf size 2. A fixed validation-only search selected a 0.6 annotation threshold.

Local inference

Explicit artifact loading and an independently trusted SHA-256 are required. A matching hash alone does not make an untrusted pickle safe.

What the model learned from.

The baseline corpus contains 6,000 balanced, invented examples grouped by source and template family. Candidate extraction finds no negatives in training or validation, so the forest trains on annotated values. That population mismatch limits what validation tells us about actual scanner candidates.

Dataset provenance and reproduction

Known failure: missed database passwords.

The historical hypothetical filter misses 210 of 300 positive database-password examples. Three already fail candidate extraction; the model rejects another 207. Provider-shaped dummy values can also look like intended positives. Shape and entropy cannot establish intent or credential validity.

Inspect the precision and recall tradeoff →

What is not established.

Scores are not calibrated probabilities. Independent review, real-repository performance, score calibration, and robustness across languages and providers remain future work. Proposed 0.4/0.8 confidence bands are not implemented. XGBoost is a separate evaluation artifact, not the CLI model; Ollama is planned.

Full model card and artifact trust policy