Skip to main content

Technical~/notes/fine-tuned-model-lost-to-regex.md

My fine-tuned model lost to regex and I'm fine

mlevaluationprivacy

My fine-tuned model did worse than the rules

I'm building a detector that finds patient info in text. It's mostly rules plus a small NER model. I added a step that fine-tunes the model on user feedback, which sounded like an obvious upgrade

Recall dropped a few points. False positives got better, but missing more patient data is the one thing you can't trade for

The gate that caught it

A fine-tuned model only gets switched on if recall doesn't drop and false positives don't rise on held-out data. This one failed, so it got rejected automatically. The worse model never shipped

Most of the real gains came from the boring rules layer. Learning ID shapes from user reports took one test from 0 of 50 caught to 50 of 50