Using AI on sensitive data without becoming a headline
Quick tips for putting AI near patient, customer or employee data. I build this stuff at work and study cyber law on the side, so you get both sides: what breaks, and what the law says when it does
For engineers
If you write the code
Mask before the call
Cleaning up the answer is too late, the prompt already left. Every call goes through the filter, even that one script from Friday
Fakes > [REDACTED]
Models get lost on "[PERSON] saw [PERSON]". Swap in realistic fakes and swap the real ones back after. Shifting every date by one offset keeps the timeline right, but shifted dates still count as dates under HIPAA's Safe Harbor rule
Fail closed
Filter down? Weird file? Block the request. An annoying error beats a model that just saw someone's SSN
Log counts, not content
Log "3 names masked", not the names. Logging the actual names just makes a second copy of the data, and debug logs tend to stick around
Test on stuff you didn't tune on
Your detector always aces its own homework. Have someone else write fresh test cases and go by that score. It'll be lower
Streaming breaks things
A fake name can show up split across two chunks. Hold back just enough text to catch it, and let tool calls finish before restoring them
For project managers
If you run the project
Draw where the data goes
App, gateway, model, logs, backups. Get it all on one diagram, because you can't secure a path nobody has written down
Accuracy is a launch blocker
Pick the minimum score before launch and agree that anything below it doesn't ship. Set it early, while nobody's on a deadline
Pay people to break it
Typos, weird formats, names in other languages, IDs pasted from a fax. Every round finds something, so plan for three
Paperwork before data
Vendor touching health data? The BAA (business associate agreement, the HIPAA contract that makes them protect it) gets signed first. "Legal's on it" doesn't count
Start read-only
AI getting access to the CRM or inbox? Start with reading. Add write access one permission at a time, each with an owner
For system designers
If you design the system
One gateway, no side doors
Every model call goes through one place that masks, checks auth and logs. Block direct calls to the provider so old apps can't go around it
Act as the user
Pass the user's own login through. A shared admin account makes your chatbot the most powerful person in the company
Keep less, for less time
Encrypt the real-to-fake map per customer and let it expire. Don't store tokens at all if you can help it
Make logs tamper-evident
Link each entry to the hash of the one before, and keep the chain somewhere admins can't rewrite. Then any edit shows up
Don't page for nothing
Ping chat on the first failure, page after two in a row, and only once per outage. Too many false alarms and people stop answering
Find the real bottleneck
Load test what you don't control, like downstream API limits and cold starts. That's usually what breaks first
For governance & compliance
If you sign off on risk
What happens when it breaks?
Only good answer: requests get blocked. If you hear "temporarily", book a follow-up
Which number, tested on what?
Ask how they tested it. You want results on data the team didn't tune on, and how many examples that was. "99% accurate" on the data they tuned on doesn't mean much
Know the 18
HIPAA Safe Harbor lists 18 identifier types. Every date part except the year goes, ZIPs get cut to 3 digits (000 for small areas), and ages over 89 become "90+". Lots of privacy tools leave dates and ZIPs in
Who did it, and as who?
Every AI action should trace back to a real person, not "the integration". Otherwise there's nothing to investigate
Who can turn it off?
Turning masking off should be admin-only and logged. If any user can flip it in settings, it's not really a policy
New rules to know about
Pulled from the Federal Register a few times a day. Recent federal rules that touch AI, privacy or data.
Oct 7, 2026 · Proposed rule · Justice
Privacy Act of 1974; ImplementationOct 6, 2026 · Final rule · Treasury
Privacy Act ExemptionsSep 29, 2026 · Final rule · Health and Human Services
Medicare Program; Hospital Inpatient Prospective Payment Systems for Acute Care Hospitals (IPPS) and the Long-Term Care Hospital Prospective Payment System and Policy Changes and Fiscal Year (FY) 2027 Rates; Requirements for Quality Programs; Other Policy Changes; and Adoption of Updated Versions of Certain Health Information Technology Standards; CorrectionSep 21, 2026 · Proposed rule · Commodity Futures Trading Commission
Privacy Act RegulationsAug 10, 2026 · Final rule · Homeland Security
9-11 Response and Biometric Entry-Exit Fee for H-1B and L-1 VisasJul 31, 2026 · Final rule · Federal Communications Commission
Modernization of the Nation's Alerting Systems; Protecting the Nation's Communications Systems From Cybersecurity Threats
stuff people actually ask
Can we use a public AI chatbot with patient data?
Only on a business plan where the vendor signs a BAA with you, or with data that's been de-identified under HIPAA's Safe Harbor or Expert Determination method. Masking helps a lot, but it doesn't make data de-identified on its own. Check with your privacy officer
Redaction vs fake values, what's the difference?
Redaction swaps a name for [NAME] and the name is gone. Surrogates are fake stand-in values: a realistic fake name goes in, and the real one gets put back in the answer. Models handle fakes way better
Is de-identified data still covered by HIPAA?
Properly de-identified data isn't PHI (protected health information) anymore. The catch is "properly": miss one date and it was never de-identified. Your real-to-fake map is only OK if it isn't built from patient info and never leaves your hands
Do we need AI rules if it's only used internally?
Yes. Internal tools touch the same CRM and records, and that's exactly where shared admin logins and old debug logs live