MaskWise: Redact, mask, and anonymize data in training files for LLMs
MaskWise: Redact, mask, and anonymize data in training files for LLMs
If you’re working with LLM training data (like I often am), you’ll know how tricky it can be to scrub out PII without breaking the dataset. I have been using MS Presidio for some time and decided to build a UI on top of it. This is a tool that scans and recognizes sensitive bits in text (eg names, emails, addresses etc), processes images to mask whats sensitive and handles structured data. Everything is written in ts + nodejs, with great help from Claude Code :) It's still early so feedback & contributions are more than welcome.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Biip – Redact PII
Labeled Training Data
Cumul – Concatenate all files in a directory for LLMs
Labelur – Text classifier without training data
GPTCache – Redis for LLMs
prompttest – pytest for LLMs
Jsonnet Training
Ear Training Exercise
kibana training, kibanaonline training, kibana tutorial
Scan IDs, redact sensitive fields, share safer copies