This document is a draft shared for review and comment. Enter the collaborator code you received to continue.
Find and remove personal information from legal documents before they are stored, logged, shared, or sent to a model.
Here is where each shared asset stands for this tool. The conformance standard is drafted and open for comment. The benchmark set is partly built. The rest follow from them.
How to evaluate a masking system. Three binary gates that make a score irrelevant when failed, five scored criteria covering what it catches and what it wrongly hides, and three result bands. Includes the entity tiers a system is scored against.
Read the standard →Three hundred samples, half clean digital text and half scanned at varying quality, stratified by document type, language mix, and the presence of tables, stamps, and signatures. Gold labels carry span, type, severity tier, and what content must be preserved.
Join the cohort to contribute →Partner templates populated with realistic but fake names, addresses, case numbers, and amounts, covering Illinois, Michigan, Minnesota, Ohio, and Oregon with two counties in each. Each courthouse appears twice, once with an ASCII name and once with a name carrying diacritics, at the same address.
Join the cohort to contribute →The repositories worth copying, with the license and the maturity of each noted. Commercial detection services, open source libraries, and self-hosted deployments are all in scope, and the standard scores any of them.
Coming as the cohort formsThe recommended architecture, the split between deterministic code and the model, the choice between cloud and local deployment, and when to buy rather than build. Which layers should survive a model swap and which are expected to be replaced.
Coming as the cohort formsCase studies from organizations running masking in production. What they chose, what the review queue costs them in staff time, and what they would do differently.
Coming as the cohort formsThe conformance standard says what good looks like. The benchmark set is what you score against. The synthetic set is what you can share and test with when real documents cannot leave the building. The reference implementations are systems the standard has been run against.
You do not need all of them. Pick the resources that match your role.
A masking system takes a legal document containing personal information and produces a version of that document safe to send somewhere else. The destinations that matter here are foundation model providers, researchers, developers, and open datasets.
Every organization doing legal help AI work hits this problem, and most of them solve it alone. The pieces are the same each time: find the identifiers, decide which ones to remove, remove them so they cannot be recovered, and leave enough of the document that a legal worker can still use it.
The last part is what makes legal documents difficult. A masker that blacks out a whole paragraph because it found one identifier has destroyed the deadline, the court location, or the amount owed. Both failures matter, and they pull in opposite directions, which is why the standard scores catching and overmasking as separate criteria.
Character recognition quality drives masking performance in scanned documents, so a system that ingests scans is evaluated end to end on its output rather than on the masking step in isolation.
Who this serves: Legal aid organizations sharing documents with researchers or vendors, courts preparing filings for analysis, teams building evaluation sets, and anyone sending legal text to a model they do not host.
The cohort is collecting this. If you are running masking in production, or you have written your own script and would rather not maintain it alone, tell us and we will connect you with the others working on it.
These are unresolved and marked so nobody mistakes them for settled. The full list sits at the end of the conformance standard.
The PII masking working group is forming now. Members settle the open decisions, contribute documents to the benchmark, and run the standard against the systems they already use.