Cohort forming

PII Masker

Find and remove personal information from legal documents before they are stored, logged, shared, or sent to a model.

← All Common Infrastructure Tools

The Package

Here is where each shared asset stands for this tool. The conformance standard is drafted and open for comment. The benchmark set is partly built. The rest follow from them.

Open for review

Conformance Standard

How to evaluate a masking system. Three binary gates that make a score irrelevant when failed, five scored criteria covering what it catches and what it wrongly hides, and three result bands. Includes the entity tiers a system is scored against.

Read the standard →
Working draft

Benchmark Set

Three hundred samples, half clean digital text and half scanned at varying quality, stratified by document type, language mix, and the presence of tables, stamps, and signatures. Gold labels carry span, type, severity tier, and what content must be preserved.

Join the cohort to contribute →
Working draft

Synthetic Document Set

Partner templates populated with realistic but fake names, addresses, case numbers, and amounts, covering Illinois, Michigan, Minnesota, Ohio, and Oregon with two counties in each. Each courthouse appears twice, once with an ASCII name and once with a name carrying diacritics, at the same address.

Join the cohort to contribute →
Planned

Reference Implementations

The repositories worth copying, with the license and the maturity of each noted. Commercial detection services, open source libraries, and self-hosted deployments are all in scope, and the standard scores any of them.

Coming as the cohort forms
Planned

Build Guidance

The recommended architecture, the split between deterministic code and the model, the choice between cloud and local deployment, and when to buy rather than build. Which layers should survive a model swap and which are expected to be replaced.

Coming as the cohort forms
Planned

Implementations

Case studies from organizations running masking in production. What they chose, what the review queue costs them in staff time, and what they would do differently.

Coming as the cohort forms

How These Resources Connect

The conformance standard says what good looks like. The benchmark set is what you score against. The synthetic set is what you can share and test with when real documents cannot leave the building. The reference implementations are systems the standard has been run against.

Conformance StandardBenchmark SetScored ResultPublished on JusticeBench

You do not need all of them. Pick the resources that match your role.

Choosing among masking toolsConformance standard and published results
Building or tuning a maskerBenchmark set and synthetic document set
Writing a data sharing policyConformance standard and the result bands
Preparing documents for researchSynthetic document set

What This Tool Is

A masking system takes a legal document containing personal information and produces a version of that document safe to send somewhere else. The destinations that matter here are foundation model providers, researchers, developers, and open datasets.

Every organization doing legal help AI work hits this problem, and most of them solve it alone. The pieces are the same each time: find the identifiers, decide which ones to remove, remove them so they cannot be recovered, and leave enough of the document that a legal worker can still use it.

The last part is what makes legal documents difficult. A masker that blacks out a whole paragraph because it found one identifier has destroyed the deadline, the court location, or the amount owed. Both failures matter, and they pull in opposite directions, which is why the standard scores catching and overmasking as separate criteria.

Character recognition quality drives masking performance in scanned documents, so a system that ingests scans is evaluated end to end on its output rather than on the masking step in isolation.

Who this serves: Legal aid organizations sharing documents with researchers or vendors, courts preparing filings for analysis, teams building evaluation sets, and anyone sending legal text to a model they do not host.

Who Is Building It

The cohort is collecting this. If you are running masking in production, or you have written your own script and would rather not maintain it alone, tell us and we will connect you with the others working on it.

Open Questions

These are unresolved and marked so nobody mistakes them for settled. The full list sits at the end of the conformance standard.

  1. Does person name belong in the high tier? The hard fail line does not currently name it, so a system can pass while leaving the tenant's full name in the caption. This is the one to settle first.
  2. What is the full set of masking modes, and are modes the same thing as the sharing contexts of internal collaboration, external expert review, and public dataset release?
  3. What does a result band permit an organization to do? The scorecard says what a reviewer is looking at and not what they may do with it.
  4. Should the five criteria carry weights, or are they scored equally?
  5. Is the paired-name design in the synthetic set a stated benchmark requirement, meant to measure recall difference by name origin while holding jurisdiction and address constant?

Join This Cohort

The PII masking working group is forming now. Members settle the open decisions, contribute documents to the benchmark, and run the standard against the systems they already use.

Join this Working Group See All Working Groups