Detection, measured
This page measures how many confidential values a configured piighost pipeline hides, and what each stage adds over calling a NER model directly. The figures come from a benchmark run on 2026-09-29 on five data sets, English and French, with piighost 1.10.0.
The question
A NER model finds names, places and organisations. piighost wraps one in a pipeline that adds chunking, regex rules, a search for repeats and entity grouping. Does that pipeline hide more than the model alone, and which stage does the work?
To answer it, the benchmark adds one stage at a time, with the same model on the same documents.
| Rung | System |
|---|---|
| A | the model called directly on the whole text, as a developer would, with the text truncated to the model's window |
| B | A, with piighost's chunking |
| C | the regex rules of the config, no model |
| D | B and C together, overlaps resolved |
| E | D, plus the word-boundary expander |
| F | E, plus the entity resolver, the full config |
Each rung is measured with two models. One is GLiNER2 (fastino/gliner2-multi-v1, the model of the fr-notarial catalog config). The other is onnx-community/gliner_multi_pii-v1 in ONNX, called through BridgeDetector, the engine that runs in the browser.
The data
| Set | Language | Documents | What it is | Caution |
|---|---|---|---|---|
| Generated deeds | French | 200 | deeds, leases, employment contracts, emails, medical reports, built from templates with fictitious values | generated text flatters the scores, because each value sits in a slot made for it |
| Long generated deeds | French | 12 | the same, from 2 to 60 pages | three documents per length, wide intervals |
| TAB | English | 127 | ECHR judgments annotated for anonymisation (Pilán et al., 2022) | public judgments a model may have seen in training |
| PARHAF | French | 101 | medical reports written by hand by medical residents for fictitious patients | almost no fixed-shape identifier, so it tests the model |
| Gretel finance | French | 443 | synthetic finance documents | written by an LLM |
No set contains a real private person's data. Only direct identifiers are scored in the headline. TAB and PARHAF also mark quasi-identifiers (a date, a profession, a family status), which are reported apart and never averaged in.
Results
Direct identifiers hidden, model alone (A) then full pipeline (F), in %.
| Set | GLiNER2 A → F | ONNX A → F |
|---|---|---|
| Generated deeds | 36 → 95 | 19 → 90 |
| Long generated deeds | 2 → 96 | 1 → 95 |
| TAB | 27 → 46 | 31 → 61 |
| PARHAF | 15 → 61 | 11 → 61 |
| Gretel finance | 66 → 79 | 55 → 75 |
The French sets run the fr-notarial config, TAB runs support-en. The published fr-notarial gets the same score as its F rung. On every set and for both models, the 95 % confidence intervals of A and F do not overlap.
Where the gain comes from
On the generated deeds with GLiNER2, rung by rung, the scores are 36 % for A, 66 % for B, 94 % for D, 95 % for E and F.
- Chunking gives 31 to 33 points on the generated deeds and 53 to 64 on the long ones. A model reads a fixed window, and without chunking it never sees past it. On the long deeds, rung A only reads page one and misses every value after it.
- The regex rules give 28 to 40 points. They handle every value with a fixed shape. Emails, IBANs, social security numbers, company numbers, phones and dates reach 100 %, where the model alone finds at most a fifth of them. The formulae of a deed ("Monsieur", "Maître", "née", "demeurant", "section") catch the names and addresses the model misses.
- The word-boundary expander adds 1 to 3 points, less than before, since the rules now find most repeats themselves.
- The entity resolver adds no recall. It groups the spellings of one person onto one token.
Rules and the model often flag the same value with different lengths. fr-notarial keeps their union with the merge overlap resolver. With the default confidence resolver, a rule's short span, at confidence 1.0, beat the model's longer span. On the finance set, the share of names hidden then fell from 89 to 77 %.
Precision, the share of masked text that was really a value, stays near 86 % on the generated deeds. On the long deeds the expander drops it to 59 % with ONNX, since a heading word masked once is then masked everywhere.
What still leaks
On the generated deeds, full GLiNER2 pipeline:
| Category | Hidden | Documents with none left in clear |
|---|---|---|
| Email, IBAN, social security number, company number, phone, date of birth | 100 % | 100 % |
| Person | 95 % | 59 % |
| Organisation | 94 % | 88 % |
| Address | 94 % | 81 % |
| Cadastral parcel | 43 % | 75 % |
43 % of the generated deeds come out with nothing left in clear, against none before the date and deed rules. A single forgotten name is still enough to spoil a deed. The cadastral parcels a table lists without the word "section" are missed.
What the benchmark changed
Each benchmark run found a defect. That defect was fixed in the library or in the catalog's fr-notarial before the next run.
piighost1.9.0. The word-boundary expander could add an occurrence inside a kept detection. The render stage then raisedOverlappingSpansError, on 163 of the 200 generated deeds. French phones typeset with no-break spaces were never matched. An email with accented letters was matched from its first ASCII run. A detector config could not setmax_chars. A config-built model therefore read a whole deed in one pass, and ran out of memory past 13,000 characters.- Dates. A date of birth is a direct identifier. Neither model of the benchmark was asked for one, so the rules find it.
fr-notarialhides every French date, because a pattern cannot tell a date of birth from the date of the deed. It spares the date of a numbered legal text ("loi n° 89-462 du 6 juillet 1989"). - Deed formulae. A name after a civility or "Maître", a maiden name after "née", an address after "demeurant" or "situé", a street address, a lieu-dit and a cadastral reference after "section".
SWIFT_BIC. It matched any run of eight or eleven capitals. It now needs a keyword or a digit, so a heading such as "DESIGNATION" stays in clear.piighost1.10.0. Adds themergeoverlap resolver. A rule's short span no longer uncovers part of the model's span.
On the generated deeds, GLiNER2, run after run:
| 1.9.0 candidate | Dates, chunking | Formulae, merge | |
|---|---|---|---|
| Direct identifiers hidden, full pipeline | 76 % | 82 % | 95 % |
| The config as published | 63 % | 82 % | 95 % |
| Deeds with nothing left in clear | 0 % | 2.5 % | 43 % |
| Precision | 77 % | 78 % | 86 % |
| Capitalised headings masked, of 1,437 | 411 | 411 | 30 |
| Legal references masked, of 452 | 14 | 250 | 14 |
These patterns belong to a document config. A chat config should not hide every date nor read "Monsieur" as the start of a name to hide.
Limits of these numbers
- The French documents are mostly generated. A hand-written medical set (PARHAF) scores lower than the generated one, and no open notarial deed exists to check the generated deeds against.
- Each figure comes from one run on CPU. The intervals account for the choice of documents, not for another model version.
- The model calls are the same across runs, replayed from a cache keyed by their exact input. A change in a figure comes from the pipeline or the config, never from a different inference.
See also
- Limitations: what detection cannot promise, whatever its score.
- Pre-built detectors: the regex groups of the catalog and the models.
- TOML reference:
max_charsand the detector keys.