Skip to content

Detection, measured

This page measures how many confidential values a configured piighost pipeline hides, and what each stage adds over calling a NER model directly. The figures come from a benchmark run on 2026-09-29 on five data sets, English and French, with piighost 1.10.0.

The question

A NER model finds names, places and organisations. piighost wraps one in a pipeline that adds chunking, regex rules, a search for repeats and entity grouping. Does that pipeline hide more than the model alone, and which stage does the work?

To answer it, the benchmark adds one stage at a time, with the same model on the same documents.

RungSystem
Athe model called directly on the whole text, as a developer would, with the text truncated to the model's window
BA, with piighost's chunking
Cthe regex rules of the config, no model
DB and C together, overlaps resolved
ED, plus the word-boundary expander
FE, plus the entity resolver, the full config

Each rung is measured with two models. One is GLiNER2 (fastino/gliner2-multi-v1, the model of the fr-notarial catalog config). The other is onnx-community/gliner_multi_pii-v1 in ONNX, called through BridgeDetector, the engine that runs in the browser.

The data

SetLanguageDocumentsWhat it isCaution
Generated deedsFrench200deeds, leases, employment contracts, emails, medical reports, built from templates with fictitious valuesgenerated text flatters the scores, because each value sits in a slot made for it
Long generated deedsFrench12the same, from 2 to 60 pagesthree documents per length, wide intervals
TABEnglish127ECHR judgments annotated for anonymisation (Pilán et al., 2022)public judgments a model may have seen in training
PARHAFFrench101medical reports written by hand by medical residents for fictitious patientsalmost no fixed-shape identifier, so it tests the model
Gretel financeFrench443synthetic finance documentswritten by an LLM

No set contains a real private person's data. Only direct identifiers are scored in the headline. TAB and PARHAF also mark quasi-identifiers (a date, a profession, a family status), which are reported apart and never averaged in.

Results

Direct identifiers hidden, model alone (A) then full pipeline (F), in %.

SetGLiNER2 A → FONNX A → F
Generated deeds36 → 9519 → 90
Long generated deeds2 → 961 → 95
TAB27 → 4631 → 61
PARHAF15 → 6111 → 61
Gretel finance66 → 7955 → 75

The French sets run the fr-notarial config, TAB runs support-en. The published fr-notarial gets the same score as its F rung. On every set and for both models, the 95 % confidence intervals of A and F do not overlap.

Where the gain comes from

On the generated deeds with GLiNER2, rung by rung, the scores are 36 % for A, 66 % for B, 94 % for D, 95 % for E and F.

  • Chunking gives 31 to 33 points on the generated deeds and 53 to 64 on the long ones. A model reads a fixed window, and without chunking it never sees past it. On the long deeds, rung A only reads page one and misses every value after it.
  • The regex rules give 28 to 40 points. They handle every value with a fixed shape. Emails, IBANs, social security numbers, company numbers, phones and dates reach 100 %, where the model alone finds at most a fifth of them. The formulae of a deed ("Monsieur", "Maître", "née", "demeurant", "section") catch the names and addresses the model misses.
  • The word-boundary expander adds 1 to 3 points, less than before, since the rules now find most repeats themselves.
  • The entity resolver adds no recall. It groups the spellings of one person onto one token.

Rules and the model often flag the same value with different lengths. fr-notarial keeps their union with the merge overlap resolver. With the default confidence resolver, a rule's short span, at confidence 1.0, beat the model's longer span. On the finance set, the share of names hidden then fell from 89 to 77 %.

Precision, the share of masked text that was really a value, stays near 86 % on the generated deeds. On the long deeds the expander drops it to 59 % with ONNX, since a heading word masked once is then masked everywhere.

What still leaks

On the generated deeds, full GLiNER2 pipeline:

CategoryHiddenDocuments with none left in clear
Email, IBAN, social security number, company number, phone, date of birth100 %100 %
Person95 %59 %
Organisation94 %88 %
Address94 %81 %
Cadastral parcel43 %75 %

43 % of the generated deeds come out with nothing left in clear, against none before the date and deed rules. A single forgotten name is still enough to spoil a deed. The cadastral parcels a table lists without the word "section" are missed.

What the benchmark changed

Each benchmark run found a defect. That defect was fixed in the library or in the catalog's fr-notarial before the next run.

  • piighost 1.9.0. The word-boundary expander could add an occurrence inside a kept detection. The render stage then raised OverlappingSpansError, on 163 of the 200 generated deeds. French phones typeset with no-break spaces were never matched. An email with accented letters was matched from its first ASCII run. A detector config could not set max_chars. A config-built model therefore read a whole deed in one pass, and ran out of memory past 13,000 characters.
  • Dates. A date of birth is a direct identifier. Neither model of the benchmark was asked for one, so the rules find it. fr-notarial hides every French date, because a pattern cannot tell a date of birth from the date of the deed. It spares the date of a numbered legal text ("loi n° 89-462 du 6 juillet 1989").
  • Deed formulae. A name after a civility or "Maître", a maiden name after "née", an address after "demeurant" or "situé", a street address, a lieu-dit and a cadastral reference after "section".
  • SWIFT_BIC. It matched any run of eight or eleven capitals. It now needs a keyword or a digit, so a heading such as "DESIGNATION" stays in clear.
  • piighost 1.10.0. Adds the merge overlap resolver. A rule's short span no longer uncovers part of the model's span.

On the generated deeds, GLiNER2, run after run:

1.9.0 candidateDates, chunkingFormulae, merge
Direct identifiers hidden, full pipeline76 %82 %95 %
The config as published63 %82 %95 %
Deeds with nothing left in clear0 %2.5 %43 %
Precision77 %78 %86 %
Capitalised headings masked, of 1,43741141130
Legal references masked, of 4521425014

These patterns belong to a document config. A chat config should not hide every date nor read "Monsieur" as the start of a name to hide.

Limits of these numbers

  • The French documents are mostly generated. A hand-written medical set (PARHAF) scores lower than the generated one, and no open notarial deed exists to check the generated deeds against.
  • Each figure comes from one run on CPU. The intervals account for the choice of documents, not for another model version.
  • The model calls are the same across runs, replayed from a cache keyed by their exact input. A change in a figure comes from the pipeline or the config, never from a different inference.

See also