Skip to content

Limitations

piighost de-identifies, it does not magically make a text safe. This page lists the known limitations, why they exist, and how to mitigate them. It extends the threat model.

Detection is not exhaustive

A detector only finds what it knows how to recognize. Two families share the work, with different blind spots.

A pattern detector (RegexDetector) recognizes strings that follow a fixed structure, such as an email, an IP, or a credit-card shape. It is deterministic on those formats and blind to the rest. A NER detector (Gliner2Detector, SpacyDetector, TransformersDetector, PresidioDetector, BridgeDetector) or LLM detector (LLMDetector) recognizes free-form entities (a name, a place, an organization), but it misses some. A rare name, an unusual spelling, an out-of-distribution entity passes in cleartext to the LLM.

Confidential data that is not detected is not de-identified. This is an engineering concern, not a conceptual flaw.

Mitigation: chain a NER detector and a RegexDetector through the CompositeDetector, to cover both free-form text and structured formats. Load a locale-specific NER model for better accuracy. See Extending piighost.

A model can truncate a text longer than its context

A NER model has a maximum input length. The model truncates a longer text. The truncated tail is never scanned, so its PII passes in cleartext. Nothing warns you by default.

The limit belongs to the model, not to the pipeline. It applies to any de-identification backed by a NER model. piighost ships the means to work around it rather than live with it.

Mitigation: set max_chars on the NER detector to the model's safe input length. With auto_chunk on (the default), a longer text is split into overlapping chunks. Each chunk is scanned separately, then the results are remapped, so the tail is covered. With auto_chunk off, an over-long text raises TextTooLongError rather than being scanned in part. For very long inputs, wrap the detector in a ChunkedDetector.

Language coverage is model-dependent

The set of languages a NER detector can cover is fixed by the model you plug in. Coverage varies from model to model, and not every language is supported equally. Before deploying on a new locale, read the model card and run a small validation set.

Here too the limit belongs to the model, not to the pipeline. A pattern detector does not have this limit, because an IBAN or an email address has the same shape in every language.

Mitigation: load a locale-specific model, or combine several detectors through the CompositeDetector.

Whole-word search assumes spaces between words

To find a value again, piighost searches it as a whole word, so Jean is not found inside Jeanne. A value must not touch a letter, a digit or a hyphen on either side. Chinese, Japanese and Thai write words without spaces between them. A value in those scripts therefore always touches a letter, and it is never found.

TextSearched valueFound
Jeanne et JeanJeanthe second Jean
田中さんは田中です田中nothing

Three components rely on this search. ExactMatchDetector finds nothing. WordBoundaryExpander finds no repetition. LLMDetector places in the text each value the LLM names, so it drops a value the LLM did find, and that value is sent as it is. The detectors that return offsets, RegexDetector and the NER detectors, are not affected.

Supporting these scripts would take a word segmenter per language, and the pipeline has none.

Mitigation: on Chinese, Japanese or Thai text, detect with a NER model or a pattern rather than with ExactMatchDetector or LLMDetector, and do not count on the expander for repetitions.

A pattern cannot cover every script at once

Python's re has no word segmentation. It also does not count combining marks as letters, for example the vowel signs of Hindi and the other Indic scripts. An email pattern therefore has to choose which letters it takes in, and each choice leaves some letters out.

TextPattern taking Unicode letters, (?u:\w)EMAIL of catalog:piighost/generic
écrire à expéditeur@exemple.frexpéditeur@exemple.frexpéditeur@exemple.fr
メールはtanaka@example.jpですthe whole sentencetanaka@example.jp
ελένη@example.grελένη@example.grnothing
राम@उदाहरण.भारतnothingnothing

A pattern that takes Unicode letters in finds an accented or a Greek address whole. In Chinese or Japanese text with no space around the address, the ideographs next to it are letters too, so the pattern takes them in. The token hides more than the address, and nothing is sent in clear. The Hindi address still escapes it, since its vowels are combining marks.

The EMAIL pattern of catalog:piighost/generic takes Latin letters only, that is the ASCII letters and digits plus the Latin range À to ɏ. It finds the accented and the Japanese examples exactly. It misses every address written in another script, and that address is sent as it is.

Mitigation: for text in non-Latin scripts, detect addresses with a NER model, or write in your config an email pattern suited to the addresses that text really holds. A pattern written inline in the config overrides the group's pattern on the same label.

No checksum validation (deliberate)

RegexDetector matches on shape alone. It verifies no checksum, no Luhn on cards, no IBAN check key, no NIR check key. This is deliberate.

A structured value can arrive mangled by OCR, one character misread. A checksum validator would then reject a real but mistranscribed IBAN or NIR, and that PII would pass in cleartext to the LLM. piighost prefers to keep a shape-level false positive rather than let a real damaged value leak. It is a security choice. When the detector is wrong, it errs on the side of detecting too much.

The trade-off is that RegexDetector can match strings that have the shape of a PII without being one (a digit run that looks like a card). The cost of such a false positive is benign, one extra token. The cost of a false negative, a real PII left undetected, would be a leak.

Mitigation: refine the patterns if shape-level false positives disturb a precise workload. Do not reintroduce a checksum filter upstream of text that may come from OCR. If your inputs are typed and never go through OCR, the trade-off flips. You can then write your own detector with checksum validation, because the AnyDetector port is open. See Extending piighost.

Placeholders can collide depending on the factory

The placeholder factory decides what distinguishes two entities. Some families produce the same output for two different inputs.

  • RedactPlaceholderFactory collapses every value to <<REDACT>>. LabelPlaceholderFactory collapses every value of one label to <<PERSON>>. Neither family distinguishes entities, so neither is reversible.
  • MaskPlaceholderFactory keeps a fragment of the value, by default its first character, so Jonathan becomes J*******. Two similarly shaped values can collide on one mask, and a mask can also collide with a real value in a tool response.
  • LabelCounterPlaceholderFactory (<<PERSON:1>>) and LabelHashPlaceholderFactory (<<PERSON:a1b2c3d4>>) give a distinct token per entity, which can be found again in text. They therefore stay reversible without ambiguity.

Mitigation: see Placeholder factories for the full taxonomy and the choice by use case.

Restoration needs a unique token per entity

Restoring a value from a placeholder assumes the placeholder identifies a unique entity. Two properties combine in the token. Typing says which kind of value it is (a person, a location, an email). Identity says which one it is among those of the same kind. Each factory carries a preservation tag that declares what its token keeps of these two properties.

FactoryToken emittedWhat the token keepsRestoration
RedactPlaceholderFactory<<REDACT>>nothingimpossible
LabelPlaceholderFactory<<PERSON>>the typeimpossible
MaskPlaceholderFactoryJ*******a fragment of the valueambiguous
LabelCounterPlaceholderFactory<<PERSON:1>>the type and the identityreliable
LabelHashPlaceholderFactory<<PERSON:a1b2c3d4>>the type and the identityreliable

The preservation tag of each factory is in Placeholder factories.

On the sentence "Patrick and Marie live in Paris", the difference shows immediately.

  • With LabelPlaceholderFactory, both people become the same <<PERSON>>. The type is there, the identity is not, so nothing says which of the two tokens was Patrick.
  • With LabelCounterPlaceholderFactory, Patrick becomes <<PERSON:1>> and Marie becomes <<PERSON:2>>. Each token maps to a single value, so restoration is unambiguous.

The PIIAnonymizationMiddleware enforces this constraint at the type level. It requires a PreservesRecognizableIdentity factory, that is a token unique per entity and findable in text. At construction, it also refuses a factory with no delimited grammar, such as a mask (UnrecognizableFactoryError). The model's reply and the tool arguments need unique tokens to stay reversible, because their restoration relies on string replacement.

Mitigation: keep LabelCounterPlaceholderFactory or LabelHashPlaceholderFactory with the middleware. See Tool-call strategies for the FULL, INPUT, OUTPUT, and PASSTHROUGH modes.

PII invented by the LLM is not in the mapping

Restoration works on values seen at the input. If the LLM hallucinates a name that never appeared in the user's messages, for instance making up a plausible client name, that PII is in no mapping. It therefore cannot be tied back to an original value.

The middleware catches a neighbouring case, the invented placeholder. If the LLM fabricates a token that looks like a placeholder but was never emitted, piighost spots it (the token has no associated value) and refuses it by default (InventedPlaceholderError, the RAISE strategy). The KEEP and DROP strategies exist for other policies.

Mitigation: run a re-detection step on the LLM output at the application layer, and decide whether to strip, flag, or re-de-identify before display. A guard rail (DetectorGuardRail, Gliner2GuardRail, LLMGuardRail, ModerationGuardRail) re-checks the de-identified output and flags residual confidential data. The pipeline then raises PIIRemainingError.

Memory is process-local by default

InMemoryConversationMemory keeps the mapping thread by thread in a process-local dictionary. Nothing survives a restart, nothing is shared across processes. As soon as you scale horizontally, two workers have two memories and two independent placeholder spaces. The same entity can therefore get two different tokens depending on which worker handles it.

Mitigation: configure RedisConversationMemory to share the mapping across workers and make it survive a restart, or SqlAlchemyConversationMemory to keep it in a SQL table. Both backends can encrypt the values and hash the keys (opt-in, all-or-nothing). The in-memory backend is bounded by default. In a long-lived process, max_threads and ttl adjust the cap on its growth. See Security and Deployment.

A thread isolates the mapping

Memory is partitioned by thread_id. Two separate conversations share no placeholder. This is intended, but the same person then gets two unrelated tokens in two threads. The middleware requires a thread_id and does not fall back to a shared default thread, to prevent one conversation from seeing another's mapping.

Mitigation: propagate a stable per-conversation thread_id. Call forget_thread to purge a conversation from memory once it no longer has reason to exist.

Latency overhead is not yet benchmarked

There is no official benchmark of the latency added by the pipeline on typical workloads. The overhead depends on the detector (NER inference), the text length, and whether values are already known in the thread's memory.

Mitigation: measure on your own workload before sizing production traffic. Keep detectors on GPU when possible for NER-heavy paths.

Minimum viable threat coverage

piighost addresses exfiltration toward the LLM and its provider. It does not replace encryption at rest, access control, or secure logging practices for the rest of your system. See Security for the full threat model.