Skip to content

FAQ

Is it really necessary to de-identify confidential data before calling an LLM?

Yes, and this holds regardless of piighost. The stakes (exfiltration to providers, legal requisition, training on conversations, GDPR compliance, data leaks) are covered in Why de-identify?. The page is library-agnostic. It explains why the problem exists before justifying a solution like piighost.

Which languages are supported?

It depends entirely on the detector you plug in. The pipeline itself is language-agnostic. With a gliner2 detector and a multilingual GLiNER2 model, you get about 100 languages out of the box. With a spacy detector, whatever spaCy supports. With a regex detector, language is irrelevant. See Extending piighost for the detector catalogue.

Which entities are detected out of the box?

None. piighost does not ship its own NER model, this is a deliberate design choice. You bring the detector. Use an exact detector for fixed dictionaries, a regex detector with a group pulled from the catalog (catalog:piighost/generic, catalog:piighost/us, catalog:piighost/eu, catalog:piighost/fr) or your own patterns, a gliner2 detector for open NER (PERSON, LOCATION, ORGANIZATION, EMAIL, any label you ask for), or compose them with a composite detector.

Does the regex detector validate checksums (Luhn, IBAN, NIR)?

No, by design. A checksum validator rejects a value whose digits do not compute. That is exactly what OCR noise or a typo produces. Rejecting such a value would leak the PII the validator was meant to catch. The regex detector matches on shape alone and errs toward over-detection, which is the safe direction for de-identification. If you need to narrow a match, add a stricter pattern rather than a validator.

How do I configure a pipeline?

Write a TOML or JSON file describing each stage, then load it. load_pipeline builds a stateless pipeline. load_thread_pipeline builds a thread pipeline with a conversation memory. The file suffix picks the parser. Every section and component type is in the configuration reference. The config extra is required (pip install "piighost[config]").

What latency does the pipeline add?

The pipeline itself is on the millisecond scale (regex and lookups). The real cost comes from the detector. GLiNER2 on CPU for a 200-token message is typically 50 to 200 ms. An LLM used as a detector, several hundred milliseconds. Resending a message inside a thread skips detection, because a thread pipeline caches each message's detections. Measuring on your actual workload remains recommended before sizing production.

Does piighost work 100% offline?

Yes. With a local detector (gliner2, spacy, regex, exact), no data leaves your process. A catalog group pinned to a commit is fetched on the first build, then read from the on-disk cache, and the fetch sends no text to the catalog. The middleware only forwards already de-identified text to the LLM. Keeping a hosted LLM under GDPR constraints without exfiltrating raw PII is the main reason teams adopt piighost. See Why de-identify? for the legal context.

Do my placeholders have to look like <<PERSON:1>>?

No. The format is driven by the placeholder factory chosen in [anonymizer.placeholder]. label_counter produces <<PERSON:1>>, label_hash produces <<PERSON:a1b2c3d4>>, label produces <<PERSON>> without a counter, mask produces P***, and you can write your own factory. See Placeholder factories.

Can I get realistic fake values instead of tokens?

No, and it is not planned. A Faker factory, which would emit a plausible name in place of Patrick, is ruled out on purpose in the roadmap. Two people could draw the same fake name, and a fake could coincide with a real value, so restoration would no longer be reliable. Today the factories emit synthetic tokens or masks, never a value that looks real.

Does the LLM see raw confidential data when it calls a tool?

It depends on the tool-call strategy. With the default (FULL), no. The middleware restores arguments right before the tool executes, then de-identifies the tool response again before it flows back to the LLM. The tool sees real values, the LLM only sees placeholders. The INPUT, OUTPUT and PASSTHROUGH modes change this behaviour, see the next question and Tool-call strategies. Full diagram in Architecture.

How do I control what a tool sees: placeholder or real value?

The tool-call strategy of PIIAnonymizationMiddleware exposes four modes (INPUT, OUTPUT, FULL, PASSTHROUGH). The right choice depends on whether the tool may emit new confidential data and how strict the privacy boundary needs to be. See Tool-call strategies for the trade-offs and the decision tree. The middleware also needs an identity-preserving, recognizable placeholder factory, see Placeholder factories for this constraint.

What happens if the LLM hallucinates confidential data that was not in the input?

It is not de-identified by piighost. Entity linking works on detections coming from the input, not on invented values. A guard for residual confidential data can re-check the output and refuse it, see the guard section of the configuration reference and Limitations.

Is the conversation memory shared across threads?

No. The memory is scoped by thread_id. Two parallel conversations never see each other's tokens. This separation prevents cross-user leaks. The thread_id is extracted automatically from the LangGraph config.

How do I run more than one worker behind a load balancer?

Use the Redis conversation memory, shared by every worker. The in-RAM memory is process-local, so two workers would number the same value differently mid-conversation. See Multi-instance deployment for the trap and the fix, and Deployment for the full setup.

Can I use piighost without LangChain?

Yes. The stateless and thread pipelines are usable standalone, without the middleware. See Basic usage.

Does piighost encrypt stored data?

The Redis conversation memory does. It encrypts every stored value with AES-GCM and hashes every key, reading its pepper and cipher key from the environment. The in-RAM memory encrypts nothing and is for development only. See Security for the at-rest threat model.

How do I trace what the pipeline does?

Through OpenTelemetry. The pipeline emits a span per stage to whatever OTel TracerProvider your application configured. It does no backend correlation itself, because that correlation belongs to the deployment's OTel configuration. See Observation. The observation extra is required.