Skip to content

How to document piighost in a DPIA

A data protection impact assessment (DPIA) describes a processing, weighs its risks to the people concerned, and lists the measures that address them. If your system sends conversations to an LLM through piighost, this page gives your DPO what they need to fill in the parts of the DPIA that concern piighost. It follows the order of Article 35(7) of the GDPR. The legal reading of pseudonymization, and what the Court of Justice held in EDPS v SRB, are on Compliance.

Check whether a DPIA is required

Article 35(1) requires a DPIA before a processing that is likely to result in a high risk, and each supervisory authority publishes a list of the processing operations that require one (Article 35(4)). In France, that is the CNIL list.

The WP248 rev.01 guidelines, adopted by the Article 29 Working Party and endorsed by the EDPB, give nine criteria. In most cases, a processing that meets two of them requires a DPIA. An LLM assistant for a law firm or a notary's office can meet several, among them "sensitive data or data of a highly personal nature" and "innovative use or applying new technological or organisational solutions".

If you conclude that a DPIA is required, gather the facts of your deployment before writing it.

  • The configuration file of the pipeline, or the code that builds it. piighost validate checks a configuration file.
  • The detectors and the labels they look for.
  • The placeholder factory.
  • The memory backend, and whether a hasher and a cipher protect it.
  • The guard rail, if any.
  • The observation redactor, if traces are exported.
  • The integration (LangChain middleware, Pydantic AI, LlamaIndex, Claude Code hooks, remote client) and, for the LangChain middleware, the tool-call strategy.

Describe the processing

Article 35(7)(a) asks for a systematic description of the processing. For the part piighost performs, it comes down to four steps.

  1. Detection. The configured detectors find the PII in each message, a name, an email, an IBAN. Only what they recognize is replaced. The regex patterns and the NER models run locally. An LLMDetector runs wherever its chat model runs. A regex group of the catalog pinned to a commit is fetched once, as patterns only, then read from the local cache. No message text is sent to the catalog.
  2. Replacement. Each detected value is replaced by a token before the text leaves for the LLM provider. Patrick becomes <<PERSON:1>>, and stays <<PERSON:1>> for the whole conversation.
  3. Retention of the mapping. The mapping from <<PERSON:1>> back to Patrick is kept in the conversation memory, partitioned by conversation (thread_id).
  4. Restoration. When the reply comes back, piighost puts Patrick back in place of <<PERSON:1>> for the user. With the LangChain middleware, the tool-call strategy decides whether tools receive real values too. The default, ToolCallStrategy.FULL, restores the arguments of a tool call and de-identifies its result.

Where the mapping lives

The mapping is what the GDPR calls additional information, and it holds the values in clear. Record where it lives.

Memory backendWhere the mapping livesSurvives a restartAt-rest protectionExtra
InMemoryConversationMemory (default)the memory of the application processnononecore
RedisConversationMemorya Redis serveryes, with an optional ttlopt-in hasher and cipherredis
SqlAlchemyConversationMemorya SQL table (SQLite, PostgreSQL)yesopt-in hasher and ciphersqlalchemy
PIIGhostClientthe piighost-api server it callsdepends on that server's backenddepends on that server's backendclient

The hasher and the cipher need the crypto extra, and Argon2Hasher the argon2 extra. The conversation identifier stays in clear in the Redis keys and the SQL table, since it is what lets a conversation be found and erased. Use an opaque thread_id, never an email or a name.

Two other places hold values in clear and belong in the description.

  • Each worker keeps a memoized copy of a conversation's tokens in its own memory. token_memo_ttl bounds how long it lives, see Multi-instance deployment.
  • With the LangChain middleware, the messages in the LangGraph state hold the restored values after the model turn. A checkpointer that persists that state therefore persists these values too. See Security.

Who can restore

Restoring takes the mapping, so it takes access to the memory backend. It also takes the cipher key when the backend encrypts. In practice, the ones who can restore are the application process and anyone who can read the store and holds PIIGHOST_CIPHER_KEY. When the store is not encrypted, reading it is enough. The LLM provider cannot restore, because it receives only tokens. Recital 29 asks the controller to indicate the authorized persons. Name them in the DPIA.

Map the data flows

The diagram follows Patrick through one turn of a conversation.

DiagramDiagram

One turn with the LangChain middleware and the default FULL tool-call strategy.

Each flow goes into the DPIA with what crosses it and who receives it.

FlowWhat crossesRecipient
User to applicationthe message in clearyou
Application to memorythe detected values, encrypted when a cipher is configuredyou, or the host of the store
Application to LLM providerthe de-identified text, with everything the detectors did not replacethe LLM provider
Application to toolthe restored values, under the FULL and INPUT strategiesthe operator of the tool
Tool result to LLM providerthe tool's response, de-identified under FULL and OUTPUT, as the tool returned it under INPUT and PASSTHROUGHthe LLM provider
Application to trace backendthe stage payloads, tokenized when an observation_redactor is set, in clear otherwisethe operator of the trace backend
Application to a remote LLMDetectorthe message in clear, since detection runs before replacementthe provider of that chat model
Application to a remote guard railthe de-identified output (LLMGuardRail, ModerationGuardRail)the provider of that model

Map the measures to the risks

Article 35(7)(d) asks for the measures envisaged to address the risks. The table lists those piighost provides, the setting to record, and the page that details it.

RiskMeasureSetting to record
The LLM provider reads the PIIthe values are replaced before the text leaves, and a counter or hash token is never computed from the value it replaces. See Placeholder factories.the detectors, the placeholder factory
The mapping reaches the providerthe mapping stays in the memory on your side and is never sent with the text. See Security.the memory backend
Theft of the persistent storethe key of each message is hashed (Sha256Hasher or Argon2Hasher) and each value encrypted (AesGcmCipher). The hasher and the cipher are configured together or not at all. A networked store built without them emits a PIIGhostSecurityWarning. See Security.the hasher, the cipher, where PIIGHOST_HASH_PEPPER and PIIGHOST_CIPHER_KEY are kept
A PII left in the outputa guard rail re-checks the de-identified text, and the pipeline raises PIIRemainingError when it flags one. See Guard rails.DetectorGuardRail, Gliner2GuardRail, LLMGuardRail or ModerationGuardRail
Logs and traces carry PIIthe library writes no PII to its loggers, and an observation_redactor tokenizes the trace payloads. See Observation.the redactor, trace_clear_text, whether PIIGHOST_HOOK_LOG is unset
One conversation sees another's valuesthe memory is partitioned by thread_id, and the integrations refuse a turn without a thread_id instead of pouring it into a shared thread. See Limitations.how thread_id is derived
A user types a token to read someone else's valuea token typed in the input is neutralized before rendering (escape_existing_tokens=True by default). See Security.left at its default
The LLM makes up a tokenan invented token is refused by default (InventedPlaceholderStrategy.RAISE). See Tool-call strategies.the strategy
Retention, and a request for erasureforget_thread erases a conversation from the memory and the local token memo, and returns how many messages and detections it dropped. See Pipeline reference.the retention rule, max_threads and ttl on the in-memory backend, ttl on Redis, token_memo_ttl

The debug log of the Claude Code hooks, written only when PIIGHOST_HOOK_LOG is set, can contain restored values. Leave it unset in production.

Record the residual risks

Article 35(7)(c) asks for an assessment of the risks. piighost lowers the exposure toward the LLM provider without removing the following risks, to record as residual.

  • Detection is not exhaustive. A PII the detectors do not recognize reaches the provider in clear. A NER model can also truncate a text longer than its context. See Limitations.
  • Context and quasi-identifiers stay in clear. "<<PERSON:1>>, the only notary in a village of 300" identifies a person without naming them. The detectors see values, not that inference.
  • The LLM can write a PII it invented. A name the model makes up is in no mapping. So nothing ties it to a person, and nothing removes it.
  • Values the assistant introduces stay in clear under the default EntityCreateByAssistantStrategy.PRESERVE. ANONYMIZE tokenizes them too.
  • The mapping store is a target. It holds the values in clear, or encrypted under a key your environment holds. The process memory and a persisted LangGraph state hold them in clear. See Security.
  • Tools receive real values under the FULL and INPUT strategies, so every tool the agent can call is a recipient. Under INPUT and PASSTHROUGH, a PII in a tool's response reaches the provider in clear.
  • Erasure has a scope. forget_thread reaches the memory and the memo of the process that runs it. It does not reach the provider's logs, your checkpointer, your traces, or the memo of another worker before its token_memo_ttl runs out.
  • The provider's position is not settled. A DPIA that treats the de-identified text as personal data for the provider holds whichever way that question is settled. See Compliance.

Fill in the template

Copy the table into your DPIA and fill in the last column for your deployment. It follows the four items of Article 35(7). You can carry it into the DPIA template of the EDPB, adopted on 14 April 2026 for public consultation, or into the CNIL method and PIA software (in French).

Article 35(7)ItemWhat to recordYour deployment
(a) descriptionpurposewhat the LLM assistant is used for
(a) descriptiondetectorsthe detectors, the labels they look for, the languages they cover
(a) descriptionplaceholderthe factory and an example token
(a) descriptionmappingthe memory backend, its host, its retention
(a) descriptionrecipientsthe LLM provider, the tools, the trace backend, any remote detector or guard rail
(a) descriptionrestorationwho can restore, and with which access
(b) necessityminimizationwhy the provider needs the de-identified text, and why the tools need real values
(c) risksresidualthe residual risks above that apply, with their likelihood and severity
(d) measurescryptothe hasher, the cipher, where the secrets live
(d) measuresguardthe guard rail, or why there is none
(d) measuresloggingthe observation redactor, the application logs, the checkpointer
(d) measuresisolationhow thread_id is derived, and that it is opaque
(d) measureserasurewhen and by whom forget_thread is called, and what it does not reach
(d) measuresproviderthe processor contract, retention and training terms, the transfer outside the EU if any

See also

  • Compliance: the GDPR provisions, the EDPB guidelines and the EDPS v SRB judgment on pseudonymization.
  • Security: the threat model, the memory backends, and the at-rest crypto.
  • Limitations: what detection misses, and how to mitigate it.
  • Deployment: bounding the memory and running piighost in production.