Where to start
In short
piighost hides the confidential data of a text before an AI model reads it, then puts the real values back in the reply. It is a Python library, with no graphical interface. It plugs into LangChain, Pydantic AI, LlamaIndex or Claude Code, or is used remotely through the piighost-api server. It is configured with a TOML or JSON file and the piighost command. The legal and technical reasons to de-identify are explained in Why de-identify?.
The path of a message:
- The user writes a message with their real data, for example their name and email.
piighostfinds the sensitive values, for example names, emails, phone numbers or secrets.- It replaces them with placeholders. A placeholder is a stand-in text, such as
<<PERSON:1>>, which stays the same over the whole conversation. - The model replies with these placeholders.
piighostputs the real values back in the displayed reply.
What this domain documentation is. It describes what piighost must do. How to use it is in the technical documentation. It defines:
- the needs of each profile, that is the compliance officer, the developer, the operator and the application user.
- the rules each process follows, each with its identifier, such as
BR-MSG-05. - the acceptance tests that check each need.
- where in the code each rule applies.
Its goal. De-identifying a conversation with an LLM is still a new practice, and its rules are written down nowhere. This documentation writes them down, so that they can be discussed, checked and improved together. Anyone can propose a need or challenge a rule.
How to read it. Start with Needs by profile to find what concerns your profile. When they disagree, the code and the tests are right. The terms are defined in the glossary.
I want to understand…
| Business need | Page to read |
|---|---|
What each profile expects from piighost, and how to check it | Needs by profile |
| What the model really sees of a message | Protect a message before it is sent to the model |
| Why a name stayed in clear text, or half of it | Protect a message before it is sent to the model |
| How a person keeps the same placeholder from one message to the next | Follow a conversation and restore the reply |
| What happens when you correct a message by hand | Follow a conversation and restore the reply |
| How to erase a conversation (right to erasure) | Follow a conversation and restore the reply |
| Keep the company name in clear text, or always mask an internal code | Impose a deny list and an allow list |
| What a tool of the agent receives, and what the model reads of its result | Let a tool act on the real values |
| Why a placeholder appears while a reply is displayed | Show a streamed reply |
| What each actor sees depending on the tool used (LangChain, Claude Code…) | Plug the protection into an agent and its tools |
| What happens when a model answers badly | Needs by profile, watch points |
| Where the conversation data is stored, and whether it is encrypted | Store conversations and protect traces |
Why piighost works this way, decision by decision | Design decisions |
| The meaning of a term or an acronym | Glossary |
| What is decided and remains to be done | Open points |
Why piighost works this way
Each rule follows from a design decision. The Design decisions page explains them in the order they arose, with an example for each.
- De-identify a text:
- DEC-01 : Replace each confidential value with a placeholder.
- DEC-02 : Find each value and its exact position.
- DEC-03 : Wrap each placeholder in
<<and>>. - DEC-04 : Say in the placeholder what type of data it stands for.
- DEC-05 : Give each entity its own identifier.
- DEC-06 : Group the detections of the same entity.
- DEC-07 : Keep a single span when two detections overlap.
- DEC-08 : Keep the mapping to restore the real values.
- DEC-09 : Let people correct the detection.
- DEC-10 : Reread the protected text before sending, as an option.
- Hold a conversation:
- Let an agent act:
- Go to production:
- The architecture:
Change the code
To change the code, the technical documentation says which pages to read and which files to open. See Change piighost's code.
The domain documentation groups
- Needs:
- Processes:
- Integrations:
- Operations:
- Architecture:
- Tests:
- Reference:
Cross-cutting watch points
- The conversation identifier decides how placeholders are shared. A call without an identifier is refused, by LangChain, the Claude Code hooks and the server. An application that names
defaultshares its placeholders between all its users. Only thepiighost anonymizecommand falls back todefault, to try out a single text. See Follow a conversation. - Correcting an old message can renumber the placeholders, and a reply of the model can then be restored with the name of another person. See Follow a conversation.
- Real values stay stored outside the model. The model only sees placeholders, but two places keep the real values.
piighost's memory keeps them in clear text if its storage is not encrypted. The history the agent records, with LangGraph or Pydantic AI, keeps the text of the messages restored. Some texts do not go throughpiighosteither. This is the case of the result of a Claude Code tool thatpiighostdoes not read, such as Grep. It is also the case of the result of a tool whose setting sends it to the model in clear text, see Let a tool act. Encrypt the memory, and protect the agent's history as personal data. See Store conversations and Plug the protection into an agent. - The technical traces carry the text in clear by default. Configure a trace redactor before you send them to a third-party service. See Store conversations and protect traces.
- Erasing a conversation does not clear the other server instances right away. The storage and the instance that receives the request are cleared. The other instances keep a copy of the values in their placeholder cache, until the lifetime of that cache ends. With no lifetime set, the copy stays until the full cache evicts it. See Store conversations.
- An LLM detector or guard rail refuses the message when it cannot read the answer of its own LLM. The message does not leave, and the application gets an error. The
fail_opensetting lets the message leave without that detection or without that check. See the watch points of Needs by profile and DEC-19 .