--- icon: lucide/shield --- # piighost `piighost` is a Python library that protects your confidential data (personal data or PII, secrets) in conversations with LLMs, through de-identification. Sensitive values are hidden before they are sent, then restored in the response. LangChain, Pydantic AI, LlamaIndex and Claude Code integrations are provided, together with an OpenAI and Anthropic API connector. This de-identification spots confidential data with pluggable detectors (regex, NER, LLM) and replaces each value with a placeholder, the token that takes its place. For example: - `John Doe`{ .pii } becomes `<>`{ .placeholder } - `john.doe@example.com`{ .pii } becomes `<>`{ .placeholder } With the conversational pipeline, this placeholder stays the same from one message to the next. That pipeline keeps the mapping between a value and its placeholder across the whole conversation. If `john.doe@example.com`{ .pii } reappears three messages later, the placeholder is still `<>`{ .placeholder }, so the LLM can follow the thread. The LLM therefore only receives de-identified text. When it returns placeholders, for example by answering "Hello `<>`{ .placeholder }", `piighost` replaces them with the real values. The user sees `John Doe`{ .pii } and never sees the de-identification. The same mechanism protects agents that call tools. With the LangChain middleware, a tool that needs the real email address receives it in clear, while the LLM that supplies it only writes `<>`{ .placeholder }. ![A user chats with an agent, confidential values are replaced by placeholders before reaching the LLM and restored afterwards for the user and for tool calls.](assets/deid-chat-light.svg#only-light) ![A user chats with an agent, confidential values are replaced by placeholders before reaching the LLM and restored afterwards for the user and for tool calls.](assets/deid-chat-dark.svg#only-dark) *Full round trip of an agent request. The user and the tool see the real values, the LLM sees only placeholders.* { .figure-caption } !!! note "Reversible de-identification" This retained mapping makes the de-identification a pseudonymization under the GDPR, not a definitive anonymization. With the conversational pipeline, the real values stay stored for the duration of the conversation and must be protected accordingly. ## Why de-identify? A cloud LLM (GPT, Claude, Gemini) receives every piece of information you send it, including your users' PII. De-identifying upstream decouples the choice of LLM from the sensitivity of the content. When confidential data never reach the LLM, the choice of provider stops being a confidentiality decision. It goes back to being a question of quality, cost, and latency. The provider spectrum, the legal detail (CLOUD Act, FISA 702, Schrems II) and the use cases are in [Why de-identify?](why-anonymize.md). The alternatives and their trade-offs are in [How piighost compares](comparison.md). ## Where to start
- :lucide-rocket: __Get started__ --- Install and take `piighost` in hand. - [Installation](getting-started/installation.md) - [Quickstart](getting-started/quickstart.md) - [First pipeline](getting-started/first-pipeline.md) - [Conversational pipeline](getting-started/conversation.md) - [Configuration file](getting-started/configuration.md) - [LangChain middleware](getting-started/langchain.md) - [API server](getting-started/api-server.md) - [Remote client](getting-started/api-client.md) - :lucide-wrench: __Recipes__ --- Solve a specific task. - [De-identify and restore a text](examples/basic.md) - [Pre-built detectors](examples/detectors.md) - [Deny and allow lists](examples/overrides.md) - [Extending piighost](extending.md) - [Testing without a model](examples/testing.md) - [Deployment](deployment.md) - [Multi-instance deployment](multi-instance.md) - :lucide-plug: __Integrations__ --- Plug `piighost` into an agent, a framework or a client. - [LangChain integration](examples/langchain.md) - [Pydantic AI integration](examples/pydantic-ai.md) - [LlamaIndex integration](examples/llama-index.md) - [Claude Code hooks](examples/claude-code.md) - [OpenAI-compatible proxy](examples/openai-proxy.md) - [Anthropic-compatible proxy](examples/anthropic-proxy.md) - :lucide-book-open: __Reference__ --- The full API documentation. - [Anonymizer](reference/anonymizer.md) - [Pipeline](reference/pipeline.md) - [Data models](reference/models.md) - [LangChain](reference/langchain.md) - [Detectors](reference/detectors.md) - [Guard rails](reference/guard-rails.md) - [Conversation memory](reference/memory.md) - [Exceptions](reference/errors.md) - [CLI](reference/cli.md) - [API endpoints](reference/api-endpoints.md) - [Server CLI](reference/api-cli.md) - [Configuration reference](configuration/toml.md) - :lucide-layers: __Concepts__ --- Understand the design choices. - [Architecture](architecture.md) - [Placeholder factories](placeholder-factories.md) - [Security](security.md)