Plug the protection into an agent and its tools
In short
- Plugged into an agent,
piighostmasks the messages before the model and puts the real values back into the reply. - By default, the agent's tools (search, sending mail, reading a file) receive the real values, and what they return is masked before the model.
- A conversation without an identifier is refused, with LangChain as with Claude Code. Otherwise, all conversations would share their placeholders.
- With Claude Code, the displayed reply keeps the placeholders, because no hook point allows rewriting it.
- The history kept by the agent holds the text of the messages with the real values. Only the tool calls stay as placeholders there. Protect this history as personal data.
The terms are defined in the glossary. The conversation mechanism is described in Follow a conversation and restore the reply.
For the business
piighost has no screen. The settings are made in the agent's code or in its configuration. This part describes what each actor sees and the choices to settle.
Who sees what
| Actor | LangChain, Pydantic AI | LlamaIndex | Claude Code |
|---|---|---|---|
| The model | placeholders | placeholders (question and indexed documents) | placeholders (request and results of the listed tools) |
| The tools | the real values (default setting) | not applicable | the real values |
| The end user | the reply with the real values | the reply with the real values | the reply with placeholders |
| The indexing service | not applicable | placeholders | not applicable |
Choose how tools are handled
Four settings decide what the tool receives and what the model reads. The choice is made for the whole agent. The table of settings, the worked example and the rules are in Let a tool act on the real values.
Rules to know
BR-AGT-01 . When a LangChain agent is called without a conversation identifier, then it stops on No thread_id in the LangGraph config; pass config={'configurable': {'thread_id': ...}} on the agent call, or 'default' if your conversations need no separation. The reason is that without an identifier, all conversations would become one and share their placeholders. An application that does not need to separate its conversations passes default itself. 1 location
BR-AGT-02 . When the model writes a placeholder that piighost never issued, then the reply is refused by default, with Deanonymized text holds tokens the pipeline never issued. The two other choices are to keep the placeholder as is or to remove it from the text. 2 locations
BR-AGT-03 . When the assistant is the first to quote a value, then it stays in clear by default. For example, the assistant answers "The head office is in Lyon". "Lyon" is not masked at the next turn, because it comes from the assistant. The two other choices are to mask it like user data or to not analyze the assistant's messages at all. 2 locations · 11 direct tests
BR-AGT-04 . When the tool setting is "Full" or "Output only", then the text returned by the tool goes through full detection and is masked before the model. With LangChain, only the text of the tool message is masked. With Pydantic AI, a structured result (list, dictionary) is walked through entirely. 2 locations · 4 direct tests
BR-AGT-05 . When a Claude Code event has no session identifier, then it is refused with The hook event carries no session_id, the thread its values belong to. No shared conversation is used. 1 location · 9 direct tests
BR-AGT-06 . When a Claude Code tool is not in the list of handled tools, then its result passes in clear. The handled tools are Bash, Read, Write, Edit, Agent, WebFetch, WebSearch, ToolSearch. Grep, in particular, is not among them. 1 location
BR-AGT-07 . When the model's reply is streamed as it is produced, then the display shows placeholders until the end of the message, unless the application plugs in the provided stream decoder. See Show a streamed reply. 2 locations · 4 direct tests
What the end user sees
- With LangChain, Pydantic AI and LlamaIndex: a readable reply, with the real values.
- With Claude Code: a reply that contains placeholders like
<<PERSON:1>>. The modified files and the commands run, however, hold the real values.
Frequently asked questions
The displayed reply contains <<PERSON:1>>. Three possible causes. You use Claude Code, which does not restore the displayed reply. Or the application streams the reply without a stream decoder (BR-AGT-07 ). Or the restoration does not take place in the same conversation as the masking. Check that the same conversation identifier is passed to both.
The agent stops with No thread_id in the LangGraph config. The call does not pass a conversation identifier. Ask the development team to pass it on each call, or default if the conversations do not need to be separated (BR-AGT-01 ).
The agent stops with Deanonymized text holds tokens the pipeline never issued. The model wrote an unknown placeholder, often by copying a placeholder from another conversation or from a document. Keep the refusal if you prefer a visible error to a doubtful text (BR-AGT-02 ).
A tool received <<EMAIL:1>> instead of the address. The tool setting is "Output only" or "None". Switch it to "Full" if the tool must act on the real address. See Let a tool act.
The result of a Grep search went out in clear in Claude Code. Grep is not in the list of handled tools (BR-AGT-06 ). Remove Grep from the session, or have the tool added to the list.
For developers
Where the rules live
| Rule | Location |
|---|---|
| BR-AGT-01 | src/piighost/integrations/langchain/middleware.py:47-66 (_thread_id) |
| BR-AGT-02 | src/piighost/integrations/_deidentify.py:133-155, default RAISE line 57 |
| BR-AGT-03 | src/piighost/integrations/langchain/middleware.py:370 (_message_role), pipeline/thread.py:322-329 |
| BR-AGT-04 | middleware.py:220-272 (LangChain), pydantic_ai/hooks.py:138-154 (Pydantic AI) |
| BR-AGT-05 | src/piighost/integrations/claude_code/hooks.py:101-105 |
| BR-AGT-06 | src/piighost/integrations/claude_code/hooks.py:22-38 |
| BR-AGT-07 | middleware.py:206-218, _deidentify.py:83-107 |
Related components:
TextDeidentifier(integrations/_deidentify.py): shared logic for masking, restoration and invented placeholders, used by LangChain, Pydantic AI and LlamaIndex. It refuses at construction a pipeline without arecognizer(UnrecognizableFactoryError).PIIAnonymizationMiddleware:abefore_model,aafter_model,awrap_tool_call.pii_hooks(pipeline, thread_id, ...): Pydantic AI capability,thread_idfixed or a function of theRunContext.PIINodeAnonymizerandPIIQueryEngine(LlamaIndex): masking of the nodes before embedding, then masking of the question and restoration of the answer in the same corpus conversation.handle_hook(event, pipeline)andrun()(Claude Code):runreads the event on stdin and callspiighost-apiatPIIGHOST_API_URL(defaulthttp://localhost:8000).PIIGhostClient: implementsAnyThreadPipelineover HTTP (/v1/anonymize,/v1/deanonymize,/v1/detect,/v1/labels,/v1/threads/{id}/tokens). It raisesRemoteErroron a non-2xx response.
Plug in the LangChain middleware
- Start from
examples/langchain_middleware.py. - Build a
ThreadAnonymizationPipelinewhose factory is delimited (by defaultLabelCounterPlaceholderFactory). - Pass it to
PIIAnonymizationMiddleware(pipeline). Addtool_strategy,invented_strategyorassistant_strategyonly to change a default. - Pass
config={"configurable": {"thread_id": "..."}}on each call of the agent. - For a streamed display, wrap the
agent.astream(..., stream_mode="messages")loop inmiddleware.deanonymize_stream(source, thread_id).
Check
uv run pytest tests/integrations/langchainIn a trace of the agent, the message received by the model must contain <<PERSON:1>> and the argument received by the tool the real value.
Pitfalls
- The LangGraph state keeps the message content in clear.
aafter_modelrestores the content in the state, and the checkpointer saves it that way. Only thetool_callsstay as placeholders (middleware.py:199-200). The same goes for the Pydantic AI history afterafter_model_request. - LangChain masks again the arguments of the
tool_callsin the history, as a precaution, except underIGNORE. Pydantic AI does not. - The Claude Code hooks tolerate an unknown output shape: they let it through. Set
PIIGHOST_HOOK_LOGto see the real shapes. This log contains restored values in clear. Keep it local and delete it afterwards. - If
piighost-apicannot be reached, the hook fails closed (claude_code/runner.py:68-88). A prompt or a tool call exits with code 2, which Claude Code reads as a block. A tool output, already produced, is replaced by a notice.PIIGHOST_HOOK_FAIL_OPEN=1lets the text through in clear (DPO-9 ). PIIQueryEnginerefuses a streaming engine (NotImplementedError). Its synchronous paths go throughasyncio.run, so callaqueryfrom asynchronous code.PIIGhostClient.anonymizereturns an empty placeholder dictionary. The mapping lives on the server. Restore withdeanonymize.
Tests
| Test | Covers |
|---|---|
tests/integrations/langchain/test_middleware.py | Conversation identifier, block content, each tool strategy, Command, invented placeholders, assistant provenance, refusal of a non-delimited factory |
tests/integrations/langchain/test_middleware_stream.py | Streamed restoration |
tests/integrations/test_pydantic_ai_hooks.py | Pydantic AI capability |
tests/integrations/llama_index/ | Node transformation, query engine |
tests/integrations/test_claude_code_hooks.py | The three events, list of fields, unknown tool let through, blocking when piighost-api cannot be reached, fail open |
tests/integrations/client/test_client.py | HTTP client |
See also Configure a pipeline.