--- icon: lucide/link --- # LangChain middleware You will wire `PIIAnonymizationMiddleware` into a LangChain agent so the LLM only ever sees tokens, while your tools receive the real values. The user asks `Where does Patrick live?`, the model reasons over `<>`{ .placeholder } and `<>`{ .placeholder }, and a lookup tool still gets the real `Patrick`{ .pii } to do its job. You build the middleware over a `ThreadAnonymizationPipeline`, register a tool, and run one turn. !!! note "Prerequisites" `piighost` installed with the middleware extra, `pip install "piighost[langchain]"`, plus an LLM provider configured for `create_agent` (here `openai:...`, so an `OPENAI_API_KEY`). The pipeline reuses the components from [Conversational pipeline](conversation.md). ## 1. Build the thread pipeline The middleware wraps a `ThreadAnonymizationPipeline`, the same one from the [Conversational pipeline](conversation.md) page. Only the detector is passed. The other components keep their defaults, and the default anonymizer uses `LabelCounterPlaceholderFactory`, a delimited token factory which emits `<>`{ .placeholder }. The middleware needs that delimited shape to find a token again. With another factory, it raises `UnrecognizableFactoryError` at construction. ```python from piighost.components.detector import ExactMatchDetector from piighost.pipeline import ThreadAnonymizationPipeline detector = ExactMatchDetector({"Patrick": "PERSON", "Paris": "LOCATION"}) pipeline = ThreadAnonymizationPipeline(detector) ``` ## 2. Declare a tool that needs the real value A tool that looks a person up by name needs `Patrick`{ .pii }, not `<>`{ .placeholder }. Write the tool as usual, against real values. The middleware restores them before the call. ```python from langchain.tools import tool @tool def lookup_city(person: str) -> str: """Return the city where a person lives.""" directory = {"Patrick": "Paris"} return directory.get(person, "unknown") ``` ## 3. Wrap the pipeline in the middleware `PIIAnonymizationMiddleware` takes the pipeline. `tool_strategy=ToolCallStrategy.FULL` restores the tool arguments on the way in and de-identifies the tool result on the way out. The tool therefore works on real values, while the model still only sees tokens. ```python from langchain.agents import create_agent from piighost.integrations.langchain import ( PIIAnonymizationMiddleware, ToolCallStrategy, ) agent = create_agent( model="openai:gpt-5.6-terra", tools=[lookup_city], middleware=[ PIIAnonymizationMiddleware( pipeline=pipeline, tool_strategy=ToolCallStrategy.FULL, ) ], ) ``` ## 4. Run one turn The `thread_id` goes in the LangGraph config, under `configurable`. The middleware reads it from there and scopes every token to that thread. ```python import asyncio async def main() -> None: result = await agent.ainvoke( {"messages": [{"role": "user", "content": "Where does Patrick live?"}]}, config={"configurable": {"thread_id": "thread-42"}}, ) print(result["messages"][-1].content) asyncio.run(main()) ``` The final message is restored for display, so the answer reads with the real values. Its wording depends on the model, for example: ```text Patrick lives in Paris. ``` ## How it works The middleware is a thin adapter around the pipeline. Before the model call, `abefore_model` sends each message through `pipeline.anonymize`, so the LLM receives `Where does <> live?` instead of the raw name. When the model calls `lookup_city` with `person="<>"`, `awrap_tool_call` under `ToolCallStrategy.FULL` restores the argument to `Patrick`{ .pii } before running the tool. It then de-identifies the tool's string result. After the model call, `aafter_model` restores the reply for the user. The `thread_id` keeps `<>`{ .placeholder } bound to `Patrick`{ .pii } across every step of the turn. Two rules are worth knowing. A call without a thread id raises. The middleware does not route every conversation into one shared thread, where tokens would leak from one conversation to another. If your conversations need no separation, pass `"default"`. `invented_strategy=InventedPlaceholderStrategy.RAISE` refuses a token that surfaces in the model's reply but was never issued by the pipeline, whether hallucinated or injected. ## See also - To pick a different tool behaviour, `INPUT` only, `OUTPUT` only, or `PASSTHROUGH`, see [Tool-call strategies](../tool-call-strategies.md). - For a complete agent with a real detector and a system prompt, see [LangChain integration](../examples/langchain.md). - To run the pipeline out of process against a shared server, see [Remote client](api-client.md).