Skip to content

Migrate from PresidioReversibleAnonymizer

To migrate from PresidioReversibleAnonymizer, wrap your Presidio AnalyzerEngine in a PresidioDetector and build a ThreadAnonymizationPipeline on it. Then call the pipeline's anonymize and deanonymize with a thread_id. Presidio keeps detecting the values, and piighost replaces and restores them.

The old class lived in langchain-experimental, which LangChain sunset on 22 May 2026. Its repository is archived, so it gets no more fixes.

Before, with PresidioReversibleAnonymizer

The old class detects with Presidio, replaces each value with a fake one drawn from Faker, and keeps the mapping inside the object.

from langchain.chat_models import init_chat_model
from langchain_experimental.data_anonymizer import PresidioReversibleAnonymizer

anonymizer = PresidioReversibleAnonymizer(analyzed_fields=["PERSON", "EMAIL_ADDRESS"])
model = init_chat_model("openai:gpt-5.6-terra")

safe_text = anonymizer.anonymize("Patrick Martin wrote from patrick@example.com.")
reply = model.invoke(safe_text)
print(anonymizer.deanonymize(reply.content))

anonymizer.save_deanonymizer_mapping("mapping.json")

The LLM reads a made-up name and a made-up email in place of Patrick Martin and patrick@example.com. deanonymize() swaps them back, and save_deanonymizer_mapping() writes the mapping to a JSON file if you want to keep it.

After, with piighost

The same round trip with piighost keeps the Presidio engine and replaces the anonymizer object with a thread pipeline.

import asyncio

from langchain.chat_models import init_chat_model
from presidio_analyzer import AnalyzerEngine

from piighost.components.detector.ner import PresidioDetector
from piighost.pipeline import ThreadAnonymizationPipeline

detector = PresidioDetector(
    AnalyzerEngine(), labels={"PERSON": "PERSON", "EMAIL": "EMAIL_ADDRESS"}
)
pipeline = ThreadAnonymizationPipeline(detector)
model = init_chat_model("openai:gpt-5.6-terra")


async def main() -> None:
    safe = await pipeline.anonymize(
        "Patrick Martin wrote from patrick@example.com.", thread_id="thread-42"
    )
    print(safe.text)

    reply = await model.ainvoke(safe.text)
    print(await pipeline.deanonymize(reply.content, thread_id="thread-42"))


asyncio.run(main())

The output should be:

<<PERSON:1>> wrote from <<EMAIL:1>>.
Thank you Patrick Martin, I will write to patrick@example.com.

The LLM reads <<PERSON:1>> and <<EMAIL:1>>. The labels map renames Presidio's EMAIL_ADDRESS to EMAIL, and it plays the role of analyzed_fields, since a type it does not list is dropped. The conversation memory keeps the mapping under thread-42, so the next message of the thread reuses the same placeholders.

In a LangChain agent

If the old class sat in a LangChain chain in front of an agent, hand the same pipeline to PIIAnonymizationMiddleware instead of calling it by hand. The middleware de-identifies each message, restores the reply, and by default gives tools the real values. See LangChain middleware.

What changes

  • Placeholders replace fake values. <<PERSON:1>> cannot collide with a real name, where a Faker name can. A Faker factory is ruled out on purpose, see the FAQ.
  • The mapping is scoped to a conversation. One old object held a single mapping for every text it saw, while the pipeline keeps one per thread_id and erases it with forget_thread.
  • The mapping is stored in a conversation memory, in RAM by default, or in Redis or SQL, which can encrypt the values. It replaces save_deanonymizer_mapping() and load_deanonymizer_mapping(). See Deployment.
  • The methods are asynchronous, so they are awaited.
  • Detection is open to other detectors. Presidio can sit next to a regex group of the catalog or a GLiNER2 model, see Pre-built detectors.
  • The allow_list argument of anonymize() becomes an allow list in the pipeline, see Deny and allow lists.
  • The LangChain middleware also restores tool-call arguments and a streamed reply, which the old class did not do.