Skip to content

LlamaIndex integration

You want a LlamaIndex RAG where neither the embedding provider nor the LLM ever sees confidential data. piighost gives you two components. PIINodeAnonymizer is an ingestion transform that de-identifies each node before it is embedded. PIIQueryEngine is a wrapper that de-identifies the query and restores the answer. Both share one thread pipeline, so a value keeps the same token across the corpus and the query.

For the same idea orchestrated by hand over a plain RAG flow, see the examples/langchain/rag.py script. This page packages that idea as reusable LlamaIndex objects.

1. Build the thread pipeline

The pipeline de-identifies and restores over a corpus thread. Here an ExactMatchDetector keeps the example deterministic. Swap in a model detector for real text.

from piighost.components.detector import ExactMatchDetector
from piighost.pipeline import ThreadAnonymizationPipeline

THREAD = "docs"
detector = ExactMatchDetector({"Patrick": "PERSON", "Paris": "LOCATION"})
pipeline = ThreadAnonymizationPipeline(detector)

2. De-identify at ingestion, before embedding

Put PIINodeAnonymizer in the transformations before the embedding model, so the index is built on tokens and the embedding provider never sees confidential data.

from llama_index.core import Document, Settings, VectorStoreIndex
from llama_index.core.node_parser import SentenceSplitter
from llama_index.embeddings.openai import OpenAIEmbedding

from piighost.integrations.llama_index import PIINodeAnonymizer

Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
anonymizer = PIINodeAnonymizer(pipeline=pipeline, thread_id=THREAD)
index = VectorStoreIndex.from_documents(
    [Document(text="Patrick lives in Paris.")],
    transformations=[SentenceSplitter(), anonymizer],
)

3. Wrap the query engine

PIIQueryEngine de-identifies the query into the same thread as the corpus, and restores the answer for the user. Retrieval therefore matches the de-identified corpus.

from llama_index.llms.openai import OpenAI

from piighost.integrations.llama_index import PIIQueryEngine

Settings.llm = OpenAI(model="gpt-5.6-terra")
engine = PIIQueryEngine(
    inner=index.as_query_engine(),
    pipeline=pipeline,
    thread_id=THREAD,
)
answer = engine.query("Where does Patrick live?")
print(answer.response)

The answer is restored for display. Its wording depends on the model, for example:

Patrick lives in Paris.

The LLM answered over <<PERSON:1>> and <<LOCATION:1>>. The user sees Patrick and Paris restored. Retrieval runs on the de-identified space. It loses some quality there, and in exchange confidential data stays out of the embedding call.

See also

  • LangChain integration: de-identify a LangChain agent with the middleware.
  • Roadmap: what else is planned.
  • The runnable script is in examples/llama_index/rag.py.