Architecture
piighost follows a hexagonal architecture, also known as ports and adapters. The core
knows only abstract contracts, the ports. Each concrete implementation, a GLiNER2
detector, a Redis backend, a LangChain middleware, is an adapter that satisfies a
port without the core knowing about it. The de-identification pipeline is assembled by
injecting the chosen adapters behind the ports it expects.
The three rings
The code reads as three rings, from the most abstract to the most concrete. The direction of the dependencies is fixed once and for all. An outer ring imports an inner ring, and an inner ring never imports an outer ring.
Three rings and the composition root. Dependencies always point toward the core.
- Core. The data models (
Detection,Entity,Span, frozen dataclasses) and the ports. No external dependency, no pydantic, no I/O. - Application. The pipeline orchestration, which depends only on the core ports.
This is where
anonymize,deanonymize, andforget_threadlive. - Adapters. The concrete implementations of the ports, that is detectors, resolvers, factories, guard rails, memory backends, observation, HTTP client, middleware. Each adapter imports the core, never the reverse.
- Config. The composition root. It is the only place allowed to know both the ports and the concrete adapters, in order to assemble them.
Ports and templates
A port is a Python Protocol marked runtime_checkable, in each component's
base.py. The typing there is structural. An object satisfies the port as soon as
it has the methods, without inheriting from it. The pipeline depends on the port, never
on a concrete class.
@runtime_checkable
class AnyDetector(Protocol):
async def detect(self, text: str) -> list[Detection]: ...
When several adapters of one port share a skeleton, that skeleton lives in a Base*
class, an abstract class that applies the Template Method pattern. The skeleton is
written once in the base class, and each subclass provides only the step that varies.
class BaseEntityLinker(ABC):
def link(self, detections: list[Detection]) -> list[Entity]:
# common skeleton: group by key
...
@abstractmethod
def _key(self, detection: Detection) -> Hashable:
# only varying step, defined by the subclass
...
Five ports have no template shared by all their adapters, the detector, override, guard rail, memory backend and cipher ports. Their adapters differ by their whole mechanism, not by a single step, so they have nothing common to factor out. This is the deliberate exception to the always-template rule.
The detector is a partial exception. The model detectors (Gliner2Detector,
SpacyDetector, TransformersDetector, PresidioDetector, BridgeDetector,
LLMDetector) share the BaseNERDetector template. It re-reads each detection's text
from the source, applies the confidence threshold and maps the labels. RegexDetector,
ExactMatchDetector, CompositeDetector and ChunkedDetector implement the port
directly.
The pipeline stages
BaseAnonymizationPipeline chains the stages from detection to de-identified text.
Only the detector is a required constructor argument. Linking, de-identification, and
overlap resolution always run. When omitted, they fall back to built-in defaults. These
defaults are an ExactEntityLinker, an Anonymizer with a LabelCounterPlaceholderFactory, and a
ConfidenceOverlapResolver. The override, expand, entity-resolve, and guard stages
are pass-throughs when not provided.
The pipeline. The stages that always run are in bold, the optional stages have a dashed border.
The Pipeline design page explains why each stage exists and why they run in this order. Here is the role and the default adapter of each.
| Port | Provided adapters | Role |
|---|---|---|
AnyDetector | Gliner2Detector, Gliner2PiiDetector, SpacyDetector, TransformersDetector, PresidioDetector, BridgeDetector, LLMDetector, RegexDetector, ExactMatchDetector, CompositeDetector, ChunkedDetector | Finds the confidential data (personal data, secrets), returns positioned and typed Detection objects. |
AnyOverlapResolver | ConfidenceOverlapResolver, MergeOverlapResolver | Arbitrates overlapping detections, keeps the highest-confidence one or their union. |
AnyDetectionExpander | WordBoundaryExpander | Catches missed occurrences of an already-detected value. |
AnyEntityLinker | ExactEntityLinker | Groups the detections of one value into an Entity. |
AnyEntityResolver | MergeEntityResolver, FuzzyEntityResolver, SeparateEntityResolver | Reconciles entities that share a detection. |
AnyAnonymizer and AnyPlaceholderFactory | Anonymizer and LabelCounterPlaceholderFactory | Replaces each entity with its token. |
AnyGuardRail | DetectorGuardRail, Gliner2GuardRail, LLMGuardRail, ModerationGuardRail | Re-checks the output, raises PIIRemainingError on residual confidential data. |
The override (AnyDetectionOverride, adapter DetectionOverride) is an optional server
component. It applies a deny list and an allow list to every detection set, right after
detection, before span resolution.
The placeholder component and its preservation tags
The anonymizer delegates the shape of the token to a placeholder factory
(AnyPlaceholderFactory). What changes between two factories is what the token
preserves of the original value.
The preservation tags, from the token that keeps nothing to the one that identifies each entity. Each arrow goes from a tag to its parent and reads "is a".
Each tag is a subclass of str. A token is therefore a real string carrying its
preservation level in its own type. These tags are phantom types, which means they
exist only for the type checker. The middleware requires a tag that preserves identity
(PreservesRecognizableIdentity). Plugging a <<PERSON>> factory into the
middleware is therefore an error caught at type-check time, not a runtime surprise.
The provided factories range from the least to the most informative.
RedactPlaceholderFactory emits <<REDACT>>, LabelPlaceholderFactory
emits <<PERSON>>, LabelCounterPlaceholderFactory emits
<<PERSON:1>>, LabelHashPlaceholderFactory emits
<<PERSON:a1b2c3d4>>. MaskPlaceholderFactory keeps the first character
and masks the rest, so Jonathan becomes J*******. The detail is in
Placeholder factories.
The single-text pipeline
AnonymizationPipeline handles an isolated text. It detects, applies the optional
stages that are present, groups into entities, de-identifies, then passes the output to
the guard rail. Its deanonymize method takes the token-to-entity mapping produced by
anonymize and restores the values.
from piighost.components.anonymizer import Anonymizer
from piighost.components.detector import ExactMatchDetector
from piighost.components.linker import ExactEntityLinker
from piighost.components.placeholder import LabelCounterPlaceholderFactory
from piighost.pipeline import AnonymizationPipeline
detector = ExactMatchDetector({"Patrick": "PERSON"})
linker = ExactEntityLinker()
factory = LabelCounterPlaceholderFactory()
anonymizer = Anonymizer(factory)
pipeline = AnonymizationPipeline(
detector=detector,
linker=linker,
anonymizer=anonymizer,
)
result = await pipeline.anonymize("Patrick habite à Paris.")
# result.text -> "<<PERSON:1>> habite à Paris."
# result.tokens -> {Entity("Patrick"): "<<PERSON:1>>"}
restored = pipeline.deanonymize(result.text, result.tokens)
# restored -> "Patrick habite à Paris."The constructor requires only the detector. The linker and anonymizer default to
ExactEntityLinker and an Anonymizer with a LabelCounterPlaceholderFactory.
The other stages come as keyword arguments.
AnonymizationPipeline(
detector,
linker,
anonymizer,
overlap_resolver=None, # AnyOverlapResolver, defaults to ConfidenceOverlapResolver
expander=None, # AnyDetectionExpander
entity_resolver=None, # AnyEntityResolver
guard=None, # AnyGuardRail
override=None, # AnyDetectionOverride
)Omitting overlap_resolver, or passing None, builds a ConfidenceOverlapResolver,
since the render stage needs disjoint spans. The expand, entity-resolve, guard, and
override stages stay disabled when None.
The conversation pipeline
ThreadAnonymizationPipeline shares the same base but adds a conversation memory
(AnyConversationMemory), passed through the memory keyword argument. Without it, the
pipeline builds an InMemoryConversationMemory. An agent chains messages, and
the same Patrick must keep the same <<PERSON:1>> from the
first to the last.
Tokens are assigned over the union of every message's detections in the thread, not over one message alone. A value seen again later therefore recovers its token instead of creating a new one. Rendering, in contrast, stays per message. Only the current message's spans are replaced, because detections from different messages do not share the same offset space.
result = await thread_pipeline.anonymize(text, thread_id="t-42")
restored = await thread_pipeline.deanonymize(reply, thread_id="t-42")
dropped = await thread_pipeline.forget_thread("t-42")- The
thread_idis mandatory. There is no shared default thread, so two callers cannot fall into the same thread and leak each other's confidential data. deanonymizerebuilds the thread's tokens from memory. It therefore restores any text carrying those tokens, including a model reply the pipeline never de-identified.forget_threaderases a thread's whole memory and reports how much was dropped, for the right to erasure.
Value provenance
A value whose first occurrence in the thread comes from a model message is not the user's confidential data.
Tokenizing it would strip the model of its world knowledge. So the memory records the
role of each value's first occurrence (MessageRole.USER or
MessageRole.ASSISTANT), and the pipeline leaves assistant-introduced values in clear.
The conversation memory and encryption
The memory is a repository, an AnyConversationMemory port with three adapters.
InMemoryConversationMemorykeeps everything in a process-local dict, bounded by default. Simple, enough for a single worker.RedisConversationMemorypersists to Redis, for a multi-worker deployment where each worker must see the others' threads.SqlAlchemyConversationMemorypersists to a SQL table, for long conversations that outlive the process.
By nature, a persistent backend stores confidential data, because it keeps the reverse
mapping, which leads each token back to its value. Two optional crypto components,
passed together, protect it on Redis as on SQL. An AnyHasher (Sha256Hasher,
Argon2Hasher) turns each message into a deterministic key without revealing the text.
An AnyCipher (AesGcmCipher) encrypts the detections at rest, so a store leak reveals
neither the message nor the values. The thread_id stays clear, a key prefix in Redis
and a column in the SQL table, so a thread can be enumerated and forgotten.
The LangChain middleware
PIIAnonymizationMiddleware wires the conversation pipeline into a LangChain agent
loop. It contains no de-identification logic, it delegates everything to the pipeline.
It is an adapter between the LangChain world and the core.
The middleware intercepts the agent loop at three points.
abefore_modelde-identifies the messages before the LLM sees them.aafter_modelrestores the model's output for the user display.awrap_tool_callhandles the tool call according to the chosen strategy (ToolCallStrategy), restoring the arguments so the tool receives real data, then de-identifying its response.
The middleware requires a factory that preserves identity, at type-check time. At
runtime, it also refuses a pipeline whose tokens have no delimited grammar, such as a
mask (UnrecognizableFactoryError). That grammar lets it
recognize the tokens the model invents (InventedPlaceholderStrategy). After
restoration, any token still following the placeholder grammar was not emitted by the
pipeline. The detail of the tool strategies is in
Tool-call strategies.
Observation
piighost emits one trace per pipeline stage through a port (AnyObservationTracer), a
seam on top of OpenTelemetry. With no backend configured, a no-op implementation traces
nothing and costs nothing. The pipeline can therefore always emit its traces without
checking whether tracing is active. An optional observation_redactor replaces the values in the traces
with tokens, for a backend not allowed to see confidential data.
The config, composition root
A TOML or JSON file describes the whole pipeline. The config subsystem reads it with
pydantic-settings and turns it into config models. These models are discriminated
unions, where each component type carries a build() method. Assembling the pipeline amounts to calling
build() on each model.
from piighost.config import load_pipeline, load_thread_pipeline
pipeline = load_pipeline("pipeline.toml")
thread_pipeline = load_thread_pipeline("thread.toml")A file without a [memory] section builds a pipeline. A file that declares a [memory] section builds a conversation pipeline. Each loader refuses the file meant for the other. load_pipeline refuses a file with [memory], and load_thread_pipeline a file without one.
The coupling is one-way. Config depends on the core and the adapters, but the core
never imports config. Adding a component means writing an adapter, a config model with
build(), and nothing else. The pipeline does not change.
Data models
All core models are frozen dataclasses, immutable so they can be shared across coroutines without risk.
| Model | Key fields |
|---|---|
Detection | text, label, span: Span, confidence |
Entity | detections: tuple[Detection, ...], label and text as properties |
Span | start, end, overlaps(), extract() |
See also
- Pipeline design: why each stage exists and in which order.
- Placeholder factories: the families of tokens and what they preserve.
- Tool-call strategies: the detail of
awrap_tool_call. - Extending piighost: plugging your own adapter behind a port.
- Data models reference: the fields, methods and validation of
Detection,Entity,SpanandChunk.