Skip to content

Architecture

piighost follows a hexagonal architecture, also known as ports and adapters. The core knows only abstract contracts, the ports. Each concrete implementation, a GLiNER2 detector, a Redis backend, a LangChain middleware, is an adapter that satisfies a port without the core knowing about it. The de-identification pipeline is assembled by injecting the chosen adapters behind the ports it expects.


The three rings

The code reads as three rings, from the most abstract to the most concrete. The direction of the dependencies is fixed once and for all. An outer ring imports an inner ring, and an inner ring never imports an outer ring.

DiagramDiagram

Three rings and the composition root. Dependencies always point toward the core.

  • Core. The data models (Detection, Entity, Span, frozen dataclasses) and the ports. No external dependency, no pydantic, no I/O.
  • Application. The pipeline orchestration, which depends only on the core ports. This is where anonymize, deanonymize, and forget_thread live.
  • Adapters. The concrete implementations of the ports, that is detectors, resolvers, factories, guard rails, memory backends, observation, HTTP client, middleware. Each adapter imports the core, never the reverse.
  • Config. The composition root. It is the only place allowed to know both the ports and the concrete adapters, in order to assemble them.

Ports and templates

A port is a Python Protocol marked runtime_checkable, in each component's base.py. The typing there is structural. An object satisfies the port as soon as it has the methods, without inheriting from it. The pipeline depends on the port, never on a concrete class.

@runtime_checkable
class AnyDetector(Protocol):
    async def detect(self, text: str) -> list[Detection]: ...

When several adapters of one port share a skeleton, that skeleton lives in a Base* class, an abstract class that applies the Template Method pattern. The skeleton is written once in the base class, and each subclass provides only the step that varies.

class BaseEntityLinker(ABC):
    def link(self, detections: list[Detection]) -> list[Entity]:
        # common skeleton: group by key
        ...

    @abstractmethod
    def _key(self, detection: Detection) -> Hashable:
        # only varying step, defined by the subclass
        ...

Five ports have no template shared by all their adapters, the detector, override, guard rail, memory backend and cipher ports. Their adapters differ by their whole mechanism, not by a single step, so they have nothing common to factor out. This is the deliberate exception to the always-template rule.

The detector is a partial exception. The model detectors (Gliner2Detector, SpacyDetector, TransformersDetector, PresidioDetector, BridgeDetector, LLMDetector) share the BaseNERDetector template. It re-reads each detection's text from the source, applies the confidence threshold and maps the labels. RegexDetector, ExactMatchDetector, CompositeDetector and ChunkedDetector implement the port directly.


The pipeline stages

BaseAnonymizationPipeline chains the stages from detection to de-identified text. Only the detector is a required constructor argument. Linking, de-identification, and overlap resolution always run. When omitted, they fall back to built-in defaults. These defaults are an ExactEntityLinker, an Anonymizer with a LabelCounterPlaceholderFactory, and a ConfidenceOverlapResolver. The override, expand, entity-resolve, and guard stages are pass-throughs when not provided.

DiagramDiagram

The pipeline. The stages that always run are in bold, the optional stages have a dashed border.

The Pipeline design page explains why each stage exists and why they run in this order. Here is the role and the default adapter of each.

PortProvided adaptersRole
AnyDetectorGliner2Detector, Gliner2PiiDetector, SpacyDetector, TransformersDetector, PresidioDetector, BridgeDetector, LLMDetector, RegexDetector, ExactMatchDetector, CompositeDetector, ChunkedDetectorFinds the confidential data (personal data, secrets), returns positioned and typed Detection objects.
AnyOverlapResolverConfidenceOverlapResolver, MergeOverlapResolverArbitrates overlapping detections, keeps the highest-confidence one or their union.
AnyDetectionExpanderWordBoundaryExpanderCatches missed occurrences of an already-detected value.
AnyEntityLinkerExactEntityLinkerGroups the detections of one value into an Entity.
AnyEntityResolverMergeEntityResolver, FuzzyEntityResolver, SeparateEntityResolverReconciles entities that share a detection.
AnyAnonymizer and AnyPlaceholderFactoryAnonymizer and LabelCounterPlaceholderFactoryReplaces each entity with its token.
AnyGuardRailDetectorGuardRail, Gliner2GuardRail, LLMGuardRail, ModerationGuardRailRe-checks the output, raises PIIRemainingError on residual confidential data.

The override (AnyDetectionOverride, adapter DetectionOverride) is an optional server component. It applies a deny list and an allow list to every detection set, right after detection, before span resolution.


The placeholder component and its preservation tags

The anonymizer delegates the shape of the token to a placeholder factory (AnyPlaceholderFactory). What changes between two factories is what the token preserves of the original value.

DiagramDiagram

The preservation tags, from the token that keeps nothing to the one that identifies each entity. Each arrow goes from a tag to its parent and reads "is a".

Each tag is a subclass of str. A token is therefore a real string carrying its preservation level in its own type. These tags are phantom types, which means they exist only for the type checker. The middleware requires a tag that preserves identity (PreservesRecognizableIdentity). Plugging a <<PERSON>> factory into the middleware is therefore an error caught at type-check time, not a runtime surprise.

The provided factories range from the least to the most informative. RedactPlaceholderFactory emits <<REDACT>>, LabelPlaceholderFactory emits <<PERSON>>, LabelCounterPlaceholderFactory emits <<PERSON:1>>, LabelHashPlaceholderFactory emits <<PERSON:a1b2c3d4>>. MaskPlaceholderFactory keeps the first character and masks the rest, so Jonathan becomes J*******. The detail is in Placeholder factories.


The single-text pipeline

AnonymizationPipeline handles an isolated text. It detects, applies the optional stages that are present, groups into entities, de-identifies, then passes the output to the guard rail. Its deanonymize method takes the token-to-entity mapping produced by anonymize and restores the values.

from piighost.components.anonymizer import Anonymizer
from piighost.components.detector import ExactMatchDetector
from piighost.components.linker import ExactEntityLinker
from piighost.components.placeholder import LabelCounterPlaceholderFactory
from piighost.pipeline import AnonymizationPipeline

detector = ExactMatchDetector({"Patrick": "PERSON"})
linker = ExactEntityLinker()
factory = LabelCounterPlaceholderFactory()
anonymizer = Anonymizer(factory)
pipeline = AnonymizationPipeline(
    detector=detector,
    linker=linker,
    anonymizer=anonymizer,
)
result = await pipeline.anonymize("Patrick habite à Paris.")
# result.text   -> "<<PERSON:1>> habite à Paris."
# result.tokens -> {Entity("Patrick"): "<<PERSON:1>>"}
restored = pipeline.deanonymize(result.text, result.tokens)
# restored -> "Patrick habite à Paris."

The constructor requires only the detector. The linker and anonymizer default to ExactEntityLinker and an Anonymizer with a LabelCounterPlaceholderFactory. The other stages come as keyword arguments.

AnonymizationPipeline(
    detector,
    linker,
    anonymizer,
    overlap_resolver=None,  # AnyOverlapResolver, defaults to ConfidenceOverlapResolver
    expander=None,  # AnyDetectionExpander
    entity_resolver=None,  # AnyEntityResolver
    guard=None,  # AnyGuardRail
    override=None,  # AnyDetectionOverride
)

Omitting overlap_resolver, or passing None, builds a ConfidenceOverlapResolver, since the render stage needs disjoint spans. The expand, entity-resolve, guard, and override stages stay disabled when None.


The conversation pipeline

ThreadAnonymizationPipeline shares the same base but adds a conversation memory (AnyConversationMemory), passed through the memory keyword argument. Without it, the pipeline builds an InMemoryConversationMemory. An agent chains messages, and the same Patrick must keep the same <<PERSON:1>> from the first to the last.

Tokens are assigned over the union of every message's detections in the thread, not over one message alone. A value seen again later therefore recovers its token instead of creating a new one. Rendering, in contrast, stays per message. Only the current message's spans are replaced, because detections from different messages do not share the same offset space.

result = await thread_pipeline.anonymize(text, thread_id="t-42")
restored = await thread_pipeline.deanonymize(reply, thread_id="t-42")
dropped = await thread_pipeline.forget_thread("t-42")
  • The thread_id is mandatory. There is no shared default thread, so two callers cannot fall into the same thread and leak each other's confidential data.
  • deanonymize rebuilds the thread's tokens from memory. It therefore restores any text carrying those tokens, including a model reply the pipeline never de-identified.
  • forget_thread erases a thread's whole memory and reports how much was dropped, for the right to erasure.

Value provenance

A value whose first occurrence in the thread comes from a model message is not the user's confidential data. Tokenizing it would strip the model of its world knowledge. So the memory records the role of each value's first occurrence (MessageRole.USER or MessageRole.ASSISTANT), and the pipeline leaves assistant-introduced values in clear.


The conversation memory and encryption

The memory is a repository, an AnyConversationMemory port with three adapters.

  • InMemoryConversationMemory keeps everything in a process-local dict, bounded by default. Simple, enough for a single worker.
  • RedisConversationMemory persists to Redis, for a multi-worker deployment where each worker must see the others' threads.
  • SqlAlchemyConversationMemory persists to a SQL table, for long conversations that outlive the process.

By nature, a persistent backend stores confidential data, because it keeps the reverse mapping, which leads each token back to its value. Two optional crypto components, passed together, protect it on Redis as on SQL. An AnyHasher (Sha256Hasher, Argon2Hasher) turns each message into a deterministic key without revealing the text. An AnyCipher (AesGcmCipher) encrypts the detections at rest, so a store leak reveals neither the message nor the values. The thread_id stays clear, a key prefix in Redis and a column in the SQL table, so a thread can be enumerated and forgotten.


The LangChain middleware

PIIAnonymizationMiddleware wires the conversation pipeline into a LangChain agent loop. It contains no de-identification logic, it delegates everything to the pipeline. It is an adapter between the LangChain world and the core.

DiagramDiagram

The middleware intercepts the agent loop at three points.

  • abefore_model de-identifies the messages before the LLM sees them.
  • aafter_model restores the model's output for the user display.
  • awrap_tool_call handles the tool call according to the chosen strategy (ToolCallStrategy), restoring the arguments so the tool receives real data, then de-identifying its response.

The middleware requires a factory that preserves identity, at type-check time. At runtime, it also refuses a pipeline whose tokens have no delimited grammar, such as a mask (UnrecognizableFactoryError). That grammar lets it recognize the tokens the model invents (InventedPlaceholderStrategy). After restoration, any token still following the placeholder grammar was not emitted by the pipeline. The detail of the tool strategies is in Tool-call strategies.


Observation

piighost emits one trace per pipeline stage through a port (AnyObservationTracer), a seam on top of OpenTelemetry. With no backend configured, a no-op implementation traces nothing and costs nothing. The pipeline can therefore always emit its traces without checking whether tracing is active. An optional observation_redactor replaces the values in the traces with tokens, for a backend not allowed to see confidential data.


The config, composition root

A TOML or JSON file describes the whole pipeline. The config subsystem reads it with pydantic-settings and turns it into config models. These models are discriminated unions, where each component type carries a build() method. Assembling the pipeline amounts to calling build() on each model.

from piighost.config import load_pipeline, load_thread_pipeline

pipeline = load_pipeline("pipeline.toml")
thread_pipeline = load_thread_pipeline("thread.toml")

A file without a [memory] section builds a pipeline. A file that declares a [memory] section builds a conversation pipeline. Each loader refuses the file meant for the other. load_pipeline refuses a file with [memory], and load_thread_pipeline a file without one.

The coupling is one-way. Config depends on the core and the adapters, but the core never imports config. Adding a component means writing an adapter, a config model with build(), and nothing else. The pipeline does not change.


Data models

All core models are frozen dataclasses, immutable so they can be shared across coroutines without risk.

ModelKey fields
Detectiontext, label, span: Span, confidence
Entitydetections: tuple[Detection, ...], label and text as properties
Spanstart, end, overlaps(), extract()

See also