--- icon: lucide/puzzle tags: - Advanced - Detector --- # Extending piighost Every pipeline stage is a **port**, a `Protocol` you satisfy by implementing its one method. There is no base class to inherit, and nothing else in the pipeline changes. Where a `Base*` template exists, you can also subclass it. That template supplies the shared skeleton and leaves you a single hook. ```mermaid flowchart LR P[AnonymizationPipeline] -->|detector| D[AnyDetector] P -->|overlap_resolver| O[AnyOverlapResolver] P -->|expander| X[AnyDetectionExpander] P -->|linker| L[AnyEntityLinker] P -->|entity_resolver| R[AnyEntityResolver] P -->|anonymizer| A[AnyAnonymizer] P -->|guard| G[AnyGuardRail] A -->|factory| F[AnyPlaceholderFactory] ``` *The pipeline injects one component per port. Only the detector is required. The linker, anonymizer, and overlap resolver default to built-ins. The expansion, entity resolution, guard rail, and deny and allow list stages are disabled by default.* { .figure-caption } The ports live in each component's `base.py`, under `piighost.components.*`. The data models they exchange live in `piighost.models`. A `Detection` is a `Span(start, end)` carrying `text`, `label`, and a `confidence` in the range 0 to 1. An `Entity` groups the detections that share a value, and derives its `label`, `text`, and `spans` from them. See the [data models reference](reference/models.md) for every field, method and validation error. ## A custom detector A detector finds confidential data (personal data, secrets) in a text. Implement one method: ```python class AnyDetector(Protocol): async def detect(self, text: str) -> list[Detection]: ... ``` `detect` is async so an implementation can await a model server or an LLM API. Return detections in any order. Overlaps and repeats are resolved by later stages, not here. ???+ example "Regex handle detector" ```python import re from piighost.models import Detection, Span class HandleDetector: """Detect @handles as USERNAME.""" async def detect(self, text: str) -> list[Detection]: detections: list[Detection] = [] for match in re.finditer(r"@\w+", text): span = Span(match.start(), match.end()) detections.append( Detection( span=span, text=match.group(), label="USERNAME", confidence=1.0, ) ) return detections ``` ### Use the detector ```python from piighost.pipeline import AnonymizationPipeline detector = HandleDetector() pipeline = AnonymizationPipeline(detector) ``` To feed a detector from a fixed value list in tests, use the built-in `ExactMatchDetector` instead. See [Testing without a model](examples/testing.md). ### For NER models, subclass `BaseNERDetector` The model-backed detectors (`Gliner2Detector`, `SpacyDetector`, `TransformersDetector`) all extend `BaseNERDetector`. `BaseNERDetector` maps the label a model emits internally to the label that appears in `Detection.label`. You can thus query a model with the strings it detects best, while producing clean labels downstream. Pass `labels` as a list to keep each label as is (identity mapping), or as an `{emitted: internal}` dict to rename: ```python from piighost.components.detector.ner import Gliner2Detector # Query GLiNER2 with "person" and "company" but emit "PERSON" / "COMPANY". detector = Gliner2Detector( model="fastino/gliner2-multi-v1", labels={"PERSON": "person", "COMPANY": "company"}, ) ``` ## A custom overlap resolver An overlap resolver takes detections whose spans overlap, and turns them into a set of detections with no overlap. The port: ```python class AnyOverlapResolver(Protocol): def resolve(self, detections: list[Detection]) -> list[Detection]: ... ``` Rather than implement `resolve` from scratch, subclass `BaseOverlapResolver`. It clusters the detections into overlap groups and hands each group to your `_reduce`, so you only decide which detections to keep from a group that overlaps. ???+ example "Longest span wins" ```python from piighost.components.overlap_resolver.base import BaseOverlapResolver from piighost.models import Detection class LongestOverlapResolver(BaseOverlapResolver): """Keep the longest detection in each overlap group.""" def _reduce(self, conflicting: list[Detection]) -> list[Detection]: return [max(conflicting, key=lambda d: d.span.length)] ``` The built-in `ConfidenceOverlapResolver` keeps the highest-confidence detection instead. The overlap resolver is always on. Omit it and the pipeline installs a `ConfidenceOverlapResolver`. Pass your own to change the rule. There is no supported way to disable it, since render assumes disjoint spans and raises `OverlappingSpansError` otherwise. ## A custom expander An expander finds occurrences a detector missed, such as a repeat of a name flagged elsewhere. The port: ```python class AnyDetectionExpander(Protocol): def expand(self, text: str, detections: list[Detection]) -> list[Detection]: ... ``` Subclass `BaseDetectionExpander`. It keeps the original detections. For each one, it adds a detection at every extra occurrence your `_find_occurrences` returns. Each added detection carries the source detection's label and confidence. An occurrence that overlaps a detection already kept is skipped, since the expander runs after the overlap resolver and the renderer refuses overlapping spans. Values are searched longest first, so a full name claims a place before its first name does. ???+ example "Whole-word repeats" ```python from collections.abc import Iterable from piighost.components.expander.base import BaseDetectionExpander from piighost.models import Detection, Span from piighost.text import find_all_word_boundary class WholeWordExpander(BaseDetectionExpander): """Find whole-word repeats of a detected value.""" def _find_occurrences(self, text: str, detection: Detection) -> Iterable[Span]: return find_all_word_boundary(text, detection.text) ``` The built-in `WordBoundaryExpander` does exactly this. The stage is optional. ## A custom entity linker A linker groups into entities the detections that refer to the same value. All occurrences of a value thus share one placeholder. The port: ```python class AnyEntityLinker(Protocol): def link(self, detections: list[Detection]) -> list[Entity]: ... ``` Subclass `BaseEntityLinker`. It groups detections by a key you compute in `_key`. It creates one entity per distinct key, in first-occurrence order. ???+ example "Group by exact value and label" ```python from collections.abc import Hashable from piighost.components.linker.base import BaseEntityLinker from piighost.models import Detection class CaseSensitiveLinker(BaseEntityLinker): """Group detections that share an exact value and label.""" def _key(self, detection: Detection) -> Hashable: return (detection.text, detection.label) ``` The built-in `ExactEntityLinker` groups on the value key. That key is the same for the same words, whatever their spaces and case. `Patrick`{ .pii } and `patrick`{ .pii } therefore become one entity. Use `piighost.text.value_key` in your own linker to follow the same rule, see [Unicode spaces](reference/detectors.md#unicode-spaces). ## A custom entity resolver An entity resolver reconciles entities that should not coexist, such as two entities sharing a detection. The port: ```python class AnyEntityResolver(Protocol): def resolve(self, entities: list[Entity]) -> list[Entity]: ... ``` Subclass `BaseEntityResolver`. It clusters entities that share a detection into groups and hands each group to your `_reduce`. Your `_reduce` returns a consistent set, whether by merging the group into one entity or by keeping the entities apart. The built-ins: - `MergeEntityResolver` merges entities that share a detection, by union-find. - `SeparateEntityResolver` keeps them apart, giving each shared detection to one entity. - `FuzzyEntityResolver` merges entities with similar values (needs the `fuzzy` extra). The stage is optional. ## A custom placeholder factory A placeholder factory turns entities into their replacement tokens. It is generic on a **preservation tag**, a phantom type stating what its tokens preserve. The type checker uses this tag to gate a consumer like the middleware. The port: ```python class AnyPlaceholderFactory(Protocol[PreservationT_co]): def create(self, entities: list[Entity]) -> Mapping[Entity, PreservationT_co]: ... ``` A token is an instance of the tag, and the tag is a `str` subclass. The token is therefore a real string, which carries its preservation level in its own type. `create` must be deterministic. The same entities yield the same tokens on every call, because the pipeline calls it more than once per run. ???+ example "Bracket label factory" ```python from collections.abc import Mapping from piighost.components.placeholder.base import AnyPlaceholderFactory from piighost.components.placeholder.tags import PreservesLabel from piighost.models import Entity class BracketLabelFactory(AnyPlaceholderFactory[PreservesLabel]): """Emit [LABEL] for every entity, collapsing each label to one token.""" def create(self, entities: list[Entity]) -> Mapping[Entity, PreservesLabel]: return {entity: PreservesLabel(f"[{entity.label}]") for entity in entities} ``` `PreservesLabel` says the token reveals the type but not a unique identity. This factory therefore suits one-shot redaction, not the middleware. For a token the middleware can restore and find again, tag it `PreservesRecognizableIdentity` (or a sub-tag such as `PreservesLabeledIdentityOpaque`) and use a delimited grammar like `<>`{ .placeholder }. To wrap an inner form in delimiters without writing the wrapping yourself, subclass `BaseDelimitedPlaceholderFactory`. See [Placeholder factories](placeholder-factories.md) for the full tag taxonomy and worked examples. ### Use the factory ```python from piighost.components.anonymizer import Anonymizer factory = BracketLabelFactory() anonymizer = Anonymizer(factory) ``` ## A custom guard rail A guard rail re-checks the de-identified output for residual confidential data. It classifies, it does not decide. It returns a `GuardVerdict` and leaves the pipeline to raise `PIIRemainingError` when a verdict is flagged. There is no `Base` template, because each guard has its own checking mechanism. The port: ```python class AnyGuardRail(Protocol): async def check(self, text: str) -> GuardVerdict: ... ``` `check` sees only the de-identified text. The placeholders in that text are clearly synthetic. A check that looks for real values therefore does not mistake them for real values. ???+ example "Flag a residual @ sign" ```python from piighost.components.guard.base import GuardVerdict class AtSignGuard: """Flag any residual @ sign as leftover PII.""" async def check(self, text: str) -> GuardVerdict: return GuardVerdict(flagged="@" in text) ``` The built-in `DetectorGuardRail` re-runs a detector and reports the residual detections. The stage is optional. Pass no `guard` and the output is returned unchecked. ### Use the guard rail ```python from piighost.pipeline import AnonymizationPipeline guard = AtSignGuard() pipeline = AnonymizationPipeline(detector, guard=guard) ``` ### A decision model behind the port A decision model does not generate text. It answers a question whose possible answers are fixed in advance, here yes or no. A guard rail does the same on the de-identified text. [`examples/guard_rail_laya.py`](https://github.com/Athroniaeth/piighost/blob/master/examples/guard_rail_laya.py) puts [Laya](https://huggingface.co/convaiinnovations/laya), an Apache 2.0 counterpart of Jev, behind the port in a dozen lines, running locally. It asks whether personal data is left and flags the text above a probability. On 24 de-identified texts, half of them leaking, it caught 11 leaks out of 12 and flagged 5 clean texts out of 12 at a 0.5 threshold. `Gliner2GuardRail` caught 4 leaks, with no false alarm. Placeholders raise its score, so it mostly flags a text dense in tokens by mistake. Its English checkpoint reads French well enough, `laya-multilingual` does not. ## Full composition The stages are independent, so a custom detector, factory, and guard combine freely with the built-ins: ```python from piighost.components.anonymizer import Anonymizer from piighost.components.entity_resolver import MergeEntityResolver from piighost.pipeline import AnonymizationPipeline pipeline = AnonymizationPipeline( HandleDetector(), anonymizer=Anonymizer(BracketLabelFactory()), entity_resolver=MergeEntityResolver(), guard=AtSignGuard(), ) ``` To unit-test a custom component deterministically, feed it through `ExactMatchDetector`. See [Testing without a model](examples/testing.md).