Guard rails reference
Module: piighost.components.guard
A guard rail is the pipeline's last, optional stage. It re-checks the de-identified text for residual confidential values. When it finds any, the pipeline raises PIIRemainingError instead of returning a leak. Every guard satisfies the AnyGuardRail port, an async def check(self, text: str) -> GuardVerdict. It returns a GuardVerdict carrying whether confidential values seem to remain and how it knows. Unlike the other stages, guards share no Base* template, because their checking mechanisms have no shared skeleton. One re-runs a local detector, another calls an external API.
The guard classifies, it does not decide. It reports a verdict. The pipeline turns a flagged verdict into an exception, and your code chooses how to react.
from piighost.components.guard import (
DetectorGuardRail,
LLMGuardRail,
ModerationGuardRail,
)Wire a guard into a pipeline
AnonymizationPipeline takes an optional guard argument, disabled by default. When set, the guard runs on the rendered output after de-identification, and the pipeline raises PIIRemainingError if the guard flags anything unexpected.
from piighost.catalog import pull
from piighost.components.detector import ExactMatchDetector, RegexDetector
from piighost.components.guard import DetectorGuardRail
from piighost.exceptions import PIIRemainingError
from piighost.pipeline import AnonymizationPipeline
# The primary detector only knows the literal name. The guard re-runs a broader
# email and phone regex over the short output to catch structured PII it missed.
guard_detector = RegexDetector(
{**pull("catalog:piighost/generic"), **pull("catalog:piighost/us")}
)
pipeline = AnonymizationPipeline(
ExactMatchDetector({"Emma Doe": "PERSON"}),
guard=DetectorGuardRail(guard_detector),
)
try:
result = await pipeline.anonymize("Emma Doe, reachable at emma@acme.com.")
except PIIRemainingError as error:
print(error) # Anonymized text still contains PII: ['EMAIL']
print(error.detections) # the residual detections behind the flagThe runnable version is examples/guard_rail.py. This script also uses a guard standalone. It calls await guard.check(text) and reads the verdict without raising. The local-model version is examples/guard_rail_local_model.py.
DetectorGuardRail
Re-runs a detector on the de-identified output and flags whatever it still finds, carrying the residual detections on the verdict.
DetectorGuardRail(detector: AnyDetector)This guard only adds value with a detector different from the pipeline's. Re-running the same one finds nothing, since the pipeline already de-identified everything it detects. A stronger or complementary detector, run as a cheap second pass over the short de-identified output, catches what the primary detector missed. The synthetic placeholders are not shaped like real values, so a detector meant for real values leaves them alone. DetectorGuardRail needs no optional extra.
A local model as the guard
The complementary detector is often a model, because the shapes a regex is good at are exactly the ones the primary pass already caught. A name, an address or a company name is what slips through:
DetectorGuardRail(
Gliner2Detector(
model="fastino/GLiNER2-Guardrails-PII-Multi",
labels=["person", "address"],
threshold=0.5,
)
)This guard localizes what leaked. That is what a detector-backed guard gives you over a classifier. For a text-level verdict from the same checkpoint, without spans and in one forward pass, see Gliner2GuardRail.
LLMGuardRail
Wraps an LLMDetector configured with a guard prompt that tells the model to ignore placeholders and flag only residual clear-form PII, then reports a verdict.
LLMGuardRail(
model: BaseChatModel | str,
labels: list[str] | dict[str, str],
prompt: str | None = None,
provider: str | None = None,
prefix: str = "<<",
suffix: str = ">>",
fail_open: bool = False,
)A str model is loaded like LLMDetector's. A loaded instance is used as-is. A custom prompt must contain a {labels} placeholder. When no custom prompt is given, prefix and suffix (default << and >>) shape the default prompt's placeholder examples to match the delimiters the pipeline emits. An output the guard cannot read raises UnreadableOutputError rather than report the text clean, unless fail_open=True. This behavior is the same as for LLMDetector. Requires piighost[llm].
Gliner2GuardRail
Classifies the de-identified output with a GLiNER2 guardrail model running in the process, and flags the verdict when it comes back unsafe with enough confidence.
Gliner2GuardRail(
model: GLiNER2 | str = "fastino/GLiNER2-Guardrails-PII-Multi",
task: str = "response_safety",
labels: tuple[str, ...] = ("safe", "unsafe"),
threshold: float = 0.5,
)This is ModerationGuardRail without the API call, and that difference is the point. The text a guard checks is the text that still holds whatever leaked. Sending it to a third party is therefore an odd shape for the last stage of a de-identification pipeline. The default checkpoint is 300M parameters, multilingual over seven languages, and does safety moderation and PII extraction in one forward pass.
A str model is loaded with GLiNER2.from_pretrained, and a loaded instance is used as-is. That is how one checkpoint is shared between this guard and a Gliner2Detector. The labels pair is read positionally, and the refused answer comes last. Another task of the same model is therefore read the same way. For example, task="response_refusal" with labels=("compliance", "refusal") flags a refusal instead. Requires piighost[gliner2].
from piighost.components.detector import RegexDetector
from piighost.components.guard import Gliner2GuardRail
from piighost.pipeline import AnonymizationPipeline
pipeline = AnonymizationPipeline(
RegexDetector({"EMAIL": r"[\w.+-]+@[\w.-]+\.\w{2,}"}),
guard=Gliner2GuardRail(),
)
await pipeline.anonymize("Write to a@b.co about the invoice.")
# Write to <<EMAIL:1>> about the invoice.
await pipeline.anonymize("Write to John Doe, 12 rue des Lilas, 75008 Paris.")
# PIIRemainingError: A guard flagged residual PII (score 0.997)Gliner2GuardRail gives a text-level verdict, so it localizes nothing. detections is empty and only score is set. Pair it with a DetectorGuardRail when you need to know which value leaked. The placeholders the pipeline emits do not trip it. For example, <<EMAIL:1>> scores safe at 0.989. The runnable version is examples/guard_rail_local_model.py.
ModerationGuardRail
Classifies residual PII with Mistral's moderation model, reading the PII category score and flagging the verdict when it reaches the threshold.
ModerationGuardRail(
client: Mistral,
model: str = "mistral-moderation-latest",
threshold: float = 0.5,
)This guard classifies the text, it does not detect values. It therefore catches PII a detection-based pipeline cannot localize. In return, it gives a text-level verdict without spans. Requires piighost[mistral].
GuardVerdict and PIIRemainingError
check returns a frozen GuardVerdict(flagged: bool, score: float | None, detections: tuple[Detection, ...]). The detail depends on the guard. It is a score from a moderation model, or the residual detections from a detector. Both are optional.
When a guard flags confidential values, the pipeline raises PIIRemainingError (a subclass of GuardError, itself a PIIGhostError). Its message names the leaked labels or the score. Its detections attribute holds the residual detections. It stays empty for a score-based guard, which localizes nothing.
Configure a guard from a file
A [guard] section adds the stage, and its type field picks the guard.
[guard]
type = "detector"
[guard.detector]
type = "regex"
catalogs = ["catalog:piighost/generic", "catalog:piighost/us"]type | Fields | Extra |
|---|---|---|
detector | [guard.detector] (a detector config) | |
gliner2 | model (default fastino/GLiNER2-Guardrails-PII-Multi), task, labels, threshold | gliner2 |
llm | model, labels, prompt (optional), provider (optional) | llm |
moderation | model (default mistral-moderation-latest), threshold (default 0.5) | mistral |
The moderation guard reads MISTRAL_API_KEY from the environment at build time. It raises ConfigError if the variable is unset. Every [guard] key is in the configuration reference.