--- icon: lucide/shield-check tags: - Guard --- # Guard rails reference Module: `piighost.components.guard` A guard rail is the pipeline's last, optional stage. It re-checks the de-identified text for residual confidential values. When it finds any, the pipeline raises `PIIRemainingError` instead of returning a leak. Every guard satisfies the `AnyGuardRail` port, an `async def check(self, text: str) -> GuardVerdict`. It returns a `GuardVerdict` carrying whether confidential values seem to remain and how it knows. Unlike the other stages, guards share no `Base*` template, because their checking mechanisms have no shared skeleton. One re-runs a local detector, another calls an external API. The guard classifies, it does not decide. It reports a verdict. The pipeline turns a flagged verdict into an exception, and your code chooses how to react. ```python from piighost.components.guard import ( DetectorGuardRail, LLMGuardRail, ModerationGuardRail, ) ``` ## Wire a guard into a pipeline `AnonymizationPipeline` takes an optional `guard` argument, disabled by default. When set, the guard runs on the rendered output after de-identification, and the pipeline raises `PIIRemainingError` if the guard flags anything unexpected. ```python from piighost.catalog import pull from piighost.components.detector import ExactMatchDetector, RegexDetector from piighost.components.guard import DetectorGuardRail from piighost.exceptions import PIIRemainingError from piighost.pipeline import AnonymizationPipeline # The primary detector only knows the literal name. The guard re-runs a broader # email and phone regex over the short output to catch structured PII it missed. guard_detector = RegexDetector( {**pull("catalog:piighost/generic"), **pull("catalog:piighost/us")} ) pipeline = AnonymizationPipeline( ExactMatchDetector({"Emma Doe": "PERSON"}), guard=DetectorGuardRail(guard_detector), ) try: result = await pipeline.anonymize("Emma Doe, reachable at emma@acme.com.") except PIIRemainingError as error: print(error) # Anonymized text still contains PII: ['EMAIL'] print(error.detections) # the residual detections behind the flag ``` The runnable version is [`examples/guard_rail.py`](https://github.com/Athroniaeth/piighost/blob/master/examples/guard_rail.py). This script also uses a guard standalone. It calls `await guard.check(text)` and reads the verdict without raising. The local-model version is [`examples/guard_rail_local_model.py`](https://github.com/Athroniaeth/piighost/blob/master/examples/guard_rail_local_model.py). ## `DetectorGuardRail` Re-runs a detector on the de-identified output and flags whatever it still finds, carrying the residual detections on the verdict. ```python DetectorGuardRail(detector: AnyDetector) ``` This guard only adds value with a detector different from the pipeline's. Re-running the same one finds nothing, since the pipeline already de-identified everything it detects. A stronger or complementary detector, run as a cheap second pass over the short de-identified output, catches what the primary detector missed. The synthetic placeholders are not shaped like real values, so a detector meant for real values leaves them alone. `DetectorGuardRail` needs no optional extra. ### A local model as the guard The complementary detector is often a model, because the shapes a regex is good at are exactly the ones the primary pass already caught. A name, an address or a company name is what slips through: ```python DetectorGuardRail( Gliner2Detector( model="fastino/GLiNER2-Guardrails-PII-Multi", labels=["person", "address"], threshold=0.5, ) ) ``` This guard localizes what leaked. That is what a detector-backed guard gives you over a classifier. For a text-level verdict from the same checkpoint, without spans and in one forward pass, see [`Gliner2GuardRail`](#gliner2guardrail). ## `LLMGuardRail` Wraps an `LLMDetector` configured with a guard prompt that tells the model to ignore placeholders and flag only residual clear-form PII, then reports a verdict. ```python LLMGuardRail( model: BaseChatModel | str, labels: list[str] | dict[str, str], prompt: str | None = None, provider: str | None = None, prefix: str = "<<", suffix: str = ">>", fail_open: bool = False, ) ``` A `str` model is loaded like `LLMDetector`'s. A loaded instance is used as-is. A custom `prompt` must contain a `{labels}` placeholder. When no custom prompt is given, `prefix` and `suffix` (default `<<` and `>>`) shape the default prompt's placeholder examples to match the delimiters the pipeline emits. An output the guard cannot read raises `UnreadableOutputError` rather than report the text clean, unless `fail_open=True`. This behavior is the same as for `LLMDetector`. Requires `piighost[llm]`. ## `Gliner2GuardRail` Classifies the de-identified output with a GLiNER2 guardrail model running in the process, and flags the verdict when it comes back unsafe with enough confidence. ```python Gliner2GuardRail( model: GLiNER2 | str = "fastino/GLiNER2-Guardrails-PII-Multi", task: str = "response_safety", labels: tuple[str, ...] = ("safe", "unsafe"), threshold: float = 0.5, ) ``` This is `ModerationGuardRail` without the API call, and that difference is the point. The text a guard checks is the text that still holds whatever leaked. Sending it to a third party is therefore an odd shape for the last stage of a de-identification pipeline. The default checkpoint is 300M parameters, multilingual over seven languages, and does safety moderation and PII extraction in one forward pass. A `str` model is loaded with `GLiNER2.from_pretrained`, and a loaded instance is used as-is. That is how one checkpoint is shared between this guard and a `Gliner2Detector`. The `labels` pair is read positionally, and the refused answer comes last. Another task of the same model is therefore read the same way. For example, `task="response_refusal"` with `labels=("compliance", "refusal")` flags a refusal instead. Requires `piighost[gliner2]`. ```python from piighost.components.detector import RegexDetector from piighost.components.guard import Gliner2GuardRail from piighost.pipeline import AnonymizationPipeline pipeline = AnonymizationPipeline( RegexDetector({"EMAIL": r"[\w.+-]+@[\w.-]+\.\w{2,}"}), guard=Gliner2GuardRail(), ) await pipeline.anonymize("Write to a@b.co about the invoice.") # Write to <> about the invoice. await pipeline.anonymize("Write to John Doe, 12 rue des Lilas, 75008 Paris.") # PIIRemainingError: A guard flagged residual PII (score 0.997) ``` `Gliner2GuardRail` gives a text-level verdict, so it localizes nothing. `detections` is empty and only `score` is set. Pair it with a `DetectorGuardRail` when you need to know which value leaked. The placeholders the pipeline emits do not trip it. For example, `<>` scores safe at 0.989. The runnable version is [`examples/guard_rail_local_model.py`](https://github.com/Athroniaeth/piighost/blob/master/examples/guard_rail_local_model.py). ## `ModerationGuardRail` Classifies residual PII with Mistral's moderation model, reading the PII category score and flagging the verdict when it reaches the threshold. ```python ModerationGuardRail( client: Mistral, model: str = "mistral-moderation-latest", threshold: float = 0.5, ) ``` This guard classifies the text, it does not detect values. It therefore catches PII a detection-based pipeline cannot localize. In return, it gives a text-level verdict without spans. Requires `piighost[mistral]`. ## `GuardVerdict` and `PIIRemainingError` `check` returns a frozen `GuardVerdict(flagged: bool, score: float | None, detections: tuple[Detection, ...])`. The detail depends on the guard. It is a score from a moderation model, or the residual detections from a detector. Both are optional. When a guard flags confidential values, the pipeline raises `PIIRemainingError` (a subclass of `GuardError`, itself a `PIIGhostError`). Its message names the leaked labels or the score. Its `detections` attribute holds the residual detections. It stays empty for a score-based guard, which localizes nothing. ## Configure a guard from a file A `[guard]` section adds the stage, and its `type` field picks the guard. ```toml [guard] type = "detector" [guard.detector] type = "regex" catalogs = ["catalog:piighost/generic", "catalog:piighost/us"] ``` | `type` | Fields | Extra | |--------|--------|-------| | `detector` | `[guard.detector]` (a detector config) | | | `gliner2` | `model` (default `fastino/GLiNER2-Guardrails-PII-Multi`), `task`, `labels`, `threshold` | `gliner2` | | `llm` | `model`, `labels`, `prompt` (optional), `provider` (optional) | `llm` | | `moderation` | `model` (default `mistral-moderation-latest`), `threshold` (default `0.5`) | `mistral` | The moderation guard reads `MISTRAL_API_KEY` from the environment at build time. It raises `ConfigError` if the variable is unset. Every `[guard]` key is in the [configuration reference](../configuration/toml.md). ## See also - [Pipeline](pipeline.md): where the guard stage sits in the run. - [Detectors](detectors.md): the detectors a `DetectorGuardRail` re-runs. - [Security](../security.md): what a guard does and does not protect against.