Skip to content

Pre-built detectors

piighost pulls ready-to-use regex pattern catalogs for structured PII (email, IP, IBAN, phone) from the piighost catalog. This guide shows how to load them, merge them, and combine several detectors, with the piighost core alone.

Four catalog groups cover the common formats. Each one is a set of label to pattern entries.

  • catalog:piighost/generic: email, URL, IPv4, credit card, country-agnostic
  • catalog:piighost/us: phone, ZIP, ITIN, SSN, prefixed US_
  • catalog:piighost/eu: pan-European ISO 13616 IBAN
  • catalog:piighost/fr: phone, IBAN, NIR, SIRET, SIREN, prefixed FR_

A reference without a suffix follows the latest version of the group, fetched each time a detector is built. To freeze a version, add its commit after a colon, as in catalog:piighost/generic:fab51b33. The group is then fetched once, and read from the on-disk cache afterwards, offline included. Secrets such as API keys are in the catalog groups piighost/secrets and piighost/secrets-extended. These groups are pulled the same way, for example with catalogs = ["catalog:piighost/secrets"] in a config.

For the label details, see the detectors reference.

Use a single group

Build a RegexDetector from the group with from_catalog, then assemble the pipeline.

import asyncio

from piighost.components.detector import RegexDetector
from piighost.pipeline import AnonymizationPipeline

detector = RegexDetector.from_catalog("catalog:piighost/generic")
pipeline = AnonymizationPipeline(detector)


async def main() -> None:
    result = await pipeline.anonymize("Email alice@example.com, server 192.168.1.42.")
    print(result.text)


asyncio.run(main())

The output should be:

Email <<EMAIL:1>>, server <<IPV4:1>>.

Merge generic and regional groups

If you want to cover both generic PII and a region's PII, pull each group with pull and merge the dictionaries you get. pull returns a label to pattern dictionary. When two dictionaries share a label, the entry from the right-hand dictionary wins.

from piighost.catalog import pull

patterns = {**pull("catalog:piighost/generic"), **pull("catalog:piighost/fr")}
detector = RegexDetector(patterns)
pipeline = AnonymizationPipeline(detector)


async def main() -> None:
    result = await pipeline.anonymize(
        "IBAN FR7630006000011234567890189, email marie@exemple.fr, tel 06 12 34 56 78."
    )
    print(result.text)


asyncio.run(main())

The output should be:

IBAN <<FR_IBAN:1>>, email <<EMAIL:1>>, tel <<FR_PHONE:1>>.

To keep only some labels, build a hand-picked dictionary.

generic = pull("catalog:piighost/generic")
french = pull("catalog:piighost/fr")
patterns = {
    "EMAIL": generic["EMAIL"],
    "FR_IBAN": french["FR_IBAN"],
}
detector = RegexDetector(patterns)

Combine several detectors

CompositeDetector runs several detectors over the same text and concatenates their detections. Overlaps are arbitrated by the pipeline's resolution stage. This is how you pair a regex detector with one that recognizes names.

from piighost.components.detector import (
    CompositeDetector,
    ExactMatchDetector,
    RegexDetector,
)

exact_detector = ExactMatchDetector({"Patrick": "PERSON"})
regex_detector = RegexDetector.from_catalog("catalog:piighost/generic")
detector = CompositeDetector([exact_detector, regex_detector])
pipeline = AnonymizationPipeline(detector)


async def main() -> None:
    result = await pipeline.anonymize("Patrick emailed alice@example.com.")
    print(result.text)


asyncio.run(main())

The output should be:

<<PERSON:1>> emailed <<EMAIL:1>>.

In production, replace ExactMatchDetector with an NER or LLM detector, see the detectors reference. ExactMatchDetector is used here to keep the example reproducible without a model.

Handle a long text

An NER detector has a bounded context window, and a long document can exceed it. ChunkedDetector wraps any detector, splits the text into overlapping chunks, detects on each, and remaps the offsets back onto the original text.

from piighost.components.detector import ChunkedDetector, RegexDetector
from piighost.text import RecursiveCharacterTextSplitter

regex_detector = RegexDetector.from_catalog("catalog:piighost/generic")
splitter = RecursiveCharacterTextSplitter(chunk_size=40, chunk_overlap=10)
detector = ChunkedDetector(regex_detector, splitter=splitter)
pipeline = AnonymizationPipeline(detector)


async def main() -> None:
    text = (
        "Filler text here. Reach alice@example.com now. "
        "More filler padding words. Then bob@example.org later."
    )
    result = await pipeline.anonymize(text)
    print(result.text)


asyncio.run(main())

The output should be:

Filler text here. Reach <<EMAIL:1>> now. More filler padding words. Then <<EMAIL:2>> later.

Leave splitter=None for a default RecursiveCharacterTextSplitter tuned for real documents. The reduced chunk_size above only forces several chunks in a short example.

Load groups from a config file

If you drive the pipeline from a config file rather than from code, a regex detector accepts a catalogs key.

[detector]
type = "regex"
catalogs = ["catalog:piighost/generic", "catalog:piighost/fr"]

The detector merges the groups first, then the inline patterns. So on a shared label, an inline pattern wins over the one from a group. See the TOML configuration.

See also