--- icon: lucide/scan-search tags: - Detector - Regex --- # Pre-built detectors `piighost` pulls ready-to-use regex pattern catalogs for structured PII (email, IP, IBAN, phone) from the [piighost catalog](https://catalog.piighost.dev). This guide shows how to load them, merge them, and combine several detectors, with the `piighost` core alone. Four catalog groups cover the common formats. Each one is a set of `label` to `pattern` entries. - `catalog:piighost/generic`: email, URL, IPv4, credit card, country-agnostic - `catalog:piighost/us`: phone, ZIP, ITIN, SSN, prefixed `US_` - `catalog:piighost/eu`: pan-European ISO 13616 IBAN - `catalog:piighost/fr`: phone, IBAN, NIR, SIRET, SIREN, prefixed `FR_` A reference without a suffix follows the latest version of the group, fetched each time a detector is built. To freeze a version, add its commit after a colon, as in `catalog:piighost/generic:fab51b33`. The group is then fetched once, and read from the on-disk cache afterwards, offline included. Secrets such as API keys are in the catalog groups `piighost/secrets` and `piighost/secrets-extended`. These groups are pulled the same way, for example with `catalogs = ["catalog:piighost/secrets"]` in a config. For the label details, see the [detectors reference](../reference/detectors.md). ## Use a single group Build a `RegexDetector` from the group with `from_catalog`, then assemble the pipeline. ```python import asyncio from piighost.components.detector import RegexDetector from piighost.pipeline import AnonymizationPipeline detector = RegexDetector.from_catalog("catalog:piighost/generic") pipeline = AnonymizationPipeline(detector) async def main() -> None: result = await pipeline.anonymize("Email alice@example.com, server 192.168.1.42.") print(result.text) asyncio.run(main()) ``` The output should be: ```text Email <>, server <>. ``` ## Merge generic and regional groups If you want to cover both generic PII and a region's PII, pull each group with `pull` and merge the dictionaries you get. `pull` returns a `label` to `pattern` dictionary. When two dictionaries share a label, the entry from the right-hand dictionary wins. ```python from piighost.catalog import pull patterns = {**pull("catalog:piighost/generic"), **pull("catalog:piighost/fr")} detector = RegexDetector(patterns) pipeline = AnonymizationPipeline(detector) async def main() -> None: result = await pipeline.anonymize( "IBAN FR7630006000011234567890189, email marie@exemple.fr, tel 06 12 34 56 78." ) print(result.text) asyncio.run(main()) ``` The output should be: ```text IBAN <>, email <>, tel <>. ``` To keep only some labels, build a hand-picked dictionary. ```python generic = pull("catalog:piighost/generic") french = pull("catalog:piighost/fr") patterns = { "EMAIL": generic["EMAIL"], "FR_IBAN": french["FR_IBAN"], } detector = RegexDetector(patterns) ``` ## Combine several detectors `CompositeDetector` runs several detectors over the same text and concatenates their detections. Overlaps are arbitrated by the pipeline's resolution stage. This is how you pair a regex detector with one that recognizes names. ```python from piighost.components.detector import ( CompositeDetector, ExactMatchDetector, RegexDetector, ) exact_detector = ExactMatchDetector({"Patrick": "PERSON"}) regex_detector = RegexDetector.from_catalog("catalog:piighost/generic") detector = CompositeDetector([exact_detector, regex_detector]) pipeline = AnonymizationPipeline(detector) async def main() -> None: result = await pipeline.anonymize("Patrick emailed alice@example.com.") print(result.text) asyncio.run(main()) ``` The output should be: ```text <> emailed <>. ``` In production, replace `ExactMatchDetector` with an NER or LLM detector, see the [detectors reference](../reference/detectors.md). `ExactMatchDetector` is used here to keep the example reproducible without a model. ## Handle a long text An NER detector has a bounded context window, and a long document can exceed it. `ChunkedDetector` wraps any detector, splits the text into overlapping chunks, detects on each, and remaps the offsets back onto the original text. ```python from piighost.components.detector import ChunkedDetector, RegexDetector from piighost.text import RecursiveCharacterTextSplitter regex_detector = RegexDetector.from_catalog("catalog:piighost/generic") splitter = RecursiveCharacterTextSplitter(chunk_size=40, chunk_overlap=10) detector = ChunkedDetector(regex_detector, splitter=splitter) pipeline = AnonymizationPipeline(detector) async def main() -> None: text = ( "Filler text here. Reach alice@example.com now. " "More filler padding words. Then bob@example.org later." ) result = await pipeline.anonymize(text) print(result.text) asyncio.run(main()) ``` The output should be: ```text Filler text here. Reach <> now. More filler padding words. Then <> later. ``` Leave `splitter=None` for a default `RecursiveCharacterTextSplitter` tuned for real documents. The reduced `chunk_size` above only forces several chunks in a short example. ## Load groups from a config file If you drive the pipeline from a config file rather than from code, a regex detector accepts a `catalogs` key. ```toml [detector] type = "regex" catalogs = ["catalog:piighost/generic", "catalog:piighost/fr"] ``` The detector merges the groups first, then the inline `patterns`. So on a shared label, an inline pattern wins over the one from a group. See the [TOML configuration](../configuration/toml.md). ## See also - [De-identify and restore a text](basic.md) for the full round-trip. - [Detectors reference](../reference/detectors.md) for the label catalog. - [Extending piighost](../extending.md) to write your own detectors.