Pre-built detectors
piighost pulls ready-to-use regex pattern catalogs for structured PII (email, IP, IBAN, phone) from the piighost catalog. This guide shows how to load them, merge them, and combine several detectors, with the piighost core alone.
Four catalog groups cover the common formats. Each one is a set of label to pattern entries.
catalog:piighost/generic: email, URL, IPv4, credit card, country-agnosticcatalog:piighost/us: phone, ZIP, ITIN, SSN, prefixedUS_catalog:piighost/eu: pan-European ISO 13616 IBANcatalog:piighost/fr: phone, IBAN, NIR, SIRET, SIREN, prefixedFR_
A reference without a suffix follows the latest version of the group, fetched each time a detector is built. To freeze a version, add its commit after a colon, as in catalog:piighost/generic:fab51b33. The group is then fetched once, and read from the on-disk cache afterwards, offline included. Secrets such as API keys are in the catalog groups piighost/secrets and piighost/secrets-extended. These groups are pulled the same way, for example with catalogs = ["catalog:piighost/secrets"] in a config.
For the label details, see the detectors reference.
Use a single group
Build a RegexDetector from the group with from_catalog, then assemble the pipeline.
import asyncio
from piighost.components.detector import RegexDetector
from piighost.pipeline import AnonymizationPipeline
detector = RegexDetector.from_catalog("catalog:piighost/generic")
pipeline = AnonymizationPipeline(detector)
async def main() -> None:
result = await pipeline.anonymize("Email alice@example.com, server 192.168.1.42.")
print(result.text)
asyncio.run(main())The output should be:
Email <<EMAIL:1>>, server <<IPV4:1>>.Merge generic and regional groups
If you want to cover both generic PII and a region's PII, pull each group with pull and merge the dictionaries you get. pull returns a label to pattern dictionary. When two dictionaries share a label, the entry from the right-hand dictionary wins.
from piighost.catalog import pull
patterns = {**pull("catalog:piighost/generic"), **pull("catalog:piighost/fr")}
detector = RegexDetector(patterns)
pipeline = AnonymizationPipeline(detector)
async def main() -> None:
result = await pipeline.anonymize(
"IBAN FR7630006000011234567890189, email marie@exemple.fr, tel 06 12 34 56 78."
)
print(result.text)
asyncio.run(main())The output should be:
IBAN <<FR_IBAN:1>>, email <<EMAIL:1>>, tel <<FR_PHONE:1>>.To keep only some labels, build a hand-picked dictionary.
generic = pull("catalog:piighost/generic")
french = pull("catalog:piighost/fr")
patterns = {
"EMAIL": generic["EMAIL"],
"FR_IBAN": french["FR_IBAN"],
}
detector = RegexDetector(patterns)Combine several detectors
CompositeDetector runs several detectors over the same text and concatenates their detections. Overlaps are arbitrated by the pipeline's resolution stage. This is how you pair a regex detector with one that recognizes names.
from piighost.components.detector import (
CompositeDetector,
ExactMatchDetector,
RegexDetector,
)
exact_detector = ExactMatchDetector({"Patrick": "PERSON"})
regex_detector = RegexDetector.from_catalog("catalog:piighost/generic")
detector = CompositeDetector([exact_detector, regex_detector])
pipeline = AnonymizationPipeline(detector)
async def main() -> None:
result = await pipeline.anonymize("Patrick emailed alice@example.com.")
print(result.text)
asyncio.run(main())The output should be:
<<PERSON:1>> emailed <<EMAIL:1>>.In production, replace ExactMatchDetector with an NER or LLM detector, see the detectors reference. ExactMatchDetector is used here to keep the example reproducible without a model.
Handle a long text
An NER detector has a bounded context window, and a long document can exceed it. ChunkedDetector wraps any detector, splits the text into overlapping chunks, detects on each, and remaps the offsets back onto the original text.
from piighost.components.detector import ChunkedDetector, RegexDetector
from piighost.text import RecursiveCharacterTextSplitter
regex_detector = RegexDetector.from_catalog("catalog:piighost/generic")
splitter = RecursiveCharacterTextSplitter(chunk_size=40, chunk_overlap=10)
detector = ChunkedDetector(regex_detector, splitter=splitter)
pipeline = AnonymizationPipeline(detector)
async def main() -> None:
text = (
"Filler text here. Reach alice@example.com now. "
"More filler padding words. Then bob@example.org later."
)
result = await pipeline.anonymize(text)
print(result.text)
asyncio.run(main())The output should be:
Filler text here. Reach <<EMAIL:1>> now. More filler padding words. Then <<EMAIL:2>> later.Leave splitter=None for a default RecursiveCharacterTextSplitter tuned for real documents. The reduced chunk_size above only forces several chunks in a short example.
Load groups from a config file
If you drive the pipeline from a config file rather than from code, a regex detector accepts a catalogs key.
[detector]
type = "regex"
catalogs = ["catalog:piighost/generic", "catalog:piighost/fr"]The detector merges the groups first, then the inline patterns. So on a shared label, an inline pattern wins over the one from a group. See the TOML configuration.
See also
- De-identify and restore a text for the full round-trip.
- Detectors reference for the label catalog.
- Extending piighost to write your own detectors.