Configuration file
You will describe a whole pipeline in a TOML file. The file starts at three lines and grows into a conversational pipeline, which keeps a token stable across the turns of a conversation. Each step changes one thing in the file, then you check the file and run it to see what changed.
1. Set up the check loop
Two commands drive every step below. Start with a pipeline.toml that is wrong on purpose, with pattern where the schema expects patterns.
[detector]
type = "regex"
pattern = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }Validate it.
piighost validate pipeline.tomlThe output should be:
invalid configuration in pipeline.toml: 1 validation error for PipelineConfig
detector.regex.pattern
Extra inputs are not permitted [type=extra_forbidden, input_value={'EMAIL': '[a-z0-9._%+-]+...a-z0-9.-]+\\.[a-z]{2,}'}, input_type=dict]
For further information visit https://errors.pydantic.dev/2.13/v/extra_forbiddenThe command names the faulty section and key, and exits 1. That exit code also makes it a CI gate. Run it after every edit below. It builds no component, so it loads no model.
Dump the schema once and point your editor at it for completion on the section and key names.
piighost schema > schema.jsonBoth commands are documented in the command-line interface.
2. Build a pipeline from three lines
Fix the key, patterns with an s. The file now carries one section, and that is enough to build a pipeline.
[detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }piighost validate pipeline.tomlThe output should be:
OK: pipeline.tomlWrite run.py next to it. It loads the file and de-identifies the text you pass on the command line. Every later step reuses run.py unchanged.
import asyncio
import sys
from piighost.config import load_pipeline
async def main() -> None:
pipeline = load_pipeline("pipeline.toml")
result = await pipeline.anonymize(sys.argv[1])
print(result.text)
asyncio.run(main())python run.py "Write to alice@corp.com from 10.0.0.7."The output should be:
Write to <<EMAIL:1>> from 10.0.0.7.The file declares no anonymizer, yet the token names the label and numbers it. It declares no linker, yet both occurrences of one address would share that token. The anonymizer and the linker each fall back to their default, and so does overlap resolution. This file is on disk as examples/config/detector_only.toml.
3. Pick the token
Ask for a plain redaction instead of the numbered token, with an [anonymizer.placeholder] section.
[detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }
[anonymizer.placeholder]
type = "redact"python run.py "Write to alice@corp.com from 10.0.0.7."The output should be:
Write to <<REDACT>> from 10.0.0.7.The address is gone and its label with it. examples/config/minimal.toml carries this file, with the default linker written out. examples/config/minimal.json carries the same file in JSON. The file suffix picks the parser. The configuration reference lists every token style.
4. Pull a group from the catalog
Your pattern covers email only, so the IP address in the sample text went through in clear. Replace the inline pattern with the generic group of the catalog, which carries email, URL, IPv4 and credit card. Without a suffix, the reference follows the latest version of the group, fetched from the catalog every time the pipeline is built. To freeze the group, pin it to a commit, as in catalog:piighost/generic:fab51b33. It is then fetched once, and read from the on-disk cache afterwards. The file no longer has an [anonymizer.placeholder] section, so the default numbered token comes back and tells the four labels apart.
[detector]
type = "regex"
catalogs = ["catalog:piighost/generic"]python run.py "Write to alice@corp.com and prénom@corp.com from 10.0.0.7."The output should be:
Write to <<EMAIL:1>> and <<EMAIL:2>> from <<IPV4:1>>.The IP address is covered now, and the accented address with it. A format of your own, an order number such as CMD-2024-0042, is in no group. Declare it inline, next to the group.
[detector]
type = "regex"
catalogs = ["catalog:piighost/generic"]
patterns = { ORDER = 'CMD-\d{4}-\d{4}' }python run.py "Order CMD-2024-0042 for alice@corp.com, from 10.0.0.7."The output should be:
Order <<ORDER:1>> for <<EMAIL:1>>, from <<IPV4:1>>.The order number is a token now. A label declared in both takes your pattern, because the groups merge first and your inline patterns after them.
5. Run two detectors at once
The generic group matches formats, and a first name has no format. Declare the names you already know in a second detector, and let a composite detector run both and merge what they return.
[detector]
type = "composite"
[[detector.detectors]]
type = "regex"
catalogs = ["catalog:piighost/generic"]
patterns = { ORDER = 'CMD-\d{4}-\d{4}' }
[[detector.detectors]]
type = "exact"
values = { Patrick = "PERSON", Patrik = "PERSON" }python run.py "Patrick writes to alice@corp.com. Patrik answers from 10.0.0.7."The output should be:
<<PERSON:1>> writes to <<EMAIL:1>>. <<PERSON:2>> answers from <<IPV4:1>>.The names and the formats are caught in one pass. One person spelled two ways still gets two tokens, <<PERSON:1>> and <<PERSON:2>>. The next step settles that duplicate.
6. Merge the near-duplicate entities
Patrick and Patrik are the same person, and a model reading two tokens follows two people. Install the fuzzy extra.
uv add "piighost[config,fuzzy]"pip install "piighost[config,fuzzy]"Append an [entity_resolver] section to the file. That section clusters the entities whose values are close enough to each other.
[entity_resolver]
type = "fuzzy"
threshold = 0.85python run.py "Patrick writes to alice@corp.com. Patrik answers from 10.0.0.7."The output should be:
<<PERSON:1>> writes to <<EMAIL:1>>. <<PERSON:1>> answers from <<IPV4:1>>.Both spellings share <<PERSON:1>>. Drop the section and the stage is gone again, as it is for every optional stage.
7. Keep the tokens across a conversation
Each run of run.py restarts the numbering, since the pipeline keeps nothing from one call to the next. Append a [memory] section. That section gives the pipeline a per-thread store, and changes the loader you call.
[memory]
type = "in_memory"piighost validate pipeline.tomlThe output should be:
OK: pipeline.tomlThe file is valid, and run.py now refuses it.
python run.py "Patrick writes to alice@corp.com."The traceback ends on:
piighost.exceptions.ConfigError: this configuration declares a memory; use load_thread_pipelineA file carrying a memory describes a thread pipeline, so it takes load_thread_pipeline. Write thread.py, which sends two messages on the thread "thread-42".
import asyncio
from piighost.config import load_thread_pipeline
async def main() -> None:
pipeline = load_thread_pipeline("pipeline.toml")
first = await pipeline.anonymize(
"Patrick writes to alice@corp.com.", thread_id="thread-42"
)
print(first.text)
second = await pipeline.anonymize(
"Patrik answers from 10.0.0.7.", thread_id="thread-42"
)
print(second.text)
asyncio.run(main())python thread.pyThe output should be:
<<PERSON:1>> writes to <<EMAIL:1>>.
<<PERSON:1>> answers from <<IPV4:1>>.The second message reuses the <<PERSON:1>> assigned by the first. Each loader refuses the other's files. load_thread_pipeline on a file without a memory therefore raises this configuration declares no memory; use load_pipeline.
See also
- Configuration reference for every section, every
typeand every key. - Deployment for a memory shared between workers, Redis or a SQL database, with the stored values encrypted at rest. The two files are
examples/config/thread_redis.tomlandexamples/config/thread_sqlalchemy.toml. - Deny and allow lists for the deny list and the allow list.