Skip to content

Configuration file

You will describe a whole pipeline in a TOML file. The file starts at three lines and grows into a conversational pipeline, which keeps a token stable across the turns of a conversation. Each step changes one thing in the file, then you check the file and run it to see what changed.

1. Set up the check loop

Two commands drive every step below. Start with a pipeline.toml that is wrong on purpose, with pattern where the schema expects patterns.

[detector]
type = "regex"
pattern = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }

Validate it.

piighost validate pipeline.toml

The output should be:

invalid configuration in pipeline.toml: 1 validation error for PipelineConfig
detector.regex.pattern
  Extra inputs are not permitted [type=extra_forbidden, input_value={'EMAIL': '[a-z0-9._%+-]+...a-z0-9.-]+\\.[a-z]{2,}'}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/extra_forbidden

The command names the faulty section and key, and exits 1. That exit code also makes it a CI gate. Run it after every edit below. It builds no component, so it loads no model.

Dump the schema once and point your editor at it for completion on the section and key names.

piighost schema > schema.json

Both commands are documented in the command-line interface.

2. Build a pipeline from three lines

Fix the key, patterns with an s. The file now carries one section, and that is enough to build a pipeline.

[detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }
piighost validate pipeline.toml

The output should be:

OK: pipeline.toml

Write run.py next to it. It loads the file and de-identifies the text you pass on the command line. Every later step reuses run.py unchanged.

import asyncio
import sys

from piighost.config import load_pipeline


async def main() -> None:
    pipeline = load_pipeline("pipeline.toml")
    result = await pipeline.anonymize(sys.argv[1])
    print(result.text)


asyncio.run(main())
python run.py "Write to alice@corp.com from 10.0.0.7."

The output should be:

Write to <<EMAIL:1>> from 10.0.0.7.

The file declares no anonymizer, yet the token names the label and numbers it. It declares no linker, yet both occurrences of one address would share that token. The anonymizer and the linker each fall back to their default, and so does overlap resolution. This file is on disk as examples/config/detector_only.toml.

3. Pick the token

Ask for a plain redaction instead of the numbered token, with an [anonymizer.placeholder] section.

[detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }

[anonymizer.placeholder]
type = "redact"
python run.py "Write to alice@corp.com from 10.0.0.7."

The output should be:

Write to <<REDACT>> from 10.0.0.7.

The address is gone and its label with it. examples/config/minimal.toml carries this file, with the default linker written out. examples/config/minimal.json carries the same file in JSON. The file suffix picks the parser. The configuration reference lists every token style.

4. Pull a group from the catalog

Your pattern covers email only, so the IP address in the sample text went through in clear. Replace the inline pattern with the generic group of the catalog, which carries email, URL, IPv4 and credit card. Without a suffix, the reference follows the latest version of the group, fetched from the catalog every time the pipeline is built. To freeze the group, pin it to a commit, as in catalog:piighost/generic:fab51b33. It is then fetched once, and read from the on-disk cache afterwards. The file no longer has an [anonymizer.placeholder] section, so the default numbered token comes back and tells the four labels apart.

[detector]
type = "regex"
catalogs = ["catalog:piighost/generic"]
python run.py "Write to alice@corp.com and prénom@corp.com from 10.0.0.7."

The output should be:

Write to <<EMAIL:1>> and <<EMAIL:2>> from <<IPV4:1>>.

The IP address is covered now, and the accented address with it. A format of your own, an order number such as CMD-2024-0042, is in no group. Declare it inline, next to the group.

[detector]
type = "regex"
catalogs = ["catalog:piighost/generic"]
patterns = { ORDER = 'CMD-\d{4}-\d{4}' }
python run.py "Order CMD-2024-0042 for alice@corp.com, from 10.0.0.7."

The output should be:

Order <<ORDER:1>> for <<EMAIL:1>>, from <<IPV4:1>>.

The order number is a token now. A label declared in both takes your pattern, because the groups merge first and your inline patterns after them.

5. Run two detectors at once

The generic group matches formats, and a first name has no format. Declare the names you already know in a second detector, and let a composite detector run both and merge what they return.

[detector]
type = "composite"

[[detector.detectors]]
type = "regex"
catalogs = ["catalog:piighost/generic"]
patterns = { ORDER = 'CMD-\d{4}-\d{4}' }

[[detector.detectors]]
type = "exact"
values = { Patrick = "PERSON", Patrik = "PERSON" }
python run.py "Patrick writes to alice@corp.com. Patrik answers from 10.0.0.7."

The output should be:

<<PERSON:1>> writes to <<EMAIL:1>>. <<PERSON:2>> answers from <<IPV4:1>>.

The names and the formats are caught in one pass. One person spelled two ways still gets two tokens, <<PERSON:1>> and <<PERSON:2>>. The next step settles that duplicate.

6. Merge the near-duplicate entities

Patrick and Patrik are the same person, and a model reading two tokens follows two people. Install the fuzzy extra.

uv add "piighost[config,fuzzy]"

Append an [entity_resolver] section to the file. That section clusters the entities whose values are close enough to each other.

[entity_resolver]
type = "fuzzy"
threshold = 0.85
python run.py "Patrick writes to alice@corp.com. Patrik answers from 10.0.0.7."

The output should be:

<<PERSON:1>> writes to <<EMAIL:1>>. <<PERSON:1>> answers from <<IPV4:1>>.

Both spellings share <<PERSON:1>>. Drop the section and the stage is gone again, as it is for every optional stage.

7. Keep the tokens across a conversation

Each run of run.py restarts the numbering, since the pipeline keeps nothing from one call to the next. Append a [memory] section. That section gives the pipeline a per-thread store, and changes the loader you call.

[memory]
type = "in_memory"
piighost validate pipeline.toml

The output should be:

OK: pipeline.toml

The file is valid, and run.py now refuses it.

python run.py "Patrick writes to alice@corp.com."

The traceback ends on:

piighost.exceptions.ConfigError: this configuration declares a memory; use load_thread_pipeline

A file carrying a memory describes a thread pipeline, so it takes load_thread_pipeline. Write thread.py, which sends two messages on the thread "thread-42".

import asyncio

from piighost.config import load_thread_pipeline


async def main() -> None:
    pipeline = load_thread_pipeline("pipeline.toml")
    first = await pipeline.anonymize(
        "Patrick writes to alice@corp.com.", thread_id="thread-42"
    )
    print(first.text)
    second = await pipeline.anonymize(
        "Patrik answers from 10.0.0.7.", thread_id="thread-42"
    )
    print(second.text)


asyncio.run(main())
python thread.py

The output should be:

<<PERSON:1>> writes to <<EMAIL:1>>.
<<PERSON:1>> answers from <<IPV4:1>>.

The second message reuses the <<PERSON:1>> assigned by the first. Each loader refuses the other's files. load_thread_pipeline on a file without a memory therefore raises this configuration declares no memory; use load_pipeline.

See also

  • Configuration reference for every section, every type and every key.
  • Deployment for a memory shared between workers, Redis or a SQL database, with the stored values encrypted at rest. The two files are examples/config/thread_redis.toml and examples/config/thread_sqlalchemy.toml.
  • Deny and allow lists for the deny list and the allow list.