# Configuration reference
Module: `piighost.config`
A configuration file describes a whole pipeline declaratively. `piighost` reads it as TOML or JSON, depending on the file suffix. It then validates it with Pydantic and builds the pipeline the file describes. This page documents every section and every component `type`.
```python
from piighost.config import load_config, load_pipeline, load_thread_pipeline
```
The `config` extra is required (`pip install "piighost[config]"`). It pulls in `pydantic-settings`. Unknown keys are rejected, so a typo fails validation rather than being ignored. A component `type` can need an extra of its own, named in the Extra column of the table that documents it.
---
## Entry points
| Function | Returns | Builds | Memory |
|----------|---------|--------|--------|
| `load_config(path)` | `PipelineConfig` | nothing, validates only | any |
| `load_pipeline(path)` | `AnonymizationPipeline` | a stateless pipeline | rejects a `[memory]` section |
| `load_thread_pipeline(path)` | `ThreadAnonymizationPipeline` | a thread pipeline | requires a `[memory]` section |
`load_config` parses and validates a file into a `PipelineConfig` without building any component, so no model loads. `load_pipeline` builds a stateless `AnonymizationPipeline` and raises `ConfigError` if the file declares a `[memory]` section, since a memory describes a thread pipeline. `load_thread_pipeline` builds a `ThreadAnonymizationPipeline` and raises `ConfigError` if the file declares no `[memory]` section.
```python
from piighost.config import load_pipeline, load_thread_pipeline
stateless = load_pipeline("pipeline.toml") # no [memory]
thread = load_thread_pipeline("thread.toml") # has [memory]
```
---
## File format
The suffix picks the parser. A `.json` suffix is read as JSON, whatever its case. Any other suffix is read as TOML. The two formats carry the same schema. A section is a TOML table or a JSON object.
```toml
[detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }
[linker]
type = "exact"
[anonymizer.placeholder]
type = "redact"
```
```json
{
"detector": { "type": "regex", "patterns": { "EMAIL": "[a-z0-9._%+-]+@[a-z0-9.-]+\\.[a-z]{2,}" } },
"linker": { "type": "exact" },
"anonymizer": { "placeholder": { "type": "redact" } }
}
```
---
## Environment overrides
Every top-level key accepts an override from an environment variable prefixed `PIIGHOST_`, whether it holds a scalar or a whole section. `PIIGHOST_NAME` overrides the `name` scalar. `PIIGHOST_DETECTOR` overrides the `[detector]` section with a JSON object. If it is not valid JSON, the variable is rejected as a validation error. Overrides layer above the file key by key. An environment value wins over the file value, and the keys it leaves out keep the file value.
```bash
export PIIGHOST_NAME="local-en"
export PIIGHOST_DETECTOR='{"type": "exact", "values": {"Patrick": "PERSON"}}'
```
No nested delimiter is configured. A variable such as `PIIGHOST_DETECTOR__TYPE` therefore names no field. It is ignored without an error rather than reaching the `type` key. A section is overridden by its JSON object only.
Secrets are never read from the file. Each is read from its own environment variable at build time, and a missing one raises `ConfigError` from `build()`.
| Secret | Variable | Format | Used by |
|--------|----------|--------|---------|
| Hash pepper | `PIIGHOST_HASH_PEPPER` | any non-empty string | `[memory.hasher]` |
| Cipher key | `PIIGHOST_CIPHER_KEY` | base64 of 16, 24, or 32 bytes | `[memory.cipher]` |
| Moderation key | `MISTRAL_API_KEY` | Mistral API key | `[guard]` type `moderation` |
| Database URL | the `url_env` value, `PIIGHOST_DATABASE_URL` by default | an async SQLAlchemy URL | `[memory]` type `sqlalchemy` |
---
## Sections
The top-level keys of a `PipelineConfig`.
| Section | Required | Meaning |
|---------|----------|---------|
| `name` | no | An optional pipeline name, a top-level scalar overridable by `PIIGHOST_NAME` |
| `token_memo_ttl` | no | The number of seconds a thread's memoized token map is kept. A top-level scalar, which needs a `[memory]` |
| `[detector]` | yes | The detect stage |
| `[linker]` | no | The entity linker, defaults to `ExactEntityLinker` |
| `[anonymizer]` | no | The render stage, defaults to an `Anonymizer` with a label-counter factory |
| `[overlap_resolver]` | no | Resolves overlapping detections, defaults to `ConfidenceOverlapResolver` |
| `[expander]` | no | Re-finds missed occurrences of a detected value |
| `[entity_resolver]` | no | Clusters entities that refer to the same thing |
| `[guard]` | no | Re-checks the output for residual confidential data |
| `[override]` | no | Forces or vetoes detections via a deny list and an allow list |
| `[observation_redactor]` | no | A placeholder factory redacting trace payloads |
| `[memory]` | no | The conversation memory, its presence makes a thread pipeline |
---
## `[detector]`
Discriminated on `type`. Required.
### `type = "regex"`
Matches confidential data by one regex per label, pulled from inline `patterns`, catalog groups in `catalogs`, or both. The groups merge first, then inline patterns. An inline pattern therefore overrides a catalog pattern on the same label. At least one inline pattern or one catalog is required. At load time, each pattern is validated as a compilable regex. It is then compiled under `re.ASCII`, so `\d` matches `0-9` and `\w` stops at the first non-ASCII character. A pattern written with `\w` therefore matches `prénom@corp.com`{ .pii } from `nom` onwards. To take every letter in, scope the Unicode flag to the class, `(?u:\w)`, or name a range. The `EMAIL` pattern of `catalog:piighost/generic` names the Latin range `À-ɏ` this way. See [Limitations](../limitations.md) for what each choice misses.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `patterns` | `dict[str, str]` | `{}` | Inline label-to-regex mapping |
| `catalogs` | `list[str]` | `[]` | Catalog references, `catalog:namespace/name` with an optional `:selector`. A 1.x `hub:` reference is still accepted. Any other entry fails validation |
```toml
[detector]
type = "regex"
catalogs = ["catalog:piighost/generic", "catalog:piighost/fr"]
patterns = { EMPLOYEE_ID = 'EMP-[0-9]{4}' }
```
A group is fetched from the catalog when the config is built, not when it is parsed. A reference pinned to a commit is fetched once, then read from the on-disk cache. `PIIGHOST_CATALOG_URL` points at a private registry, and `PIIGHOST_HUB_URL` from 1.x is still read when it is unset. The names `generic`, `us`, `eu` and `fr` are refused. See [Catalog groups](../reference/detectors.md#catalog-groups) for the groups that replace them.
### `type = "composite"`
Runs child detectors together and merges their detections.
| Key | Type | Meaning |
|-----|------|---------|
| `detectors` | `list[detector]` | The child detector configs, at least one, as `[[detector.detectors]]` |
```toml
[detector]
type = "composite"
[[detector.detectors]]
type = "regex"
catalogs = ["catalog:piighost/generic"]
[[detector.detectors]]
type = "exact"
values = { Patrick = "PERSON" }
```
### `type = "exact"`
Finds occurrences of literal values, each mapped to a label.
| Key | Type | Meaning |
|-----|------|---------|
| `values` | `dict[str, str]` | Literal value to label mapping, at least one |
```toml
[detector]
type = "exact"
values = { Patrick = "PERSON", Lyon = "LOCATION" }
```
### `type = "chunked"`
Wraps a detector with a splitter that cuts long text into overlapping chunks.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `detector` | `detector` | | The detector run on each chunk, as `[detector.detector]` |
| `chunk_size` | `int` | `1000` | Maximum chunk size, greater than 0 |
| `chunk_overlap` | `int` | `100` | Overlap between chunks, below `chunk_size` |
```toml
[detector]
type = "chunked"
chunk_size = 2000
chunk_overlap = 200
[detector.detector]
type = "spacy"
model = "en_core_web_sm"
```
### Model-backed detectors
Each needs its own extra, and every one but `presidio` needs a model. `labels` accepts a list or an `{emitted: internal}` map. `max_concurrency` caps concurrent inferences. `None` leaves them unbounded.
| `type` | Extra | Keys |
|--------|-------|------|
| `gliner2` | `gliner2` | `model` (required), `labels` (required), `threshold` (default `0.5`), `max_concurrency`, `max_chars` |
| `spacy` | `spacy` | `model` (required), `labels`, `max_concurrency` |
| `transformers` | `transformers` | `model` (required), `labels`, `threshold` (default `0.0`), `aggregation_strategy` (default `simple`), `max_concurrency`, `max_chars` |
| `presidio` | `presidio` | `labels`, `language` (default `en`), `threshold` (default `0.0`) |
| `llm` | `llm` | `model` (required), `labels` (required), `prompt`, `provider` |
```toml
[detector]
type = "gliner2"
model = "fastino/gliner2-multi-v1"
labels = ["PERSON", "LOCATION"]
threshold = 0.5
max_chars = 2000
```
The `gliner2` and `transformers` detectors take `max_chars`, the longest text one inference sees. A longer text is split into overlapping chunks. Each chunk is scanned separately, then the spans are mapped back. Without this key, the whole text goes to the model in one pass. A model with a short window then truncates the text, and a long document can exhaust memory.
The `transformers` detector passes `aggregation_strategy` to its token-classification pipeline, which groups sub-word tokens into whole entities.
The `presidio` detector takes no `model` key, since the config path builds Presidio's default English `AnalyzerEngine` with its default recognizers. Another language, a custom recognizer, or a custom NLP engine is the programmatic path, constructing the engine and passing it to `PresidioDetector`.
The `llm` detector reads its provider credential from the provider's own environment variable, never from the file.
---
## `[linker]`
Optional. Defaults to `ExactEntityLinker`. One linker exists, so `type` names it rather than discriminating a union.
| `type` | Meaning |
|--------|---------|
| `exact` | Groups detections by value, the same words whatever their spaces and case |
```toml
[linker]
type = "exact"
```
---
## `[anonymizer]`
Optional. Defaults to an `Anonymizer` with a label-counter factory. When present it carries one `[anonymizer.placeholder]` table selecting the placeholder factory, discriminated on `type`.
| `type` | Token | Keys |
|--------|-------|------|
| `redact` | `<
>`{ .placeholder } | |
| `label` | `<>`{ .placeholder } | |
| `label_counter` | `<>`{ .placeholder } | |
| `label_hash` | `<>`{ .placeholder } | `hash_length` (default `8`, at least 1) |
| `mask` | `P***`{ .placeholder } | `visible` (default `1`, 0 or more), `mask_char` (default `*`, exactly one character) |
```toml
[anonymizer.placeholder]
type = "label_counter"
```
The middleware needs a delimited factory, meaning `redact`, `label`, `label_counter`, or `label_hash`. The `mask` factory produces `P***`{ .placeholder }, which keeps no delimiters and has no recognizer.
---
## `[overlap_resolver]`
Optional in the file, but the stage runs either way. Omitting the section builds a `ConfidenceOverlapResolver`. There is no supported way to disable the stage, since the render stage assumes disjoint spans.
| `type` | Meaning |
|--------|---------|
| `confidence` | Keeps the highest-confidence detection when two overlap |
| `merge` | Keeps the union of overlapping detections, with the label of the most confident. At equal confidence, the label of the widest |
```toml
[overlap_resolver]
type = "merge"
```
`merge` hides every character a detector flagged. With `confidence`, a regex at confidence 1.0 that found `Wirth`{ .pii } beats a model that found `Loni M. Wirth`{ .pii }. `Loni M.`{ .pii } is then sent in clear. Choose `merge` when a leak costs more than a masked neighbour word, as in a document with rules and a model together.
---
## `[expander]`
Optional, and disabled when omitted. One expander exists, so `type` names it rather than discriminating a union.
| `type` | Keys | Meaning |
|--------|------|---------|
| `word_boundary` | `case_sensitive` (default `false`) | Re-finds a detected value's other whole-word occurrences, whatever spaces separate its words |
```toml
[expander]
type = "word_boundary"
case_sensitive = false
```
---
## `[entity_resolver]`
Optional. Discriminated on `type`.
| `type` | Extra | Keys | Meaning |
|--------|-------|------|---------|
| `merge` | | | Unions entities that share detections |
| `separate` | | | Keeps every entity distinct |
| `fuzzy` | `fuzzy` | `threshold` (default `0.85`) | Clusters entities at or above a Jaro-Winkler similarity |
```toml
[entity_resolver]
type = "fuzzy"
threshold = 0.85
```
---
## `[guard]`
Optional. Discriminated on `type`. Re-checks the de-identified output for residual confidential data and refuses it when any remains.
| `type` | Extra | Re-checks with |
|--------|-------|----------------|
| `detector` | | A detector re-run on the output |
| `llm` | `llm` | A chat model prompted to find residual PII |
| `moderation` | `mistral` | A Mistral moderation model scoring the output |
| `gliner2` | `gliner2` | A local GLiNER2 guardrail model classifying the output |
### `type = "detector"`
Re-runs a detector on the output. Carries a nested `[guard.detector]` config.
```toml
[guard]
type = "detector"
[guard.detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }
```
### `type = "llm"`
Prompts a chat model to find residual PII.
| Key | Type | Meaning |
|-----|------|---------|
| `model` | `str` | The chat model identifier (required) |
| `labels` | `list` or `dict` | The labels to look for (required) |
| `prompt` | `str` | A prompt overriding the default, or omitted |
| `provider` | `str` | The provider, or omitted to infer from the model |
### `type = "moderation"`
Scores the output with a Mistral moderation model. The credential is read from `MISTRAL_API_KEY` at build time, and `build()` raises `ConfigError` when it is unset.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `model` | `str` | `mistral-moderation-latest` | The moderation model |
| `threshold` | `float` | `0.5` | The category score at or above which the text is flagged |
### `type = "gliner2"`
Classifies the output with a GLiNER2 guardrail model that runs in the process, so it needs no credential. The checkpoint is downloaded on the first build, then read from the Hugging Face cache.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `model` | `str` | `fastino/GLiNER2-Guardrails-PII-Multi` | The GLiNER2 checkpoint the guard classifies with |
| `task` | `str` | `response_safety` | The classification task read from the model's answer |
| `labels` | `list` | `["safe", "unsafe"]` | The answers the task chooses between, two at least. The unsafe one comes last |
| `threshold` | `float` | `0.5` | The confidence at or above which an unsafe answer flags the output |
```toml
[guard]
type = "gliner2"
threshold = 0.5
```
---
## `[override]`
Optional. Forces detections through a deny list, whose values are always masked, and vetoes them through an allow list, whose values are always left in clear. Each list is a detector config, `[override.deny_list]` and `[override.allow_list]`, and both are optional. The 1.x keys, `whitelist`, `blacklist` and their strategies, are refused at load time with the key that replaces them, see [Upgrading to 2.0](../community/upgrading.md#the-override-lists-are-renamed).
| Key | Values | Default | Meaning |
|-----|--------|---------|---------|
| `[override.deny_list]` | detector | | A detector whose hits are always masked, forced into the set |
| `[override.allow_list]` | detector | | A detector whose hits are always left in clear, invalidating the detections they match |
| `allow_list_strategy` | `exact`, `value`, `overlap` | `value` | How an allow list hit invalidates a detection. `value` needs the same value, whatever its spaces and case. `exact` needs the same span and label. `overlap` invalidates any overlapping span |
| `deny_list_strategy` | `respect_provenance`, `force` | `respect_provenance` | Whether a deny list hit leaves an assistant-introduced value in clear, or tokenizes it regardless |
| `conflict_strategy` | `deny_list_wins`, `allow_list_wins`, `raise` | `deny_list_wins` | Who wins when the two lists contradict. `raise` refuses the collision with `ConflictingOverrideError` |
```toml
[override]
allow_list_strategy = "value"
[override.deny_list]
type = "regex"
patterns = { CODENAME = 'ACME-[A-Z]+' }
[override.allow_list]
type = "exact"
values = { "public@corp.com" = "EMAIL" }
```
---
## `[observation_redactor]`
Optional. A placeholder factory config, with the same `type` values as `[anonymizer.placeholder]`. It redacts the payloads sent to a tracing backend, so a trace holds tokens, not raw values.
```toml
[observation_redactor]
type = "label"
```
Without this section, the clear text and the detection values are traced. A live tracer then emits a `PIIGhostSecurityWarning`. The pipeline's `trace_clear_text` flag silences that warning, but it has no key in a configuration file. A file-built pipeline therefore cannot acknowledge clear-text tracing. Passing `trace_clear_text=True` to the pipeline is the programmatic path.
---
## `[memory]`
Optional. Its presence makes the pipeline a `ThreadAnonymizationPipeline` keeping per-thread state. Discriminated on `type`.
The `token_memo_ttl` scalar goes with this section, but it sits at the top level, because it bounds the pipeline's own memoized token map and not the store. Setting it without a `[memory]` raises, because a stateless pipeline memoizes nothing. [Multi-instance deployment](../multi-instance.md) explains why it matters on a multi-worker deployment.
| `type` | Extra | Store |
|--------|-------|-------|
| `in_memory` | | Process-local, lost on restart |
| `redis` | `redis` | Persistent, shared across workers |
| `sqlalchemy` | `sqlalchemy` | Durable, in a SQL database |
### `type = "in_memory"`
A process-local store, lost on restart and not shared across workers.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `max_threads` | `int` | `10000` | Cap on kept threads, LRU eviction beyond it (at least 1) |
| `ttl` | `float` | `86400` | Idle time, in seconds, after which a thread expires (greater than 0). It is only dropped on the next access |
```toml
[memory]
type = "in_memory"
```
### `type = "redis"`
A persistent, multi-worker store. Optionally, it keys each stored message with a hasher and encrypts each stored value with a cipher.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `url` | `str` | | The Redis connection URL (required) |
| `namespace` | `str` | `piighost` | The key prefix isolating this library's keys |
| `ttl` | `int` | `None` | Seconds a stored message lives, or omitted to keep until eviction |
| `[memory.hasher]` | hasher | | Optional (both or neither). The hasher keying each message |
| `[memory.cipher]` | cipher | | Optional (both or neither). The cipher encrypting each value |
Configure both `[memory.hasher]` and `[memory.cipher]`, or neither. With neither, the backend stores the mapping in clear and warns. With exactly one, `build()` raises `ConfigError`.
The hasher, `[memory.hasher]`, is discriminated on `type`.
| `type` | Extra | Keys | Meaning |
|--------|-------|------|---------|
| `sha256` | | | HMAC-SHA256, a fast keyed digest |
| `argon2` | `argon2` | `time_cost` (default `2`), `memory_cost` (default `19456`), `parallelism` (default `1`), `hash_length` (default `32`) | Argon2id, a slow memory-hard digest |
The cipher, `[memory.cipher]`, has one type.
| `type` | Extra | Meaning |
|--------|-------|---------|
| `aesgcm` | `crypto` | AES-GCM authenticated encryption of stored values |
The hasher reads its pepper from `PIIGHOST_HASH_PEPPER` and the cipher reads its base64 key from `PIIGHOST_CIPHER_KEY`, both at build time. A missing or malformed value raises `ConfigError`.
```toml
[memory]
type = "redis"
url = "redis://localhost:6379/0"
namespace = "piighost"
ttl = 3600
[memory.hasher]
type = "argon2"
[memory.cipher]
type = "aesgcm"
```
### `type = "sqlalchemy"`
A durable, multi-worker store backed by any SQLAlchemy-supported database (SQLite, PostgreSQL, ...). It reads the database URL from an environment variable rather than the config file, so the URL and its password stay out of version control. An optional hasher and cipher protect the stored values exactly as they do for Redis.
| Key | Type | Default | Meaning |
|-----|------|---------|---------|
| `url_env` | `str` | `PIIGHOST_DATABASE_URL` | The environment variable holding the async database URL |
| `table_name` | `str` | `piighost_conversation_messages` | The table storing per-thread messages |
| `[memory.hasher]` | hasher | | Optional (both or neither). The hasher keying each message |
| `[memory.cipher]` | cipher | | Optional (both or neither). The cipher encrypting each value |
Configure both `[memory.hasher]` and `[memory.cipher]`, or neither, exactly as for Redis. With neither, the backend stores the mapping in clear and warns. With exactly one, `build()` raises `ConfigError`.
The URL must use an async driver, for example `postgresql+asyncpg://...` or `sqlite+aiosqlite://...`. A missing environment variable raises `ConfigError` at build time. Call `await memory.create_schema()` once at startup to create the table.
```toml
[memory]
type = "sqlalchemy"
url_env = "PIIGHOST_DATABASE_URL"
table_name = "piighost_conversation_messages"
[memory.hasher]
type = "argon2"
[memory.cipher]
type = "aesgcm"
```
---
## Full example
The keys of `examples/config/pipeline.toml`, a stateless pipeline pulling a catalog group, adding one inline pattern, and enabling several optional stages. The file itself carries the same keys with a comment on each stage.
```toml
[detector]
type = "regex"
catalogs = ["catalog:piighost/generic"]
patterns = { EMPLOYEE_ID = 'EMP-[0-9]{4}' }
[overlap_resolver]
type = "confidence"
[expander]
type = "word_boundary"
[entity_resolver]
type = "fuzzy"
threshold = 0.85
[linker]
type = "exact"
[anonymizer.placeholder]
type = "label_counter"
[override.deny_list]
type = "regex"
patterns = { CODENAME = 'ACME-[A-Z]+' }
[guard]
type = "detector"
[guard.detector]
type = "regex"
patterns = { EMAIL = '[a-z0-9._%+-]+@[a-z0-9.-]+\.[a-z]{2,}' }
[observation_redactor]
type = "label"
```
The same content in JSON, chosen by a `.json` suffix, is equivalent. A table becomes an object, an inline table becomes a nested object, and an array of tables becomes an array of objects.
---
## Errors
| Error | Raised when |
|-------|-------------|
| `ConfigFileError` | The file is missing, unreadable, or invalid TOML or JSON |
| `ConfigValidationError` | The parsed data fails schema validation |
| `ConfigError` | A secret is missing at build time, or the wrong entry point is used for the memory declared |
`ConfigFileError` and `ConfigValidationError` are subclasses of `ConfigError`, so catching `ConfigError` covers all three. The classes live in `piighost.exceptions`, so a caller can catch them without the `config` extra.
---
## See also
- `examples/config/` in the repository for six runnable files, `detector_only.toml`, `minimal.toml`, `minimal.json`, `pipeline.toml`, `thread_redis.toml` and `thread_sqlalchemy.toml`, all six loaded by `examples/config/run.py`.
- [Command-line interface](../reference/cli.md) for validating a file from the shell.
- [Detectors reference](../reference/detectors.md) for the detector each `type` builds.
- [LangChain middleware reference](../reference/langchain.md) for driving a thread pipeline in an agent.
- [Configure a pipeline by file, catalog and command line](../../../openwiki/en/operations/configuration-and-catalog.md), for the configuration rules `BR-CFG-01` to `BR-CFG-09` and where each lives in the code.