Skip to content

Placeholder factories

A placeholder is the synthetic token that takes the place of a detected value before the text reaches the LLM. Instead of sending Patrick lives in Paris to the LLM, the pipeline sends <<PERSON:1>> lives in <<LOCATION:1>>. The original values stay in the conversation memory. The LLM never sees them.

A placeholder factory decides what those tokens look like and how much information they carry. Two questions structure the choice.

  1. Is the token unique per entity? Patrick and Marie should not both collapse onto a generic <<PERSON>>, otherwise the LLM cannot tell them apart. A unique token per entity lets the model reason about relations. The question is the manager the same person as Patrick? becomes is <<PERSON:1>> the same as <<PERSON:2>>?, and it gets a clear answer.

  2. Is the token reversible and findable? Does the token denote a single value in the conversation memory, and can it be relocated in a text the pipeline never produced? Restoration needs both properties, whether it runs on the model's reply or on a tool's arguments. If two entities collapse onto the same <<PERSON>>, there is no way to know which original to restore.

Six families of factories sit at different points on that spectrum, and the choice has direct consequences on which ToolCallStrategy you can use safely. See Tool-call strategies for the runtime side.

  • No information (<<REDACT>>): a constant token that reveals nothing to the LLM. Classic redaction. No reasoning is possible on entities. For example, the model cannot tell that the value was a city and decide to call the get_weather tool.

  • Type only (<<PERSON>>, <<EMAIL>>): the type is revealed, not the identity. Multiple persons in the same conversation collapse onto the same <<PERSON>>, so cross-references break.

  • Type + id (opaque) (<<PERSON:1>>, <<PERSON:a1b2c3d4>>): type revealed, stable identity, clearly synthetic token. The LLM can tell that <<PERSON:1>> and <<PERSON:2>> are two different people. Unique, so reversible by string replacement.

  • Id only (<<REDACT:a1b2c3d4>>): a unique hash per entity, without revealing the type. The LLM sees that two distinct entities exist but cannot tell whether they are persons, emails, or cards. Keeps reversibility on the tool side without giving any semantic hint to the model.

  • Partial value (J******* for Jonathan): part of the real content stays visible, here the first letter and the length. The LLM sees the start of the value, not the full value. Riskier on privacy (real fragments) and on reversibility (collisions possible).


Family details

No information, total destruction

The token is a fixed marker, e.g. <<REDACT>>. The LLM learns that something was removed but nothing about its type, count, or relations. The conversation loses every internal reference. An agent trying to act on send the invoice to the client cannot tell whether the client is the one mentioned earlier or someone new.

Useful for archival redaction, useless once an agent has to reason.

  • Built-in: RedactPlaceholderFactory (output <<REDACT>>, delimiters configurable).
  • Preservation tag: PreservesNothing.

Type only, identities collapsed

<<PERSON>>, <<EMAIL>>. The LLM knows that something is a person, an email, a card, and can answer questions that depend on the type alone. But two different persons in the same conversation collapse onto the same token.

The classic failure mode is cross-reference. The question is Patrick the same person as the manager mentioned earlier? becomes is <<PERSON>> the same as <<PERSON>>?, and that question has no answer.

  • Built-in: LabelPlaceholderFactory (output <<PERSON>>).
  • Preservation tag: PreservesLabel.

Type + id (opaque)

<<PERSON:1>>, <<PERSON:a1b2c3d4>>. The string clearly is not a person, an email, or a card number, it is a token. The LLM cannot mistake it for real data, audit logs are easy to scan, and there is zero chance of collision with a real value.

Its delimiters also make it findable. A consumer can then spot a token the model invented.

In return, a strict downstream prompt or tool that requires the argument must look like an email will reject these tokens.

  • Built-in: LabelCounterPlaceholderFactory (<<PERSON:1>>) and LabelHashPlaceholderFactory (<<PERSON:a1b2c3d4>>).
  • Preservation tag: PreservesLabeledIdentityOpaque.

Both number the entities per label, in order. The first person becomes ordinal 1, the second 2, while an email starts its own count at 1.

LabelHashPlaceholderFactory renders that ordinal as a hash. The hash is a sha256 of the string label:ordinal, never of the value. It only gives an opaque look, so that two consecutive entities look unrelated.

Id only, identity without type

<<REDACT:a1b2c3d4>>. The token keeps the synthetic <<...>> shape but does not reveal the label, while carrying a unique hash per entity. The LLM cannot tell whether the entity is a person, an email, or a card, but it can see that <<REDACT:a1b2c3d4>> and <<REDACT:ef98abcd>> are two distinct entities.

It is one of the most protective levels that stays usable on the tool side. The string replacement works, because the hash is unique.

  • Built-in: none for this branch.
  • Preservation tag: PreservesIdentityOnly, meant for a factory you write, a hashed redaction with no label prefix. See Writing your own below.

Type + id (realistic hashed)

A custom factory can produce values that look like the original format but whose content is driven by a hash, e.g. a1b2c3d4@anonymized.local for an email, or Patient_a1b2c3d4 for a name.

The token passes basic format validation (email regex, length, allowed characters), so downstream tools and prompt templates that expect a real-looking value still work. Because the content is a hash, the token is unique and cannot coincidentally match an existing real value.

  • Built-in: none. See Writing your own below for a complete example.
  • Preservation tag: PreservesLabeledIdentityHashed.

Partial value, a fragment leaks

J*******, j***@mail.com, ****4567. The token keeps part of the original value, for example the email domain, the last four digits of a card, the first letter of a name. The LLM can reason on more than the type, the email is on the company domain, the card ends in 4567, the name starts with J. Two trade-offs come with this.

  1. Real fragments of the value reach the LLM. It cannot reconstruct the full value, but j***@mail.com already places the user inside a known mail provider.
  2. Collisions are possible. Two different cards ending in 4567 collapse onto ****4567, two emails sharing the first letter and domain end up identical. The token is mostly unique, with no guarantee.
  • Built-in: MaskPlaceholderFactory, which by default keeps the first character of the value and masks the rest with *, so Jonathan becomes J******* and jean@mail.com becomes j************. The j***@mail.com and ****4567 forms need a factory you write.
  • Preservation tag: PreservesShape.

The middleware refuses it, at type-check time and at runtime. An ambiguous token cannot be restored through string replacement, and a mask has no grammar the middleware can find again.


Preservation tags

Every factory carries a phantom type that summarises the preservation level of its tokens. A phantom type is a generic parameter that exists only at type-check time, it does not affect execution. The type-checker reads this tag to validate a factory against its consumers.

The table below gives an example token and the tag of each family.

FamilyExampleTag
No information<<REDACT>>PreservesNothing
Type only<<PERSON>>PreservesLabel
Type + id (opaque)<<PERSON:1>>, <<PERSON:a1b2c3d4>>PreservesLabeledIdentityOpaque
Id only<<REDACT:a1b2c3d4>>PreservesIdentityOnly
Type + id (realistic hashed)a1b2c3d4@anonymized.local, Patient_a1b2c3d4PreservesLabeledIdentityHashed
Partial valueJ*******, ****4567PreservesShape

Two tables read these families from two angles. The Confidentiality table shows what leaks to the LLM, from the attacker and privacy point of view. The Exploitation table shows what the agent and the system can do with the token, from the point of view of functional capabilities. The same answer can be good in one and problematic in the other, and the two tables make this tension explicit.

Both tables share the same colour code, from best to problematic, explained in the legend under the second table.

Confidentiality (what leaks to the LLM)

FamilyType seen?Values distinguished?Real-value leak?Collision with a real value?
No informationnonononeno
Type onlyyesnononeno
Type + id (opaque)yesyesnoneno
Id onlynoyesnoneno
Type + id (realistic hashed)yesyesnoneno
Partial valueyesyespartialrisk

Exploitation by the LLM and the agent

FamilyReason about the typeTrack cross-referencesReversible at the tool boundaryToken findable
No informationnononoyes
Type onlyyesnonoyes
Type + id (opaque)yesyesyesyes
Id onlynoyesyesyes
Type + id (realistic hashed)yesyesyesno
Partial valueyesmostlyyes (collisions)no
Legend: best acceptable partial problematic

Tags form an inheritance hierarchy that the type-checker exploits through the covariance of AnyPlaceholderFactory[PreservationT_co]. A factory tagged more specifically therefore satisfies a consumer asking for a looser one.

Three independent axes structure the taxonomy:

  • Label: the token reveals the type.
  • Identity: the token is unique per entity.
  • Recognizable: the factory can find its token again in arbitrary text. A delimited token allows this, a realistic one does not.

PreservesLabeledIdentity combines label and identity via multiple inheritance. A <<PERSON:1>> factory is therefore both a PreservesLabel and a PreservesIdentity.

PreservesRecognizableIdentity crosses identity with findability. The middleware accepts only this intersection. A consumer typed against PreservesRecognizableIdentity sorts the tags as follows:

  • Accepts: PreservesIdentityOnly and PreservesLabeledIdentityOpaque.
  • Rejects: PreservesLabel, PreservesShape and PreservesNothing, which lack the uniqueness guarantee, along with PreservesLabeledIdentityHashed, which is not findable.
DiagramDiagram

Preservation tag hierarchy. Each node carries an example token, the abstract nodes are intersections between axes. Each arrow goes from a tag to its parent and reads "is a".

PreservesLabeledIdentity inherits from both PreservesLabel and PreservesIdentity. This inheritance expresses the A is a B but not every B is an A relation. Every PreservesLabeledIdentity is also a PreservesLabel and a PreservesIdentity, but a PreservesLabel is not necessarily a PreservesLabeledIdentity.

PreservesShape extends PreservesLabel, because a masked token implies the label through its format. It does not guarantee uniqueness, so it does not descend from PreservesIdentity.

Each tag is a subclass of str, so a token is a real string that carries its preservation level in its own type.

A factory declares the most specific tag that matches its guarantees.

class LabelCounterPlaceholderFactory(
    BaseCounterPlaceholderFactory
): ...  # PreservesLabeledIdentityOpaque


class LabelHashPlaceholderFactory(
    BaseCounterPlaceholderFactory
): ...  # PreservesLabeledIdentityOpaque


class LabelPlaceholderFactory(AnyPlaceholderFactory[PreservesLabel]): ...


class MaskPlaceholderFactory(AnyPlaceholderFactory[PreservesShape]): ...


class RedactPlaceholderFactory(AnyPlaceholderFactory[PreservesNothing]): ...


# No built-in for the id-only branch nor the realistic hashed one,
# implement your own with PreservesIdentityOnly or PreservesLabeledIdentityHashed.

Built-in factories

FactoryStyleMechanismOutput example
RedactPlaceholderFactoryRedactnone<<REDACT>>
LabelPlaceholderFactoryLabelnone<<PERSON>>
LabelCounterPlaceholderFactory (default)LabelCounter<<PERSON:1>>
LabelHashPlaceholderFactoryLabelHash<<PERSON:a1b2c3d4>>
MaskPlaceholderFactoryMaskpartialJ*******

The tag of each factory is in the family table, above. The naming follows a <Style><Mechanism>PlaceholderFactory schema.

  • Style: what the token preserves. Redact = nothing, Label = type, Mask = partial value.
  • Mechanism: how uniqueness is achieved. Counter = sequential per-label count, Hash = sha256 of label:ordinal rendered as hex. Absent when not relevant.

LabelCounterPlaceholderFactory and LabelHashPlaceholderFactory are the safe defaults, reversible and findable. RedactPlaceholderFactory, LabelPlaceholderFactory and MaskPlaceholderFactory are non-reversible redaction tools. The type checker rejects them under the middleware, and the middleware also refuses the mask at construction. The id-only and realistic-hashed branches have no built-in. You write them with the matching tag.


Which placeholder to pick?

The placeholder factory is the place where the privacy / agent-capability trade-off is made explicit. The right choice depends on the use case. Two scenarios cover most needs.

Case 1, one-off de-identification (archival, compliance)

The goal is to produce a sanitised version of a document, for example redacting a court ruling, scrubbing an HR record before archival, exporting a dataset. No agent, no tools, sometimes not even reversibility.

NeedRecommended familyWhy
Erase every trace, no reversibility neededNo information (<<REDACT>>)The most protective, no semantic leak. The document stays readable but the LLM cannot infer anything. Built-in RedactPlaceholderFactory.
Keep the text readable, a human reader sees <<EMAIL>> rather than <<REDACT>>Type only (<<PERSON>>, <<EMAIL>>)The type aids human reading without leaking the value. Built-in LabelPlaceholderFactory.
Allow server-side restorationType + id (opaque) (<<PERSON:1>>)Reversible, trivial to audit, no collisions. Built-in LabelCounterPlaceholderFactory or LabelHashPlaceholderFactory.
Track who is who without revealing the type (medical, HR)Id only (<<REDACT:a1b2c3d4>>)Distinguishes entities without a semantic hint. Custom factory, no built-in.

Case 2, de-identification for an LLM or an agent with tools

The LLM reasons about the conversation, and tools (CRM, DB, mail) need real values at call time. The middleware restores through string replacement, on the model's reply as on tool arguments. It therefore requires a unique and findable token per entity.

As a direct consequence, only families with preserved identity and a findable grammar are compatible, that is id only and type + id opaque. The no-information, type-only and partial-value families are rejected at type-check time. Realistic hashed preserves identity but is not findable, so it fails the middleware constraint.

NeedRecommended familyWhy
DefaultType + id (opaque) (<<PERSON:1>>, <<PERSON:a1b2c3d4>>)Reversible, findable, opaque, zero collision. The safe default. Built-in LabelCounterPlaceholderFactory (per-thread counter) or LabelHashPlaceholderFactory (hash of the ordinal).
Bias reduction (CV screening, hiring)Id only (<<REDACT:a1b2c3d4>>)The LLM does not see the type, so gender or origin inferable from a first name vanishes. Distinguishes candidates without biasing reasoning. Custom factory.
Sensitive type (medical category, clearance level)Id only (<<REDACT:a1b2c3d4>>)Same reason, the type itself is a PII and must not reach the LLM. Custom factory.

To avoid in an agent under the middleware.

  • LabelPlaceholderFactory and MaskPlaceholderFactory are rejected by the middleware, whatever the ToolCallStrategy, because they do not guarantee uniqueness. The type checker rejects both, and the middleware also refuses the mask at construction. Use them with the bare pipeline, outside the middleware.
  • A realistic-hashed factory (PreservesLabeledIdentityHashed) preserves identity but stays not findable, so the middleware cannot spot a token the model would invent. Reserve it for de-identification outside an agent, or for a flow where a placeholder invented by the model is not a concern.

The preservation tag exists so this choice is visible to the type-checker, not buried in placeholder-format trivia. A factory tagged PreservesShape cannot be plugged into the middleware by accident, the error falls at type-check time, not on the first tool call in production.


Why PIIAnonymizationMiddleware requires a findable identity

The middleware operates on three boundaries, input messages (LLM in), output messages (LLM out), and tool calls. All three rely on the conversation memory, which keeps each message's detections.

Input and output messages. When abefore_model de-identifies a message, the pipeline records its detections in the memory. When the LLM replies, aafter_model restores the reply through string replacement. It looks for every known token of the thread and replaces it with the value of its entity. The model's reply is a new text the pipeline never produced, so there is no other way to restore it.

Tool calls. The LLM produces tool arguments by combining and paraphrasing the tokens it just saw. The middleware restores them the same way, scanning the args for known tokens. The tool response, for its part, goes through the thread's pipeline like a user message, detection included.

On both channels, that replacement is unambiguous only if every entity maps to a unique token. If two entities collapse onto <<PERSON>>, there is no way to know which original to restore.

The middleware also requires a findable grammar, a token shape it can spot in a text. Once every issued token has been replaced, any token still matching the grammar was invented by the model and can be refused (see Tool-call strategies).

The middleware therefore narrows its accepted type to a pipeline whose tokens are PreservesRecognizableIdentity. Through covariance, this type encompasses PreservesIdentityOnly (hashed redact, no label) and PreservesLabeledIdentityOpaque (with label). pyrefly catches a PreservesLabel, PreservesShape, PreservesNothing or PreservesLabeledIdentityHashed factory before the program runs.

PIIAnonymizationMiddleware mirrors part of that constraint at runtime. At construction, it asks the pipeline for a recognizer, the object that knows how to find its own tokens. A delimited factory is its own recognizer. A factory with no grammar, such as a mask, has none, and the middleware then raises UnrecognizableFactoryError. This runtime check catches untyped or remote pipelines that bypassed the type checker. It only checks the grammar, so a delimited factory without identity, such as LabelPlaceholderFactory, passes at runtime. Only the type checker rejects it.

The recognizer's grammar is bounded, not "anything between the delimiters". It reads as follows:

  • Inner form: a label, then an optional colon and identifier, such as <<PERSON>>, <<PERSON:1>> or <<PERSON:a1b2c3d4>>.
  • Label: a letter or underscore, then letters, digits, underscores, spaces, or hyphens, so a multi-word label a detector emits, such as date of birth, still fits.
  • Identifier: after the colon, alphanumeric, an ordinal or a hex digest.

Arbitrary delimited content is not a token. A C++ shift cout << x >> y or a markdown run therefore never trips the invented-token guard.

A streaming reply that opens << without closing it is released rather than buffered indefinitely.

The choice of ToolCallStrategy does not lift this constraint. Even under PASSTHROUGH, the middleware restores the model's reply for the user, so each token must denote a single entity. See Tool-call strategies.


Writing your own

Subclass AnyPlaceholderFactory[<tag>] with the right preservation tag for your guarantees, then implement create().

Id-only factory (id without label), PreservesIdentityOnly
import uuid
from collections.abc import Mapping

from piighost.components.placeholder import AnyPlaceholderFactory
from piighost.components.placeholder.tags import PreservesIdentityOnly
from piighost.models import Entity


class UUIDPlaceholderFactory(AnyPlaceholderFactory[PreservesIdentityOnly]):
    """Generate opaque delimited ids, e.g. <<a3f21b4c>>, no label revealed."""

    def create(self, entities: list[Entity]) -> Mapping[Entity, PreservesIdentityOnly]:
        tokens: dict[Entity, PreservesIdentityOnly] = {}
        seen: dict[str, PreservesIdentityOnly] = {}  # canonical value -> token

        for entity in entities:
            canonical = entity.text.lower()
            if canonical not in seen:
                seen[canonical] = PreservesIdentityOnly(f"<<{uuid.uuid4().hex[:8]}>>")
            tokens[entity] = seen[canonical]

        return tokens


# Not shown: the factory, put to work.
import asyncio

from piighost.components.anonymizer import Anonymizer
from piighost.components.detector import ExactMatchDetector
from piighost.pipeline import AnonymizationPipeline

pipeline = AnonymizationPipeline(
    ExactMatchDetector({"Patrick": "PERSON", "office@example.com": "EMAIL"}),
    anonymizer=Anonymizer(UUIDPlaceholderFactory()),
)
text = asyncio.run(pipeline.anonymize("Patrick writes from office@example.com.")).text
print("<<" in text and "Patrick" not in text)

The token is delimited, hence findable, and unique per entity. This factory can be used under PIIAnonymizationMiddleware.

Bracket format factory (label + id), PreservesLabeledIdentityOpaque
from collections import defaultdict
from collections.abc import Mapping

from piighost.components.placeholder import AnyPlaceholderFactory
from piighost.components.placeholder.tags import PreservesLabeledIdentityOpaque
from piighost.models import Entity


class BracketPlaceholderFactory(AnyPlaceholderFactory[PreservesLabeledIdentityOpaque]):
    """Generate tokens in the format [PERSON:1], [LOCATION:2], etc."""

    def create(
        self, entities: list[Entity]
    ) -> Mapping[Entity, PreservesLabeledIdentityOpaque]:
        tokens: dict[Entity, PreservesLabeledIdentityOpaque] = {}
        counters: dict[str, int] = defaultdict(int)

        for entity in entities:
            counters[entity.label] += 1
            inner = f"{entity.label}:{counters[entity.label]}"
            tokens[entity] = PreservesLabeledIdentityOpaque(f"[{inner}]")

        return tokens


# Not shown: the factory, put to work.
import asyncio

from piighost.components.anonymizer import Anonymizer
from piighost.components.detector import ExactMatchDetector
from piighost.pipeline import AnonymizationPipeline

pipeline = AnonymizationPipeline(
    ExactMatchDetector({"Patrick": "PERSON", "office@example.com": "EMAIL"}),
    anonymizer=Anonymizer(BracketPlaceholderFactory()),
)
text = asyncio.run(pipeline.anonymize("Patrick writes from office@example.com.")).text
print(text)
Realistic hashed factory, PreservesLabeledIdentityHashed

This factory produces a real-looking value whose content comes from a hash of the original value, hence unique and collision-free. The token has no delimited grammar, so it is not findable. Keep it out of the middleware.

import hashlib
from collections.abc import Mapping

from piighost.components.placeholder import AnyPlaceholderFactory
from piighost.components.placeholder.tags import PreservesLabeledIdentityHashed
from piighost.models import Entity


class HashedEmailPlaceholderFactory(
    AnyPlaceholderFactory[PreservesLabeledIdentityHashed]
):
    """Generate realistic emails like a1b2c3d4@anonymized.local."""

    def create(
        self, entities: list[Entity]
    ) -> Mapping[Entity, PreservesLabeledIdentityHashed]:
        tokens: dict[Entity, PreservesLabeledIdentityHashed] = {}

        for entity in entities:
            digest = hashlib.sha256(entity.text.encode()).hexdigest()[:8]
            tokens[entity] = PreservesLabeledIdentityHashed(
                f"{digest}@anonymized.local"
            )

        return tokens


# Not shown: the factory, put to work.
import asyncio

from piighost.components.anonymizer import Anonymizer
from piighost.components.detector import ExactMatchDetector
from piighost.pipeline import AnonymizationPipeline

pipeline = AnonymizationPipeline(
    ExactMatchDetector({"Patrick": "PERSON", "office@example.com": "EMAIL"}),
    anonymizer=Anonymizer(HashedEmailPlaceholderFactory()),
)
text = asyncio.run(pipeline.anonymize("Patrick writes from office@example.com.")).text
print(text)

See also