--- icon: lucide/replace --- # Placeholder factories A *placeholder* is the synthetic token that takes the place of a detected value before the text reaches the LLM. Instead of sending `Patrick`{ .pii } lives in `Paris`{ .pii } to the LLM, the pipeline sends `<>`{ .placeholder } lives in `<>`{ .placeholder }. The original values stay in the conversation memory. The LLM never sees them. !!! note "Why the name placeholder factory" *Placeholder* because the token holds the place of the original value. We could have said *token*, but that word is already overloaded in the LLM context (language tokens). *Factory* because the component builds these tokens on the fly, based on the entities detected in each message. A **placeholder factory** decides what those tokens look like and how much information they carry. Two questions structure the choice. 1. *Is the token unique per entity?* `Patrick`{ .pii } and `Marie`{ .pii } should not both collapse onto a generic `<>`{ .placeholder }, otherwise the LLM cannot tell them apart. A unique token per entity lets the model reason about relations. The question *is the manager the same person as `Patrick`{ .pii }?* becomes *is `<>`{ .placeholder } the same as `<>`{ .placeholder }?*, and it gets a clear answer. 2. *Is the token reversible and findable?* Does the token denote a single value in the conversation memory, and can it be relocated in a text the pipeline never produced? Restoration needs both properties, whether it runs on the model's reply or on a tool's arguments. If two entities collapse onto the same `<>`{ .placeholder }, there is no way to know which original to restore. Six families of factories sit at different points on that spectrum, and the choice has direct consequences on which `ToolCallStrategy` you can use safely. See [Tool-call strategies](tool-call-strategies.md) for the runtime side. - **No information** (`<>`{ .placeholder }): a constant token that reveals nothing to the LLM. Classic redaction. No reasoning is possible on entities. For example, the model cannot tell that the value was a city and decide to call the `get_weather` tool. - **Type only** (`<>`{ .placeholder }, `<>`{ .placeholder }): the type is revealed, not the identity. Multiple persons in the same conversation collapse onto the same `<>`{ .placeholder }, so cross-references break. - **Type + id (opaque)** (`<>`{ .placeholder }, `<>`{ .placeholder }): type revealed, stable identity, clearly synthetic token. The LLM can tell that `<>`{ .placeholder } and `<>`{ .placeholder } are two different people. Unique, so reversible by string replacement. - **Id only** (`<>`{ .placeholder }): a unique hash per entity, without revealing the type. The LLM sees that two distinct entities exist but cannot tell whether they are persons, emails, or cards. Keeps reversibility on the tool side without giving any semantic hint to the model. - **Partial value** (`J*******`{ .placeholder } for `Jonathan`{ .pii }): part of the real content stays visible, here the first letter and the length. The LLM sees the start of the value, not the full value. Riskier on privacy (real fragments) and on reversibility (collisions possible). !!! note "Token format convention" Tokens in this documentation follow a simple rule. - **Synthetic token** (does not look like any real value), wrapped in `<<` and `>>`. For example `<>`{ .placeholder }, `<>`{ .placeholder }, `<>`{ .placeholder }, `<>`{ .placeholder }, `<>`{ .placeholder }. The delimiters serve two purposes. An LLM or a human re-reading never mistakes the token for a regular word or for an HTML/XML tag the model might emit. And the middleware can find the token again to run its string replacement, including spotting a token the model invented. - **Token that replicates a real value's format** (realistic hashed, masked), no delimiters. For example `a1b2c3d4@anonymized.local`{ .placeholder }, `Patient_a1b2c3d4`{ .placeholder }, `j***@mail.com`{ .placeholder }. The absence of delimiters is deliberate. The token has to look natural, so that a downstream tool that validates a format (email regex, card length) still accepts it. The rule also applies to any factory you write. Purely opaque token, wrap it. Token that mimics a real value, leave it raw. --- ## Family details ### No information, total destruction The token is a fixed marker, e.g. `<>`{ .placeholder }. The LLM learns *that* something was removed but nothing about its type, count, or relations. The conversation loses every internal reference. An agent trying to act on *send the invoice to the client* cannot tell whether the client is the one mentioned earlier or someone new. Useful for archival redaction, useless once an agent has to reason. - Built-in: `RedactPlaceholderFactory` (output `<>`{ .placeholder }, delimiters configurable). - Preservation tag: `PreservesNothing`. ### Type only, identities collapsed `<>`{ .placeholder }, `<>`{ .placeholder }. The LLM knows that something is a person, an email, a card, and can answer questions that depend on the type alone. But two different persons in the same conversation collapse onto the same token. The classic failure mode is cross-reference. The question *is `Patrick`{ .pii } the same person as the manager mentioned earlier?* becomes *is `<>`{ .placeholder } the same as `<>`{ .placeholder }?*, and that question has no answer. - Built-in: `LabelPlaceholderFactory` (output `<>`{ .placeholder }). - Preservation tag: `PreservesLabel`. ### Type + id (opaque) `<>`{ .placeholder }, `<>`{ .placeholder }. The string clearly is *not* a person, an email, or a card number, it is a token. The LLM cannot mistake it for real data, audit logs are easy to scan, and there is **zero chance** of collision with a real value. Its delimiters also make it findable. A consumer can then spot a token the model invented. In return, a strict downstream prompt or tool that requires *the argument must look like an email* will reject these tokens. - Built-in: `LabelCounterPlaceholderFactory` (`<>`{ .placeholder }) and `LabelHashPlaceholderFactory` (`<>`{ .placeholder }). - Preservation tag: `PreservesLabeledIdentityOpaque`. Both number the entities per label, in order. The first person becomes ordinal 1, the second 2, while an email starts its own count at 1. `LabelHashPlaceholderFactory` renders that ordinal as a hash. The hash is a sha256 of the string `label:ordinal`, never of the value. It only gives an opaque look, so that two consecutive entities look unrelated. ### Id only, identity without type `<>`{ .placeholder }. The token keeps the synthetic `<<...>>` shape but does not reveal the label, while carrying a unique hash per entity. The LLM cannot tell whether the entity is a person, an email, or a card, but it can see that `<>`{ .placeholder } and `<>`{ .placeholder } are two distinct entities. It is one of the most protective levels that stays usable on the tool side. The string replacement works, because the hash is unique. - Built-in: none for this branch. - Preservation tag: `PreservesIdentityOnly`, meant for a factory you write, a hashed redaction with no label prefix. See *Writing your own* below. ### Type + id (realistic hashed) A custom factory can produce values that **look like the original format** but whose content is driven by a hash, e.g. `a1b2c3d4@anonymized.local`{ .placeholder } for an email, or `Patient_a1b2c3d4`{ .placeholder } for a name. The token passes basic format validation (email regex, length, allowed characters), so downstream tools and prompt templates that expect a real-looking value still work. Because the content is a hash, the token is **unique and cannot coincidentally match** an existing real value. - Built-in: none. See *Writing your own* below for a complete example. - Preservation tag: `PreservesLabeledIdentityHashed`. !!! warning "Token not findable" This tag is not findable. The middleware therefore cannot spot an invented token in this form. Weigh this before using it under the middleware. ### Partial value, a fragment leaks `J*******`{ .placeholder }, `j***@mail.com`{ .placeholder }, `****4567`{ .placeholder }. The token keeps *part* of the original value, for example the email domain, the last four digits of a card, the first letter of a name. The LLM can reason on more than the type, *the email is on the company domain*, *the card ends in 4567*, *the name starts with J*. Two trade-offs come with this. 1. **Real fragments of the value reach the LLM.** It cannot reconstruct the full value, but `j***@mail.com`{ .placeholder } already places the user inside a known mail provider. 2. **Collisions are possible.** Two different cards ending in `4567` collapse onto `****4567`{ .placeholder }, two emails sharing the first letter and domain end up identical. The token is *mostly* unique, with no guarantee. - Built-in: `MaskPlaceholderFactory`, which by default keeps the first character of the value and masks the rest with `*`, so `Jonathan`{ .pii } becomes `J*******`{ .placeholder } and `jean@mail.com`{ .pii } becomes `j************`{ .placeholder }. The `j***@mail.com`{ .placeholder } and `****4567`{ .placeholder } forms need a factory you write. - Preservation tag: `PreservesShape`. The middleware refuses it, at type-check time and at runtime. An ambiguous token cannot be restored through string replacement, and a mask has no grammar the middleware can find again. --- ## Preservation tags Every factory carries a **phantom type** that summarises the preservation level of its tokens. A phantom type is a generic parameter that exists only at type-check time, it does not affect execution. The type-checker reads this tag to validate a factory against its consumers. The table below gives an example token and the tag of each family. | Family | Example | Tag | |---|---|---| | No information | `<>`{ .placeholder } | `PreservesNothing` | | Type only | `<>`{ .placeholder } | `PreservesLabel` | | Type + id (opaque) | `<>`{ .placeholder }, `<>`{ .placeholder } | `PreservesLabeledIdentityOpaque` | | Id only | `<>`{ .placeholder } | `PreservesIdentityOnly` | | Type + id (realistic hashed) | `a1b2c3d4@anonymized.local`{ .placeholder }, `Patient_a1b2c3d4`{ .placeholder } | `PreservesLabeledIdentityHashed` | | Partial value | `J*******`{ .placeholder }, `****4567`{ .placeholder } | `PreservesShape` | Two tables read these families from two angles. The **Confidentiality** table shows what leaks to the LLM, from the attacker and privacy point of view. The **Exploitation** table shows what the agent and the system can do with the token, from the point of view of functional capabilities. The same answer can be good in one and problematic in the other, and the two tables make this tension explicit. Both tables share the same colour code, from best to problematic, explained in the legend under the second table. #### Confidentiality (what leaks to the LLM)
FamilyType seen?Values distinguished?Real-value leak?Collision with a real value?
No informationnonononeno
Type onlyyesnononeno
Type + id (opaque)yesyesnoneno
Id onlynoyesnoneno
Type + id (realistic hashed)yesyesnoneno
Partial valueyesyespartialrisk
#### Exploitation by the LLM and the agent
FamilyReason about the typeTrack cross-referencesReversible at the tool boundaryToken findable
No informationnononoyes
Type onlyyesnonoyes
Type + id (opaque)yesyesyesyes
Id onlynoyesyesyes
Type + id (realistic hashed)yesyesyesno
Partial valueyesmostlyyes (collisions)no
Legend: best acceptable partial problematic Tags form an **inheritance hierarchy** that the type-checker exploits through the covariance of `AnyPlaceholderFactory[PreservationT_co]`. A factory tagged more specifically therefore satisfies a consumer asking for a looser one. Three independent axes structure the taxonomy: - *Label*: the token reveals the type. - *Identity*: the token is unique per entity. - *Recognizable*: the factory can find its token again in arbitrary text. A delimited token allows this, a realistic one does not. `PreservesLabeledIdentity` combines label and identity via multiple inheritance. A `<>`{ .placeholder } factory is therefore both a `PreservesLabel` *and* a `PreservesIdentity`. `PreservesRecognizableIdentity` crosses identity with findability. The middleware accepts only this intersection. A consumer typed against `PreservesRecognizableIdentity` sorts the tags as follows: - Accepts: `PreservesIdentityOnly` and `PreservesLabeledIdentityOpaque`. - Rejects: `PreservesLabel`, `PreservesShape` and `PreservesNothing`, which lack the uniqueness guarantee, along with `PreservesLabeledIdentityHashed`, which is not findable. ```mermaid classDiagram class PlaceholderPreservation { root } class PreservesNothing { <<REDACT>> } class PreservesLabel { <<PERSON>> } class PreservesShape { "J*******" } class Recognizable { abstraction } class PreservesIdentity { abstraction } class PreservesRecognizableIdentity { abstraction } class PreservesIdentityOnly { <<REDACT:a1b2c3d4>> } class PreservesLabeledIdentity { abstraction } class PreservesLabeledIdentityOpaque { <<PERSON:1>> <<PERSON:a1b2c3d4>> } class PreservesLabeledIdentityRealistic { abstraction } class PreservesLabeledIdentityHashed { a1b2c3d4@anonymized.local Patient_a1b2c3d4 } PlaceholderPreservation <|-- PreservesNothing PlaceholderPreservation <|-- PreservesLabel PlaceholderPreservation <|-- Recognizable PlaceholderPreservation <|-- PreservesIdentity PreservesLabel <|-- PreservesShape PreservesIdentity <|-- PreservesRecognizableIdentity Recognizable <|-- PreservesRecognizableIdentity PreservesRecognizableIdentity <|-- PreservesIdentityOnly PreservesLabel <|-- PreservesLabeledIdentity PreservesIdentity <|-- PreservesLabeledIdentity PreservesLabeledIdentity <|-- PreservesLabeledIdentityOpaque PreservesRecognizableIdentity <|-- PreservesLabeledIdentityOpaque PreservesLabeledIdentity <|-- PreservesLabeledIdentityRealistic PreservesLabeledIdentityRealistic <|-- PreservesLabeledIdentityHashed ``` *Preservation tag hierarchy. Each node carries an example token, the abstract nodes are intersections between axes. Each arrow goes from a tag to its parent and reads "is a".* { .figure-caption } `PreservesLabeledIdentity` inherits from both `PreservesLabel` and `PreservesIdentity`. This inheritance expresses the *A is a B but not every B is an A* relation. Every `PreservesLabeledIdentity` is also a `PreservesLabel` and a `PreservesIdentity`, but a `PreservesLabel` is not necessarily a `PreservesLabeledIdentity`. `PreservesShape` extends `PreservesLabel`, because a masked token implies the label through its format. It does not guarantee uniqueness, so it does not descend from `PreservesIdentity`. Each tag is a subclass of `str`, so a token is a real string that carries its preservation level in its own type. A factory declares the **most specific** tag that matches its guarantees. ```python class LabelCounterPlaceholderFactory( BaseCounterPlaceholderFactory ): ... # PreservesLabeledIdentityOpaque class LabelHashPlaceholderFactory( BaseCounterPlaceholderFactory ): ... # PreservesLabeledIdentityOpaque class LabelPlaceholderFactory(AnyPlaceholderFactory[PreservesLabel]): ... class MaskPlaceholderFactory(AnyPlaceholderFactory[PreservesShape]): ... class RedactPlaceholderFactory(AnyPlaceholderFactory[PreservesNothing]): ... # No built-in for the id-only branch nor the realistic hashed one, # implement your own with PreservesIdentityOnly or PreservesLabeledIdentityHashed. ``` --- ## Built-in factories | Factory | Style | Mechanism | Output example | |---|---|---|---| | `RedactPlaceholderFactory` | Redact | none | `<>`{ .placeholder } | | `LabelPlaceholderFactory` | Label | none | `<>`{ .placeholder } | | `LabelCounterPlaceholderFactory` (default) | Label | Counter | `<>`{ .placeholder } | | `LabelHashPlaceholderFactory` | Label | Hash | `<>`{ .placeholder } | | `MaskPlaceholderFactory` | Mask | partial | `J*******`{ .placeholder } | The tag of each factory is in the family table, above. The naming follows a `