First pipeline
You will build a pipeline one component at a time, then run it on a sentence. You start from a detector, add the linker and the anonymizer, then put them together. The detector depends on what you look for. An NER model such as GLiNER2 detects names and locations it has never seen. A regex only recognizes fixed formats, such as an email address.
1. Pick a detector
The detector reads the text and returns detections, one per value found. The rest of the pipeline is the same whatever the detector, so pick the one that matches your text.
A RegexDetector recognizes patterns, that is strings of characters following a fixed structure. You pass it a dictionary mapping a label to a pattern. A first name has no fixed structure, so the two patterns below simply list the values of the example sentence. They show how the pipeline works, they detect no other name.
from piighost.components.detector import RegexDetector
patterns = {
"PERSON": r"\b(?:Patrick|Marie)\b",
"LOCATION": r"\bParis\b",
}
detector = RegexDetector(patterns)For fixed formats that do not depend on a language, such as email and URL, the piighost catalog publishes ready-made groups. The generic group below holds no pattern for a name or a location. It is fetched from the catalog every time the detector is built.
from piighost.components.detector import RegexDetector
detector = RegexDetector.from_catalog("catalog:piighost/generic")An NER is an AI model that sorts the words of a text into categories decided in advance (name, first name, location, organization). Unlike the regex, it does not need to know the values in advance. It detects a first name it has never seen.
from piighost.components.detector.ner import Gliner2Detector
detector = Gliner2Detector(
model="fastino/gliner2-multi-v1",
labels=["PERSON", "LOCATION"],
threshold=0.5,
)The first argument is a model name loaded by GLiNER2, or an already loaded instance. labels sets the queried categories. threshold is the minimum confidence above which a detection is kept.
2. Group detections into entities
One first name can appear several times. The linker groups the detections of the same value and the same label into a single entity, so every occurrence later receives the same token.
from piighost.components.linker import ExactEntityLinker
linker = ExactEntityLinker()3. Assign a token to each entity
The anonymizer replaces each entity with a placeholder, that is the token that takes its place in the text. The token depends on the chosen factory. LabelCounterPlaceholderFactory numbers the tokens per label. This gives <<PERSON:1>>, <<PERSON:2>>, <<LOCATION:1>>.
from piighost.components.anonymizer import Anonymizer
from piighost.components.placeholder import LabelCounterPlaceholderFactory
factory = LabelCounterPlaceholderFactory()
anonymizer = Anonymizer(factory)4. Assemble and run
AnonymizationPipeline chains the three components in order. It detects, groups, then replaces. Its anonymize method is asynchronous. It returns a result whose text attribute carries the de-identified sentence.
import asyncio
from piighost.pipeline import AnonymizationPipeline
pipeline = AnonymizationPipeline(detector, linker, anonymizer)
async def main() -> None:
text = "Patrick lives in Paris. Patrick loves Paris. Marie too."
result = await pipeline.anonymize(text)
print(result.text)
asyncio.run(main())The output should be:
<<PERSON:1>> lives in <<LOCATION:1>>. <<PERSON:1>> loves <<LOCATION:1>>. <<PERSON:2>> too.Each occurrence of Patrick receives the same <<PERSON:1>>. Paris keeps <<LOCATION:1>> at both appearances. Marie receives the next number, <<PERSON:2>>. The linker from step 2 is what makes this consistency possible.
How it works
AnonymizationPipeline runs three mandatory stages. The detector finds the confidential data. The linker groups the occurrences of the same value into one entity. The anonymizer replaces each entity with the token from its factory. Optional stages (missed-occurrence expansion, entity merging) exist, disabled by default. Overlap resolution, in contrast, runs by default. Only the detector is strictly required to construct the pipeline, and that minimum is enough for a first pipeline.
See also
- To describe this pipeline in a file rather than in Python, see the configuration reference. A regex detector takes its catalogs there with
catalogs = ["catalog:piighost/generic"]. - To de-identify across a conversation with tokens stable between messages, see the Conversational pipeline.