--- icon: lucide/play --- # First pipeline You will build a pipeline one component at a time, then run it on a sentence. You start from a detector, add the linker and the anonymizer, then put them together. The detector depends on what you look for. An NER model such as GLiNER2 detects names and locations it has never seen. A regex only recognizes fixed formats, such as an email address. !!! note "Prerequisites" `piighost` installed, see [Installation](installation.md). The regex path uses only the core, no extra. The GLiNER2 path needs the `gliner2` extra and downloads a model on first load. ## 1. Pick a detector The detector reads the text and returns detections, one per value found. The rest of the pipeline is the same whatever the detector, so pick the one that matches your text. === "Regex (catalog)" A `RegexDetector` recognizes patterns, that is strings of characters following a fixed structure. You pass it a dictionary mapping a label to a pattern. A first name has no fixed structure, so the two patterns below simply list the values of the example sentence. They show how the pipeline works, they detect no other name. ```python from piighost.components.detector import RegexDetector patterns = { "PERSON": r"\b(?:Patrick|Marie)\b", "LOCATION": r"\bParis\b", } detector = RegexDetector(patterns) ``` For fixed formats that do not depend on a language, such as email and URL, the [piighost catalog](https://catalog.piighost.dev) publishes ready-made groups. The `generic` group below holds no pattern for a name or a location. It is fetched from the catalog every time the detector is built. ```python from piighost.components.detector import RegexDetector detector = RegexDetector.from_catalog("catalog:piighost/generic") ``` === "GLiNER2 (NER)" An NER is an AI model that sorts the words of a text into categories decided in advance (name, first name, location, organization). Unlike the regex, it does not need to know the values in advance. It detects a first name it has never seen. ```python from piighost.components.detector.ner import Gliner2Detector detector = Gliner2Detector( model="fastino/gliner2-multi-v1", labels=["PERSON", "LOCATION"], threshold=0.5, ) ``` The first argument is a model name loaded by GLiNER2, or an already loaded instance. `labels` sets the queried categories. `threshold` is the minimum confidence above which a detection is kept. ## 2. Group detections into entities One first name can appear several times. The linker groups the detections of the same value and the same label into a single entity, so every occurrence later receives the same token. ```python from piighost.components.linker import ExactEntityLinker linker = ExactEntityLinker() ``` ## 3. Assign a token to each entity The anonymizer replaces each entity with a placeholder, that is the token that takes its place in the text. The token depends on the chosen factory. `LabelCounterPlaceholderFactory` numbers the tokens per label. This gives `<>`{ .placeholder }, `<>`{ .placeholder }, `<>`{ .placeholder }. ```python from piighost.components.anonymizer import Anonymizer from piighost.components.placeholder import LabelCounterPlaceholderFactory factory = LabelCounterPlaceholderFactory() anonymizer = Anonymizer(factory) ``` ## 4. Assemble and run `AnonymizationPipeline` chains the three components in order. It detects, groups, then replaces. Its `anonymize` method is asynchronous. It returns a result whose `text` attribute carries the de-identified sentence. ```python import asyncio from piighost.pipeline import AnonymizationPipeline pipeline = AnonymizationPipeline(detector, linker, anonymizer) async def main() -> None: text = "Patrick lives in Paris. Patrick loves Paris. Marie too." result = await pipeline.anonymize(text) print(result.text) asyncio.run(main()) ``` The output should be: ```text <> lives in <>. <> loves <>. <> too. ``` Each occurrence of `Patrick`{ .pii } receives the same `<>`{ .placeholder }. `Paris`{ .pii } keeps `<>`{ .placeholder } at both appearances. `Marie`{ .pii } receives the next number, `<>`{ .placeholder }. The linker from step 2 is what makes this consistency possible. ## How it works `AnonymizationPipeline` runs three mandatory stages. The detector finds the confidential data. The linker groups the occurrences of the same value into one entity. The anonymizer replaces each entity with the token from its factory. Optional stages (missed-occurrence expansion, entity merging) exist, disabled by default. Overlap resolution, in contrast, runs by default. Only the detector is strictly required to construct the pipeline, and that minimum is enough for a first pipeline. ## See also - To describe this pipeline in a file rather than in Python, see the [configuration reference](../configuration/toml.md). A regex detector takes its catalogs there with `catalogs = ["catalog:piighost/generic"]`. - To de-identify across a conversation with tokens stable between messages, see the [Conversational pipeline](conversation.md).