Acceptance tests
In short
- An acceptance test checks one criterion of a need from Needs by profile, with literal values and through the public entry points.
- Each test carries the identifier
AT-<need>-<n>, the same in every language. - The page compares each test with the existing suite and marks it covered, partial or missing.
- The tests written to fill the gaps are in
tests/acceptance/.
The decisions still to make are in Open points. The terms are defined in the glossary.
How to read the tables
- "Existing tests" cites pytest nodes relative to the root of the
piighostrepository. A node prefixedpiighost-api:comes from thepiighost-apirepository. - The LlamaIndex tests are skipped without the
llama-indexextra. They pass withuv run --with llama-index-core pytest tests/integrations/llama_index. - Status: covered (the behavior is checked, sometimes with values other than those of the inventory, and those values are then flagged), partial (what is missing is stated), missing.
- The library tests use
ExactMatchDetectoror a simulated catalog. The detection quality of a real model or of a real catalog group stays outside these tests.
DPO (compliance officer)
| Id | Need | Scenario | Expected result | Existing tests | Status |
|---|---|---|---|---|---|
| AT-DPO-1-1 | DPO-1 | Given a protected agent (LangChain middleware, Pydantic AI capability or proxy), when the user writes "Write to Jean Dupont, jean.dupont@exemple.fr", then the model receives the de-identified text | The model receives "Write to <<PERSON:1>>, <<EMAIL:1>>", without "Jean Dupont" or "jean.dupont@exemple.fr" in any message | tests/integrations/langchain/test_middleware_e2e.py::test_second_model_call_never_sees_clear_tool_argstests/integrations/test_pydantic_ai_hooks.py::TestAroundTheModel::test_model_sees_placeholders_and_reply_is_restoredpiighost-api:tests/routes/test_openai_chat.py::test_upstream_sees_tokens_reply_is_restored | covered (values "Patrick Dupont" and "patrick@acme.com") |
| AT-DPO-1-2 | DPO-1 | Given a config that pulls a secrets group from the catalog, when the message contains an API key, then the key leaves as a placeholder | The de-identified text contains a placeholder with the group's label (for example "<<API_KEY:1>>") and no longer the key | tests/acceptance/test_dpo.py::TestSecretsLeaveAsTokens::test_an_api_key_leaves_as_a_token | covered (patterns OPENAI_API_KEY and AWS_ACCESS_KEY copied from catalog:piighost/logs:b635d867, simulated catalog, placeholders "<<OPENAI_API_KEY:1>>" and "<<AWS_ACCESS_KEY:1>>") |
| AT-DPO-2-1 | DPO-2 | Given a config whose catalogs contains catalog:piighost/fr-default:<commit>, when the text contains the IBAN "FR76 3000 6000 0112 3456 7890 189" and a NIR, then both are masked | Two placeholders, neither the IBAN nor the NIR in clear | tests/config/test_detectors.py::TestRegexCatalogs::test_a_catalog_is_pulled_at_buildtests/test_catalog.py::TestFromCatalog::test_builds_a_detector_that_detectstests/pipeline/test_pipeline.py::TestUnicodeSpaces::test_a_regex_value_typed_with_no_break_spaces_is_hidden | partial (loading a group is checked on a simulated catalog, the IBAN is masked only by an inline pattern, the NIR never) |
| AT-DPO-2-2 | DPO-2 | Given the same config without the group that carries the IBAN, when the same text goes through, then the IBAN stays in clear | "FR76 3000 6000 0112 3456 7890 189" appears as is in the output | tests/acceptance/test_dpo.py::TestChosenGroups::test_a_removed_group_leaves_its_values_clear | covered (simulated catalog) |
| AT-DPO-2-3 | DPO-2 | Given a config with an inline pattern specific to the company, when the text contains a value that matches it, then it is masked | "reach a@b.co now" becomes "reach <<REDACT>> now" | tests/config/test_settings.py::TestLoadPipeline::test_builds_a_working_pipelinetests/config/test_detectors.py::TestRegexCatalogs::test_catalogs_merge_in_order_then_inline_patterns | covered |
| AT-DPO-3-1 | DPO-3 | Given a deny list (deny_list) containing "Acme" and a detector that sees nothing, when "Acme rocks" goes through, then "Acme" is masked | "<<ORG:1>> rocks" | tests/pipeline/test_override_integration.py::TestBasePipelineOverride::test_deny_list_forces_a_value_the_detector_missedtests/config/test_settings.py::TestOverrideEffect::test_deny_list_forces_a_detection | covered |
| AT-DPO-3-2 | DPO-3 | Given an allow list (allow_list) containing "Paris" and a detector that sees it, when "Visit Paris" goes through, then "Paris" stays in clear | "Visit Paris" | tests/pipeline/test_override_integration.py::TestBasePipelineOverride::test_allow_list_keeps_a_false_positive_in_clear | covered |
| AT-DPO-4-1 | DPO-4 | Given a guard rail that detects e-mail addresses and a main detector that misses them, when "Emma wrote leak@x.com" goes through, then nothing is returned | PIIRemainingError is raised | tests/pipeline/test_override_integration.py::TestBasePipelineOverride::test_a_real_leak_still_trips_the_guardtests/config/test_settings.py::TestGuardEffect::test_detector_guard_raises_on_residualtests/pipeline/test_pipeline.py::TestGuard::test_a_flagged_guard_raises | covered |
| AT-DPO-5-1 | DPO-5 | Given a Redis memory with a SHA-256 hasher and AES-GCM encryption, when the message "I am Emma" is stored, then the database contains neither the message nor the value | No key contains "I am Emma", no stored value contains "Emma" | tests/conversation_memory/test_redis.py::TestAtRestProtection::test_message_text_is_not_stored_in_any_keytests/conversation_memory/test_redis.py::TestAtRestProtection::test_detections_are_encrypted_at_rest | covered (on RedisConversationMemory and fakeredis, not through a config file) |
| AT-DPO-5-2 | DPO-5 | Given a Redis config file with hasher and cipher, when PIIGHOST_HASH_PEPPER or PIIGHOST_CIPHER_KEY is missing, then load_thread_pipeline refuses to build | ConfigError, no pipeline returned | tests/acceptance/test_ops.py::TestSecretsFromTheEnvironment::test_a_full_redis_config_without_its_secrets_does_not_buildtests/config/test_crypto_models.py::TestHasherConfig::test_missing_pepper_is_rejectedtests/config/test_crypto_models.py::TestCipherConfig::test_missing_key_is_rejectedtests/config/test_redis_memory.py::TestRedisMemoryConfig::test_half_configured_crypto_raises_config_error | covered |
| AT-DPO-6-1 | DPO-6 | Given a conversation "t1" that de-identified "Hi Emma", when forget_thread("t1") is called and then deanonymize("<<PERSON:1>>", "t1"), then the value is no longer returned | The restoration returns "<<PERSON:1>>", never "Emma" | tests/acceptance/test_dpo.py::TestErasure::test_a_forgotten_thread_restores_nothingtests/pipeline/test_thread.py::TestForget::test_forget_thread_clears_the_memorytests/pipeline/test_thread.py::TestForget::test_forget_thread_drops_the_memoized_tokenstests/conversation_memory/test_redis.py::TestForget::test_forget_purges_and_reports | covered |
| AT-DPO-6-2 | DPO-6 | Given the piighost-api server, when the client sends DELETE /v1/threads/t1, then the conversation is erased | 200 and {"messages": 2, "detections": 3} | piighost-api:tests/test_app.py::test_forget_thread_returns_countstests/integrations/client/test_client.py::TestForgetThread::test_deletes_and_returns_a_forgotten | covered (simulated pipeline on the server side) |
| AT-DPO-7-1 | DPO-7 | Given a pipeline with observation_redactor, when "Hi Emma!" goes through, then the exported traces contain only placeholders | No span attribute contains "Emma", the piighost.detect span carries a placeholder | tests/observation/test_pipeline_spans.py::TestRedaction::test_redactor_replaces_values_with_its_tokenstests/observation/test_pipeline_spans.py::TestRedaction::test_redactor_removes_clear_values_from_payloads | covered (placeholder "<<REDACT>>" and not "<<PERSON:1>>", the test uses RedactPlaceholderFactory) |
| AT-DPO-8-1 | DPO-8 | Given the documentation, when you open docs/fr/dpia.md and docs/en/dpia.md, then the page lists risks, measures and settings | Both pages exist and carry a section that maps each risk to a setting | none | missing (documentation criterion, make lint checks only the style of the pages) |
| AT-DPO-9-1 | DPO-9 | Given an LLMDetector whose model returns an output without an entities field, when a message goes through, then the pipeline raises an error | An error raised, no text returned | tests/components/detector/test_llm.py::TestDetect::test_an_unreadable_output_refuses_the_message | covered |
| AT-DPO-9-2 | DPO-9 | Given the same detector with fail open explicitly requested, when the message goes through, then it leaves without detection and a warning is logged | The text unchanged, a warning | tests/components/detector/test_llm.py::TestDetect::test_fail_open_lets_an_unreadable_output_through | covered |
| AT-DPO-10-1 | DPO-10 | Given a human correction exported in the "placeholders" form, when the dataset is read, then it contains no real value | The dataset carries "<<PERSON:1>>", never "Emma" | none | missing (feature to design) |
Developer (who integrates piighost)
| Id | Need | Scenario | Expected result | Existing tests | Status |
|---|---|---|---|---|---|
| AT-DEV-1-1 | DEV-1 | Given an existing create_agent agent, when you add middleware=[PIIAnonymizationMiddleware(pipeline)], then the agent runs with no other change and the model sees only placeholders | The ainvoke call succeeds, no call to the model contains the value in clear | tests/integrations/langchain/test_middleware_e2e.py::test_second_model_call_never_sees_clear_tool_args | covered |
| AT-DEV-1-2 | DEV-1 | Given an OpenAI client and a piighost-api with a default upstream, when only base_url points to /openai/v1, then the upstream receives placeholders and the reply comes back restored | The upstream receives "<<PERSON:1>>", the client reads "Hi Patrick" | piighost-api:tests/routes/test_openai_chat.py::test_upstream_sees_tokens_reply_is_restoredpiighost-api:tests/routes/test_upstream.py::test_upstream_default_used_when_header_absentpiighost-api:tests/routes/test_anthropic_messages.py::test_default_upstream_used_without_header | partial (the end-to-end test passes the x-piighost-upstream header, the default upstream is tested separately, no test goes through the openai SDK) |
| AT-DEV-2-1 | DEV-2 | Given a conversation where "Hi Emma" was de-identified, when the model replies "Hello <<PERSON:1>>", then the application receives the restored reply | "Hello Emma" | tests/integrations/langchain/test_middleware.py::TestWhenInstalled::test_after_model_deanonymizes_for_displaytests/integrations/test_pydantic_ai_hooks.py::TestAroundTheModel::test_model_sees_placeholders_and_reply_is_restoredtests/pipeline/test_thread.py::TestDeanonymize::test_restores_tokens_the_pipeline_never_anonymized | covered |
| AT-DEV-3-1 | DEV-3 | Given a conversation "t1", when "Hello Emma" then "Bye Emma" go through, then the value keeps its placeholder | "Hello <<PERSON:1>>" then "Bye <<PERSON:1>>" | tests/pipeline/test_thread.py::TestThreadConsistency::test_a_value_keeps_its_token_across_messagestests/config/test_thread_pipeline.py::TestLoadThreadPipeline::test_memory_shares_placeholder_across_messages | covered (two messages, not three) |
| AT-DEV-3-2 | DEV-3 | Given "Liam" in conversation "a" and "Emma" in conversation "b", when you restore "<<PERSON:1>>" in conversation "b", then you get the value of conversation "b" | Each conversation numbers from "<<PERSON:1>>", the restoration in "b" returns "Emma", never "Liam" | tests/acceptance/test_dev.py::TestThreadIsolation::test_a_token_restores_in_its_own_thread_onlytests/pipeline/test_thread.py::TestThreadConsistency::test_threads_are_isolatedtests/config/test_thread_pipeline.py::TestLoadThreadPipeline::test_threads_are_isolated | covered (a third, empty thread returns the placeholder as is) |
| AT-DEV-4-1 | DEV-4 | Given the FULL strategy (the default), when the model calls send_mail(to="<<EMAIL:1>>"), then the tool receives the real value and its result leaves again as placeholders | The tool receives "patrick@acme.com", the next call to the model contains neither "patrick@acme.com" nor "Patrick Dupont" | tests/integrations/langchain/test_middleware_e2e.py::test_second_model_call_never_sees_clear_tool_argstests/integrations/langchain/test_middleware.py::TestToolCalls::test_strategy_routes_each_directiontests/integrations/test_pydantic_ai_hooks.py::TestTools::test_tool_gets_the_value_and_its_result_is_reanonymized | covered |
| AT-DEV-5-1 | DEV-5 | Given a config with patterns = { ORDER = "CMD-\d{4}-\d{4}" }, when "CMD-2024-0042" goes through, then the number is masked | "<<ORDER:1>>" in place of the number | tests/config/test_settings.py::TestLoadPipeline::test_builds_a_working_pipeline | covered (inline EMAIL pattern instead of the order number) |
| AT-DEV-5-2 | DEV-5 | Given a class that inherits from nothing but exposes async detect(text) -> list[Detection], when you pass it to AnonymizationPipeline, then its detections become placeholders | The output carries the placeholders of the class's detections | tests/pipeline/test_pipeline.py::TestMergeOverlap::test_the_merge_resolver_leaves_no_fragment_of_a_longer_spantests/pipeline/test_pipeline.py::TestUnicodeSpaces::test_a_value_keeps_one_token_whatever_its_spaces | covered (incidentally, through the _FixedDetector of these tests) |
| AT-DEV-6-1 | DEV-6 | Given a valid config, when you run piighost validate config.toml, then the command succeeds | Exit code 0, output that starts with "OK:" | tests/cli/test_cli.py::TestValidate::test_valid_config_exits_zero | covered |
| AT-DEV-6-2 | DEV-6 | Given a config where patterns is written pattern, when you run piighost validate, then the command fails and names the error | Exit code 1, stderr that names the faulty key | tests/cli/test_cli.py::TestValidate::test_invalid_schema_exits_onetests/config/test_catalog_config.py::TestLoadConfigFromTheCatalog::test_an_invalid_configuration_names_the_referencetests/config/test_models.py::TestDetectorConfig::test_unknown_key_is_rejected | partial (the CLI is tested on a missing type and checks only "invalid configuration", the misspelled key is tested only in the loader) |
| AT-DEV-7-1 | DEV-7 | Given ExactMatchDetector({"Emma": "PERSON", "Liam": "PERSON"}), when "Emma met Liam" goes through, then the text is de-identified without a model | "<<PERSON:1>> met <<PERSON:2>>" | tests/pipeline/test_pipeline.py::TestAnonymize::test_replaces_detected_piitests/components/detector/test_exact.py::TestDetect::test_finds_multiple_values | covered (the absence of network is not enforced by a socket block) |
| AT-DEV-8-1 | DEV-8 | Given the RAISE strategy, when the model replies with "<<PERSON:9>>", never issued, then the reply is refused | InventedPlaceholderError is raised | tests/integrations/langchain/test_middleware.py::TestInventedPlaceholders::test_raise_refuses_an_invented_tokentests/integrations/test_pydantic_ai_hooks.py::TestInventedStrategy::test_an_invented_token_in_the_reply_is_refused | covered |
| AT-DEV-8-2 | DEV-8 | Given the DROP strategy, when the model replies "Hi <<PERSON:1>>, cc <<PERSON:9>>", then the invented placeholder disappears | "Hi Emma, cc " | tests/integrations/langchain/test_middleware.py::TestInventedPlaceholders::test_kept_or_dropped | covered |
| AT-DEV-9-1 | DEV-9 | Given a conversation where "<<PERSON:1>>" stands for "Patrick", when deanonymize_stream receives "Hi <<PER" then "SON:1>>", then the value appears only once | "Hi Patrick" | tests/integrations/langchain/test_middleware_stream.py::test_middleware_deanonymize_stream_restores_split_tokentests/integrations/test_deidentify_stream.py::test_token_split_across_chunks_is_restored | covered |
| AT-DEV-10-1 | DEV-10 | Given the LangChain middleware, the Claude Code hooks or the API server, when a call names no conversation, then it is refused | MissingThreadIdError, or 400 on the server | tests/integrations/langchain/test_middleware.py::TestThreadId::test_a_missing_id_raisestests/integrations/test_claude_code_hooks.py::test_an_event_without_a_session_id_is_refusedpiighost-api:tests/test_app.py::test_a_thread_route_without_a_thread_id_is_refused | covered |
| AT-DEV-10-2 | DEV-10 | Given a call that names "default", when it goes through the middleware, then it is accepted | The conversation used is "default" | tests/integrations/langchain/test_middleware.py::TestThreadId::test_the_shared_thread_is_named_explicitly | covered |
| AT-DEV-11-1 | DEV-11 | Given the PRESERVE strategy (the default), when the assistant writes "It was Napoleon." and then the user asks "What did Napoleon do?", then the value stays in clear | "Napoleon" in clear in both messages | tests/pipeline/test_thread.py::TestProvenance::test_assistant_introduced_value_stays_cleartests/pipeline/test_thread.py::TestProvenance::test_user_reference_after_assistant_stays_cleartests/integrations/langchain/test_middleware.py::TestAssistantProvenance::test_preserve_keeps_assistant_value_clear | covered |
| AT-DEV-11-2 | DEV-11 | Given the ANONYMIZE or IGNORE strategy, when the assistant introduces the same value, then it becomes a placeholder, or the assistant's message is not analyzed | "<<PERSON:1>>", or no detector call for the assistant | tests/integrations/langchain/test_middleware.py::TestAssistantProvenance::test_anonymize_treats_assistant_value_as_piitests/integrations/langchain/test_middleware.py::TestAssistantProvenance::test_ignore_does_not_analyze_assistant_messages | covered |
Operator (who deploys and runs it)
| Id | Need | Scenario | Expected result | Existing tests | Status |
|---|---|---|---|---|---|
| AT-OPS-1-1 | OPS-1 | Given piighost-api serve --config hub:piighost/fr-default:e6990159, when a client sends POST /v1/anonymize, then the server responds | 201 and an anonymized_text that carries placeholders | piighost-api:tests/test_cli_serve.py::test_serve_takes_a_hub_referencepiighost-api:tests/test_app.py::test_anonymizepiighost-api:tests/test_routes_labels.py::test_the_labels_of_a_hub_catalog_come_from_the_hub | partial (serve is tested with a simulated uvicorn, the route with a simulated pipeline, no test links the two) |
| AT-OPS-2-1 | OPS-2 | Given two ThreadAnonymizationPipeline on the same Redis database, when instance A processes "Hi Emma" and instance B "Bye Emma" in conversation "t1", then they agree | Both return "<<PERSON:1>>", B restores "<<PERSON:1>>" to "Emma" | tests/acceptance/test_ops.py::TestSharedMemory::test_two_instances_agree_on_a_thread | covered (shared fakeredis, a new value seen by B also restores from A) |
| AT-OPS-3-1 | OPS-3 | Given a Redis config with hasher and cipher, when the server starts without PIIGHOST_HASH_PEPPER or PIIGHOST_CIPHER_KEY, then it refuses to start | ConfigError at startup, no clear-text startup | tests/acceptance/test_ops.py::TestSecretsFromTheEnvironment::test_a_full_redis_config_without_its_secrets_does_not_buildtests/config/test_crypto_models.py::TestHasherConfig::test_missing_pepper_is_rejectedtests/config/test_crypto_models.py::TestCipherConfig::test_missing_key_is_rejectedtests/config/test_redis_memory.py::TestRedisMemoryConfig::test_half_configured_crypto_raises_config_error | partial (checked through load_thread_pipeline on a complete file, not at server startup) |
| AT-OPS-4-1 | OPS-4 | Given no API_KEY_<name> and no PIIGHOST_ALLOW_ANONYMOUS, when the server starts, then it refuses | RuntimeError that names PIIGHOST_ALLOW_ANONYMOUS | piighost-api:tests/test_auth.py::test_startup_fails_without_keys_by_default | covered |
| AT-OPS-4-2 | OPS-4 | Given the server, when a client sends a text of 2,000,000 characters, then the request is rejected | 413 | piighost-api:tests/test_app.py::test_oversized_body_is_rejected | covered |
| AT-OPS-4-3 | OPS-4 | Given PIIGHOST_RATE_LIMIT=minute:1, when a client sends two requests within the minute, then the second is refused | 200 then 429, /health stays at 200 | piighost-api:tests/test_app.py::test_rate_limit_throttles_second_request | covered |
| AT-OPS-5-1 | OPS-5 | Given an NER detector with max_chars=20 and auto_chunk=True, when "SECRET" comes after 96 characters, then it is detected | A detection whose span reads "SECRET" in the original text | tests/components/detector/ner/test_base.py::TestMaxChars::test_long_text_is_chunked_and_offsets_remapped | covered (fake NER detector derived from BaseNERDetector) |
| AT-OPS-5-2 | OPS-5 | Given max_chars=20 and auto_chunk=False, when a 200-character text goes through, then the detector refuses | TextTooLongError | tests/components/detector/ner/test_base.py::TestMaxChars::test_long_text_raises_when_auto_chunk_is_off | covered |
| AT-OPS-6-1 | OPS-6 | Given the reference catalog:piighost/notarial:2f602547, when you call load_pipeline on it, then the pipeline builds and runs | "write a@b.co" becomes "write <<EMAIL:1>>" | tests/config/test_catalog_config.py::TestLoadConfigFromTheCatalog::test_load_pipeline_builds_and_runs_it | covered (simulated catalog) |
| AT-OPS-6-2 | OPS-6 | Given the same pinned reference, when you call load_pipeline twice, then the second call reads the disk cache | A single HTTP request to the catalog | tests/test_catalog.py::TestPullConfig::test_a_pinned_configuration_is_fetched_once | partial (the cache is checked on pull_config, not end to end through load_pipeline) |
| AT-OPS-6-3 | OPS-6 | Given a reference without a commit, when you load it twice, then the catalog is queried twice | Two requests, nothing in cache | tests/test_catalog.py::TestCache::test_a_moving_selector_is_never_cached | covered |
| AT-OPS-6-4 | OPS-6 | Given a file:// catalog, or a catalog that serves a model detector, when you pull a reference, then it is refused | CatalogUrlError or CatalogPayloadError | tests/test_catalog.py::TestPull::test_a_catalog_that_is_not_http_is_refusedtests/test_catalog.py::TestPull::test_a_model_detector_is_refused | covered |
| AT-OPS-7-1 | OPS-7 | Given an in-process memory built with no setting, when 10,001 conversations are written, or a conversation stays a day without a message, then the oldest or the idle one is forgotten | The first conversation no longer returns its detections | tests/conversation_memory/test_in_memory.py::TestBounding::test_the_default_store_caps_its_thread_counttests/conversation_memory/test_in_memory.py::TestBounding::test_the_default_store_expires_an_idle_threadtests/config/test_thread_pipeline.py::TestInMemoryConfig::test_a_bare_section_is_bounded_like_the_store | covered |
User (of the application built with piighost)
| Id | Need | Scenario | Expected result | Existing tests | Status |
|---|---|---|---|---|---|
| AT-USER-1-1 | USER-1 | Given a LangChain, Pydantic AI or LlamaIndex application, or one that goes through a proxy, when the model replies "Hello <<PERSON:1>>", then the user reads the real value | "Hello Jean Dupont", never "Hello <<PERSON:1>>" | tests/integrations/langchain/test_middleware.py::TestWhenInstalled::test_after_model_deanonymizes_for_displaytests/integrations/test_pydantic_ai_hooks.py::TestAroundTheModel::test_model_sees_placeholders_and_reply_is_restoredtests/integrations/llama_index/test_query_engine.py::TestQuery::test_anonymizes_query_and_restores_answerpiighost-api:tests/routes/test_openai_chat.py::test_upstream_sees_tokens_reply_is_restoredpiighost-api:tests/routes/test_anthropic_messages.py::test_upstream_sees_tokens_reply_is_restored | covered (the LlamaIndex test is skipped without the extra) |
| AT-USER-1-2 | USER-1 | Given Claude Code with the piighost hooks, when the assistant replies with a placeholder, then the user reads the real value | The displayed reply contains no placeholder | none | missing (known limit, handle_hook handles only UserPromptSubmit, PreToolUse and PostToolUse) |
| AT-USER-2-1 | USER-2 | Given a conversation that mentions "Emma" then "Emma and Liam", when the model replies "Thanks <<PERSON:1>> and <<PERSON:2>>.", then each person keeps their placeholder and their value | "<<PERSON:1>> and <<PERSON:2>>" then "Thanks Emma and Liam." | tests/pipeline/test_thread.py::TestThreadConsistency::test_a_new_value_gets_the_next_numbertests/pipeline/test_thread.py::TestDeanonymize::test_restores_tokens_the_pipeline_never_anonymized | covered (no deterministic test can check that the model does not confuse two placeholders) |
| AT-USER-3-1 | USER-3 | Given a LangChain or Pydantic AI agent, when the assistant sends an e-mail to "<<EMAIL:1>>", then the tool receives the real address | The tool receives "patrick@acme.com" | tests/integrations/langchain/test_middleware_e2e.py::test_second_model_call_never_sees_clear_tool_argstests/integrations/test_pydantic_ai_hooks.py::TestTools::test_tool_gets_the_value_and_its_result_is_reanonymized | covered |
| AT-USER-3-2 | USER-3 | Given an OpenAI or Anthropic proxy, when the model returns a tool call with {"who": "<<PERSON:1>>"}, streamed or not, then the client receives the restored argument | {"who": "Patrick"} in all four cases (OpenAI and Anthropic, streamed and not) | piighost-api:tests/routes/test_rewrite.py::test_deanonymize_chat_response_restores_content_and_tool_argspiighost-api:tests/routes/test_anthropic_shape.py::test_restorer_restores_tool_input_json_safelypiighost-api:tests/routes/test_anthropic_messages.py::test_tool_result_anonymized_and_tool_use_restored | partial (the streamed OpenAI case is neither restored nor tested) |
| AT-USER-4-1 | USER-4 | Given the stream decoder, when "Hi <<PER" arrives, then only "Hi " is emitted, then "SON:1>>!" emits "Emma!" | No emitted chunk contains "<<PER" | tests/components/placeholder/test_streaming_async.py::TestFeed::test_a_split_token_is_held_then_completedtests/components/placeholder/test_streaming.py::TestFeed::test_split_delimiter_is_held | covered (a stream cut in the middle of a placeholder returns "<<PER" at the end, by choice, see tests/integrations/test_deidentify_stream.py::test_trailing_incomplete_token_is_flushed) |
| AT-USER-4-2 | USER-4 | Given a streaming proxy, when the upstream cuts "<<PERSON:1>>" into two events, then the client never receives a fragment | The received stream contains "Patrick" and not "<<PERSON" | piighost-api:tests/routes/test_openai_stream.py::test_stream_restores_a_token_split_across_chunkspiighost-api:tests/routes/test_anthropic_messages.py::test_stream_restores_token_split_across_deltas | covered |
| AT-USER-5-1 | USER-5 | Given "Paris" in the allow list, when "Visit Paris" goes through, then the city stays readable | "Visit Paris" | tests/pipeline/test_override_integration.py::TestBasePipelineOverride::test_allow_list_keeps_a_false_positive_in_clear | covered |
| AT-USER-5-2 | USER-5 | Given a generic catalog group, when "Meeting on 12/03/2026" goes through, then the date stays in clear | "Meeting on 12/03/2026" | none | missing (depends on the content of the catalog groups, never loaded in the library tests) |
| AT-USER-6-1 | USER-6 | Given a detector that misses "Acme", when the user adds the detection, then "Acme rocks" leaves as a placeholder and is restored | "<<ORG:1>> rocks", restored to "Acme" | tests/pipeline/test_thread_hitl.py::TestAnonymizeCorrected::test_adds_a_missed_valuetests/pipeline/test_thread_hitl.py::TestAnonymizeCorrected::test_added_value_is_deanonymizable_thread_wide | covered |
| AT-USER-6-2 | USER-6 | Given "Paris" wrongly masked, when the user removes the detection, then this message leaves in clear and the following ones do not change | "Visit Paris", then "Go to <<LOCATION:1>>" | tests/pipeline/test_thread_hitl.py::TestAnonymizeCorrected::test_drops_a_false_positivetests/pipeline/test_thread_hitl.py::TestAnonymizeCorrected::test_correction_is_local_to_the_message | covered |
| AT-USER-6-3 | USER-6 | Given a deny list that contains the value, when the user removes it, then it stays masked | A placeholder, not the value | tests/pipeline/test_override_integration.py::TestThreadPipelineOverride::test_deny_list_trumps_a_hitl_droptests/pipeline/test_override_integration.py::TestThreadPipelineOverride::test_allow_list_trumps_a_hitl_add | covered |
Summary
| Profile | Covered | Partial | Missing | Total |
|---|---|---|---|---|
| DPO | 14 | 1 | 2 | 17 |
| DEV | 16 | 2 | 0 | 18 |
| OPS | 10 | 3 | 0 | 13 |
| USER | 9 | 1 | 2 | 12 |
| Total | 49 | 7 | 4 | 60 |
Missing
Ranked by risk, starting with those that can cause a leak.
- AT-USER-3-2 (partial): tool arguments restored in the OpenAI stream. Known functional defect, the tool receives a placeholder.
- AT-DPO-2-1 (partial): a real catalog group (IBAN, NIR) masks its values. Only a simulated catalog is loaded.
- AT-DPO-10-1 (missing): the storage form of human corrections, feature to design.
- AT-OPS-3-1 (partial): the refusal without secrets is checked through
load_thread_pipeline, not atpiighost-apistartup. - AT-OPS-1-1 (partial):
serveon a catalog reference thenPOST /v1/anonymizein a single test. - AT-DEV-6-2 (partial):
piighost validateon a misspelled key, with the key named in the message. - AT-DEV-1-2 (partial): the
openaiSDK with onlybase_urlchanged. - AT-OPS-6-2 (partial): the disk cache seen through
load_pipeline. - AT-USER-5-2 (missing): a date is not masked by a generic group. Depends on the content of the groups, to test in the catalog's repository,
piighost-hub. - AT-USER-1-2 (missing): displayed reply restored under Claude Code, impossible as long as no hook rewrites the reply.
- AT-DPO-8-1 (missing): content of the DPIA page, checkable by a documentation test.
Tested behaviors missing from the inventory
| Behavior | Tests that guard it | Proposed need or criterion |
|---|---|---|
| A placeholder typed by the user does not restore another person's value | tests/components/anonymizer/test_span_anonymizer.py::TestAnonymize::test_user_typed_token_cannot_hijack_a_restoretests/components/anonymizer/test_span_anonymizer.py::TestDeanonymize::test_no_prefix_collision_when_one_token_prefixes_another | Criterion of DPO-1 : "a message that contains <<PERSON:2>> does not reveal the second person's value" |
| A value typed with no-break spaces stays masked and keeps its placeholder | tests/pipeline/test_pipeline.py::TestUnicodeSpaces::test_a_regex_value_typed_with_no_break_spaces_is_hiddentests/pipeline/test_thread.py::TestUnicodeSpaces::test_a_value_keeps_its_token_across_spacings_and_messages | Criterion of DPO-1 |
| Two overlapping detections leave no fragment in clear | tests/pipeline/test_pipeline.py::TestAnonymize::test_overlapping_detections_are_resolved_by_defaulttests/pipeline/test_pipeline.py::TestMergeOverlap::test_the_merge_resolver_leaves_no_fragment_of_a_longer_span | Criterion of DPO-1 |
| The history, block content and tool arguments already passed leave again as placeholders | tests/integrations/test_pydantic_ai_hooks.py::TestMultiTurn::test_prior_turn_pii_is_not_leaked_to_the_modeltests/integrations/langchain/test_middleware.py::TestWhenInstalled::test_before_model_anonymizes_block_contenttests/integrations/langchain/test_middleware.py::TestToolCalls::test_before_model_reanonymizes_clear_tool_call_args | Criterion of DPO-1 : "no later turn sends a value in clear" |
The text analyzed by LLMDetector cannot leave its data region (injection) | tests/components/detector/test_llm.py::TestDetect::test_a_data_tag_in_the_text_cannot_close_the_data_region | Watch point of DPO-1 |
Tracing in clear, or storing in clear in Redis or a networked SQL database, emits PIIGhostSecurityWarning | tests/observation/test_pipeline_spans.py::TestClearTextTracingWarning::test_warns_without_a_redactor_or_acknowledgmenttests/conversation_memory/test_redis.py::TestRedisPlaintext::test_round_trips_without_cryptotests/conversation_memory/test_sqlalchemy.py::TestWarning::test_networked_dialect_without_crypto_warns | Criterion of OPS-3 |
| With keys configured, a protected route without a Bearer token responds 401 | piighost-api:tests/test_auth.py::test_protected_route_401s_without_bearer_when_auth_enabled | Criterion of OPS-4 |
| The OpenAI proxy forgets the ephemeral conversation of a request without a conversation header | piighost-api:tests/routes/test_openai_chat.py::test_ephemeral_thread_is_forgotten | Criterion of DPO-6 , for data minimization |
piighost anonymize de-identifies an argument or standard input without a config file | tests/cli/test_cli.py::TestAnonymize::test_default_detector_anonymizes_an_argumenttests/cli/test_cli.py::TestAnonymize::test_reads_stdin_on_dash | a developer need to create (de-identify without a configuration file) |
| The middleware refuses a pipeline whose placeholders are not recognizable | tests/integrations/langchain/test_middleware.py::TestFactoryContract::test_a_pipeline_without_a_recognizer_is_refused | Criterion of DEV-1 |
BridgeDetector reads the positions in the runner's unit (UTF-16 on the browser side) | tests/components/detector/ner/test_bridge.py::TestDetect::test_offsets_are_read_in_the_runner_unit | a developer need to create (local detection in the browser) |