--- type: workflow title: Show a streamed reply description: How piighost restores a model reply streamed as it arrives, holds back a placeholder cut between two chunks until it is whole, and what stays on screen when the stream is interrupted. tags: [streaming, deanonymize, stream-decoder, invented-placeholder, langchain] sources: - id: openwiki-source-219ef8159700bea2d8181beb resource: repo://src/piighost/components/placeholder/streaming.py - id: openwiki-source-60f405cf9fd8c0cba8a61889 resource: repo://src/piighost/integrations/_deidentify.py - id: openwiki-source-85881a85af445f438a8d7d5f resource: repo://src/piighost/integrations/langchain/middleware.py generated: { by: "claude-code", at: "2026-10-02T18:00:00.000Z" } --- # Show a streamed reply ## In short - The model sends its reply in chunks, which the application shows as they arrive. - A placeholder can arrive cut between two chunks, for example "`<>`". - The `piighost` stream decoder holds back the start of the placeholder until it is whole, then restores it once. - Without this decoder, the user reads the placeholder on screen. - If the stream is cut in the middle of a placeholder, the start of that placeholder stays on screen. It contains no real value. Needs covered, described in [Needs by profile](../needs-by-profile.md): - Developer: DEV-8, DEV-9 - Application user: USER-1, USER-4 The terms are defined in the [glossary](../glossary.md). The restoration of a whole reply is described in [Follow a conversation and restore the reply](follow-a-conversation.md). ## For the business `piighost` has no screen. What you can observe is the text that appears while the model replies. The development team must plug in the stream decoder. ### Who is involved | Actor | Role | |---|---| | The end user | reads the reply while it is being written | | The model | sends its reply in chunks | | The application | reads the stream and passes each chunk through the decoder | | `piighost` | restores the placeholders chunk by chunk | ### The path of a streamed reply ```mermaid flowchart TD A["Chunk sent by the model"] --> B{"Unclosed start of placeholder?"} B -- no --> C["Text shown at once"] B -- yes --> D["Start held back"] D --> E["Next chunk"] E --> F["Complete placeholder restored"] F --> C ``` Example: the conversation maps `<>` to Jean Dupont and `<>` to jean.dupont@exemple.fr. | Chunk received | Text shown | |---|---| | "Hello <>, I am writing" | "Jean Dupont, I am writing" | | "to you at <>." | "jean.dupont@exemple.fr." | The user read "Hello Jean Dupont, I am writing to you at jean.dupont@exemple.fr." without waiting for the end of the stream. **How to check**: have the model reply with a name known to the conversation. The screen must never show "`<>." gives | |---|---| | Refuse (default) | "Hello", then the stream is interrupted by an error | | Drop | "Hello ." | | Keep | "Hello `<>`." | **BR-STREAM-07.** When the application restores each chunk separately, without the decoder, then a cut placeholder is never recognized, and the user reads "Hello `<>`.". **BR-STREAM-08.** When the reply goes through a proxy of the `piighost-api` server, then the proxy also restores the stream with this decoder. The OpenAI proxy restores only the text, not the tool arguments. Neither proxy, OpenAI nor Anthropic, applies the invented placeholder setting. ### What the end user sees A reply that is written as it arrives, with the real values. A slight delay appears when a placeholder or a "`<<`" is in progress. Only an interrupted stream leaves a fragment of a placeholder at the end. ### Frequently asked questions **The screen shows `<>` during the stream.** The application does not pass the chunks through the decoder (BR-STREAM-07). Ask the development team to plug it in. **The reply ends with "`<