OpenAI-compatible proxy
piighost-api serves an OpenAI-compatible proxy under /openai/v1. Point a client's base_url at it, and the proxy de-identifies each request, forwards it to the real provider, and restores the reply. The provider receives <<PERSON:1>>, never Jane Doe.
Point the client at the proxy
Change only the base_url. The api_key stays the provider's key.
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8000/openai/v1",
api_key="sk-...",
)
response = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[
{
"role": "user",
"content": "Write a short greeting to Jane Doe, jane.doe@example.com.",
}
],
)
print(response.choices[0].message.content)The provider receives <<PERSON:1>> and <<EMAIL:1>>, and the printed reply carries Jane Doe again. The server's API_KEY_ keys do not apply to /openai/v1, because the proxy asks for no server key of its own. It relays the Authorization header to the provider as is.
The same call with curl:
curl http://127.0.0.1:8000/openai/v1/chat/completions \
-H "Authorization: Bearer sk-..." \
-H "Content-Type: application/json" \
-d '{"model": "gpt-5.6-terra", "messages": [{"role": "user", "content": "I am Jane Doe"}]}'Choose the provider
Without a header, the proxy forwards to https://api.openai.com/v1. If you want another OpenAI-compatible provider for every client, set PIIGHOST_OPENAI_UPSTREAM before starting the server:
export PIIGHOST_OPENAI_UPSTREAM="http://vllm.internal:8000/v1"If you want it for one client only, name the provider's base URL in the X-PIIGhost-Upstream header:
client = OpenAI(
base_url="http://127.0.0.1:8000/openai/v1",
api_key="sk-...",
default_headers={"X-PIIGhost-Upstream": "http://vllm.internal:8000/v1"},
)The proxy strips every X-PIIGhost-* header before forwarding, so the provider never sees it.
Keep a thread across requests
Each request runs in a fresh thread, forgotten as soon as the reply is restored. A chat client resends the whole history every turn, so the placeholder numbering stays consistent within each request. If you want the thread to outlive the request, for example to restore a stored reply later through /v1/deanonymize, pin it with X-PIIGhost-Thread-Id:
response = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "I am Jane Doe"}],
extra_headers={"X-PIIGhost-Thread-Id": "user-42"},
)A pinned thread stays in the server memory until DELETE /v1/threads/user-42 erases it.
Stream the reply
stream=True works unchanged. The proxy restores each placeholder as the chunks arrive, even when the provider splits <<PERSON:1>> across two chunks.
stream = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "I am Jane Doe"}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="")The routes the proxy serves, and the fields it de-identifies on each, are listed in API endpoints.
See also
- Anthropic-compatible proxy: the same relay for the Messages API and Claude Code.
- Server CLI: every environment variable of the server.
- Tool-call strategies: how the library handles tool arguments when it runs inside the agent.