Open Source · AI Data Privacy

Dataiku Launches Kiji Privacy Proxy™ — An Open-Source Privacy Layer to Safeguard Sensitive Data in the Age of Generative AI

iTech360Hub | 5 min read | New Launch

Every time a team member types a prompt into a generative AI service — ChatGPT, Claude, or any other large language model — sensitive data travels to external servers. For casual queries that's fine, but in enterprise settings those prompts frequently contain customer names, email addresses, social security numbers, medical records, financial details, and internal business data that should never leave the organisation's environment. Dataiku's answer is Kiji Privacy Proxy™, a newly launched open-source privacy layer that ensures personally identifiable information never leaves an organisation's control — even when using third-party AI services.

This is not a hypothetical risk. Regulations like GDPR, HIPAA, and CCPA impose real penalties on organisations that fail to protect personal data. A 2026 study of 600 CIOs found that 85% have seen AI projects delayed or blocked entirely due to gaps in traceability, explainability, and privacy — making data protection one of the single biggest blockers to enterprise AI adoption today.

85%
of CIOs report AI projects blocked by privacy concerns
16+
PII types automatically detected and masked
Zero
code changes needed — works as a transparent proxy

"Open source isn't just a distribution model — it's a trust model. As AI systems become more autonomous and more consequential, enterprises need tools they can inspect, verify, and adapt. By building these foundations in the open, we're helping teams to manage risk and use AI responsibly."

— Hannes Hapke, Director, 575 Lab at Dataiku

The Problem: Generative AI's Enterprise Privacy Blind Spot

Until now, enterprises facing this challenge have had three deeply unsatisfying options: avoid external AI services entirely and forfeit the productivity gains; invest in expensive, bespoke privacy infrastructure; or accept the exposure risk and proceed without proper safeguards. None of those options is good enough for organisations operating in regulated industries or handling sensitive customer data at scale.

Kiji Privacy Proxy™ is built to make all three compromises obsolete. Operating as a transparent gateway sitting directly within the organisation's network, it automatically identifies and redacts personally identifiable information before any data is transmitted — allowing teams to leverage the full power of generative AI without trusting third-party servers with sensitive information.

How Kiji Privacy Proxy™ Works

The mechanism is elegant in its simplicity. Kiji sits between local applications and external AI APIs, intercepting every outbound request and running it through an ML-powered PII detection model before anything leaves the network. The full four-step flow works as follows:

1

Your app sends a request to the Kiji Privacy Proxy. Kiji intercepts it and runs the content through its ML-powered PII detection engine before forwarding anything externally.

2

Sensitive data is replaced with realistic dummy values — emails, phone numbers, SSNs, credit card numbers, IP addresses, and 16+ other PII types are masked before the request leaves your network.

3

The masked request goes to the AI API — OpenAI, Anthropic, or any other external service. The model processes fully anonymised data and returns a response.

4

Original values are restored in the response before it reaches your application. The AI model never saw the real data, but your application behaves exactly as though nothing changed.

Built by the 575 Lab — Dataiku's Open Source Office

Kiji Privacy Proxy™ is a flagship release from the 575 Lab, Dataiku's dedicated Open Source Office, established to translate a decade of enterprise AI experience into reusable, community-driven governance tools. Both the training data and the underlying PII detection model are fully open on GitHub and HuggingFace, making the entire stack inspectable and auditable by any organisation.

Critically, Kiji is also designed for deep customisation. Organisations requiring domain-specific PII detection — medical record formats, industry-specific identifiers, or non-English PII patterns — can train their own models using Label Studio for data labelling and Outerbounds' Metaflow infrastructure for orchestrating training pipelines at scale. The full customisation workflow was demonstrated at a cost of just $50 using open-source inference, illustrating the practical economics of the open ecosystem approach.

On macOS, Kiji runs as a native desktop application with automatic proxy configuration — no environment variables or code changes required. A Chrome extension is also available to route browser-based AI requests through the proxy automatically, making adoption frictionless for everyday enterprise users as well as developers.

"Enterprises are building increasingly complex agentic ecosystems. To make them safer to use, they need reusable building blocks that can become the standards for how agentic systems are controlled and inspected. The 575 Lab is contributing to open source to foster the community from which those standards will emerge."

— Florian Douetteau, CEO & Co-Founder, Dataiku

Part of a Broader Open Trust Infrastructure for Enterprise AI

Kiji Privacy Proxy™ sits alongside Kiji Inspector™, another 575 Lab release — one of the first open-source explainability frameworks purpose-built for enterprise AI agents, with initial support for NVIDIA Nemotron open models. Together, these tools address the two most pressing enterprise AI trust concerns: keeping data private and keeping AI decisions explainable and auditable.

Dataiku is a member of both the Linux Foundation and the newly formed Agentic AI Foundation, signalling a deliberate strategy of building AI governance standards through open community collaboration rather than proprietary lock-in. With over 750 enterprise customers — trusted by 1 in 4 of the world's top companies — Dataiku is positioning this open infrastructure as the foundation on which responsible, governed AI can be scaled reliably.

PII Types Automatically Detected
Email Addresses Phone Numbers Social Security Numbers Credit Card Numbers IP Addresses Medical Record Identifiers + 10 More PII Types

Key Takeaways

1

Kiji Privacy Proxy™ ensures PII never leaves an organisation's control — even when using external AI services like OpenAI or Anthropic — with zero code changes required.

2

An ML-powered detection engine automatically identifies and replaces 16+ PII types with realistic dummy values before any data is transmitted externally.

3

The tool is fully open-source — model, training data, and architecture available on GitHub — with support for custom domain-specific PII models for regulated industries.

4

Kiji is part of Dataiku's broader open trust infrastructure strategy via the 575 Lab, alongside Kiji Inspector™ for AI agent explainability. Learn more at dataiku.com.

The launch of Kiji Privacy Proxy™ represents a meaningful shift in how enterprises can approach generative AI adoption. Rather than choosing between data privacy and capability, organisations now have an open, inspectable, and customisable tool that allows them to protect both. As AI systems grow more autonomous and deeply embedded in critical business operations, foundational trust infrastructure like Kiji will be essential to scaling responsibly.

Developers, data scientists, and AI specialists can get involved with the project and join the contributor community at the official 575 Lab open-source page.

Tags
Data Privacy Open Source AI Generative AI PII Protection Enterprise AI AI Governance GDPR Compliance 575 Lab