Veritone Data Refinery Helps Safeguard Personal Data with Veritone Redact
5 min read
Veritone, Inc. (NASDAQ: VERI), a leader in building enterprise AI and data solutions, has announced a significant advancement in its commitment to privacy-first AI. By deploying Veritone Redact alongside Veritone Data Refinery (VDR), personally identifiable information (PII) and other sensitive data are now automatically removed before processing and refinement — enabling VDR to convert unstructured data into AI-ready assets while protecting intellectual property and the rights of data owners throughout the pipeline.
The Problem: AI Training Data Is Scaling Faster Than Governance
According to the Stanford HAI 2025 AI Index Report, citing Epoch AI research, AI training datasets are doubling every eight months as model scale continues to grow — increasing the volume of data flowing into AI systems at an unprecedented rate. This acceleration has raised pressing concerns about the legal and ethical risks embedded in AI training pipelines.
The scale of the problem is underscored by a revealing audit conducted by arXiv — a platform supported by Cornell University — of more than 1,800 text datasets. It found frequent miscategorisation of licences on widely used dataset hosting sites, with licence omission exceeding 70%. The race to train models on vast datasets has outpaced the governance frameworks needed to ensure that data is clean, properly licensed, and free of sensitive personal information before it enters AI systems.
This governance gap is precisely what the combination of Veritone Data Refinery and Veritone Redact is designed to close — ensuring AI-ready data is clean from the outset, meeting strict industry compliance and privacy standards while enabling a broader range of companies to innovate responsibly.
How Veritone Redact Works Within the Data Refinery Pipeline
Veritone Redact is an industry-leading automated redaction application with deep roots in public safety and law enforcement — where it is used by agencies including the Department of Justice and state and local police departments to redact sensitive information from audio, video, and image-based evidence. The application significantly reduces manual redaction processes while increasing accuracy, minimising errors, and helping agencies meet critical deadlines.
Recent enhancements to Redact — including AI-powered voice masking, inverse blur, and transcription capabilities in 64 languages — have significantly broadened the application's reach, addressing critical privacy, compliance, and productivity needs across legal, law enforcement, and corporate environments simultaneously.
Integrated into the Veritone Data Refinery pipeline, Redact automatically strips PII and other sensitive data from unstructured content before that content is refined and packaged as AI-ready training assets. This sequence — redact first, refine second — ensures that the downstream datasets flowing to enterprise AI systems and hyperscalers are compliant, ethically sourced, and legally defensible from the point of origin.
"We are committed to helping data-driven organizations protect their valuable assets and help ensure that their data is used cleanly and ethically. We're proud to use our own proprietary, AI-enabled tool, Redact, which is traditionally used by our public sector customers, including the Department of Justice, and state and local police agencies, to help ensure PII is safeguarded before going through refinement. This is a prime example of our dedication to providing innovative solutions that the market needs while fostering a more responsible and ethical AI ecosystem for everyone."
— Ryan Steelberg, CEO, Veritone
Surging Demand: 3.5x Volume Growth in H2 2025
Market demand for compliant, ethically sourced AI-ready datasets is accelerating sharply. Veritone is seeing significant demand for VDR from both content owners and hyperscalers — with the volume of data processed through the platform increasing by 3.5 times in the second half of 2025 compared to the first half. This growth trajectory reinforces that the need for clean, privacy-preserving AI training data is not a niche concern — it is a mainstream enterprise and hyperscaler requirement that is only intensifying.
As regulatory scrutiny of AI training data continues to tighten globally — from the EU AI Act to emerging data governance frameworks across North America and Asia-Pacific — the ability to demonstrate that AI systems are trained on properly licensed, PII-free data will increasingly become a baseline compliance requirement rather than a differentiator. Platforms that embed this capability natively into the data pipeline, as Veritone has done, are well positioned to serve both sides of the market: content owners seeking to protect their IP, and AI developers seeking clean, compliant training data at scale.
A Model for Ethical AI Data Infrastructure
The integration of Veritone Redact into Veritone Data Refinery represents something broader than a product update — it is a demonstration of how ethical AI infrastructure can be engineered rather than simply declared. By taking a tool proven in some of the most legally sensitive environments imaginable — public safety, law enforcement, justice — and embedding it at the entry point of commercial AI data pipelines, Veritone is applying a standard of data hygiene to enterprise AI that the industry's rapid growth has made urgently necessary.
For enterprises and hyperscalers building or scaling AI systems, the message is clear: clean data is not an optional quality characteristic — it is the foundation on which responsible, legally defensible, and competitively durable AI is built.
Key Takeaways
- Veritone Redact is now integrated into the Veritone Data Refinery pipeline, automatically removing PII from unstructured data before it is refined into AI-ready training assets.
- AI training datasets are doubling every eight months according to Stanford HAI research, while a Cornell/arXiv audit found licence omission exceeding 70% across 1,800+ text datasets.
- Veritone Redact supports AI-powered voice masking, inverse blur, and transcription in 64 languages — capabilities proven with the Department of Justice and law enforcement agencies.
- VDR data processing volume grew 3.5x in H2 2025 vs H1 2025, reflecting surging demand from content owners and hyperscalers for clean, compliant AI training data.
- The integration positions Veritone as a privacy-first infrastructure provider for AI pipelines, helping enterprises and hyperscalers meet tightening global compliance standards from the point of data origin.
