ShengShu Technology, a global leader in multimodal generative AI, today introduced Vidu Q2 “Reference-to-Video”, a breakthrough feature that moves AI video creation from simple motion generation to realistic, expressive performance.
The new Reference-to-Video feature allows creators to generate consistent, lifelike videos using up to seven reference images for faces, gestures, scenes, or props. Multiple unrelated elements—such as characters, items, or backgrounds—can be blended seamlessly into a single video. Using a text prompt and the platform’s Multiple-Entity Consistency capability, each element remains visually accurate, even in complex or changing scenes. With faster generation times and more affordable pricing, Vidu Q2 now competes with leading global video platforms, ushering in the “AI performance era.”
“Vidu Q2 ‘Reference-to-Video’ marks a new chapter in AI video creation,” said Yihang Luo, CEO of ShengShu Technology. “This launch is about teaching AI to act and tell stories alongside creators, capturing human expression and cinematic flair.”
Realism and Cinematic Quality
Vidu Q2 captures subtle emotions—hesitant smiles, curious looks, or tense anticipation—with natural motion. Movements are smooth and alive, replacing robotic motions with vibrant energy. Cinematic techniques like camera panning, depth of field, and seamless transitions between wide shots and close-ups allow creators to produce engaging, professional-quality storytelling.
The platform also understands prompts more accurately, letting creators focus on storytelling rather than corrections. Longer, expressive videos for film, animation, advertising, and commercial content can now be produced with minimal adjustments.
Global Adoption and Commercial Value
Alongside Vidu Q2, the Vidu Q2 MaaS API is now globally available, enabling businesses to integrate Reference-to-Video capabilities into workflows. Advertising and e-commerce companies have already adopted the technology, benefiting from higher efficiency, reduced costs, and improved creative output.
The platform ensures high consistency in commercial video production. Even with complex camera movements or multi-character interactions, subjects and product details remain stable, delivering lifelike 360-degree visuals. AI-generated models now display natural gestures and micro-expressions, producing realistic and engaging footage for advertising and marketing.
Innovation and Platform Evolution
Since its founding, ShengShu Technology has led creative AI innovation. Previous breakthroughs include the U-ViT architecture, the DiT architecture adopted by top competitors, UniDiffuser for joint text-image generation, and the Analytic-DPM framework for faster AI processing. The Vidu series evolved as follows:
- Vidu 1.5: First consistent multi-character scenes.
- Vidu 2.0: 10-second videos at half industry cost.
- Vidu Q1: Cinematic transitions with realistic sound.
- Vidu Q2: Combines all advancements for expressive AI performance.
Rapid Global Growth
Since Vidu’s launch in April 2024, ShengShu Technology has expanded to over 200 countries, reached 30 million users, and produced over 400 million videos. “With each release, we blend technology and creativity more closely,” said Luo. “Our goal is not to replace creativity but to expand it, making imagination visible and emotions limitless.”
For the latest cybersecurity and technology updates, visit Itech360Hub.
