2026年9月10日

Going Head-to-Head with Sora 2: Google’s Veo 3.1 Brings Some Surprising Upgrades

On October 16 (Beijing time), Google officially unveiled Veo 3.1 and its Veo 3.1 Fast paid preview v...

On October 16 (Beijing time), Google officially unveiled Veo 3.1 and its Veo 3.1 Fast paid preview version within the Gemini API, instantly drawing attention across the tech world. Much like OpenAI’s recently released Sora 2, Veo 3.1 introduces an exciting new capability — audio integration, marking another leap in AI-generated video realism.

Compared to its predecessor, Veo 3.1 focuses on three key upgrades.

First, AI-generated videos can now move beyond silent imagery into the world of synchronized sound. Veo 3.1 allows creators to align visuals and audio seamlessly — not just syncing sound with motion, but also ensuring that the background music fits the emotional tone and rhythm of the scene.

Second, Veo 3.1 introduces the ability to define opening and closing frames, enabling smoother transitions between short clips and tighter control over video continuity. Even more impressively, it can generate new clips that continue directly from the final frame of the previous one — effectively simulating AI-generated long-form videos through continuous chaining.

Third, Veo 3.1 now supports character creation from just three reference images — a headshot, a clothing example, and a scene setup. Using these, the AI can construct a consistent character who moves naturally and even speaks lines according to the prompt.

In essence, Veo 3.1 represents Google’s ongoing effort to refine the cinematic experience of AI-generated video — focusing on visual stability, storytelling continuity, and synchronized soundscapes.

Veo 3.1 is now accessible for free trials through Google’s Gemini app and Flow platform, though with limited availability. Unsurprisingly, Chinese AI platforms such as Imagine.art, Fal-ai, and Lovart were quick to integrate the model. Our tests on Lovart provided firsthand insight into what’s new — and what still needs work.

When prompted with: “It’s raining on the streets of New York. Suddenly, a lightning bolt strikes with thunder,” Veo 3.1 successfully synchronized the thunderclap with the lightning flash. Even the subtle sound of cars splashing through puddles shifted naturally from near to far.

However, while the generation process was quick (about one minute), the resulting clips were short — roughly six seconds, compared to Sora 2’s ten to twenty seconds, limiting creative flexibility. We also noticed limited motion within scenes — vehicles, rain, and lightning were dynamic, but pedestrians and trees remained static, making the AI origin rather obvious.

Next, we tested Veo 3.1’s frame control. By providing two photos as the start and end of a video — featuring a tabby cat jumping onto a desk — Veo generated a mostly smooth motion. Yet, midway, a second “magical” cat appeared, disrupting realism. Even so, when we stitched two related clips together, the transition between them remained decently consistent, showing Veo’s growing potential for scene continuity.

Lastly, we tried generating a character using three reference images: one for the woman’s face, one for clothing, and one for the environment. While the idea was promising, the results showed heavy “AI modeling” artifacts — characters appeared stiff, and visual details failed to fully match the reference images.

Overall, Veo 3.1 showed meaningful progress in sound synchronization and frame stability but fell short in realistic character creation.

Google, however, remains confident. The company’s official site boldly claimed that Veo 3.1 outperformed competitors like Sora 2 Pro, Hailuo 2.0, Seedance 1.0 Pro, and Renway Gen 3 in overall visual quality and alignment. Interestingly, Google also “subtly” pointed out that Sora 2 Pro was excluded from comparison due to its lack of portrait generation capability.

Yet, many AI experts disagreed. Otherside AI’s founder Matt Shumer expressed disappointment, saying Veo 3.1 lagged behind Sora 2 in performance while being significantly more expensive — especially since Sora 2 remains free to use. 3D artist Travis David noted that Veo 3.1 failed to break the “eight-second limit” of AI video generation and that users couldn’t customize audio, a letdown for many creators. Others lamented the absence of the long-awaited “automated storyboarding” feature.

Pricing reveals Google’s strategic direction. While Veo 3.1 maintains the same price as Veo 3, it remains one of the costlier models on the market, second only to Sora 2 Pro. The faster Veo 3.1 Fast version can generate clips more quickly and at lower cost — $0.15 per second without audio, or $0.40 per second with audio.

Google also acknowledged potential instability, noting that users would only be charged for successfully generated videos, implying that audio rendering still has reliability issues.

Compared with Sora 2’s social and entertainment focus, Veo 3.1 clearly aims for professional-grade video production. Google highlights its potential use in studios and narrative storytelling. Promise Studios, for example, is already using Veo 3.1 in its MUSE platform to achieve cinematic-level AI filmmaking, while Latitude is testing it for instant story visualization.

In short, Veo 3.1 represents Google’s push toward professional AI video creation — offering better sound synchronization, frame continuity, and storytelling tools. Yet despite these improvements, it’s fair to say that after five months, Google’s Veo has advanced just “0.1 steps” forward.

接著讀