UlazAI - AI Image & Video Tools
How Veo 3 works, step by step
How Google Veo 3 works: the public flow, controls, and limits
The public Gemini API docs describe an async video generation flow, not a fully published internal architecture. You send a prompt, optionally add images or frames, poll the operation until it finishes, and download an 8-second 720p, 1080p, or 4k video with native audio. Not sure which route you have yet? See how to access Veo 3.
Quick answer
The current public docs confirm the generation flow and controls: prompt input, async job polling, 8-second output, native audio, portrait support, video extension, frame-specific generation, and up to three reference images. They do not publish the full internal model diagram, training pipeline, or a detailed transformer stack explanation in the API docs.
Need the public source first? Use the official Gemini API video docs.
Veo 3 guides
Choose the Veo guide you need
Find practical help with access, video length, prompts, pricing, API setup, and how Veo works.
Generation shape
8 seconds
The current public Veo 3.1 docs describe high-fidelity 8-second generation.
Output modes
720p, 1080p, 4k
The public docs list 720p, 1080p, and 4k output support, with native audio generation.
Control layer
Up to 3 images
Public image-based direction supports up to three reference images, plus first/last-frame workflows.
Job model
Async polling
You submit a generation request, poll the operation, and download the generated file when it is ready.
Veo 3 in Flow: the model, the interface and the scenes
Searches for “veo 3 flow” usually mix two different things. Veo is the video model that turns a prompt into a clip. Flow is Google's creative interface built on top of it, where you manage shots as scene cards instead of one-off generations.
The model
Veo 3 renders one clip per request, with native audio, through an async generate-and-poll API or a product UI. Everything technical in the public docs describes this layer.
The interface
Flow adds project-style tooling: prompt boxes per shot, extending a scene, reusing assets, and downloading finished clips. Flow calls the model under the hood; it does not change what the model can render.
What it means for you
If a feature is advertised in Flow, check whether it is a UI convenience (saved scenes, asset reuse) or a model behavior (duration, audio, reference images). The API only exposes the model behaviors.
On UlazAI the same split applies: the Video Studio is the interface layer, and the Veo docs describe the model layer.
Veo 3 model versions at a glance
The API accepts several model names. Usage differs mainly in speed, resolution ceiling and price, not in the request shape: every version takes a prompt, runs async, and returns a downloadable file.
| Model name | Positioning | Typical usage |
|---|---|---|
| veo-3 | The original Veo 3 model with native audio. | One-shot cinematic clips where quality beats latency. |
| veo-3-fast | Speed- and cost-optimized variant of Veo 3. | Drafting, prompt iteration and volume social clips. |
| veo-3.1 | Current flagship: better reference-image and first/last-frame handling. | Character work, multi-shot sequences and extensions. |
| veo-3.1-fast | Faster, cheaper variant of Veo 3.1. | Everyday generation where the 3.1 features matter but speed wins. |
Exact per-model credits and resolution options on UlazAI are listed in the pricing guide; the official model surface is documented by Google at the Gemini API video docs.
Public generation flow
This is the part of “how Veo 3 works” that the public docs actually show.
Step 1
Write the prompt and optional controls
Start with a text prompt, then optionally add reference images, first/last frames, portrait orientation, or extension settings depending on the workflow.
Step 2
Submit a long-running generation request
The official examples use async video generation calls. The generation does not return instantly as a finished inline response.
Step 3
Poll the operation until it is done
The public code samples repeatedly check the operation status. That is the documented flow for waiting on the finished video.
Step 4
Download the generated file
Once the job is complete, the examples download the resulting video file rather than treating the response as a synchronous final asset.
What the public docs confirm vs what they do not
| Topic | Publicly confirmed | Not publicly documented in detail |
|---|---|---|
| Generation flow | Prompt in, long-running operation, poll status, download output. | Internal scheduler design, cluster orchestration, or exact serving stack. |
| Capabilities | 8-second video, native audio, portrait mode, extension, frame-specific generation, up to three reference images. | A complete internal breakdown of which submodels handle audio, motion, or image conditioning. |
| Model internals | Google describes Veo publicly as a state-of-the-art video generation model. | The full transformer topology, weight layout, training corpus, and exact pipeline internals. |
| Prompt control | Prompting, reference images, first/last frames, and extension are publicly described control layers. | A complete public spec for every latent control or motion-planning subsystem. |
What “how it works” means in practice
Practical signal 1
Native audio is first-class
The public docs frame Veo 3.1 as a video model with natively generated audio, not as a silent clip generator that always needs a separate audio pass.
Practical signal 2
Reference images shape the result
Up to three reference images, plus first and last frame control, show that Veo is not just prompt-only text-to-video in the current public workflow.
Practical signal 3
Extension is built into the workflow
Video extension means the generation process can continue an earlier output instead of starting from zero every time.
Practical signal 4
The controls matter more than the model internals
For a usable result, focus on the prompt, reference image, aspect ratio, duration and audio settings. Google does not publish every internal implementation detail.
Generate a clip or open the guide you need
Use the API docs for integration, the prompt guide for better shot descriptions, pricing for a cost estimate, or the access guide to create an account.
Veo is the video model; Flow is one creation interface
Veo generates the video, while Flow organises filmmaking and scene work around supported Google models. They are related but not interchangeable product names, so access and controls can differ.
In UlazAI, choose the visible Veo mode, provide the supported prompt or references and review duration, audio and cost before generating. Treat unofficial “Flow character” field names as unverified unless they appear in the API contract.