UlazAI - AI Image & Video Tools
Now in Video Studio
Kling O3 gives you four ways into the same video model.
Kling O3, in full Kling 3.0 Omni, generates video from a prompt, from images, from references, or from an existing clip you want to rebuild. Name up to three subjects, call them in your prompt with @name, and write your own shots if you want the cutting to follow your script instead of the model's.
- Modes
- 4
- Duration
- 3-15s
- Resolution
- 720p / 1080p / 4K
- From
- 84 credits
84 credits buys a 5 second 720p clip without audio. Full rates are in the pricing table below.
Pick the mode by what you already have
Most video models give you one door: type something, get a clip. Kling O3 has four, and the difference between them is not quality but input. Nothing but an idea, a still image, a face you need to stay the same, or a clip you want rebuilt in another style. Below is each door, what it takes, and when it is the right one.
Text to video
A prompt and nothing else. You set duration, resolution, aspect ratio and whether the model writes audio, and Kling O3 builds the shot from scratch.
Choose it when the scene does not exist yet and you have no still or clip worth starting from. It is also the mode where audio is least restricted, since there is no reference video in play.
Open text to videoImage to video
One image becomes the first frame. Give it two images instead and the second one becomes the last frame, so the model has to get from A to B inside the duration you set.
Choose it when you already have the frame you want: a product shot, a still from a shoot, a rendered scene. Two frames is the cheapest way to control where the motion ends up.
Aspect ratio is auto by default here, so the output follows your input image.
Reference to video
This mode comes in three flavours, and you use one of them per job:
- Reference images only — up to 7 of them.
- A reference video only — exactly 1.
- A reference video plus images — 1 video and up to 4 images.
Choose it when something has to carry over: a face, a garment, a packaging design, the light and grade of an existing clip.
Both video flavours generate silently. A reference video means audio has to be off.
Open reference to videoTransformation
Exactly one source video goes in, optionally with up to 4 images alongside it, and comes back in a different style or a different filling.
Choose it when the take is already right. You shot the movement, the timing works, and only the look needs to change.
Transformation always runs on a source video, so audio is off here too.
Open transformationNamed subjects, called with @name
A character who changes face halfway through is the usual reason an AI clip falls apart. Named subjects fix that. You define up to 3 of them, and each one gets:
- A name you use in the prompt.
- A description in your own words.
- 1 to 4 images, or a single video of the character.
You can attach audio clips to a subject as well. And if you gave it a video, you can trim the usable part with a start and an end time, so a two second stretch of a longer clip is enough.
Subjects
@sarah— dark curly hair, olive coat. 3 photos.@kettle— matte black kettle, copper handle. 2 photos.
Prompt
@sarah walks into the kitchen, picks up @kettle and fills it at the tap. Morning light from the left, handheld camera.
Same Sarah, same kettle, every shot in the video.
Several shots in one video
You have two ways to handle cutting, and they exclude each other. Pick one.
Write the shots yourself
Up to 6 shots. Every shot gets its own prompt of at most 512 characters and its own duration between 1 and 15 seconds.
Use this when the order matters: a hook, a demo, a close-up, a logo. You decide where each cut lands and how long it holds.
Let the model divide them
One prompt, and Kling O3 works out how many shots the scene needs and how long each one runs.
Use this for a first pass, or when the clip is one continuous moment and you only care about the outcome.
Switching on your own shot list turns off the automatic split, and the other way round. There is no mode where both run.
What every mode accepts
The settings below are the full set. Anything not listed for a mode is not available in that mode.
| Setting | Text to video | Image to video | Reference to video | Transformation |
|---|---|---|---|---|
| Prompt | Required, 1-3072 characters | 1-3072 characters | 1-3072 characters | 1-3072 characters |
| Images in | None | 1 (first frame) or 2 (first and last frame) | Up to 7 on their own, or up to 4 next to a video | Up to 4, optional |
| Video in | None | None | Exactly 1, on its own or with images | Exactly 1, required |
| Duration | 3-15s, default 5 | 3-15s, default 5 | 3-15s, default 5 | 3-15s, default 5 |
| Resolution | 720p, 1080p, 4K | 720p, 1080p, 4K | 720p, 1080p, 4K | 720p, 1080p, 4K |
| Aspect ratio | 16:9, 9:16, 1:1 | auto by default, or 16:9, 9:16, 1:1 | auto with a video, otherwise 16:9, 9:16, 1:1 | auto, since there is always a video |
| Audio | Yes | Yes | Images only: yes. Any reference video: no | No |
| Named subjects | Up to 3 | Up to 3 | Up to 3 | Up to 3 |
| Shots | Up to 6 of your own, or automatic | Up to 6 of your own, or automatic | Up to 6 of your own, or automatic | Up to 6 of your own, or automatic |
Each named subject holds a name, a description and 1 to 4 images or one character video, plus optional audio clips and a start and end time for that video.
Pricing
Credits are charged per second of finished video. Audio costs extra, and any job with a video going in sits in its own column because that work is heavier.
| Resolution | Without audio | With audio | With video input |
|---|---|---|---|
| 720p | 16.8 credits/s | 21.6 credits/s | 24 credits/s |
| 1080p | 21.6 credits/s | 27.6 credits/s | 32.4 credits/s |
| 4K | 67 credits/s | 67 credits/s | 67 credits/s |
4K is one flat rate, sold at cost. Audio or video input changes nothing there.
720p, 5s, no audio
84 credits
720p, 5s, with audio
108 credits
720p, 10s, no audio
168 credits
1080p, 5s, no audio
108 credits
1080p, 5s, with audio
138 credits
4K, 5s
335 credits
Test at 720p, then rerun the prompt that worked at 1080p or 4K. A 5 second test costs 84 credits against 335 for the 4K version of the same mistake.
Kling O3 next to Kling 3.0
Kling 3.0 is not going anywhere. It stays in Video Studio and keeps its own page with its own modes and rates. O3 is the newer model in the same family, and the honest summary of the difference is: more ways in, and more control once you are in.
| What you get | Kling 3.0 | Kling O3 |
|---|---|---|
| Studio entries | One | Four, one per mode |
| Reference video in | Not available | Yes, in reference to video and transformation |
| Named subjects with @name | Not available | Up to 3, each with images or a character video |
| Shot control | Multi-shot storytelling from one prompt | Up to 6 shots you write yourself, each with its own prompt and duration |
| Output | See the Kling 3.0 page for its own rates | 3-15s at 720p, 1080p or 4K |
One thing O3 gives up: audio. Kling 3.0 can score any clip it makes. In O3 the moment a reference video is involved, audio is off, so a transformation always comes back silent.
Questions people ask first
What are the four Kling O3 modes?
Text to video works from a prompt alone. Image to video takes one image as the first frame, or two as the first and last frame. Reference to video takes up to 7 reference images, or exactly one reference video, or one reference video plus up to 4 images. Transformation takes exactly one source video, optionally with up to 4 images, and rebuilds it in another style or filling.
Which mode should I pick?
Pick by what you already have. Nothing but an idea: text to video. A still you want to move: image to video. A look, a product or a face you want carried over: reference to video. An existing clip whose motion you want to keep while the style changes: transformation.
Can I keep the same character or product across the whole video?
Yes. Add up to 3 named subjects. Each gets a name, a description and 1 to 4 images, or a single video of the character. You call it in your prompt with an at sign plus the name, for example @sarah. You can attach audio clips to a subject, and for a video character you can trim the usable part with a start and end time.
Why can I not turn on audio for my generation?
As soon as a reference video is part of the job, audio cannot be switched on. That covers reference to video with a video, reference video plus images, and transformation. Text to video, image to video, and reference to video from images only can generate audio.
How long can a Kling O3 video be?
3 to 15 seconds, and 5 seconds is the default. If you write your own shots, you get up to 6 of them and each shot has its own duration between 1 and 15 seconds.
Can I write the shot list myself?
Yes, up to 6 shots, each with its own prompt of at most 512 characters and its own duration of 1 to 15 seconds. The alternative is letting the model divide the shots. The two exclude each other, so you pick one.
What does a Kling O3 video cost?
Credits go per second. A 5 second 720p clip without audio is 84 credits, the same clip with audio is 108, and 10 seconds at 720p without audio is 168. At 1080p a 5 second clip is 108 credits without audio and 138 with audio. At 4K a 5 second clip is 335 credits.
Is Kling 3.0 still available?
Yes. Kling 3.0 stays in Video Studio and keeps its own page. Kling O3 is the newer member of the same family, split into four modes with named subjects and your own shot list on top.
Start at 720p and 5 seconds.
84 credits tells you whether the prompt holds up. Move the one that works to 1080p or 4K after that.