Documentos para desarrolladores UlazAI
Matriz de modelo de vídeo
Matriz de capacidades del modelo de vídeo
Esta matriz proviene del registro de modelos de Video Studio. Úselo para validar las restricciones específicas del modelo antes de enviar solicitudes de generación.
| modelo | motor | Entradas | Relaciones de aspecto | Duraciones | Modos de calidad | Estimación de créditos |
|---|---|---|---|---|---|---|
|
Seedance 2.5
seedance_2_5
4-30 second generation in 480p/720p with text, first-frame, first+last-frame, or multimodal mode. Supports up to 30 image, 10 video, and 10 audio references plus MP4/MOV output.
|
Seedance 2.5 | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, 21:9, adaptive | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30s | 480p, 720p, 1080p | 480p 4 no video input: 112, 480p 5 no video input: 140, 480p 6 no video input: 168, 480p 7 no video input: 196, +77 more |
|
Veo 3.1 Lite
veo31_lite
Most cost-effective Veo 3.1 mode (40 credits) for text-to-video and image-to-video.
|
Veo 3.1 | texto , imagen | 16:9, 9:16, Auto | 8s | - | base: 40 |
|
Veo 3.1 Fast
veo31_fast
Fast 8-second generation with text-to-video and image-to-video.
|
Veo 3.1 | texto , imagen | 16:9, 9:16, Auto | 8s | - | base: 100 |
|
Veo 3.1 Quality
veo31_quality
Higher-fidelity Veo 3.1 output with the same 8-second duration.
|
Veo 3.1 | texto , imagen | 16:9, 9:16, Auto | 8s | - | base: 220 |
|
Kling 3.0
kling_3_0
Supports text, image, and frame-based generation with optional elements.
|
Kling 3.0 | texto , imagen | 16:9, 9:16, 1:1 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | std, pro, 4K | std no audio per second: 20, std audio per second: 30, pro no audio per second: 27, pro audio per second: 40, +1 more |
|
Kling 3.0 Motion Control
kling_3_0_motion_control
Requires exactly one image URL plus one motion reference video URL.
|
Kling 3.0 Motion Control | imagen , vídeo | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p per second: 20, 1080p per second: 27 | |
|
Kling V3 Turbo
kling_v3_turbo_t2v
Fast text-to-video route with 5s/10s durations and 720p/1080p output.
|
Kling V3 Turbo | texto | 16:9, 9:16, 1:1 | 5, 10s | 720p, 1080p | 720p per second: 18, 720p 5: 90, 720p 10: 180, 1080p per second: 22.5, +2 more |
|
Kling V3 Turbo Image to Video
kling_v3_turbo_i2v
Fast image-to-video route that requires one image URL and supports 5s/10s durations.
|
Kling V3 Turbo | texto , imagen | 5, 10s | 720p, 1080p | 720p per second: 18, 720p 5: 90, 720p 10: 180, 1080p per second: 22.5, +2 more | |
|
Kling 2.6
kling_2_6
Stable 5s/10s generation with optional audio.
|
Kling 2.6 | texto , imagen | 16:9, 9:16, 1:1 | 5, 10s | - | 5 no audio: 55, 10 no audio: 110, 5 audio: 110, 10 audio: 220 |
|
MiniMax H3
minimax_h3_t2v
2K text-to-video with native stereo audio and 4-15s durations.
|
MiniMax H3 | texto | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | - | per second: 36.5 |
|
MiniMax H3 Image to Video
minimax_h3_i2v
2K image-to-video from a first frame or first/last-frame pair.
|
MiniMax H3 | texto , imagen | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | - | per second: 36.5 | |
|
MiniMax H3 Reference to Video
minimax_h3_r2v
2K reference-to-video with image, video, and optional audio references.
|
MiniMax H3 | texto , imagen | adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | - | per second: 36.5, per input video second: 36.5, per additional image: 11, included images: 5 |
|
PixVerse V6
pixverse_v6_t2v
Text-to-video with optional native audio, multi-clip, seed, 1-15s duration and 360p-1080p output.
|
PixVerse V6 | texto | 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9 | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 360p, 540p, 720p, 1080p | 360p silent per second: 4.0, 360p audio per second: 5.6, 540p silent per second: 5.6, 540p audio per second: 7.2, +4 more |
|
PixVerse V6 Image to Video
pixverse_v6_i2v
Image-to-video from one or two images, with optional audio, multi-clip and seed.
|
PixVerse V6 | texto , imagen | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 360p, 540p, 720p, 1080p | 360p silent per second: 4.0, 360p audio per second: 5.6, 540p silent per second: 5.6, 540p audio per second: 7.2, +4 more | |
|
PixVerse V6 Transition
pixverse_v6_transition
First-to-last-frame transition from exactly two images, with optional audio and seed.
|
PixVerse V6 | texto , imagen | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 360p, 540p, 720p, 1080p | 360p silent per second: 4.0, 360p audio per second: 5.6, 540p silent per second: 5.6, 540p audio per second: 7.2, +4 more | |
|
PixVerse V6 Extend
pixverse_v6_extend
Extend a completed task or public source video by 1-15 seconds.
|
PixVerse V6 | , vídeo | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 360p, 540p, 720p, 1080p | 360p silent per second: 4.0, 360p audio per second: 5.6, 540p silent per second: 5.6, 540p audio per second: 7.2, +4 more | |
|
PixVerse V6 Reference to Video
pixverse_v6_r2v
Reference-to-video with 1-7 named subject or background images.
|
PixVerse V6 | texto , imagen | 16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2, 21:9 | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 360p, 540p, 720p, 1080p | 360p silent per second: 4.0, 360p audio per second: 5.6, 540p silent per second: 5.6, 540p audio per second: 7.2, +4 more |
|
Seedance 2.0 Mini
seedance_2_mini
Lower-cost Seedance 2 workflow with text, first-frame, first+last-frame, multimodal references, optional audio, 480p/720p output, and 4-15s durations.
|
Seedance 2.0 Mini | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, 21:9, adaptive | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 480p, 720p | 480p 4 with video input: 24, 480p 4 no video input: 38, 480p 5 with video input: 30, 480p 5 no video input: 48, +44 more |
|
Seedance 2.0
seedance_2
Supports text, first-frame, first+last-frame, and multimodal image/video/audio references. For real-person footage, use pre-registered asset:// IDs.
|
Seedance 2.0 | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, 21:9, adaptive | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 480p, 720p, 1080p | 480p 4 with video input: 86, 480p 4 no video input: 76, 480p 5 with video input: 108, 480p 5 no video input: 95, +68 more |
|
Seedance 2.0 Fast
seedance_2_fast
Seedance 2 Fast with the same first/last frame and multimodal reference options. For real-person footage, use pre-registered asset:// IDs.
|
Seedance 2.0 Fast | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, 21:9, adaptive | 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 480p, 720p | 480p 4 with video input: 71, 480p 4 no video input: 62, 480p 5 with video input: 88, 480p 5 no video input: 78, +44 more |
|
Seedance 1.5 Pro
seedance_1_5_pro
Audio-video model with fixed lens, multi aspect ratios, and 4/8/12s.
|
Seedance 1.5 Pro | texto , imagen | 1:1, 21:9, 4:3, 3:4, 16:9, 9:16 | 4, 8, 12s | 480p, 720p, 1080p | 480p 4 silent: 10, 480p 4 audio: 20, 480p 8 silent: 20, 480p 8 audio: 30, +14 more |
|
Wan 2.7 Video
wan_2_7_t2v
Text-to-video generation with 720p/1080p quality modes and 2-15s durations.
|
Wan 2.7 | texto | 16:9, 9:16, 1:1, 4:3, 3:4 | 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 2: 32, 720p 3: 48, 720p 4: 64, 720p 5: 80, +24 more |
|
Wan 2.7 Image to Video
wan_2_7_i2v
Image-to-video generation with first-frame reference and 2-15s durations.
|
Wan 2.7 | texto , imagen | 16:9, 9:16, 1:1, 4:3, 3:4 | 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 2: 32, 720p 3: 48, 720p 4: 64, 720p 5: 80, +24 more |
|
Wan 2.7 Video Edit
wan_2_7_videoedit
Prompt-based video editing that requires one source video URL.
|
Wan 2.7 | imagen , vídeo | 16:9, 9:16, 1:1, 4:3, 3:4 | 0, 2, 3, 4, 5, 6, 7, 8, 9, 10s | 720p, 1080p | 720p 2: 32, 720p 3: 48, 720p 4: 64, 720p 5: 80, +16 more |
|
Wan 2.7 R2V
wan_2_7_r2v
Reference-to-video flow for image/video references and 2-10s durations.
|
Wan 2.7 | texto , imagen | 16:9, 9:16, 1:1, 4:3, 3:4 | 2, 3, 4, 5, 6, 7, 8, 9, 10s | 720p, 1080p | 720p 2: 32, 720p 3: 48, 720p 4: 64, 720p 5: 80, +14 more |
|
HappyHorse 1.1 Video
happyhorse_t2v
HappyHorse 1.1 text-to-video with 3-15s durations and 720p/1080p quality modes.
|
HappyHorse | texto | 16:9, 9:16, 1:1, 4:3, 3:4 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 3: 99, 720p 4: 132, 720p 5: 165, 720p 6: 198, +22 more |
|
HappyHorse 1.1 Image to Video
happyhorse_i2v
HappyHorse 1.1 image-to-video with up to two image references.
|
HappyHorse | texto , imagen | 16:9, 9:16, 1:1, 4:3, 3:4 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 3: 99, 720p 4: 132, 720p 5: 165, 720p 6: 198, +22 more |
|
HappyHorse 1.1 Reference to Video
happyhorse_r2v
HappyHorse 1.1 reference-to-video with up to nine image references.
|
HappyHorse | texto , imagen | 16:9, 9:16, 1:1, 4:3, 3:4 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 3: 99, 720p 4: 132, 720p 5: 165, 720p 6: 198, +22 more |
|
HappyHorse Video Edit
happyhorse_videoedit
HappyHorse prompt-based video editing that requires one source video URL.
|
HappyHorse | imagen , vídeo | 16:9, 9:16, 1:1, 4:3, 3:4 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p | 720p 3: 93, 720p 4: 124, 720p 5: 155, 720p 6: 186, +22 more |
|
AI Video Lip Sync
video_lip_sync
Video-to-video lip sync route that requires one source video URL and one target audio URL.
|
AI Video Lip Sync | imagen , vídeo | 5, 10, 15, 30, 60s | - | per second: 8, 5 seconds: 40, 10 seconds: 80, 15 seconds: 120, +2 more | |
|
OmniHuman 1.5
omnihuman_1_5
Talking portrait generation from one image URL and one speech audio URL.
|
OmniHuman 1.5 | texto , imagen | 5, 10, 15, 30, 60s | 1080 | per second: 27, 5 seconds: 135, 10 seconds: 270, 15 seconds: 405, +2 more | |
|
Grok Imagine Video
grok_imagine_video
Mode + resolution quality selector (for example normal|720p).
|
Grok Imagine Video | texto , imagen | 1:1, 2:3, 3:2, 9:16, 16:9 | 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30s | normal|480p, normal|720p, fun|480p, fun|720p, spicy|480p, spicy|720p | 6 480p: 10, 7 480p: 13, 8 480p: 15, 9 480p: 18, +46 more |
|
Hailuo 2.3 Standard
hailuo_2_3_standard
Image-to-video model that requires one reference image URL.
|
Hailuo 2.3 | imagen | 16:9 | 6, 10s | 768P, 1080P | 768P 6: 30, 768P 10: 50, 1080P 6: 50 |
|
Hailuo 2.3 Pro
hailuo_2_3_pro
Higher-cost Hailuo tier with improved quality presets.
|
Hailuo 2.3 | imagen | 16:9 | 6, 10s | 768P, 1080P | 768P 6: 45, 768P 10: 90, 1080P 6: 80 |
|
Sora 2
sora_2
Sora 2 generation with 10s or 15s durations.
|
Sora 2 | texto , imagen | landscape, portrait | 10, 15s | - | per second: 8 |
|
Sora 2 Pro
sora_2_pro
Sora 2 Pro with high and standard quality modes.
|
Sora 2 | texto , imagen | landscape, portrait | 10, 15s | high, standard | per second: 8 |
|
Sora 2 Pro Storyboard
sora_2_pro_storyboard
Storyboard workflow with optional prompt and longer duration mode.
|
Sora 2 | texto , imagen | landscape, portrait | 10, 15, 25s | - | 10 seconds: 150, 15 25 seconds: 270 |
|
Gemini Omni Flash 1.1
gemini_omni_flash_1_1
|
Video engine | imagen , vídeo | 16:9, 9:16, 1:1, 4:3, 3:4 | 4, 6, 8, 10s | 360p, 720p, 1080p, 4k | 360p 4 no video input: 90, 360p 6 no video input: 120, 360p 8 no video input: 150, 360p 10 no video input: 180, +16 more |
|
Gemini Omni Video
gemini_omni_video
|
Video engine | imagen , vídeo | 16:9, 9:16, 1:1, 4:3, 3:4 | 4, 6, 8, 10s | 720p, 1080p, 4k | 360p 4 no video input: 90, 360p 6 no video input: 120, 360p 8 no video input: 150, 360p 10 no video input: 180, +16 more |
|
Gemini Omni Audio
gemini_omni_audio
|
Video engine | texto | 4s | - | base: 50 | |
|
Gemini Omni Character
gemini_omni_character
|
Video engine | texto , imagen | 4s | - | base: 50 | |
|
Kling O3 Text to Video
kling_o3_t2v
|
Video engine | texto | 16:9, 9:16, 1:1 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p, 4k | 720p 3 silent: 51, 720p 3 audio: 65, 720p 3 with video input: 72, 720p 4 silent: 68, +114 more |
|
Kling O3 Image to Video
kling_o3_i2v
|
Video engine | texto , imagen | auto, 16:9, 9:16, 1:1 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p, 4k | 720p 3 silent: 51, 720p 3 audio: 65, 720p 3 with video input: 72, 720p 4 silent: 68, +114 more |
|
Kling O3 Reference to Video
kling_o3_r2v
|
Video engine | texto , imagen | 16:9, 9:16, 1:1, auto | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p, 4k | 720p 3 silent: 51, 720p 3 audio: 65, 720p 3 with video input: 72, 720p 4 silent: 68, +114 more |
|
Kling O3 Transformation
kling_o3_transformation
|
Video engine | imagen , vídeo | auto, 16:9, 9:16, 1:1 | 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15s | 720p, 1080p, 4k | 720p 3 silent: 51, 720p 3 audio: 65, 720p 3 with video input: 72, 720p 4 silent: 68, +114 more |
|
Wan 3.0 Video
wan_3_0_video
|
Video engine | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, adaptive | 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30s | 480p, 720p, 1080p | 480p 2: 20, 480p 3: 29, 480p 4: 39, 480p 5: 48, +84 more |
|
Wan 3.0 Video Prime
wan_3_0_video_prime
|
Video engine | texto , imagen , vídeo | 1:1, 4:3, 3:4, 16:9, 9:16, adaptive | 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30s | 480p, 720p, 1080p | 480p 2: 30, 480p 3: 44, 480p 4: 59, 480p 5: 74, +84 more |
Guía de selección de modelo
- uso
veo31_litepara el perfil de costo más bajo Veo 3.1 (40 créditos) en flujos de texto a video e imagen a video. - uso
veo31_fastcuando la velocidad y la salida predecible de 8s son lo más importante. - uso
kling_3_0para duraciones flexibles, modos de calidad y controles de fotogramas. - uso
kling_v3_turbo_t2vokling_v3_turbo_i2vpara clips rápidos de 5s/10s en 720p o 1080p. - uso
seedance_2_minipara clips Seedance de 4 a 15 segundos de menor costo en 480p o 720p con modos de texto, fotograma y referencia multimodal. - uso
video_lip_synccuando ya tienes un clip fuente y solo necesitas el tiempo de la boca para seguir el nuevo audio. - uso
omnihuman_1_5para clips de retratos parlantes con audio a partir de una imagen de retrato. - uso
happyhorse_t2v,happyhorse_i2v,happyhorse_r2v, ohappyhorse_videoeditpara flujos de trabajo de vídeos cortos de 3 a 15 s 720p/1080p. - uso
wan_2_6_v2vpara flujos de trabajo de remezcla de vídeo fuente. - uso
hailuo_2_3_*Sólo cuando puedas proporcionar una imagen de referencia. - uso
sora_2_pro_storyboardpara flujos de planificación de guiones gráficos de múltiples tomas.