Models

Five AI video models. One production context.

GATA renders shots on Veo 3.1, Kling o3, Kling v3, Seedance 2.0, and Grok Imagine 1.5. The model is a per-shot choice, so you can put the expensive one on the hero shot and the cheap one on the cutaway without re-doing script, cast, or look. Here is what each is good at, what it costs, and where it falls down.

Side by side

Credit rates are the base cost per second of finished video. Uplifts are listed below the table. Verified against application code on 2026-07-30.

Model Maker Credits / sec 4K Locked-cast references Best for
Kling o3 Pro Kuaishou 20 Yes Yes Reference-anchored, stable identity.
Kling v3 Pro Kuaishou 25 Yes Yes Frame-to-frame interpolation.
Seedance 2.0 Pro ByteDance 45 No No Strong realism, vivid motion.
Veo 3.1 Pro Google DeepMind 45 Yes No Top quality, native audio.
Grok Imagine 1.5 xAI 20 No No Animates a single first frame.

A locked-cast reference means the model accepts an approved character composite as an input, so identity is bound structurally instead of re-described in each prompt. See locked cast and character drift for why that distinction decides whether a multi-shot sequence holds together.

Optional uplifts

Added to the base per-second rate when the matching toggle is on.

Uplift Credits / sec Applies to Notes
4K output +30 Kling o3, Kling v3, Veo 3.1 Hidden in the UI for models that cannot render 4K.
Lip-sync +15 Any model Covers dialogue voice generation and the lip-sync pass. Same rate for original and localised versions.
Background noise +5 Any model Per-shot ambience toggle.
Topaz upscale +10 Finished scene clips Master-stage upscale, charged per second of the source clip.

Each model in detail

Kling o3 Pro

Kuaishou · 20 credits/sec · 4K capable

Best for. Character-driven work: anything where the same person has to appear in shot after shot. This is GATA's default.

Where it falls down. Less punchy on fast, chaotic action than Seedance. Trades a little motion drama for identity stability.

Read the Kling o3 Pro deep dive

Kling v3 Pro

Kuaishou · 25 credits/sec · 4K capable

Best for. Shots where you have both a first and last frame and want controlled movement between two known compositions.

Where it falls down. Costs 5 credits/second more than o3 for the same duration, and needs a usable last frame to earn that premium.

Read the Kling v3 Pro deep dive

Seedance 2.0 Pro

ByteDance · 45 credits/sec

Best for. Product, texture, landscape, and abstract motion where realism and energy matter more than a recurring face.

Where it falls down. Rejects generations that reference faces. If your script has human characters, pick a different model — GATA warns you before the project starts. No 4K.

Read the Seedance 2.0 Pro deep dive

Veo 3.1 Pro

Google DeepMind · 45 credits/sec · 4K capable

Best for. Hero shots and dialogue beats where prompt adherence and native audio justify the highest per-second rate.

Where it falls down. Durations snap to 4, 6, or 8 seconds — a 5-second shot is billed and rendered as 6. No locked-cast reference binding.

Read the Veo 3.1 Pro deep dive

Grok Imagine 1.5

xAI · 20 credits/sec

Best for. Cheap motion passes over a first frame you already like, and quick alternates while you are still finding a shot.

Where it falls down. First frame only — it cannot use a last frame. Caps at 720p. xAI charges for every generation, including ones it blocks, so blocked clips are not refunded.

Billing caveat: blocked generations on this model are not refunded.

How to choose

The decision is usually not "which model is best" but "which model is right for this shot". Three questions settle almost every case.

Does a recurring person appear in it? If yes, use Kling o3. It accepts the locked-cast composite, which is the only mechanism here that binds identity structurally. Seedance 2.0 will refuse the shot outright, and Veo 3.1 and Grok will render a plausible stranger.

Is it a hero shot? Veo 3.1 is the highest-fidelity option and brings native audio, at 45 credits per second and durations snapped to 4, 6, or 8 seconds. On a six-second hero shot that is 270 credits — worth it once per film, rarely worth it twelve times.

Are you still exploring? Grok Imagine 1.5 at 20 credits per second is the cheapest way to see a first frame move, as long as you accept 720p and the no-refund caveat. Promote the shot to a better model once the composition is settled.

What a real project costs

A 30-second ad with eight shots, averaging six seconds each, is 48 seconds of finished video. On Kling o3 at 20 credits per second that is around 960 credits. The same 48 seconds on Veo 3.1 is 2,160. Mixing them — Veo on the two hero shots, o3 on the other six — lands near 1,300, which is the practical reason per-shot model choice exists.

Add lip-sync to the four shots with dialogue and you add 15 credits per second across roughly 24 seconds: 360 credits. Every plan on the pricing page lists its monthly credit inclusion, and the project math there works through a 90-second film and a localised variant.

Why the model is not the product

Every model on this page is available to everyone. A subscription to any one of them gets you a prompt box and a clip. What it does not get you is the thing that makes eight clips cut together: a locked cast, an approved look, locations that persist, and a script the shot list actually reads from.

That is the part GATA is. Models are interchangeable renderers inside it — see script to video, in parallel for the workspace the renderers plug into, and shot inheritance for the mechanism that lets you swap a shot's model without redoing anything around it.

Frequently asked

Which AI video model is best for character consistency?

Kling o3 Pro. It is the only pair of models — o3 and v3 — that accept locked-cast composites as references, so the approved face is bound to the generation rather than described in a prompt. It is also GATA's default model for that reason. Veo 3.1, Seedance 2.0, and Grok Imagine 1.5 do not take character references.

How much does a second of AI video cost in GATA?

Between 20 and 45 credits per second depending on the model: Kling o3 and Grok Imagine 1.5 are 20, Kling v3 is 25, and Seedance 2.0 and Veo 3.1 are 45. Optional uplifts add per second on top — 4K adds 30, lip-sync adds 15, background noise adds 5.

Can I use a different model for each shot?

Yes. The model is a per-shot choice, not a per-project one. Because script, cast, look, and locations live in project state, switching a shot's model does not re-do any of them — the shot inherits the same context and only the renderer changes.

Why does Seedance 2.0 refuse shots with people in them?

Seedance 2.0 rejects generations that reference faces. That is a model-side restriction, not a GATA setting. If your script has human characters, GATA warns you before the project starts and you should pick Kling o3 instead.

Which models can render 4K?

Kling o3, Kling v3, and Veo 3.1. The 4K toggle adds 30 credits per second and is hidden for Seedance 2.0 and Grok Imagine 1.5, which cannot render it. Finished scene clips can also be upscaled at the master stage for 10 credits per second.

What happens to my credits if a generation fails?

Failed generations are refunded, with one disclosed exception: xAI charges for every Grok Imagine 1.5 request including ones it blocks for policy reasons, so a blocked Grok clip is not refunded. Every other model refunds a content-policy failure.