One model instead of a toolchain
You have a clip that's almost right. A stranger walks through frame. The jacket is the wrong color. Kling O1 takes the clip plus a sentence and hands back the fixed shot. Kling O1 reads images, video clips, subjects and text as executable prompts. So the reference photo and the instruction go in the same box, and you skip the round trip to an editor.
Seven references, one cast
Load up to 7 images. Faces stay put across camera moves. Props keep their shape. Group scenes hold together, because each subject is tracked on its own. That's what makes this model useful for series work: same character, new shot, next day.
Length, size, sound
Clips run 3–10s at 720p or 1080p, in 16:9, 9:16 or 1:1. A render takes ~2 min, so you can iterate over a coffee. There's no native audio here. Score it afterwards, or send the idea to Kling 3.0, which brings 4K and sound in the same composer.
What it costs
Video is paid per render, not per month. Kling O1 starts at ⚡77 and tops out at ⚡343 for the longest 1080p take. The composer prices your exact settings before you commit, so you always see the number first. New accounts get ⚡50 on signup. $10/1,000⚡
Where to go next
Browse every version on the Kling family page. Compare vendors side by side in the AI Video Generator. Run the same brief on two models and keep the better take. One wallet, one library, all of it.