Unified creation and editing
Kling O1 is the first video model to unify generation and editing in a single system — you can create a new video from scratch and then edit specific sections, restyle footage, extend shots, or swap elements within the same model, without exporting to a separate editing tool. The Multi-modal Visual Language (MVL) architecture accepts six input types simultaneously: text, images, keyframes, reference videos, motion references, and video editing instructions. This makes Kling O1 uniquely capable for production pipelines that need a single model to handle multiple stages.Capabilities
Unified generation and editing
The first model to handle both video creation and video editing in one system — generate footage and edit it within the same generation pipeline.
6 input types
Accepts text, images, keyframes, reference videos, motion references, and editing instructions as simultaneous inputs.
Up to 7 reference images
Anchor character appearance, visual style, and scene composition with up to 7 reference images in a single generation.
Up to 6 camera cuts
Generates up to 6 distinct shots per generation — structured multi-shot output from a single model invocation.
Video restyling
Transform the visual style of existing footage — apply new aesthetics, change time of day, or retheme content while preserving the underlying motion.
Shot extension
Extend existing shots seamlessly — continue the motion and scene from the end of an existing clip.
Input types supported
Specifications
How to use
1
Open the AI Video Generator
Log into ImagineArt and go to the AI Video Generator.
2
Select Kling O1
Choose Kling O1 from the model dropdown.
3
Choose your input combination
Select the combination of input types that fits your use case — text only, text + images, keyframes + motion reference, or video editing mode.
4
Upload references
Upload up to 7 reference images, a reference video, or motion reference as needed.
5
Describe your multi-shot structure
For multi-shot output, structure your prompt with explicit shot descriptions — up to 6 shots per generation.
6
Generate
Click Generate. Generation typically completes in 1–2 minutes for complex multi-input requests.
Prompting tips
- Describe edit targets precisely — In editing mode: “Change the background from day to night while keeping the subject unchanged” is more accurate than “make it darker.”
- Use keyframes for transitions — Define your start and end keyframes; let Kling O1 fill in the motion between them consistently.
- Combine input types — “Based on this reference image [image], in this visual style [image 2], with this camera movement [motion ref]…” — the MVL architecture processes all inputs cohesively.
Example prompts
SHOT 1 (wide, 3s): A detective walks into a rain-soaked alley at night. SHOT 2 (close-up, 2s): Detective looks at a clue on the ground, rain drops visible. SHOT 3 (medium, 3s): Detective turns and exits the alley. Reference image for detective character appearance attached.
Restyle the provided footage to a vintage 1970s Super 8 film look. Keep all motion and subjects identical; change only the visual aesthetic.