How to Transfer Motion From Video to Character With AI
2026/07/28

How to Transfer Motion From Video to Character With AI

Learn how to transfer motion from video to character with AI. Prepare the image and reference clip, choose real settings, write prompts, and fix common failures.

To transfer motion from video to character, upload a character image and a reference video that clearly shows the movement you want the character to perform. The image defines the character's visible appearance. The video defines the pose sequence, timing, gestures, and performance rhythm.

Quick answer: Open AI Image to Video Motion Control, upload one JPG or PNG character image, add a 3–30 second MP4 or MOV reference clip, choose whether character orientation follows the video or image, select 720P or 1080P, and generate a short test. Kling 2.6 also accepts MKV video; its Image Orientation mode is limited to 10 seconds.

This workflow creates a new rendered video. It does not export an animation skeleton, rig, FBX file, BVH file, or editable motion-capture data.

What motion transfer does

Motion transfer separates two parts of a performance:

  • Character appearance: face, hair, clothing, proportions, illustration style, or creature design from the image
  • Performance guidance: pose changes, movement path, gesture order, timing, and rhythm from the reference video

The model analyzes the reference clip and generates new frames in which the image character attempts to follow that performance. The result is more directed than ordinary image-to-video generation because the action already exists in the reference.

Motion transfer is still generative. It does not copy every joint coordinate exactly, and it can reinterpret difficult anatomy, occlusion, props, camera changes, or background motion.

WorkflowInputOutputBest used when
Motion transferCharacter image + reference videoGenerated character videoYou need a specific visible performance
Image-to-videoImage + promptGenerated videoThe model can invent the movement
Character swapExisting video + replacement characterRecast generated videoReplacing the performer is the main goal
Traditional motion captureTracked performer or mocap dataSkeleton or animation dataYou need an editable 3D animation pipeline

Use AI Image to Video Motion Control when the action is the instruction. Use Character Swap AI when replacing the person in an existing scene is the central goal. Use AI Photo to Video when you do not have a reference performance and only need prompt-directed animation.

Verified input limits in the current editor

The site currently exposes Kling 2.6 Motion Control and Kling 3.0 Motion Control. The following values come from the implemented form controls, upload routes, request validation, and model-provider requirements.

ParameterKling 2.6 Motion ControlKling 3.0 Motion Control
Character image typeJPG or PNGJPG or PNG
Maximum image upload10MB10MB
Minimum image dimension shown in the editorGreater than 300pxGreater than 340px
Reference video typeMP4, MOV, or MKVMP4 or MOV
Maximum video upload100MB100MB
Reference duration3–30s in Video Orientation; 3–10s in Image Orientation3–30s
Character OrientationVideo or ImageVideo or Image
Background SourceNot exposed as a separate controlInput Video or Input Image
QualityStandard 720P or Pro 1080PStandard 720P or Pro 1080P
PromptOptionalOptional; provider validation caps it at 2,500 characters

Kling 3.0's provider requirements also specify an aspect ratio between 2:5 and 5:2 for both the image and motion video.

The editor clips and transcodes the selected reference range. It passes the applied duration to the generation request and uses that duration when calculating the live credit estimate.

How to transfer motion from video to character step by step

Step 1: define the movement you actually need

Do not start by looking for a visually impressive reference. Start by writing down the action:

  • A five-second greeting with one wave
  • A seated explanation with hand gestures
  • A continuous dance phrase
  • A short walk and turn
  • A martial-arts strike
  • A product demonstration

One clear action is easier to evaluate than a montage. If the intended output is a 20-second social video, it can be safer to design it as several controlled shots rather than one reference containing multiple cuts and unrelated movements.

Step 2: choose a readable reference video

The reference video is the motion instruction. A useful source makes the motion easy to see:

  • One main performer
  • Head, torso, and relevant limbs visible
  • Stable camera or intentional, smooth camera movement
  • No rapid edits inside the selected range
  • Limited overlap with other people or foreground objects
  • A clear first pose
  • Enough contrast to separate the performer from the background

Avoid using a group performance when only one character will be generated. Avoid a close crop if the important action happens in the legs. Avoid starting halfway through a turn or jump because the first generated frame must reconcile the still image with that pose.

The model accepts up to 30 seconds in supported modes, but the maximum is not automatically the best starting point. Use a short diagnostic range first.

Step 3: prepare the character image

Match the image crop to the reference action.

For full-body motion, the image should show the full body, including hands and feet. For a talking or gesturing reference, an upper-body image can work if the reference uses similar framing.

Use an image with:

  • One clearly identifiable character
  • A face large enough to read
  • Even lighting and visible facial features
  • A pose reasonably close to the reference's first frame
  • Clean separation between the character and background
  • Clothing and limbs that are not cut off

JPG and PNG are supported. The app upload route limits images to 10MB. Kling 2.6's interface asks for an image larger than 300 pixels, while Kling 3.0 asks for dimensions greater than 340 pixels.

Higher pixel count alone does not guarantee a better source. A sharp, naturally lit image is more useful than a heavily upscaled image containing compression artifacts or invented detail.

Step 4: choose Kling 2.6 or Kling 3.0

The two current models share the same basic inputs but expose different controls.

Kling 2.6 Motion Control

  • Accepts MP4, MOV, and MKV references
  • Offers Video Orientation up to 30 seconds
  • Offers Image Orientation up to 10 seconds
  • Provides 720P Standard and 1080P Pro output choices

Kling 3.0 Motion Control

  • Accepts MP4 and MOV references
  • Supports 3–30 second references in the current editor
  • Adds a Background Source setting
  • Provides 720P Standard and 1080P Pro output choices

If both models fit the file and duration, compare them using the same short clip. Keep the image, selected range, orientation, quality, and prompt unchanged so the comparison isolates the model.

Step 5: set Character Orientation

Character Orientation tells the model whether the generated character should be guided more by the orientation in the reference video or the uploaded image.

Choose Video when following the reference pose direction and camera-facing behavior is the priority. This is the default in the current editor and supports references up to 30 seconds.

Choose Image when preserving the source image's orientation is more important. In Kling 2.6, this limits the reference to 3–10 seconds. Kling 3.0 currently allows up to 30 seconds for either orientation.

This setting is not the aspect ratio control. A vertical 9:16 clip can use either Character Orientation option.

Step 6: choose Background Source in Kling 3.0

Kling 3.0 can guide the generated background from:

  • Input Video: appropriate when the reference environment, camera, and scene should remain central
  • Input Image: appropriate when the character image's environment should influence the result

If the image is a transparent-looking character on a plain background and the goal is to place it into the recorded performance, Input Video is the logical first test. If the character is already inside a carefully designed environment that should remain visible, test Input Image.

The result remains generative; the setting is guidance, not a guarantee that every background pixel will remain unchanged.

Step 7: choose Standard 720P or Pro 1080P

Use Standard 720P for setup tests. It costs fewer credits and is sufficient for judging:

  • Whether the character follows the action
  • Whether the starting pose is compatible
  • Whether orientation is correct
  • Whether the background source works
  • Whether the face and body remain recognizable

Use Pro 1080P after the input pairing is working and output detail matters. A failed composition at 720P usually remains a failed composition at 1080P.

Step 8: add an optional prompt

The prompt should not repeat every movement visible in the reference video. Use it to reinforce the character, scene, and visual constraints.

Animate the uploaded character using the reference performance.
Keep the same face, hair, outfit, and body proportions throughout.
Preserve the gesture timing and camera-facing direction.
Natural movement, stable hands, consistent clothing, no extra people.

For Kling 3.0, you can also make the intended background treatment explicit:

Use the movement and timing from the reference video.
Keep the uploaded character's appearance consistent.
Retain the reference-video environment and camera position.

The provider accepts a long prompt, but a 2,500-character limit is not a target. Extra instructions can conflict with one another. Begin with the minimum needed to explain what the image and video do not already communicate.

Step 9: generate and compare against the source

Review the output beside the reference video. Do not judge only by whether the new video looks attractive.

Check:

  1. Does the same action happen in the same order?
  2. Are important beats early or late?
  3. Does the character face the expected direction?
  4. Are the face, body proportions, clothing, and style stable?
  5. Do hands and feet remain readable?
  6. Are contact moments with the floor, chair, prop, or body plausible?
  7. Does the background stay consistent with the selected source?
  8. Is the final frame still usable?

If motion fidelity is good but identity is weak, change the character image. If identity is good but the action is weak, change or trim the reference video. If both fail, the framing or initial pose may be incompatible.

How to choose the right framing

Framing mismatch forces the model to invent missing visual information.

Reference framingRecommended character image
Full-body dance, run, or walkFull body with hands and feet visible
Waist-up presentationWaist-up or wider
Seated conversationSeated or similarly framed upper body
Close facial performanceClear face and shoulders
Side-facing actionSimilar side or three-quarter starting pose

A front-facing portrait can sometimes be animated with a side-facing reference, but the model must reconstruct unseen parts of the head, clothing, and body. Use a more compatible image when consistency matters.

Prompt examples for common motion-transfer tasks

Dance or choreography

The uploaded character performs the reference choreography with matching
timing and body direction. Keep the exact character identity, outfit, and
proportions stable. Clean full-body silhouette and consistent hands and feet.

Talking avatar

Apply the reference speaker's upper-body gestures and posture to the uploaded
character. Preserve the character's face and clothing. Stable seated framing,
natural hand motion, and no background changes.

Mascot performance

Animate the mascot using the reference movement and gesture timing. Preserve
the mascot design, colors, proportions, and logo placement. Keep the motion
readable without adding human facial features.

Action reference

Transfer the reference action to the uploaded character. Preserve the action
order, direction, and pacing. Keep the character anatomy and costume stable,
with clear contact between the feet and ground.

Treat these as starting structures, not magic words. The image and video quality have more influence than decorative prompt language.

Common motion-transfer failures and what to change

SymptomLikely causeBest first test
Character changes identityFace is unclear or source is heavily stylizedUse a sharper, larger, evenly lit character image
Legs or arms deformReference limbs are cropped or occludedChoose a wider, cleaner reference
Output starts with a visual jumpImage pose conflicts with the first video frameTrim to a compatible first pose
Movement is incompleteAction is too fast, complex, or hiddenUse a shorter range with one readable action
Motion timing driftsCuts or camera changes interrupt trackingUse a continuous shot
Background flickersBusy scene or wrong Kling 3.0 background sourceTest the other source or simplify the reference
Character faces the wrong wayOrientation setting conflicts with the goalSwitch Character Orientation
Clothing morphsFine pattern, loose fabric, or repeated occlusionUse a simpler image and shorter clip
Hands fail at one momentFast overlap or contact with an objectTrim around the moment or choose another take
Longer clip degrades over timeMore opportunities for generative driftSplit the performance into controlled shots

Change one variable at a time and keep a record of the settings. Randomly rewriting the prompt after every result makes the workflow harder to diagnose.

Building a longer video from short motion transfers

For a longer project, divide the performance into shots with clear start and end poses. Reuse:

  • The same character image
  • The same model and quality
  • The same Character Orientation
  • The same background strategy
  • The same identity description in the prompt
  • Compatible aspect ratios and framing

Generate and approve each shot before editing them together. A single 30-second reference is supported in some settings, but shot-based production gives you more control over failures and credit use.

Do not assume the last frame of one generated clip will automatically match the first frame of the next. Plan cuts at pauses, camera changes, or moments where a transition can hide small differences.

Motion transfer is not traditional motion capture

Traditional motion capture records or estimates movement and maps it to an editable skeleton or rig. That data can be adjusted in 3D software.

The workflow described here generates a finished video from an image and reference clip. It does not provide:

  • Joint coordinates
  • A reusable character rig
  • FBX or BVH motion files
  • Editable keyframes
  • Guaranteed frame-exact retargeting

If your production requires animation data for Blender, Maya, Unreal Engine, or another 3D pipeline, use a motion-capture or pose-estimation system designed to export that data.

Rights and source-video quality

Use a reference video you recorded or are licensed to use. A clip found online may still be protected even if you only intend to copy its movement. Obtain permission before using a person's likeness, distinctive performance, choreography, costume, or branded character in public or commercial work.

Keep realistic AI outputs clearly contextualized when viewers could mistake them for authentic footage.

Frequently asked questions

Can AI copy movement from a video to a character?

Yes. Motion-control video models can use a character image for appearance and a reference video for performance guidance. They generate a new character video that attempts to follow the source poses, gestures, direction, and timing.

What image and video formats are supported?

The current Kling 2.6 workflow accepts a JPG or PNG image and an MP4, MOV, or MKV reference. Kling 3.0 accepts a JPG or PNG image and an MP4 or MOV reference. The app limits images to 10MB and videos to 100MB.

How long can the reference video be?

References must be at least 3 seconds. Kling 2.6 supports up to 30 seconds with Video Orientation and up to 10 seconds with Image Orientation. Kling 3.0 supports up to 30 seconds in the current editor.

What is the difference between Video and Image Orientation?

Video Orientation follows orientation cues from the reference performance. Image Orientation gives more weight to the character image's existing orientation. It is separate from landscape or portrait aspect ratio.

Should I choose 720P or 1080P?

Use 720P to test the image, video, orientation, background, and prompt. Use 1080P after the setup works and you need more output detail. The editor shows the current credit estimate before generation.

Does the prompt control the movement?

The reference video is the primary movement guide. The optional prompt refines the character, visual treatment, environment, and consistency requirements. Avoid writing motion instructions that contradict the source clip.

Can motion transfer produce an FBX or BVH file?

No. This workflow produces a generated video, not an editable skeleton, rig, FBX, BVH, or keyframe sequence.

Why does the character drift during a long clip?

Longer clips contain more pose changes, occlusions, and opportunities for the model to reinterpret the character. Test a shorter range, simplify the action, improve the source image, or divide the performance into separate shots.

Verified implementation references

The formats, file sizes, durations, orientation rules, background control, and quality options above were checked against the current site code and the current Kling 2.6 Motion Control API requirements and Kling 3.0 Motion Control API requirements.

When your image and reference are ready, open AI Image to Video Motion Control and begin with a short 720P test.