Skip to content
AI Kiss

What Is an AI Kiss Video? A Plain Explanation

An AI kiss video animates two still photos into a short clip of two people kissing. Here is how the models work, what the output looks like, and where it fails.

AI Kiss TeamPublished Updated 6 min read

You have probably seen one already: two photos of a couple, side by side, that suddenly move and lean into a kiss. No camera was there. No editor keyframed anything. A model looked at a still image and predicted what the next hundred and twenty frames would look like if those people had moved.

This guide explains what is really happening underneath, what these clips look like when they work, and where they still come apart. If you would rather just try it, the app runs in your browser, but the background is worth five minutes because it explains most of the reasons a generation turns out badly.

What an AI kiss video actually is

An AI kiss video is a short generated clip, usually around five seconds, in which two people in still photographs appear to move and kiss. It belongs to a family of tools called image-to-video generators. You supply a starting image. The model produces a sequence of frames that could plausibly follow from it.

The important part of that sentence is could plausibly follow. The model is not retrieving footage of your face from anywhere. It has learned, from an enormous amount of video, how human faces and bodies tend to move: how a head tilts before contact, how eyelids close, how hair shifts, how light rolls across a cheek when the angle changes. Given your photo as frame one, it generates a version of that motion using your faces.

That is also why the output is a genuine video file rather than a slideshow with a transition. Every frame after the first is new pixels.

How a still photo becomes a moving clip

Three things happen in order.

The starting frame is assembled. If you upload two separate photos, they are not sent to the model separately. They are composed into one image first: two halves of a single canvas, each photo fitted into its half. If you upload one photo that already contains two people, that photo is the canvas. Either way the model receives exactly one image to start from. In our app you choose whether that canvas is 4:3 or 9:16, and the finished video inherits that shape.

The model is given an instruction. Alongside the image goes a short text prompt describing what should happen: the people are kissing, or hugging. That instruction is what separates a kiss clip from a wave or a hug, and it is why picking a different style gives you different motion from the same two photos.

Frames are predicted. The model generates the sequence, conditioned on the starting image so the faces stay recognisable, and conditioned on the prompt so the motion goes where you asked. It runs on a queue of graphics hardware, which is why the result takes a few seconds rather than arriving instantly. In our app the usual wait is under a minute, and the finished MP4 lands in your gallery.

There is a longer, step-by-step version of this on the how it works page.

Why two photos become one video

People often expect the two uploads to stay separate, like a split screen. They do not. The composition step is the whole trick.

Because the model sees one image containing two faces, it treats them as two people in one scene, and it can move them toward each other. That is why the guidance is one person per photo. Upload a group shot as one half of the pair and the model has four faces to reason about in a scene that is supposed to hold two, and it will usually merge or discard someone.

It also explains why lighting matters more than resolution. Two photos taken in wildly different light, one in warm evening sun and one under blue office strip lights, become a single frame with an obvious seam down the middle, and the model tries to animate across that seam. Photos that look like they came from the same afternoon blend far better. We collected the full list of what helps and what breaks in the photo requirements guide.

What the results actually look like

Set your expectations at convincing short clip, not found footage.

When it works, the two faces stay recognisably themselves, the heads turn with believable weight, eyes close near the contact, and there is small ambient motion in hair and shoulders that sells the whole thing. On a phone screen, at the speed people scroll, it reads as real.

When you look closely, the tells are there. Fine detail around the mouth softens during the closest moment. Earrings, glasses frames and stray hair strands can wobble. Backgrounds sometimes ripple slightly. Hands, when they are in the frame, are still the least reliable part of any generative video model.

The output itself is a standard MP4 you can download and post anywhere. Five seconds is short enough to loop cleanly, which is a large part of why the format works on social feeds. You can see a range of finished clips on the examples page before spending anything.

Kiss, French kiss and hug

The same pipeline produces different motion depending on the instruction. In our app there are three styles.

Kiss is the default and the most reliable: a lean in, contact, a short hold. It is the one to use if you are unsure.

French kiss asks for something more intimate and longer in contact. It is a heavier ask of the model, so it rewards good, front-facing source photos. There is a dedicated page for that style.

Hug moves shoulders and arms rather than just heads, which is why it is the style people use for the reunion and younger-self clips rather than romantic ones. That is covered on the hug video page.

Same two photos, three different clips. If the first result is not what you pictured, changing the style often does more than changing the photos.

Where AI kiss videos fall short

Being honest about the limits saves you money.

  • Profile shots struggle. The model needs to see enough of both faces to rotate them. A photo taken from the side gives it very little to work with.
  • Small faces lose detail. If a face occupies a tiny part of the frame, it has few pixels to preserve, and identity drifts during the clip.
  • Heavy filters confuse it. Strong beauty filters, thick grain and dramatic colour grading all fight against the model's sense of what a face is.
  • Five seconds is the working length. Longer generations from a single image tend to lose the person somewhere in the middle. Short is not a limitation of one tool; it is where the current generation of models is good.
  • Sunglasses and hats crop the signal. Anything hiding the eyes removes most of the information the model uses to animate expression.

Generations do fail. When one does in our app, the credit goes back to your balance automatically, so a bad run costs you time rather than money.

Whose face you may animate

This technology is easy to misuse, so the rule is simple and worth stating plainly: use photos of yourself, and photos of people who have agreed to it.

That means no strangers pulled from social media, no public figures, and no minors, ever. Those are the terms here and they are not negotiable. Beyond the rules, a fair test is whether the person in the photo would be pleased to be shown the result. If you would hesitate to send it to them, do not make it.

Where the photo goes and how long it stays is a separate question, and a reasonable one to ask before uploading anything of your face. We wrote it up in the privacy and safety guide.

Trying one yourself

There is no sign-up and no subscription. You open the app in a browser on a phone or a laptop, add two photos, choose a style and an aspect ratio, and pay for that video. One payment, one clip. If the generation fails you are credited back and can retry.

The most useful advice is the least technical: spend your effort on the photos. Two clear, well-lit, front-facing portraits taken in similar conditions will beat any amount of retrying with bad source material.

Frequently asked questions

What is an AI kiss video?
It is a short video clip generated from still photographs. You give a tool two portraits, or one photo containing two people, and an image-to-video model invents the frames in between so the two faces move toward each other and kiss. Nothing was filmed. The whole clip is predicted from that single starting image.
Is an AI kiss video a deepfake?
It uses related technology, but the intent is what matters. A deepfake usually means putting someone into a scene to deceive viewers. An AI kiss video of you and your partner, made from your own photos and posted as an AI edit, deceives nobody. Animating a stranger or a public figure without consent is a different thing entirely, and it is not allowed here.
How long is an AI kiss video?
About five seconds. Current image-to-video models hold a face together well for a few seconds and then start to drift, so short clips are where the quality is. Five seconds is also close to the ideal length for a social post that loops.
Do I need video editing skills to make one?
No. The work is choosing two good photos, not editing. You upload them in a browser, pick a style and an aspect ratio, and wait about a minute. If you can post a photo to Instagram, you can make one of these clips.

Make your own AI kiss video

Upload two photos in your browser, pick Kiss, French kiss or Hug, and get a 5-second MP4 with sound back in under a minute. No sign-up, and a failed generation is credited back automatically.

6 min read

AI Kiss Trends on TikTok: What Works in 2026

The AI kiss formats people actually post on TikTok in 2026, the hooks that hold a scroll, caption and sound choices, and the mistakes that flatten a good clip.

6 min read

The Hug Your Younger Self AI Trend, Explained

How the hug your younger self AI trend works: one childhood photo, one recent photo, the Hug style, and the tips that make old scanned prints animate well.