AI Kiss Photo Requirements: Why Generations Fail
Face size, angle, lighting, resolution, glasses and group shots decide whether an AI kiss video works. What to look for, why it matters, and a checklist.
AI Kiss TeamPublished 5 min read
If a clip comes out wrong, the tool is rarely the problem. Nearly every disappointing generation traces back to one of seven things about the source photographs, and all seven are checkable in under a minute.
This page is the checklist, with the reasoning behind each item, because the reasoning is what lets you judge a photo you have never seen before.
Why an AI kiss generation fails
The model receives one image and one instruction. From that image it has to work out where the faces are, what those people look like from angles it cannot see, and how to move them believably.
Everything it needs comes from your photographs. If the face is small, it has few pixels to preserve and identity drifts. If the face is turned away, it has to invent the other side. If the lighting between the two halves disagrees, it has to animate across a seam. Nothing in the pipeline can recover information that was not in the picture.
So the useful mental model is not how good is this photo but how much does this photo tell the model. A slightly imperfect close-up beats a beautiful photograph where the face is a hundred pixels across.
Face size in the frame
This is the most common cause of a weak result.
Aim for head and shoulders, with the face occupying a good share of its half of the frame. A portrait crop is ideal. A full-body shot on a beach is not, and neither is a wide holiday photo where you can only just tell who it is.
Crop before uploading rather than hoping the tool will zoom for you. Every phone gallery has a crop tool, and thirty seconds there does more for the result than anything else on this page.
The same applies to both photos equally. If one face fills its half and the other is distant and small, the size mismatch shows up in the video even when the motion is fine.
Front-facing beats profile
The model animates a head turning toward another head. To do that it needs to see enough of the face to understand its structure.
A photo taken straight on, or at a slight three-quarter angle, gives it everything. A pure profile shot gives it one eye, half a mouth and no information about the other side, so it has to invent that half, and inventions drift.
Look for both eyes visible, the nose roughly centred, and the chin not tipped far up or down. Slightly relaxed expressions work better than wide, fixed smiles, because a lot of dramatic mouth shape at frame one constrains what the motion can do.
This matters most for the French kiss style, which asks for more movement and a longer moment of contact, and therefore leans harder on the source.
Lighting, and why the two photos must agree
Your two photos become one image before anything is generated: a single canvas with one person in each half. That frame is what the model sees.
Which means lighting has to work across both halves. A warm sunset selfie next to a photo lit by blue office strips produces a video with an obvious join down the middle, because that is literally what the model was given.
Even, soft light is the friendliest. Daylight from a window, open shade outdoors, or a diffuse indoor light. What to avoid: hard direct sun casting deep shadows across a face, strong backlighting that leaves the face dark, and on-camera flash that flattens everything and blows out highlights.
If you are shooting new photos, take both in the same place within a few minutes of each other. That solves lighting, angle and colour in one move.
Resolution, focus and filters
Three related things.
Focus. A soft photo produces a soft video. Motion blur, camera shake and misfocused shots all lose the fine facial detail the model preserves, and it cannot invent detail that was never captured.
Resolution. No strict minimum, since images are resized before upload. What matters is pixels on the face rather than pixels in the file. A cropped phone photo is plenty.
Filters. Beauty filters, heavy smoothing, thick grain and strong colour grading all work against you. Smoothing in particular removes the texture the model uses to keep a face looking like that specific person. Use the original photo rather than the filtered version wherever you have the choice.
Screenshots of photos are worth avoiding for the same reason: you are handing over a compressed copy of a compressed copy.
Glasses, hats and anything over the eyes
Sunglasses are the clearest failure case. The eyes carry most of the information about expression and head orientation, and covering them removes it. Take them off.
Clear glasses are usually fine, though the frames can wobble slightly during motion and reflections across the lenses can cause odd artefacts. If you have a similar photo without them, prefer it.
Hats pulled low, hoods and heavy fringes over the eyes all cause the same problem to a lesser degree. Face paint, masks and anything that obscures the mouth also make the contact moment unpredictable.
One person per photo
In two-photo mode, each upload should contain exactly one face.
The reason is the same combined frame described above. Upload a group shot and the model sees a scene with four or five people in it while being asked to make two of them kiss. It does not know which two you meant, so it merges faces, drops someone, or animates the wrong pair.
Crop each photo down to a single person. If your only picture of someone is a group shot, crop tightly around them and check that no second face has survived at the edge of the frame.
If you have one photograph that already contains both people, switch to the single-photo mode rather than trying to split it. That path is described in the two-photo workflow guide, which covers when each mode is the right choice.
The checklist
Run through this before you generate. It takes under a minute.
- One person per photo, or one photo with exactly two people.
- Head and shoulders, face large in its half of the frame.
- Both eyes visible, face roughly front-on.
- Sharp and in focus, not a screenshot.
- Soft, even light, no hard shadows or flash.
- Similar lighting, angle and distance across both photos.
- No sunglasses, no low hats, no beauty filters.
- Nobody else at the edges of the crop.
- Everyone in the photos has agreed to this, no minors, no strangers.
Get eight of those nine right and the result will usually be good. The ninth is not optional at all.
If a generation still fails, the credit returns to your balance automatically and you can retry, so the cost of a bad run is time rather than money. The pricing page covers how that works. To see what strong source photographs produce, the examples page has finished clips, the how it works page explains the pipeline, and the app is where you try your own. The full step-by-step is in the walkthrough.
Frequently asked questions
What resolution do my photos need to be?
Can I use a selfie?
Do the two photos need to be taken on the same day?
Why did my generation fail rather than produce a bad video?
Make your own AI kiss video
Upload two photos in your browser, pick Kiss, French kiss or Hug, and get a 5-second MP4 with sound back in under a minute. No sign-up, and a failed generation is credited back automatically.
Related guides
6 min read
AI Kiss Videos on Instagram: Best Practices
Where an AI kiss video belongs on Instagram, why 9:16 wins, how to write the caption, how to handle AI disclosure and consent, and the mistakes to avoid.
6 min read
AI Kiss Privacy and Safety: What Happens to Your Photos
Where your uploads go, what the Privacy Policy says about keeping them, how to ask for deletion, and the consent rules that decide whose face you may animate.
6 min read
AI Kiss App Features: An Honest Tour
Every feature in the AI Kiss app, plainly: the two upload modes, three styles, two aspect ratios, the gallery, retries and refunds, and what it cannot do.