Hotel Lobby AI Prompt: The Real One, Line by Line
A good Hotel Lobby AI prompt does four jobs: it fixes the orange booth, hangs one mic between the pair, ties each photo to a side, and says who raps and who reacts. Below are the two prompts the Hotel Lobby Video (hotel-lobby-video.com) generator actually builds, not a guess at one: the Wan 3.0 version, which keeps the template track, and the Seedance version, which writes a new rap. Paste them into your own tool, rework them, or let the generator handle it. If your tool takes two photos and no reference clip, skip to the five photo-only prompts further down.
Last updated
The Wan 3.0 prompt
This text covers two photos in any frame shape, because the aspect ratio is sent as a setting rather than written into the prompt. First, GPT Image 2 turns each upload into a full-length studio portrait of the same person or animal. The model then gets four attachments, always in this order: Image 1 (left portrait), Image 2 (right portrait), Video 1 (the Hotel Lobby reference clip, muted) and Audio 1 (the Hotel Lobby track, trimmed to match).
Follow Video 1 from start to finish: copy its framing, its timing and every movement, gesture, head turn, facial expression and mouth shape, along with the plain orange background and the silver microphone hanging from the ceiling between the two performers. Keep the camera framing of Video 1: no push-in, no pull-back, no pan. Do not add any movement or scene element that is not in Video 1.
Use Audio 1 as the background music for the whole video, exactly as it is: it is the entire sound of the video. Do not write new lyrics, do not re-sing it, do not add other music or voices.
The performer on the left is the subject in Image 1; the performer on the right is the subject in Image 2. They completely replace the two people in Video 1: face, hair or fur, build and clothes all come from the photos, nothing from the people in Video 1. Neither performer wears gloves (unless already wearing them in the photo): people show their own hands, animals keep their own paws. Ignore the photo backgrounds.
The left performer raps the vocals of Audio 1 into the microphone, lips precisely in sync with Audio 1. The right performer follows the moves of the person on the right in Video 1: reacting and gesturing on the beat, mouthing only short ad-libs.
Orange fills the whole frame, with no black areas or dark corners. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.The Seedance version with an AI rap
Seedance never keeps the original recording; its soundtrack is always a rap it writes. For the Hotel Lobby booth it still gets to hear the template while writing: the reference clip is uploaded muted, the track is uploaded separately as reference audio 1, and the prompt asks for a new beat on the same rhythm, with fresh lyrics and none of the original melody or voices. Seedance refers to attachments in words (“reference image 1”). The version shown uses the sample topic below.
Use reference video 1 for the whole performance: copy its framing, timing and every movement, gesture, head turn, facial expression and mouth movement, the plain orange background and the single silver microphone hanging between the performers. Keep the camera framing of reference video 1: no zoom in, no zoom out, no pan. Do not add any movement or scene element that is not in reference video 1.
Performer on the left: the person in reference image 1. Performer on the right: the person in reference image 2. They replace the two people in reference video 1 completely: face, hair, body and clothes come from the photos, nothing from the people in reference video 1. If a photo shows none of a person's clothes (a face-only crop or bare shoulders), dress that person in plain everyday clothes that suit them, such as a simple T-shirt and jeans; never the costume, cape, coat, belts, gloves or blindfold of the people in reference video 1. Animals stay as they are, with no clothes added. Ignore the photo backgrounds.
Soundtrack: an original dark Atlanta trap song made for this video that follows reference audio 1 for its beat and rhythm: the same tempo, groove, drum pattern and bounce, timed to the moves in reference video 1. Make a new beat in that style with new lyrics; do not reuse the lyrics, melody or voices of reference audio 1. Sparse but hard-hitting: a thumping, gritty 808 that slides between notes, a piercing snare-clap with sharp hi-hats, and one eerie, Eastern-tinged melody loop. Dark, menacing, confident mood. The rap is a fast, bouncy triplet flow riding on top of the slow beat. Mix: vocals clear on top, the beat right behind them at nearly the same level, loud and punchy, never quiet and never dropping out between lines. The left performer raps an original, clean verse into the hanging microphone about: "Maya’s 30th birthday and her terrible parallel parking". Rap in the language of that topic. Clear rap vocals, lips in sync with every word, natural head nods and expressive mouth movement. The right performer is the hype partner, pointing at the rapper, reacting, smiling and adding short ad-libs on the beat. Do not use any existing song, existing lyrics or the voice of any real artist.
9:16. The orange fills the whole frame edge to edge, with no black or dark areas. No split screen, no extra people, no captions or text on screen. Faces stay consistent from start to end.Why each line is there
- Movement comes from the clip. The opening paragraph tells the model to take framing, timing, gestures, head turns, expressions and mouth shapes from Video 1, along with the bare orange backdrop and the single silver mic on its cable. It locks the camera (no push-in, pull-back or pan) and bans anything the clip does not contain.
- Sound stays untouched. Audio 1 is the entire soundtrack: no rewritten lyrics, no new vocals, no extra music or voices.
- Faces and bodies come from the photos. Each picture belongs to one side, and the two subjects fully replace the people in the clip, taking face, hair or fur, body and clothing from the pictures while ignoring their backgrounds. The wording is “the subject in Image 1” rather than man or woman, so it suits a person or a pet and never argues with the photo. A line also forbids gloves: the performers in the reference clip wear them, and models tend to copy them onto hands and paws.
- Each side gets a job. The left subject raps the Audio 1 vocals with matching lip movement; the right subject mirrors the clip’s hype partner, reacting on the beat with only brief ad-libs.
- Guard rails close the gaps. Orange fills the frame to every edge with no dark corners, no split screen, no extra people, no captions, and faces stay the same from first frame to last. These lines target the failures that come up most.
Moving the prompt to another AI video tool
- Your model must accept reference images, a reference video and a reference audio track, and must keep that track. Wan 3.0 does. ByteDance Seedance 2.x accepts all three inputs but generates its own music, which is why the Hotel Lobby Video generator uses it only for AI rap.
- Use your tool’s naming for attachments. Here they are numbered by upload order (Image 1, Image 2, Video 1, Audio 1); other tools expect @Image1 or their own tags. Whatever the syntax, upload the left photo first and the right one second.
- Leave the subjects generic. “The subject in Image 1” fits anyone, human or animal. The only things the Hotel Lobby Video generator changes are the photo layout (two photos or one of both) and, when you upload your own clip, the moves.
- Frame shape lives in the settings, not the text: choose 9:16, 16:9 or 1:1 there.
- No reference clip? Then you must describe the booth and the movement in words instead of pointing at Video 1. Dreamina and similar tools publish their own text-only prompts for this case. Only the two prompts above are what the generator sends.
Photo-only Hotel Lobby AI prompts (no reference clip)
Plenty of image-to-video tools accept reference photos but not a reference clip. The five prompts below spell out the room and the choreography in words, so they paste straight into a tool that takes two pictures. They hold on to what makes the format read: one photo per side, a single hanging mic, one lead and one reactor, and a camera that stays put. They were written for this page and are not what the Hotel Lobby Video generator sends; each tool interprets prompts its own way, so plan on a retry or two.
Hotel Lobby AI prompt for a couple
Create a 12-second vertical 9:16 video from two reference photos. The person in image 1 stands on the LEFT and the person in image 2 on the RIGHT, both full body, each keeping their own face, hair, clothes and height. They are a couple performing a rap duet in a plain, evenly lit orange studio: orange wall and floor, nothing else in the room. One silver condenser microphone hangs from the ceiling on a thin cable, centered between them. The left performer raps into the microphone with confident hand gestures; the right performer laughs, points and dances on the beat, then leans in for the last line. One continuous shot from a fixed camera at chest height: no cuts, zooms or pans. No other people, no captions, no logos.Two best friends
Create a 12-second vertical 9:16 video from two reference photos of two adult best friends. Image 1 is on the LEFT, image 2 on the RIGHT, both full body, keeping each face, hairstyle, outfit and body shape exactly. Setting: a plain orange studio with a seamless orange wall and floor and one silver microphone hanging from the ceiling between them. For the first half the left friend raps into the microphone while the right friend nods, grins and hypes with big gestures; for the second half they swap, and the left friend reacts. Their moves stay independent, never mirrored. Fixed camera, one continuous take, soft even light, both bodies and feet visible. No extra people, props, text or logos.Grandparent and grown grandchild
Create a 12-second vertical 9:16 video from two reference photos. Image 1 is an older adult on the LEFT and image 2 is their adult grandchild on the RIGHT, both full body, each keeping their own face, hair, clothes and height. They perform a rap duet in a plain orange studio with one silver microphone hanging from the ceiling between them. The grandparent raps into the microphone with slow, deliberate, confident gestures; the grandchild looks amazed, laughs and dances on the beat, then joins in for the last line. Natural, comfortable movement for both. One fixed, continuous shot with even studio light: no cuts, zooms or pans, no other people, no captions.Cartoon or anime characters
Stick to characters you created or hold rights to; famous characters are someone else’s property.
Create a 12-second vertical 9:16 video from two reference images of original illustrated characters. The character in image 1 stands on the LEFT and the character in image 2 on the RIGHT, full body, each keeping its art style, colors, proportions and outfit exactly; render both in that same style. They perform a rap duet in a plain orange studio with one silver microphone hanging from the ceiling between them. The left character raps into the microphone with expressive hand gestures; the right character reacts, points and dances on the beat. One fixed camera, one continuous take, no cuts, zooms or pans. No extra characters, no text, no logos.Hotel Lobby AI prompt for pets
Dogs and cats were among the earliest Hotel Lobby clips to spread, all made in other tools (watch a few). Give each animal one sharp photo with its whole body visible.
Create a 12-second vertical 9:16 video from two reference photos of pets. The animal in image 1 stands upright on its hind legs on the LEFT and the animal in image 2 on the RIGHT, each keeping its breed, face, eye color, fur pattern, size and collar. They are in a plain orange studio with one silver microphone hanging from the ceiling between them. The left animal moves its mouth to the beat and gestures with its front paws as if rapping into the microphone; the right animal bobs its head, looks at the camera and dances on the beat. Natural anatomy: four legs and one tail each, both fully visible from ears to paws. One fixed, continuous shot. No people, no extra animals, no text.The Hotel Lobby Video generator has its prompt built in, for people and animals alike, and the dog and cat rap page opens it set up for a pet plus a partner. A dog beside a person is the most predictable pairing; cats and two animals together are less certain, so start with your clearest photos.
When the render goes wrong
| What you see | Probable reason | What to change |
|---|---|---|
| A face drifts into a stranger | Profile shot, sunglasses or a face too small in frame | Swap in a front-facing, well-lit photo from the waist up |
| Left and right trade places | Nothing in the prompt links each photo to a side | State which image is the left performer and which the right |
| An extra person shows up | Someone from the reference or a photo background got copied | Keep “no extra people” and “ignore the photo backgrounds” |
| Text appears on screen | The model added captions on its own | Keep “no captions or text on screen” |
| Mouth and audio drift apart | Audio missing, or it starts on silence | Attach the track and trim it to begin on a word |
| Dark borders around the orange | A camera move was requested that the clip does not have | Hold the clip’s framing and ask for orange edge to edge |
| Different music replaces the track | The model composes its own soundtrack, as Seedance 2.x does | Switch to a model that keeps the reference track, such as Wan 3.0 |
Rather not write prompts?
Hotel Lobby Video ships with this prompt, the reference clip and the reference track already wired in. Add two photos and press create. 12 seconds on Wan 3.0 at 480p costs 6 credits.
FAQ
What should a Hotel Lobby AI prompt include?
The orange booth, one mic hanging between the pair, which photo is left and which is right, and what each person does. The Wan 3.0 prompt above is the one the Hotel Lobby Video generator sends with every standard video.
Which video models can run this prompt?
The first prompt goes to Wan 3.0, which keeps the reference track; the AI rap version goes to Seedance 2.x, which composes its own music. Both accept reference images, a reference video and reference audio. Other models use their own attachment syntax and may not take all three inputs.
I only have two photos and no reference clip. Which prompt do I use?
One of the five photo-only prompts: couple, best friends, grandparent and grandchild, characters or pets. They describe the booth and moves in words, so no clip is needed. Name the photos the way your tool expects and keep left and right fixed.
Is there a Hotel Lobby prompt for my dog or cat?
Yes, the pet prompt above: one photo per animal, each tied to a side, standing upright with natural anatomy. On Hotel Lobby Video you can skip it; the [dog and cat rap page](/dog-and-cat-rap-video) is preset for a pet and a partner. A dog next to a person is the safest pairing.
Will these prompts work in Kling or Dreamina?
The photo-only ones are written for any image-to-video tool that takes two reference photos, Kling and Dreamina included. They were not written against those tools, so expect to adjust naming and retry.
Can I make a Hotel Lobby AI video with no prompt at all?
Yes, with a two-photo maker. On Hotel Lobby Video the prompt is built in and you only upload the pictures.
References
Up next
Hotel Lobby Trend Tutorial: Pick Your Route, Then Post
A decision-first tutorial for the Hotel Lobby trend: which of three routes suits your time and budget, the steps for each, the free options and how to post
Hotel Lobby AI Video Generators Compared (Oct 4, 2026)
A dated, sourced comparison of Hotel Lobby AI video generators (Hotel Lobby Video, Rap Duo AI, LightX, Dreamina, Media.io, Summrs, Fuzana, Kapwing, plus Higgsfield, Starrd and phone apps): input, audio, price and free options
Best Photos for a Hotel Lobby AI Video (and Ones to Skip)
Choosing photos for Hotel Lobby AI: the cropped group-chat selfie mistake, what the portrait redraw does, two photos versus one, and what no photo can fix
Hotel Lobby Video is an independent product with no ties to Quavo, Takeoff, Migos, Quality Control Music, Motown or COLORS.
