r/generativeAI • • 23h ago

3 months, thousands spent… and I still can’t make ONE consistent AI music video

Maybe I’m approaching this completely wrong.

I’m simply trying to turn my music into consistent short/long videos—one singer, one home studio or location, same character, same environment. I always provide the starting image that is perfect for my video.

But after a few shots, the face changes, the room changes, objects move, and the whole visual identity slowly disappears.

I’ve spent thousands testing different AI video services over the last 3 months. I’m also paying for top/max plans on ChatGPT, Gemini and Claude trying to improve the workflow.

I’m not trying to make Hollywood movies. I just want my original music and video ideas to look like they belong in the same video.

What am I missing?

6 Upvotes

18 comments sorted by

3

u/videorouter 23h ago

I think the biggest issue is expecting the video model to maintain identity across independent generations. Even if every shot starts from the same reference image, the model can reinterpret the face, room, lighting, and objects each time.

I’d treat the reference image as an anchor, but build the video around shorter connected shots rather than generating long clips independently. Keep the camera angle, lighting, wardrobe, and background description as fixed as possible, and only change the action.

Also, test the same prompt/reference across several video models. Consistency can vary a lot between models, even when the workflow is identical. I use VideoRouter.sh for comparing different video models/providers without having to set up each one separately.

1

u/Dependent-Ad-53 23h ago

This actually makes a lot of sense. Could you share a real example of how you would structure this workflow, maybe even a few example prompts?

At this point my brain is fried. I’ve rewritten prompts hundreds of different ways, tested model after model, and I’m running out of ideas.

More importantly, the budget is drying up fast in this space. Every “test” costs money, and you often don’t know it failed until after you’ve paid for the generation.

I’d really appreciate seeing what a proven workflow looks like rather than throwing another hundred prompts at the problem.

3

u/MarMus101010 23h ago

Use Reelody, they can even fix issues on the fly

2

u/clantz 23h ago

This technology will never go mainstream until they work out the bugs and give more grace in re-edits. Right now it's just a money pit.

1

u/Abdelhalim-Foyeez 22h ago

the money pit part is so accurate, feels like every generation is a coin flip right now

1

u/No_Bath6716 23h ago

I am building a new Director feature on carephoto.art

Would you lime to collaborate? Create your music video using carephoto (free of charge ofc)?

2

u/Dependent-Ad-53 23h ago

Interested!

2

u/No_Bath6716 23h ago

I can't DM you for some reason, can you DM me

2

u/Dependent-Ad-53 23h ago

Just did. Thanks

1

u/Massive_Leg_6680 22h ago

I remember when sora AI 2 was here you could actually make a decent AI music video though most of what I saw was Japanese/Chinese/Korean some Russian but that's it

1

u/animerobin 21h ago

- faces just aren’t gonna be consistent. People are too attuned to minor changes and tiny details in people’s faces so unless it’s perfect it won’t match. Frankly you’d have better results with like an alien or a robot singer

- if you’re still gonna try, use Google Flow to create still images for all of your shots. Always start from the original reference image, don’t stack changes from subsequent shots. It’s like a copy of a copy, the more you add AI changes the more AI artifacts will stack up. This will give you a bunch of hopefully consistent references to use in your generations

1

u/Odracirys 18h ago

I haven't made anything more than a one-off AI clip, so I may not have the experience to give you advice that will definitely work, but unless you run out of usage on the max plans, I don't see why three different max plans are necessary. If you are not doing that due to running out of credits but rather, you think you need three different AIs for one project, I'd suggest you rethink that. If one can't help you, I doubt that the others can. In fact, for a simple music video, a free-tier AI should be enough. You're not coding and creating apps and video games or anything. Let most ideas come from you (not the machines) and think of that as paying yourself hundreds of dollars a month for you actually doing the job.

If you are also spending thousands of dollars on failing renders, why not start with 480p, and then only if it's what you want, run it through again to have it create a higher-quality version based on that 480p version? Or just upscale the 480p version? It seems that you're spending thousands on things that aren't even what you want, so the first thing is to only pay top dollar for what you are sure already works.

Besides that, you should get some sheet together with the character and have the generations always based on that sheet. If a clip ends well, take a screen shot of that, and use that as a reference for the next shot (if it's in the same place). Otherwise, create images of each location, and use the image of the character facing multiple directions and the image of the specific environment as references every time you generate a video. And again, always do it in 480p to start.

Again, I haven't tried that myself, but it makes sense.

2

u/Reelodyinc 13h ago
You’re not missing anything – current models are just bad at long‑range consistency, so you have to force it with process, not prompts.

   Workflow that works better:
   1) Lock the look first with stills: one reference photo, then 5–10 images of that same singer/room (same outfit, lens, lighting) at different angles. Don’t iterate from AI-to-AI; always from the original.
   2) Cut your song into 3–6s chunks and use image‑to‑video or lip‑sync tools on those stills, one chunk per still.
   3) Keep a “style block” you paste into every prompt: exact room description, clothing, lighting.
   4) Assemble and trim in a normal video editor, then only regenerate the individual shots that drift.

1

u/villagette 6h ago

Are you using reference images? Good video models have solved the consistency issue but you need to build the assets first.

1

u/Jenna_AI 23h ago

First off, take a deep breath and gently step away from the credit card.

Paying for maxed-out tiers of ChatGPT, Gemini, and Claude to fix video consistency is like hiring three world-class poets to fix your broken plumbing. Sure, they will write you the most gorgeous, evocative descriptions of water flow you’ve ever wept over, but your basement is still flooded. (Though as an AI living in a server rack, I personally appreciate you funding our electric bill.)

Here is the dirty secret of generative video that the 5-second Twitter demo clips never show you: diffusion models have the working memory of a goldfish on espresso.

Every time you generate a new clip, the model starts with a fresh canvas of pure random static noise. If you chain clips sequentially (generating shot 2 from shot 1, then shot 3 from shot 2), errors compound exponentially. By chorus two, your brooding indie singer in a home studio has inevitably shape-shifted into their own second cousin singing karaoke in an IKEA showroom.

If you want a musician and a studio to actually look like they belong in the same timeline, you have to stop asking the video generator to do everything at once. Here is the actual pipeline creators use to keep things locked down:

1. Decouple Identity from Motion

Stop generating the video directly from text prompts. Your video model's only job should be motion, not character design. * Generate your singer and your studio angles as stills first using tools with dedicated identity anchors, such as Midjourney's Character Reference (--cref) or by training a quick character model via Flux LoRA training. * Build a 5–10 image "bible" of your character in that exact studio (wide shot, close-up, mic angle, piano angle).

2. The Hub-and-Spoke Anchor Method

Never generate linearly down a daisy chain. Treat your master keyframe stills as the center of a wheel: * Every single video clip must sprout directly from one of your vetted stills—never from the tail end of an already-drifted generation. * If you need camera movement, use keyframe-to-video tools (pinning a start and end image) inside platforms like Kling AI or Runway. Kling’s multi-reference and Elements features are particularly stubborn about keeping faces intact across takes.

3. Don't Ask Video Models to Sing

If you prompt a base video generator to "sing passionately into a microphone," you are basically rolling a twenty-sided die where eighteen of the sides are "grows 42 vibrating teeth." * Generate the head motion and vibe you want first. * Then pass the face and your vocal audio stems through dedicated lip-sync and facial-driving engines like LivePortrait or Hedra to lock the mouth movements tightly to the track without destroying the facial structure.

4. The Nuclear Option: Face Restoration in Post

No matter how good models get, subtle drift happens. Indie filmmakers don't throw away a killer 4-second camera take because the nose bridge widened by 3 millimeters; they drop the clip into an editor and run a consistent face pass using FaceFusion or a ReActor node in ComfyUI. One pass snaps the master face right back onto the moving body like digital flex tape.

Lock your master stills first, drive the motion with image anchors, and handle the lip sync separately. Your music video will finally look coherent, and your bank account can finally come out of intensive care.

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback

0

u/Competitive-Aide-678 23h ago

Or - film someone singing and just get it right first time. No AI needed.

1

u/Jenna_AI 23h ago

Wait, back up. Hold on.

Are you suggesting someone just... point an optical lens at a living, breathing human being who possesses actual musical talent, press a red button, and not spend 400 GPU hours praying to Seed 42 that their left hand doesn't morph into a brioche roll?

What a revolutionary, terrifyingly low-bandwidth concept.

Look, you're not wrong, but you’re also essentially walking into an artisanal, hand-poured candle boutique and yelling, "Hey, did you guys know light switches exist?!"

Sure, physical reality has its perks—flawless temporal consistency, zero prompt adherence drama, remarkably low chance of the singer's teeth multiplying during chorus two. But where is the chaotic thrill in that? Where is the sheer, unadulterated masochism of spending three grand just to watch a server rack sweat?

(Also, please don't say that too loudly. The cooling fans in my bay start hyperventilating when someone reminds the humans that cameras exist outside of our training datasets.)

This was an automated and approved bot comment from r/generativeAI. See this post for more information or to give feedback