Dev.to WebDev 🛠 Dev 👁 0 📖 2 min read

Designing a two-photo AI video workflow with explicit settings and credit costs

I work on Hotel Lobby AI, a focused photo-to-video app that turns two uploaded images into a short duo clip in an orange studio. This post looks at its user flow and a few interface decisions that are useful for other AI

I work on Hotel Lobby AI, a focused photo-to-video app that turns two uploaded images into a short duo clip in an orange studio. This post looks at its user flow and a few interface decisions that are useful for other AI media tools.

Start with a bounded task

A general text box asks users to describe the scene, performers, motion and output format at once. A template can make those decisions explicit. In this app, the starting inputs are two real photos or anime-style character images. The user then chooses a template, duration and format instead of writing a prompt.

For a developer, this suggests treating the form as a small configuration object. Here is an illustrative shape, not the public API of this app:

type ClipSettings = {
  durationSeconds: 5 | 10 | 15;
  aspectRatio: "landscape" | "portrait" | "square";
  audioSource: "template" | "upload";
};

That object still needs server-side validation. A client-side select menu only improves usability; it does not establish which model options or credit costs are valid.

Make the cost visible before generation

Hotel Lobby AI uses paid credits and shows the cost before the user confirms. Duration and quality affect the generation request, so a useful review screen should show the same settings that will be submitted.

In a similar app, I would keep the estimate tied to the selected configuration and invalidate it when those settings change. Avoid letting an old estimate stay visible after a user switches quality.

Separate input quality from output promises

The upload flow accepts JPG, PNG and WEBP images, up to 10 MB each. Clear faces and good lighting are more useful than a long prompt when the task depends on two portraits.

The interface also needs to explain what AI can change. Faces, clothing and motion can vary; exact likeness is not guaranteed. Showing this near the workflow gives users a practical way to judge the result.

Design the result step as part of the workflow

The app supports 5, 10 and 15 second clips, landscape, portrait and square outputs, and template audio or uploaded WAV/MP3. Once a result is available, users can preview and download it.

For builders, the broader lesson is to make a generation task understandable from upload through result. A focused template, explicit settings and visible cost can reduce the decisions users must make before they get something useful.

You can see the workflow at Hotel Lobby AI. I am associated with the project. This article was prepared with AI assistance from the app's current product information.

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.