The general shape
A UGC ad is usually three to six short clips stitched together with voiceover and on-screen copy. The Melius workflow:- Define the character — generate a character sheet so the same presenter and visual style appear across every cut.
- Generate each clip — image-to-video from a starting frame, using the character sheet for consistency.
- Add voiceover — use an audio node to generate the voice.
- Stitch the clips — use a stitch node to concatenate them into the finished video.
Step 1: Character sheet
A character sheet shows the same presenter from multiple angles with consistent features, proportions, clothing, and materials. Your presenter can be a live-action human or a fictional character in a style such as claymation, stop-motion, 2D illustration, or stylized 3D. The reference sheet keeps both the identity and the visual medium consistent across cuts and locations.For a speaking presenter, use a design with a visible, unobstructed mouth.
If the presenter holds or demonstrates a product, give the design usable,
visible hands. Voiceover-only videos can use a wider range of character
designs because the presenter does not need to lip-sync on camera.
1
Drop in reference art or a photo
Use artwork for an existing fictional design, a moodboard image, or a
custom-generated portrait. Try not to use a real identifiable person
without permission.
2
Brief the agent
“Help me build a character sheet from this reference. I want four or
five angles — front-facing, three-quarter left, three-quarter right,
profile, and a casual everyday shot. Use GPT Image 2 (Medium).
Image-to-image, keep the identity, body plan, surface materials, and
visual medium consistent.”
3
The agent generates the angles
You’ll get four to five image nodes. If the identity, anatomy, or visual
style drifts from the reference, regenerate just that node.
4
Unified-group the angles
Select the character sheet images and create a unified group. You now
have a single reference port you can drag into every downstream video
node.
Step 2: Generate the clips
For each clip in your script:1
Create a video node
Right-click → New video node. Pick Seedance 2.0 (best for character work
— we have un-gated face access) at 9x16, 720p (faster) or 1080p
(production).
2
Connect inputs
- Character sheet → reference image input (keeps face consistent). - Product pack shot → reference image input (if the product appears in the clip). - Brand anchor / script context → context input.
3
Write the clip prompt
Describe what’s happening in the clip and what the character is
doing/saying. Example:
9 seconds. Character (from @character-sheet) sitting at a kitchen counter, holding @product, looking directly at the phone camera UGC-style. Natural morning light. Slight head movement, conversational energy. Setting: casual home kitchen.4
Run with 2–3 variations
Video generations are slower and more expensive than image gens, so
don’t overdo variations. Three is usually plenty for the model to give
you a usable take.
5
Repeat for each clip
Duplicate the video node (
Cmd+C / Cmd+V) and change only the prompt
for each clip. The character sheet and product references stay
connected.Step 3: Voiceover
1
Create an audio node
Right-click → New audio node. ElevenLabs is the current voice provider.
2
Pick a voice
Roughly 20 voices are curated in the picker. We have access to the full
ElevenLabs catalog — if you need a specific voice that isn’t in the
picker, ping us in Slack.
3
Paste your VO script
The audio node will generate a clip from the script. Keep clips short —
under 30 seconds each — for quality and consistency.
Custom voice cloning at the team level is in active development. Once it
ships, you’ll be able to upload a voice sample and reuse it across canvases
for consistent persona-driven UGC. We’ll update this page when it’s live.
Step 4: Stitch the clips
1
Create a stitch node
Right-click → New stitch node. Or brief the agent: “Use a stitch node
to combine the video clips in order: opening, body 1, body 2, body 3,
closer.”
2
Connect each video clip
Drag from each video node into the stitch node in the order they should
appear.
3
Set output settings
Aspect ratio, resolution, fit-to-screen vs stretch-to-fill. Match the
platform you’re shipping to.
4
Run
The stitch node outputs a single video file with the clips concatenated.
Step 5: Finish in your editor
Download the stitched video and the voiceover audio (as separate files). Hand both to your video editor — they’ll add captions, sync the voiceover, drop in music, and handle the last-mile polish.Captions inside Melius are coming with the timeline editor. Until then,
captioning is the one part of UGC video work that still leaves the canvas.
Most teams handle it in their video editor of choice.
A common pattern: agency UGC at scale
Agencies producing UGC for clients typically structure work like this:- One project per client. Brand anchor lives in a text node at the top.
- One canvas per campaign. Holds the character sheet, product references, and all the clips.
- Claude drives the build. Via MCP, Claude reads the campaign brief and creative concept, then builds out the canvas — character sheet, clip nodes, voiceover, stitch — automatically. The designer or marketer comes in afterwards to pick winners and polish.
Common pitfalls
- Character drift across clips. Without a character sheet, every video generation can change the face or head design, anatomy, materials, or clothing. Always build the character sheet first.
- One clip per take, no variations. Video generations are probabilistic too. Run 2–3 variations per clip — the cost is real but small, and the time saved on bad takes is large.
- Trying to render a 30-second clip in one node. Models cap at 9–12 seconds. Plan your script around 3–4 short cuts.
- Forgetting product references in the clips that show the product. The model will invent a product that isn’t yours. Connect the pack shot to every video node that needs the product visible.