Voice and video are trust media — generate with consent, ship with labels
AI video, voice & music: Veo, Runway, Higgsfield, ElevenLabs, Suno
The question: How do you produce professional motion, voice and music with AI — without stepping on the trust landmines?
Video models (Veo/Sora/Runway/Higgsfield) now produce usable cinematic shots; ElevenLabs makes studio voice from text; Suno writes songs. The production economics are historic — and so are the consent and authenticity stakes.
What the lesson covers
Video generation thinks in SHOTS, not films: 4–10 second clips from a text/image brief, assembled in an editor. The craft is shot design (subject, camera move, lighting, mood), continuity via reference frames/character consistency features (Higgsfield's specialty), and knowing the failure modes — hands, physics, text — so you design around them. Storyboard first; generate second; edit like a human.
Voice: text-to-speech is a solved commodity; voice CLONING is the powerful, dangerous edge. The professional rules are bright lines: written consent for any real person's voice, disclosure when synthetic voice fronts your brand, and internal-use watermarks. ElevenLabs-class dubbing translates a speaker into other languages in their own voice — transformative for training content, radioactive without consent.
Music: Suno/Udio produce broadcast-ready tracks from prompts. Commercial terms hinge on your plan tier; style mimicry of living artists carries the same line as Lesson 27; and generated music is fantastic for scratch tracks, internal media and prototypes even where you'd license the final.
The assembly discipline: multimodal projects (your Team Battle "Multimedia" round) succeed on pre-production — script, storyboard, asset list — exactly like real production. AI collapses the cost of each asset; it does not collapse the need for a director. Label synthetic media (EU AI Act), attach provenance where the channel supports it, and keep the consent file with the project.
The strategic question behind all of this is what happens to trust as the cost of a convincing fake goes to zero, and the answer is not that audiences become sceptical — it is that they become dependent on signals. Which signals they use is being decided now, and an organisation can either supply them deliberately or have them assigned by default. Supplying them is unglamorous and cheap: a consistent named presenter, a stated policy on synthetic voice, visible labelling before anybody asks, and a provenance record you could produce on request. Every one of those costs almost nothing while trust is intact and cannot be manufactured after it is not, which is the entire argument for doing it early.
Key points
- Generate shots, assemble films; storyboard first and design around failure modes (hands, physics, text).
- Voice cloning: written consent + disclosure + revocation clause — bright lines, no exceptions.
- Music: great for scratch/internal; check plan terms for commercial release; same style-mimicry line.
- AI collapses asset cost, not the need for a director; label synthetic media and archive consent.
- As convincing fakes get free, audiences do not become sceptical — they become **dependent on signals**. Supply them deliberately: a named presenter, a stated synthetic-voice policy, labels before anyone asks, a provenance record.
Framework — Script → Storyboard → Generate per shot → Assemble → Consent & label → Publish
Pre-production first, generation second, ethics gate before publish. The director's judgement is the scarce input now.
The lab
Produce a 30–60s multimodal piece with a clean ethics trail.
Open this lesson, its lab and its quiz
Sources and further reading
- ElevenLabs — voice AI (consent & safety pages) — ElevenLabs
- EU AI Act — synthetic content transparency — AI Act Explorer