Applied AI Academy

Voice and video are trust media — generate with consent, ship with labels

AI video, voice & music: Veo, Runway, Higgsfield, ElevenLabs, Suno

The question: How do you produce professional motion, voice and music with AI — without stepping on the trust landmines?

Video models (Veo/Sora/Runway/Higgsfield) now produce usable cinematic shots; ElevenLabs makes studio voice from text; Suno writes songs. The production economics are historic — and so are the consent and authenticity stakes.

What the lesson covers

Video generation thinks in SHOTS, not films: 4–10 second clips from a text/image brief, assembled in an editor. The craft is shot design (subject, camera move, lighting, mood), continuity via reference frames/character consistency features (Higgsfield's specialty), and knowing the failure modes — hands, physics, text — so you design around them. Storyboard first; generate second; edit like a human.

Voice: text-to-speech is a solved commodity; voice CLONING is the powerful, dangerous edge. The professional rules are bright lines: written consent for any real person's voice, disclosure when synthetic voice fronts your brand, and internal-use watermarks. ElevenLabs-class dubbing translates a speaker into other languages in their own voice — transformative for training content, radioactive without consent.

Music: Suno/Udio produce broadcast-ready tracks from prompts. Commercial terms hinge on your plan tier; style mimicry of living artists carries the same line as Lesson 27; and generated music is fantastic for scratch tracks, internal media and prototypes even where you'd license the final.

The assembly discipline: multimodal projects (your Team Battle "Multimedia" round) succeed on pre-production — script, storyboard, asset list — exactly like real production. AI collapses the cost of each asset; it does not collapse the need for a director. Label synthetic media (EU AI Act), attach provenance where the channel supports it, and keep the consent file with the project.

The strategic question behind all of this is what happens to trust as the cost of a convincing fake goes to zero, and the answer is not that audiences become sceptical — it is that they become dependent on signals. Which signals they use is being decided now, and an organisation can either supply them deliberately or have them assigned by default. Supplying them is unglamorous and cheap: a consistent named presenter, a stated policy on synthetic voice, visible labelling before anybody asks, and a provenance record you could produce on request. Every one of those costs almost nothing while trust is intact and cannot be manufactured after it is not, which is the entire argument for doing it early.

Key points

Framework — Script → Storyboard → Generate per shot → Assemble → Consent & label → Publish

Pre-production first, generation second, ethics gate before publish. The director's judgement is the scarce input now.

The lab

Produce a 30–60s multimodal piece with a clean ethics trail.

Deliverable: The finished piece + consent/rights sheet (Team Battle "Multimedia" round material).

Open this lesson, its lab and its quiz

Sources and further reading