Facilitator notes — Brand Voice Studio
Presenter-only. Generated from the Engage demo kit.
Talk track
- One campaign brief becomes a full asset set — brand voice, a 30-second ad script, an SFX bed, a music track, and dubs into six markets — from a single Creative Platform run. The pitch is speed and breadth from one brief, not any single API; say that up front, the same framing as scenario 3, just for a campaign brief instead of a lifecycle moment.
- Point at the audio tags in the ad script — [warm], [pause], [confident] — as the difference between v3 TTS and flat narration. This is expressive delivery control written directly into the script, not a post-processing effect, and it's worth reading a line of the script out loud from the brief before playing the generated audio so the audience hears the intent land.
- The jingle is the one deliberate exception to D4's 'everything pre-rendered' rule in the whole kit — the presenter picks a genre and it generates live via Music v2 on stage, with a fallback take already loaded in case it's slow. Naming that trade-off explicitly, rather than letting it look like every other pre-rendered asset, is itself a mini lesson in how to demo AI reliably: know exactly which moment is live, and always have a fallback ready for it.
- Same honesty point as scenario 3, worth repeating rather than assuming it carries over: the brand voice here is built via TTS against an existing voice, not the Voice Design creation flow itself, since that isn't exposed through the tools used to generate these assets. Say this before a technical buyer asks.
Discovery questions
- How many markets or languages does your creative team produce ad content for today, and what's the current cost and turnaround per market?
- Who owns brand voice consistency today across agencies, freelance VO artists, or regional teams — and how is that actually enforced?
- Would your brand team want live, on-stage control over something like a genre-selectable jingle, or does a sonic identity like that need to be a single locked, approved asset?
- Realistically, how many campaign briefs like this would move through your team in a quarter, and what's the bottleneck today — ideation, production, or localisation?
Objection handling
Six-market dubbing sounds efficient, but can we really trust it for brand-voice consistency?
Flip the comparison: the alternative most brands run today is six different local VO artists interpreting the same brief six different ways, which is actually less consistent, not more. Dubbing from one script and one brand voice keeps the same voice across every market — a human review pass before shipping is still the normal step, same as it would be with six separate human recordings.
Generating a jingle live on stage feels risky for a demo.
It is, deliberately, and that's exactly why a fallback take is already loaded before the call — if the live generation is slow or fails, switch to the pre-rendered take and say so plainly rather than stalling on stage. Every other asset in this scenario is pre-rendered for exactly this reason; the jingle is the one place the kit chose to show a live capability anyway, with a safety net built in rather than hoping it works.
If a tool call fails or times out
- The live jingle generation is slow or fails: switch to the pre-rendered fallback take immediately and say so out loud — this is the expected, planned-for outcome for the one live-generation moment in the scenario, not an error to recover from silently.
- A market's dub doesn't play in the side-by-side player: don't skip it silently — since every dub is pre-rendered per D4, a missing one means the asset row or Storage upload needs checking, not that dubbing itself failed live. Move to the next market and flag it for a fix before the next run-through.
- Asked to create a brand-new voice live, on the spot: say plainly that this specific integration surface only does TTS against an existing voice, not voice creation, and point at the dashboard's own Voice Design feature as the real path — don't attempt to fake it with a different stock voice and imply it's new.