ElevenLabs will take your existing ads, localize them into 50+ languages, and push them back into Google, Meta, and LinkedIn. The only quality gate in that pipeline is an optional human approval step. Generation got cheap. Checking the output is still a person with headphones.
In August I built an MVP of that check. It’s called Soundcheck.
Somebody used to listen to the spot before it aired
At Crispin Porter + Bogusky, a radio spot went through a producer, a copywriter, the account team, and legal before a client heard it. Someone confirmed the brand name was pronounced the way the client says it. Someone timed the read. Someone checked the disclaimer made it in. That worked because a campaign was six spots. Now one person can render six hundred by Friday, and nothing checks six hundred.
Every vendor checks something. None of them listen to the ad.
DoubleVerify launched AI-powered brand suitability for audio in June; it analyzes the podcast or playlist your ad runs inside, not the ad. AudioStack partnered with Adclear to screen scripts for regulatory problems before the audio exists. Adobe GenStudio’s brand compliance scores generated work against your guidelines, and its page doesn’t mention audio once. CreativeX and Vidmob do the same for images and video. Nobody on that list listens to the rendered file that goes on air.
What the MVP does
It starts with the brand kit a client would already have: brand bible, pronunciation lexicon, campaign brief, media plan, talent consent scope. Soundcheck compiles those into one machine-readable standard, generates a 12-spot, four-market campaign for a fictional cold-brew brand through six ElevenLabs APIs (Voice Design, text-to-speech, Music, sound effects, Dubbing, and Scribe v2 to transcribe every render back with word-level timestamps), then runs six checks on each spot: script fidelity, pronunciation against the lexicon, pace, loudness, whether the legal and AI-disclosure lines made it in, and voice drift.
The disclosure check got more important on August 2, when Article 50 of the EU AI Act started requiring deployers to disclose AI-generated audio. Every asset also gets a provenance entry (voice ID, model version, prompt, consent scope, file hash, timestamp), which is the audit trail a brand’s counsel and the SAG-AFTRA digital-replica paperwork both ask for.
Every finding rolls up against the media plan, because “asset 7 has a pace violation” gets ignored and “$650,000 of the plan is sitting behind a mispronounced brand name” gets a meeting.
What the live run caught
I seeded four defects into the twelve scripts: an invented retail claim, a mispronounced product term, a 107-word script jammed into a :30, and a missing AI-disclosure line. On August 13 I ran it against the live APIs. Seven passed, five failed. All four seeded defects were caught, and $1.82M of the fictional brand’s $2.8M plan sat behind creative that would have shipped with a problem in it.
The fifth failure was the one I hadn’t planted. The brand voice couldn’t say the founder’s name: “Renata Oyelaran” came back as “renata ollerenshaw,” in the hero spot, on the placement with the most money behind it. The script was fine, so nobody reading it would have caught that.
There was one miss. Scribe quietly corrected two of the staged mispronunciations before my check could see them. That’s the ceiling on checking pronunciation through a transcript; a real version needs phoneme-level alignment.
What to ask before you buy
If a vendor is pitching you on volume, ask what happens between the render and the ad server. Who or what checks each file against your brand standard, and what does it cost per asset? If the answer is a person, ask how that works at a thousand assets. If the answer is nothing, that’s the risk you’re carrying, and it’s biggest in the markets where you can’t judge the audio yourself.
This is an MVP, not a product. Thresholds are hand-set, drift is a loudness proxy, and pronunciation runs off transcripts. I’m building it out from here, starting with phoneme-level scoring. And if you know of something that already solves this that I’m not aware of, please let me know.
