Skip to content

Glossary

Audio generation

Audio generation produces music, atmosphere or effects from a written description, which is a different job from reading text aloud and carries a rights question that speech does not.

In plain terms

Describing a sound and getting it back: a piece of music, a background atmosphere, a door closing. It is not the same as having text read aloud in a voice, which is a separate thing. The output is often startlingly good, and the question that actually decides whether you can use it is who owns what came out.

01

Why it matters

Because quality stopped being the constraint and rights did not. A generated track can be entirely usable and still be something you hold under a licence granted by the vendor rather than something you own, and that distinction only surfaces when a piece of work is commercial, published and hard to unpick. It is a procurement question wearing creative clothing.

02

How it works

The output arrives complete rather than as parts, which shapes what you can do with it. A track comes back with its arrangement already decided, so revising it usually means generating again rather than adjusting a layer, and a piece that is nearly right is often further from usable than one that is obviously wrong. Iteration is regeneration.

Rights are granted rather than transferred, and the wording differs sharply between tiers. Free tiers commonly restrict output to personal use, paid tiers commonly grant commercial rights as a licence from the vendor, and none of that is the same as owning what came out. Anything going into an advertisement or a product deserves the terms read rather than assumed.

The training question sits underneath the licensing one and is unsettled. Where the material a system learnt from came from is the subject of continuing dispute in this modality more than in most, and vendors have begun moving towards licensed catalogues in response. That direction is a real change and it is not yet a settled position.

It is strongest where the audio is a bed rather than the point. Atmosphere behind a video, a placeholder while a piece is being cut, effects nobody will listen to closely and drafts that let a decision be made: all of these are served extremely well. Anything where the music is what the audience came for is a much harder ask.

Two questions, and the order they arrive in

Two questions, and the order they arrive inThe reason the rights question lands late is that nothing about generating audio feels like procurement. Somebody describes a mood, gets back something genuinely good, drops it under a video, and the whole interaction takes less time than reading a paragraph of terms would. There is no moment in that flow where a licensing question naturally announces itself, and the output carries no marking that would raise one. So the discovery usually happens much later and from the wrong direction: a piece of work is going out to a real audience, somebody in a legal or brand role asks where the music came from, and the honest answer turns out to be a licence granted under a tier that may since have changed. None of that makes generated audio a bad idea; it makes the terms worth reading once, early, for the tier you actually use, rather than each time under pressure.Is it good?Answered in the first minute.Increasingly yes.Easy to demonstrate.Not the constraint any more.May we use it?Answered in the terms.Depends on tier and on vendor.Granted, rarely transferred.The constraint that actuallybites.Almost every evaluation runs theleft column and almost everyproblem comes from the right.The gap is widest exactly whenthe work is commercial andalready published.
The reason the rights question lands late is that nothing about generating audio feels like procurement. Somebody describes a mood, gets back something genuinely good, drops it under a video, and the whole interaction takes less time than reading a paragraph of terms would. There is no moment in that flow where a licensing question naturally announces itself, and the output carries no marking that would raise one. So the discovery usually happens much later and from the wrong direction: a piece of work is going out to a real audience, somebody in a legal or brand role asks where the music came from, and the honest answer turns out to be a licence granted under a tier that may since have changed. None of that makes generated audio a bad idea; it makes the terms worth reading once, early, for the tier you actually use, rather than each time under pressure.
03

Seen in the wild

  • A flagship video model whose dialogue, ambient sound and effects are generated with the pictures rather than laid on afterwards.

    Veo (Google)
  • An audio platform whose range spans generated effects alongside speech work and dubbing finished material into other languages.

    ElevenLabs
  • A tool taking text and image to video with generated sound effects matched to the action, so the audio arrives with the pictures.

    Pika
04

Common misconceptions

People assume

If I generated it, I own it.

In fact

Commonly you hold it under a licence the vendor grants, which is a weaker and more conditional thing. Free tiers frequently restrict output to personal use and paid tiers frequently grant commercial rights without transferring ownership, so anything going into published work needs the terms read rather than assumed.

People assume

It is the same capability as reading text aloud.

In fact

They are separate jobs that often share a vendor. Speech turns words into a voice, and this turns a description into music or sound, so a product excellent at one may not offer the other at all. The rights position also differs, because a generated composition raises questions a spoken sentence does not.

05

Telling them apart

Audio generation vs Text to speech

Audio generation

A description becomes music or sound.

Text to speech

Words become a voice saying them.

Vendors sell both and the licensing differs, because a generated composition raises ownership questions that reading a sentence aloud does not.

06

Questions

Can we use this in something commercial?
Read the tier's terms before assuming, because they vary and they change. Free tiers commonly restrict to personal use, paid tiers commonly grant commercial rights as a licence rather than as ownership, and the practical consequence is that your position depends on a vendor's continuing grant. For anything expensive to unpick later, that is worth checking first.
Why can I not just tweak the part I dislike?
Because the output arrives as a finished piece rather than as separable layers. Adjusting one element usually means generating again and accepting a different whole, which is why a nearly-right result can be more frustrating than an obviously wrong one. Treat each attempt as a draft rather than as something to refine.
Where is it genuinely reliable?
Where the audio supports something else. Atmosphere under a video, placeholders during an edit, effects that will not be listened to closely and drafts that let a decision get made are all handled well and cheaply. Music that an audience is attending to on its own terms is a much harder ask and usually still a human one.
07

Key takeaways

  • Output arrives complete, so iteration means regenerating rather than adjusting.
  • Rights are usually granted by licence, not transferred as ownership.
  • The training-source question is genuinely unsettled in this modality.
  • It is strongest where the audio is a bed rather than the point.
09

Tools that use this

  • Veo (Google)

    Dialogue, ambience and effects generated alongside the pictures.

  • ElevenLabs

    Generated effects alongside speech work and dubbing.

  • Pika

    Sound effects generated to match the action in the video.

Last checked July 2026

All glossary terms