Seed Audio 1.0 by ByteDance Seed: An Honest Explainer (Text-to-Audio, Not Just TTS)
Every page about Seed Audio 1.0 recycles the same vendor bullet list. This one does three things those pages skip: it separates API-doc-verified specs from marketing, normalizes the scattered reseller prices into a like-for-like table, and states plainly that no independent benchmark exists. That honesty is the whole point.
What Seed Audio 1.0 actually is
Seed Audio 1.0 is ByteDance Seed's all-in-one text-to-audio model. It was announced on June 23, 2026 at Volcano Engine FORCE 2026; some API docs list a June 29, 2026 availability date, which likely reflects the API go-live rather than the event announcement. The core idea, consistent across every source we checked, is that it renders a full sound scene from a single prompt: spoken dialogue, background music, ambience, and foley-style sound effects, all composed in one generation pass.
That framing matters. Most tools you have heard of are text-to-speech only. They turn words into a voice. Seed Audio is positioned as text-to-any-audio: describe a scene and it produces the whole soundstage at once. That single-pass composition is the honest differentiator, and it is corroborated by three independent API-provider docs that actually host the model: fal.ai, Segmind, and WaveSpeed.
Before we go further, a disclosure that no competing page seems to make: we did not benchmark this model ourselves. This is an explainer, not a first-party test. Where a claim is verified against rendered API documentation we say so; where a claim is vendor or reseller marketing we flag it; and where evidence simply does not exist, we say that too.
Verified specs vs marketing claims
Here is the separation the SEO pages never draw. The left column below is what two or three independent provider API docs agree on. The right column is what appears only in vendor or reseller marketing and is not confirmed in the rendered first-party docs.
| Attribute | Status | Detail |
|---|---|---|
| Output formats | API-doc verified | wav, mp3, pcm, ogg_opus (mp3 default) — fal.ai + WaveSpeed + Segmind agree |
| Sample rates | API-doc verified | 8000 / 16000 / 24000 / 32000 / 44100 / 48000 Hz; 24000 Hz default |
| Voice cloning | API-doc verified | Up to 3 reference clips, each up to 30s / 10MB; referenced as @Audio1/@Audio2/@Audio3; zero-shot, no fine-tuning |
| Delivery controls | API-doc verified | Speed 0.5–2.0, volume 0.5–2.0, pitch -12 to +12 semitones (defaults 1.0 / 1.0 / 0) |
| Prompt length | API-doc verified | Max 2048 characters |
| Clip length | Provider-consistent | Up to ~2 minutes (roughly 120s) per generation, voice held consistent; extendable |
| Multi-character dialogue | Provider-consistent | Distinct voice, emotion, pacing per speaker in one pass; no independent test verifies quality |
| Six preset languages | Marketing only | EN/ZH/ES/JA/ID/PT claimed by resellers; rendered fal.ai page does not list languages; EN+ZH implied as functional core |
| Independent benchmark | Does not exist | No arXiv paper, no third-party benchmark, no major-press launch coverage found |
A note on the audio reference rule that matters if you plan to build with it: per the fal.ai and Segmind docs, audio references cannot be combined with an image input. If your workflow needs both, plan around it.
How it differs from TTS-only tools
The lazy version of this comparison says "ElevenLabs can't do X." That is not true and we will not write it. ElevenLabs has added multi-speaker dialogue and sound-effects features too. The fair, defensible difference is architecture and access: Seed Audio generates voice, music and SFX together in a single pass, and it is delivered primarily through ByteDance's enterprise clouds. That is a genuine distinction in the model and the access model, not a capability gap you can wave around.
If you are mapping the broader generative-media landscape, our companions on best text-to-video 2026 and best text-to-image 2026 cover the visual side of the same shift toward single-prompt scene generation.
One key, 300+ models
DataLLM Lab is an OpenAI-compatible gateway. Route text, reasoning and multimodal models through a single key while you decide which audio channel fits your stack.
Normalized reseller pricing
This is the part every other page gets wrong or omits. There is no confirmed first-party Volcano Engine list price. Every number below is a third-party reseller rate, as of research date, and they differ. Do not present any of these as ByteDance's official rate.
| Provider (reseller) | Quoted price | Normalized to /min | Verification |
|---|---|---|---|
| fal.ai | $0.1875 / min | $0.1875 | Model page (verified) — cleanest anchor |
| AIMLAPI (aggregated) | $0.00325 / sec | ~$0.195 | Secondary aggregation |
| PiAPI (aggregated) | $0.002 / sec | ~$0.12 | Secondary aggregation |
| WaveSpeed | from $0.30 / run | varies | API (verified); scales with params |
| Atlas Cloud | from $0.015 / 1M tokens | not directly comparable | Aggregation, less verified |
| Volcano Engine (first-party) | no confirmed list price | — | Not found |
The clean takeaway: on a per-minute basis the verified rates cluster loosely around $0.12–$0.20 per minute, with fal.ai's $0.1875/min as the most defensible anchor because it is a flat, rendered figure. WaveSpeed and Atlas Cloud price on different units (per-run, per-token), so they are not directly comparable and we do not force them into a false equivalence.
What a real 2-minute scene costs
Numbers in isolation do not help you budget. So here is the concrete reframe no competing page offers. Suppose you generate a 2-minute multi-character scene — the model's stated per-generation ceiling — at the cleanest verified anchor of $0.1875/min:
2 min × $0.1875 = $0.375 per finished scene (fal.ai, modeled from the verified per-minute rate).
At the low end of the aggregated spread ($0.002/sec = $0.12/min) the same scene is about $0.24; at the high end ($0.00325/sec = $0.195/min) it is about $0.39. So a full two-minute dialogue-plus-music-plus-SFX scene lands roughly in the $0.24–$0.39 band depending on provider. These are modeled figures derived from published reseller per-second/per-minute rates, not a first-party quote and not a bill we ran.
This is the same cost-per-task lens we apply to visual models in Sora 2 cost and Nano Banana pricing: don't compare unit prices, compare what one finished deliverable actually costs.
How to access it
There are three official channels and several third-party APIs. Be careful here: many top search results (seedaudio.im, seedaudio.co, seedaudio.run and similar) are affiliate/SEO domains, not ByteDance. The only official channels are below.
- Volcano Engine — ByteDance's China enterprise cloud, via API. This is the enterprise access reality.
- BytePlus — for international developers; an official doc exists at docs.byteplus.com (the product is listed as "Audio 1.0 / Seed Speech"). The page is a single-page app that did not fully render for us, so treat the listing as confirmed-to-exist rather than fully read.
- Doubao app — consumer access. Integrations into CapCut/Jianying, Jimeng and Fanqie have been reported, but only in secondary sources.
- Third-party APIs — fal.ai, WaveSpeed, Atlas Cloud and Segmind host the model at the differing prices covered above.
If the "gateway" model of access is new to you, our primer on what an LLM gateway is explains why routing many models through one endpoint beats juggling a dozen provider keys — the same logic that makes multi-provider audio access manageable.
Honest verdict
Seed Audio 1.0 is real, hosted on multiple live APIs, and its core capability — one prompt to a full sound scene in a single pass — is well corroborated. If you want dialogue, music and effects composed together and you are comfortable with ByteDance-cloud access, it is worth a trial.
But be clear-eyed about what is not established. There is no technical paper, no independent benchmark, and no first-party press coverage. Quality claims — foley realism, "better than SeeDance 2" — trace to vendor demos or a single social-media anecdote (an X post from @TomLikesRobots claiming a ~15-second test cost a couple of cents). The six-language preset list is marketing, not confirmed in rendered docs. And no official price has been published. Treat the enthusiasm accordingly, run your own listening test on a real script, and price it per finished scene, not per unit.
FAQ
What is Seed Audio 1.0?
Seed Audio 1.0 is ByteDance Seed's all-in-one text-to-audio model, announced June 23, 2026 at Volcano Engine FORCE 2026. Unlike text-to-speech-only tools, it renders a full sound scene from a single prompt: spoken dialogue, background music, ambience and foley-style sound effects in one generation pass.
How is Seed Audio 1.0 different from ElevenLabs?
The honest differentiator is single-pass full-scene generation plus enterprise-cloud access. Seed Audio composes voice, music and SFX together in one call and is delivered through ByteDance clouds (Volcano Engine in China, BytePlus internationally). ElevenLabs has also added multi-speaker dialogue and sound effects, so this is a difference in the generation model and access channel, not a blanket claim that ElevenLabs cannot do something.
How much does Seed Audio 1.0 cost?
There is no confirmed first-party Volcano Engine list price. Verified third-party reseller pricing as of research date: fal.ai charges about $0.1875 per minute (roughly $0.003125 per second), WaveSpeed starts around $0.30 per run scaling with parameters, and secondary aggregation cites a per-second spread of about $0.002 to $0.00325. Treat fal.ai as the cleanest verified anchor and expect prices to vary by provider.
Does Seed Audio 1.0 support voice cloning?
Yes. It offers zero-shot voice cloning from up to three reference clips, each up to 30 seconds and 10MB, in wav, mp3, pcm or ogg_opus. Clips are referenced in the prompt as @Audio1, @Audio2 and @Audio3 with no training or fine-tuning. Audio references cannot be combined with an image input, per the fal.ai and Segmind API docs.
How many languages does Seed Audio 1.0 support?
Vendor and reseller pages claim preset voices for English, Chinese, Spanish, Japanese, Indonesian and Portuguese with cross-lingual synthesis. This is marketing, not first-party confirmed: the rendered fal.ai model page does not list supported languages, and an earlier fal snippet implied English and Chinese are the functional core. Treat the six-language list as unconfirmed.
Are there independent benchmarks for Seed Audio 1.0?
No. As of July 2026 there is no arXiv paper, no independent benchmark, and no first-party major-tech-press launch coverage. Evidence of quality is limited to vendor demos and a couple of anecdotal social-media user tests. Any quality ranking versus ElevenLabs should be read as vendor demo or single-user anecdote, not tested fact.
DataLLM Lab