Model Review

Seed Audio 1.0 by ByteDance Seed: An Honest Explainer (Text-to-Audio, Not Just TTS)

Every page about Seed Audio 1.0 recycles the same vendor bullet list. This one does three things those pages skip: it separates API-doc-verified specs from marketing, normalizes the scattered reseller prices into a like-for-like table, and states plainly that no independent benchmark exists. That honesty is the whole point.

Diagram of Seed Audio 1.0 turning one text prompt into a full sound scene

What Seed Audio 1.0 actually is

Seed Audio 1.0 is ByteDance Seed's all-in-one text-to-audio model. It was announced on June 23, 2026 at Volcano Engine FORCE 2026; some API docs list a June 29, 2026 availability date, which likely reflects the API go-live rather than the event announcement. The core idea, consistent across every source we checked, is that it renders a full sound scene from a single prompt: spoken dialogue, background music, ambience, and foley-style sound effects, all composed in one generation pass.

That framing matters. Most tools you have heard of are text-to-speech only. They turn words into a voice. Seed Audio is positioned as text-to-any-audio: describe a scene and it produces the whole soundstage at once. That single-pass composition is the honest differentiator, and it is corroborated by three independent API-provider docs that actually host the model: fal.ai, Segmind, and WaveSpeed.

Before we go further, a disclosure that no competing page seems to make: we did not benchmark this model ourselves. This is an explainer, not a first-party test. Where a claim is verified against rendered API documentation we say so; where a claim is vendor or reseller marketing we flag it; and where evidence simply does not exist, we say that too.

Verified specs vs marketing claims

Here is the separation the SEO pages never draw. The left column below is what two or three independent provider API docs agree on. The right column is what appears only in vendor or reseller marketing and is not confirmed in the rendered first-party docs.

AttributeStatusDetail
Output formatsAPI-doc verifiedwav, mp3, pcm, ogg_opus (mp3 default) — fal.ai + WaveSpeed + Segmind agree
Sample ratesAPI-doc verified8000 / 16000 / 24000 / 32000 / 44100 / 48000 Hz; 24000 Hz default
Voice cloningAPI-doc verifiedUp to 3 reference clips, each up to 30s / 10MB; referenced as @Audio1/@Audio2/@Audio3; zero-shot, no fine-tuning
Delivery controlsAPI-doc verifiedSpeed 0.5–2.0, volume 0.5–2.0, pitch -12 to +12 semitones (defaults 1.0 / 1.0 / 0)
Prompt lengthAPI-doc verifiedMax 2048 characters
Clip lengthProvider-consistentUp to ~2 minutes (roughly 120s) per generation, voice held consistent; extendable
Multi-character dialogueProvider-consistentDistinct voice, emotion, pacing per speaker in one pass; no independent test verifies quality
Six preset languagesMarketing onlyEN/ZH/ES/JA/ID/PT claimed by resellers; rendered fal.ai page does not list languages; EN+ZH implied as functional core
Independent benchmarkDoes not existNo arXiv paper, no third-party benchmark, no major-press launch coverage found

A note on the audio reference rule that matters if you plan to build with it: per the fal.ai and Segmind docs, audio references cannot be combined with an image input. If your workflow needs both, plan around it.

How it differs from TTS-only tools

The lazy version of this comparison says "ElevenLabs can't do X." That is not true and we will not write it. ElevenLabs has added multi-speaker dialogue and sound-effects features too. The fair, defensible difference is architecture and access: Seed Audio generates voice, music and SFX together in a single pass, and it is delivered primarily through ByteDance's enterprise clouds. That is a genuine distinction in the model and the access model, not a capability gap you can wave around.

One text prompt + up to 3 @Audio refs Seed Audio 1.0 single generation pass ~120s, voice consistent Dialogue (multi-voice) Background music Ambience Foley SFX
One prompt, one pass, a full soundstage — the structural difference from TTS-only tools. Illustrative structure, not measured output. Chart: DataLLM Lab

If you are mapping the broader generative-media landscape, our companions on best text-to-video 2026 and best text-to-image 2026 cover the visual side of the same shift toward single-prompt scene generation.

One key, 300+ models

DataLLM Lab is an OpenAI-compatible gateway. Route text, reasoning and multimodal models through a single key while you decide which audio channel fits your stack.

Normalized reseller pricing

This is the part every other page gets wrong or omits. There is no confirmed first-party Volcano Engine list price. Every number below is a third-party reseller rate, as of research date, and they differ. Do not present any of these as ByteDance's official rate.

Provider (reseller)Quoted priceNormalized to /minVerification
fal.ai$0.1875 / min$0.1875Model page (verified) — cleanest anchor
AIMLAPI (aggregated)$0.00325 / sec~$0.195Secondary aggregation
PiAPI (aggregated)$0.002 / sec~$0.12Secondary aggregation
WaveSpeedfrom $0.30 / runvariesAPI (verified); scales with params
Atlas Cloudfrom $0.015 / 1M tokensnot directly comparableAggregation, less verified
Volcano Engine (first-party)no confirmed list priceNot found

The clean takeaway: on a per-minute basis the verified rates cluster loosely around $0.12–$0.20 per minute, with fal.ai's $0.1875/min as the most defensible anchor because it is a flat, rendered figure. WaveSpeed and Atlas Cloud price on different units (per-run, per-token), so they are not directly comparable and we do not force them into a false equivalence.

What a real 2-minute scene costs

Numbers in isolation do not help you budget. So here is the concrete reframe no competing page offers. Suppose you generate a 2-minute multi-character scene — the model's stated per-generation ceiling — at the cleanest verified anchor of $0.1875/min:

2 min × $0.1875 = $0.375 per finished scene (fal.ai, modeled from the verified per-minute rate).

At the low end of the aggregated spread ($0.002/sec = $0.12/min) the same scene is about $0.24; at the high end ($0.00325/sec = $0.195/min) it is about $0.39. So a full two-minute dialogue-plus-music-plus-SFX scene lands roughly in the $0.24–$0.39 band depending on provider. These are modeled figures derived from published reseller per-second/per-minute rates, not a first-party quote and not a bill we ran.

This is the same cost-per-task lens we apply to visual models in Sora 2 cost and Nano Banana pricing: don't compare unit prices, compare what one finished deliverable actually costs.

How to access it

There are three official channels and several third-party APIs. Be careful here: many top search results (seedaudio.im, seedaudio.co, seedaudio.run and similar) are affiliate/SEO domains, not ByteDance. The only official channels are below.

If the "gateway" model of access is new to you, our primer on what an LLM gateway is explains why routing many models through one endpoint beats juggling a dozen provider keys — the same logic that makes multi-provider audio access manageable.

Honest verdict

Seed Audio 1.0 is real, hosted on multiple live APIs, and its core capability — one prompt to a full sound scene in a single pass — is well corroborated. If you want dialogue, music and effects composed together and you are comfortable with ByteDance-cloud access, it is worth a trial.

But be clear-eyed about what is not established. There is no technical paper, no independent benchmark, and no first-party press coverage. Quality claims — foley realism, "better than SeeDance 2" — trace to vendor demos or a single social-media anecdote (an X post from @TomLikesRobots claiming a ~15-second test cost a couple of cents). The six-language preset list is marketing, not confirmed in rendered docs. And no official price has been published. Treat the enthusiasm accordingly, run your own listening test on a real script, and price it per finished scene, not per unit.

FAQ

What is Seed Audio 1.0?

Seed Audio 1.0 is ByteDance Seed's all-in-one text-to-audio model, announced June 23, 2026 at Volcano Engine FORCE 2026. Unlike text-to-speech-only tools, it renders a full sound scene from a single prompt: spoken dialogue, background music, ambience and foley-style sound effects in one generation pass.

How is Seed Audio 1.0 different from ElevenLabs?

The honest differentiator is single-pass full-scene generation plus enterprise-cloud access. Seed Audio composes voice, music and SFX together in one call and is delivered through ByteDance clouds (Volcano Engine in China, BytePlus internationally). ElevenLabs has also added multi-speaker dialogue and sound effects, so this is a difference in the generation model and access channel, not a blanket claim that ElevenLabs cannot do something.

How much does Seed Audio 1.0 cost?

There is no confirmed first-party Volcano Engine list price. Verified third-party reseller pricing as of research date: fal.ai charges about $0.1875 per minute (roughly $0.003125 per second), WaveSpeed starts around $0.30 per run scaling with parameters, and secondary aggregation cites a per-second spread of about $0.002 to $0.00325. Treat fal.ai as the cleanest verified anchor and expect prices to vary by provider.

Does Seed Audio 1.0 support voice cloning?

Yes. It offers zero-shot voice cloning from up to three reference clips, each up to 30 seconds and 10MB, in wav, mp3, pcm or ogg_opus. Clips are referenced in the prompt as @Audio1, @Audio2 and @Audio3 with no training or fine-tuning. Audio references cannot be combined with an image input, per the fal.ai and Segmind API docs.

How many languages does Seed Audio 1.0 support?

Vendor and reseller pages claim preset voices for English, Chinese, Spanish, Japanese, Indonesian and Portuguese with cross-lingual synthesis. This is marketing, not first-party confirmed: the rendered fal.ai model page does not list supported languages, and an earlier fal snippet implied English and Chinese are the functional core. Treat the six-language list as unconfirmed.

Are there independent benchmarks for Seed Audio 1.0?

No. As of July 2026 there is no arXiv paper, no independent benchmark, and no first-party major-tech-press launch coverage. Evidence of quality is limited to vendor demos and a couple of anecdotal social-media user tests. Any quality ranking versus ElevenLabs should be read as vendor demo or single-user anecdote, not tested fact.

Written by
Kevin Fan

Founder of DataLLM Lab, the unified LLM gateway. Kevin tests models the boring way — same prompts, real costs, unedited outputs — and writes up what the runs actually show.

One API for every model

One API, every model.

Get a single API key for Claude Opus 4.7, GPT-5.4, and 300+ more — with automatic price comparison and routing to the best model for every request.