SeeDance 2.5
What is SeeDance 2.5?
SeeDance 2.5 is ByteDance's August 2026 audio-video model for longer stories, precise reference control, and video editing. It is live on Unifically as bytedance/seedance-2.5. The callable route generates 4 to 30 second videos at 480p or 720p, with synchronized audio on by default.
Three workflows share the same model ID: text-to-video, first-and-last-frame generation, and multimodal reference generation. Reference mode accepts images, videos, and audio in one request. This is the biggest practical change from the older SeeDance line: one generation can carry much more source material across a longer timeline.
What's new in SeeDance 2.5
- Up to 30 seconds. Duration is any whole second from 4 through 30, twice the 15-second ceiling on SeeDance 2.0.
- A 50-asset reference budget. Send up to 30 images, 10 videos, and 10 audio files, capped at 50 assets total.
- Longer reference media. Reference videos can total 30 seconds. Reference audio can also total 30 seconds.
- Audio-only reference runs. A prompt plus audio references is valid; an image or video reference is not required.
- First and last frame control. Pin both ends of the video while the model creates the motion between them.
- Native audio by default. Voice, effects, and background music generate with the picture in the same run.
Best for
Longer narrative videos
A 30-second ceiling gives a short story room for multiple beats without joining separate generations.
Reference-heavy campaigns
Combine product images, motion references, voice, and music cues in one request.
Pinned opening and closing shots
Set the first and last frame, then generate the transition between them.
Dialogue and sound-led scenes
Native audio is on by default, and audio-only reference requests are supported.
Storyboard-driven production
Use numbered image, video, and audio references directly inside the prompt.
Use cases
Build a 30-second product story that moves from a wide establishing shot to a close product reveal without joining separate videos. Feed a campaign's product angles, style frames, camera reference, and music cue into one generation. Pin a first frame and final call-to-action frame for an ad with a fixed open and close. For character work, use reference images plus a motion video and address each asset as [Image1] or [Video1] in the prompt. Audio-led workflows can start from a voice or rhythm reference without adding an image.
Limitations
The current API route exposes 480p and 720p only. The 4K output and region-editing tools shown in ByteDance's consumer product are not callable parameters here. First-and-last-frame mode requires both images and cannot be mixed with reference arrays. Reference videos and audio each have a 30-second combined limit. Arena has not added Seedance 2.5 to its video boards yet, so there is no independent rank or Elo score for this version as of August 7, 2026.
SeeDance 2.5 vs SeeDance 2.0
SeeDance 2.5 wins on duration and reference capacity: 30 seconds against 15, and up to 50 assets against 15. It also accepts audio-only reference requests. SeeDance 2.0 Pro still wins on output size because its Unifically route reaches 1080p and 4K, while 2.5 currently stops at 720p. SeeDance 2.0 also has the stronger public quality record today. On Arena's August 2 boards, its 720p build ranks #1 for image-to-video at 1478±10 and #2 for text-to-video at 1479±11.
When to use SeeDance 2.5
Use SeeDance 2.5 when a single video needs more than 15 seconds, a large reference set, or audio-only guidance. Use SeeDance 2.0 Pro when 1080p or 4K output matters more than duration. For early drafts, compare the full job price: 2.5 starts at $0.109 per second with a reference video and $0.176 per second without one at 480p.
Feature Stair
Native dialogue and cinematic scene progression
A close-up radio detail expands into a storm-lit rescue scene, then lands on a spoken emotional beat. This tests coherent camera progression, weather, a human face, and synchronized dialogue in one short generation.
Product detail, glass, and liquid physics
A macro product orbit combines transparent glass, refracted light, moving water, and fine surface detail. It is a compact stress test for material realism and controlled commercial camera movement.
Fast action and camera handoff
The shot moves from wheel-level tracking to a whip-pan and aerial chase while the rider, bicycle, leaves, and landing stay in motion. This tests speed, articulated movement, continuity, and layered environmental sound.
FAQs
People also ask
Yes. The model is live on Unifically as bytedance/seedance-2.5, using the shared /v1/tasks generation flow.
Any whole number of seconds from 4 through 30. The default is 4 seconds.
The current Unifically route supports 480p and 720p. ByteDance shows higher-resolution output in its consumer product, but 1080p and 4K are not callable on this API route.
Yes. Synchronized voice, sound effects, and background music are enabled by default. Set generate_audio to false for silent output.
Up to 30 images, 10 videos, and 10 audio files, with no more than 50 assets in one request. Video and audio references can each total up to 30 seconds.
Yes. Both frame URLs are required in frame mode, and the output follows the first frame's aspect ratio. Frame URLs cannot be mixed with reference arrays.
Current Unifically rates run from $0.109 to $0.378 per output second. Resolution and the presence of a reference video determine the rate; the pricing page shows the live matrix.
