Official BytePlus partnerBuilt on BytePlus, ByteDance's official AI platform

Guide · Audio

How to write timestamped audio prompts

Write Audio Generation prompts with time cues like [2s:5s]: one scene line, one cue per line, two speakers, and what the @Audio 1 tag does in KreatorFlow today.

  • [start s:end s] syntax
  • 24 languages
  • Up to 3,000 characters

Short answer

In Audio Generation, describe the voice and scene on the first line, then put each spoken line on its own row with a time cue like [2s:5s] (start second : end second). The / command inserts one for you. Prompts take up to 3,000 characters in 24 languages and cost 1 credit per started 10 seconds of audio (minimum 2 credits).

Updated:

Step by step

  1. 1

    Describe the scene

    First line: who speaks, in what tone and where. For example: "A crisp, calm news anchor opens the bulletin".

  2. 2

    One line per cue

    Put each line on a new row that starts with [start s:end s], such as [0s:4s]. Use the / command and pick the timestamp to insert it.

  3. 3

    Size each cue to its line

    Give every cue enough time for its words: about two to three words per second. Keep the cues in order and without overlaps.

  4. 4

    Add a second voice if needed

    Start the line with "Speaker A:" or "Speaker B:" after the cue for a dialogue.

  5. 5

    Pick the language and generate

    Choose one of 24 languages in the menu, set tone, speed and volume if you like, and generate. If the timing is off, move the cues and generate again.

Example prompts

Copy them into Audio Generation and swap in your own lines.

  • Product launch

    A confident promo voice introduces a new product. [0s:3s] The new charger is here. [3s:7s] One cable for your phone, your watch and your earbuds. [7s:10s] Available from today.

  • Two-voice dialogue

    Two friends plan a trip, relaxed and natural. [0s:4s] Speaker A: How about the coast this weekend? [4s:8s] Speaker B: Only if we leave early and stop for breakfast. [8s:11s] Speaker A: Deal.

  • Guided breath

    A calm guide leads a breathing exercise with long pauses. [0s:6s] Breathe in slowly through the nose. [6s:12s] And let it go, shoulders dropping.

The [2s:5s] syntax

A cue gives the second a line starts and ends: [2s:5s] means from second 2 to second 5. The studio's own Timing templates use the same format.

Cues are timing hints the model reads, not exact edits: the result can drift a little. For dependable timing, leave some room in each cue and generate two takes.

What @Audio 1 does

The / command also inserts @Audio 1, a tag the studio's Reference templates use for "the voice in reference clip 1". Audio Generation in KreatorFlow sends your text prompt only and has no clip upload, so describe the voice you want in words as well, for example "a warm, measured narrator".

The video studio is different: audio clips you attach there are numbered Audio 1, Audio 2 and so on up to 3, and the prompt can refer to them.

Frequently asked questions

Are the timestamps exact?

No. They are timing hints the model follows as closely as it can. Leave some room and generate two takes when timing matters.

Can I use a reference clip for the voice?

Audio Generation sends text only and takes no clip upload. Describe the voice in words. The video studio does take audio references.

What does it cost?

1 credit per started 10 seconds of audio (minimum 2 credits). The studio shows an estimate from your prompt before you generate, and the final charge follows the length of the audio it makes.

Try it in the studio

New accounts get 75 starter credits for images and video. Free accounts can make video on Seedance 2 Fast and Seedance 2.0 Mini at 480p, queued and rate limited. A plan adds every video model, 720p and above, and no queue. The cost is shown before every generation.

More guides