Songs and beds

A track with your words in it, at the length you set.

A theme for a show, a loop for a game, a bed under a video, or a song for the sake of it. Describe what you want and the words are written before anything renders, or hand it a lyric you already wrote and those exact words are sung. Either way it is written to the length you asked for rather than trimmed into it afterwards.
:10 to 10:00
the engine's envelope for one render. Your plan sets the ceiling inside it.
24
full length tracks a month on Creator, at its own 4:00 ceiling
10
saved style recipes on Creator, and 200 on Group
Three ways in

Describe it, write it yourself, or take the words out.

Which one you get is decided by what you send rather than by a mode switch you have to find first. All three run on the same engine and are billed the same way, by the second of finished audio.

Described

Give it a sentence about what the track is for and the words are written first, then sung. Writing words runs a language model rather than a GPU, so it costs a flat 5 seconds of the allowance and it needs a language model configured on the deployment. Without one, the other two ways in still render.

Custom

Hand it a lyric and those exact words are sung, verbatim. Section markers the engine understands survive into the arrangement, so a chorus you marked as a chorus is arranged as one rather than being read as a line of text.

Instrumental

No words at all, as a flag on the render rather than a hopeful line in the caption, which is the difference between no vocals and probably no vocals. It is also the cheaper rate: an instrumental second costs 1 second of the allowance where a sung second costs 1.25.

Style words, tempo, key and a seed are fields on the same request rather than adjectives hidden in a sentence, and every one of them comes back on the finished track. A caption that produced something you liked is a caption you can send again.

What the writer is given

A syllable budget, not a word count.

Below is one written lyric exactly as the engine receives it: the syllables counted line by line in the margin, the section markers in the form that goes on the wire, and the whole thing measured against the length it was asked for. Every figure on it is produced by the same functions the studio's lyric editor calls.
The words, as the engine receives them
Cross Street, opening theme

Opening theme for a two person podcast about cycling across a city at night. Warm indie folk, female lead, acoustic guitar and hand claps, ends on the show name.

  • Verse[verse]
  • 9Wheels on the wet road, quarter past five
  • 12Everybody's heading home, we're heading out
  • 10Two helmets, one map and a bad idea
  • Chorus[chorus]
  • 4Cross Street, Cross Street
  • 9Two of us and a city to cross
  • 8Cross Street, every Thursday night
  • Tag[outro]
  • 2Cross Street
Budget at :30
56
syllables the length holds
Written
54
syllables in the body above
Sings in about
28.8s
Fits the :30

Why syllables and not words

A word count is wrong by a factor of two between one word and another, and a line count says nothing at all. A digit is its own beat, so a phone number is counted digit by digit rather than as one word, and a run of capitals is spelled rather than read: FM is two beats, and a W on its own is three.

A :10 holds about 19 of them and a :60 about 112, and neither figure is the whole length: a track has an instrumental lift at the front, a breath between sections and a button on the end, and a lyric written for every second leaves nowhere for any of it.

Why a tag is written [outro]

The engine knows a small set of section markers and SINGS anything it does not recognise, so a body containing the word tag in brackets puts the word tag in the vocal. Tag is what a producer calls the sung logo at the end, so that is the word in the interface; on the wire it is [outro].

The sheet shows both, because that mismatch is the one thing about writing for this engine nobody can guess. The markers it does understand are [verse], [chorus], [bridge], [outro], [inst], and anything else in square brackets is a line of the lyric.

The estimate is reported before anything renders, and past 15 percent either side of the target the studio warns rather than just printing the number. That is the honest moment to find out a lyric is long: on the page, not after you have listened to it run over the end of the slot.

What the length changes

A short track is not a long one with the end cut off.

The same brief is written to a different shape at every length, because a very short piece has one musical idea in it and giving it a verse and a chorus means two ideas competing for a few seconds each and neither landing. This is the section plan the writer is handed at each mark, run rather than drawn.
The shape at each length
Each section shows the syllables it may spend. The figure after the bar is the whole row.
  • :10
    19Tag19
  • :15
    28Chorus28
  • :30
    56Verse22Chorus26Tag8
  • 1:00
    112Verse24Chorus28Verse24Chorus28Tag8
  • 4:00
    450Verse72Chorus83Verse72Chorus83Bridge57Chorus83

The budget is split by section, not evenly

A chorus carries more than its share because it is the part that has to be remembered, and a tag carries far less because it is two or three words. At :30 it is given 8 syllables where an even split across the plan would hand it 19, which is the difference between a sung logo and a mouthful.

Above the top mark, the shape stops changing

The longest shape is the last one on the strip, and a render longer than 4:00 gets the same plan with more room inside each section rather than more sections. The engine goes to 10:00, and which plan you are on decides how far up that envelope you can go.

Make another one of those

The second version is what makes the first one worth it.

Most of the work on a track happens after the first render lands. These are the six things that make the next one cheap, and every one of them is a real surface rather than a plan we have.

Personas: a saved style recipe

A set of style words and an optional reference track, saved and reused, so the fourth episode theme sounds like the first. NOT a cloned voice: this engine has no speaker embedding for singing, and nothing here promises a particular person’s voice.

Seeds, so a take can be made again

Pin the sampler and the same take comes back. A seed reproduces a render only while everything else is unchanged, which is why the caption, the duration and the model are all on the record beside it: a track made a year ago must not claim to have come from today’s checkpoint.

Takes, under one title

Ask for alternates and they arrive as versions of one track rather than as separate entries: one title, one place in a playlist, several things to listen to. Each take is a render and is billed as one, and the cost of the whole set is stated before you press it.

Cover: the same words, another style

Re-render a track you own in a different style, keeping the words and the shape. One approved lyric as country, as pop and as rock is three finished tracks off one set of words. It covers only what your workspace owns, which is a licensing boundary rather than a missing feature.

Extend, billed on the added seconds

Continue a finished track past its ending. You are charged for the seconds you added and never for the whole result, so a :30 extended by another :30 claims 30 seconds of audio rather than 60. The continuation is a new track with the original on its record, not an edit that overwrote it.

Lyrics that outlive a session

A draft is kept in a library rather than inside one render, and writing it runs a language model rather than a GPU: a flat 5 seconds of the allowance however long the words, with no audio made. Redeeming one copies the words onto the track, so deleting the draft next month leaves the finished piece holding what it was actually sung from.

Rights and files

What you are allowed to do with it, per plan.

The free tier is a real product with a real ceiling rather than a crippled one: it renders sung vocals and it renders full tracks. What it does not carry is the right to sell what comes out, and the file it hands you says so on the download rather than in a clause nobody read.
Longest single render
10:00

The engine's ceiling on one track. Every plan sits somewhere inside it, from 1:00 upwards.

Tracks a month on Creator
24

At the 4:00 ceiling and the sung rate, which is the dearer of the two. Shorter tracks go further.

Free, with no card
10 minutes

About 8 tracks at the free tier's own 1:00 ceiling, for evaluation rather than for release.

PlanLongest single trackStemsCommercial useSaved style recipes
Free1:00Not on this planEvaluation only1
Creator4:00IncludedIncluded10
Station5:00IncludedIncluded40
Group10:00IncludedIncluded200

Stems, for when the track has to go under something

Vocals, drums, bass and the rest, pulled out of a finished track in one separation pass however many files come back. It costs 0.75 of the allowance per second of the track it separates, and re-exporting overwrites rather than accumulating.

What happens when the allowance runs out

Renders stop and we ask you to move up a plan. There is no overage line and nothing you do can run up a bill overnight. The whole of that promise, including the refund rule for a render that failed on our side, is stated once on the pricing page.

Start without a card

Hear your own words come back sung.

The free tier is 10 minutes of audio a month with no card, which is about 8 full length tracks. It sings real lyrics at real lengths, so what you evaluate is the thing you would buy rather than a watermarked demonstration of it.