Described
Give it a sentence about what the track is for and the words are written first, then sung. Writing words runs a language model rather than a GPU, so it costs a flat 5 seconds of the allowance and it needs a language model configured on the deployment. Without one, the other two ways in still render.