Language Tags, Voices, and Denoising Steps
Supertonic accepts Unicode text together with an explicit language selection. The integration wraps the text in the corresponding language tag before inference, so choosing the matching language is part of the model input rather than a display-only filter. Its ten M1–M5 and F1–F5 choices are built-in style files; the labels indicate catalog categories and do not identify real speakers.
Denoising steps trade computation for another refinement pass, but more steps do not guarantee that every sentence sounds better. Begin around the interface default, compare a fixed sample, and increase steps only when the audible result justifies the extra time. Speed and steps interact with device performance, text length, punctuation, and backend, so retain those settings with an important export.