GLM-TTS clones the speaker identity from a 3–10 second reference clip.
Clear speech with minimal background noise gives the best results.
Both Chinese and English references are supported.
Reference preview:
Leave blank if unknown. Providing the correct transcript significantly improves voice cloning quality.