LatentSync
π
611
Audio Conditioned LipSync with Latent Diffusion Models
Generate 3D models from images
Generate spokenβready scripts from documents for podcasts, lectures, or summaries
WebGPU text-to-Speech powered by OuteTTS and Transformers.js
Scalable and Versatile 3D Generation from images
Fill image areas with AI using a text prompt
Chat with an AI that understands images and text