2,357 tools found
Text-to-Audio Generation with Latent Diffusion Models - Speech Research
EPUB to audiobook converter, optimized for Audiobookshelf
3D character lib synced
AI Voiceover & Text to Speech Platform with human-like voices
transform vocals to match the style of a list of singers
Implementation of MusicLM, Google's new SOTA model for music generation using attention networks, in Pytorch
architecture of voice AI, from speech recognition to emotional intelligence, and learn how to build, scale, and evaluate them
this AI system generates singing voice for literally any text as input
high-resolution video synthesis with latent diffusion models [[arxiv]](https://arxiv.org/abs/2304.08818)
collaborative video editor with a collection of AI plugins
"a multi-modal AI system that can generate novel videos with text, images, or video clips" [[arxiv]](https://arxiv.org/abs/2302.03011)
"transforming the future of music creation"
"Toward Controllable Text-to-Music Generation"
a video foundation model that allows you to craft characters and animate them
Open Diffusion Models for High-Quality Video Generation
Faceless Video Generator
a text-to-animation tool for developers by Stability AI [[dev platform]](https://platform.stability.ai/docs/features/animation)
video editing with cross-attention control [[tweet]](https://twitter.com/akhaliq/status/1637838648463749120)
](https://emu-video.metademolab.com/demo#/demo): state-of-the-art text-to-video generation
vocal removal using AI
a ControlNet model designed to enhance the temporal consistency of generated outputs [[tweet]](https://twitter.com/ciararowles1/status/1639321818581303310)
OpenAI's text-to-video model [[technical report]](https://openai.com/research/video-generation-models-as-world-simulators)