KittenTTS logo

KittenTTS

KittenTTS is an open-source, lightweight text-to-speech library built on ONNX that ships state-of-the-art voice synthesis in models ranging from 15M to 80M parameters, only 25 to 80 MB on disk, and runs entirely on CPU without a GPU. It exposes a clean Python API for synthesis, writes audio directly to a file, normalises numbers and currencies per locale, switches voices by name, and adjusts speed with a multiplier, all without external services or accounts. The v0.8 release added 15M, 40M, and 80M parameter variants so you can trade size for fidelity on anything from an edge device to a server batch job. It is built for developers who want a dependency-minimal, offline TTS for embedded apps, agents, accessibility tooling, and speech output where shipping a multi-gigabyte model is not feasible. KittenTTS matters now because high-quality, CPU-only speech under 25 MB makes on-device voice practical at the long tail of constrained hardware.

Reader rating

No ratings yet

Visit website

You might also like

Related tools

View all
Demon favicon
Demon
No ratings yet

Demon (Diffusion Engine for Musical Orchestrated Noise) is an open-source real-time music generation system that runs locally on consumer GPUs at 25Hz. It is built for musicians, sound designers, music producers, and AI researchers who want to generate, iterate, and perform with AI music in real time without relying on cloud APIs. The system uses diffusion-based synthesis to produce musical audio streams with low latency, enabling live experimentation and performance workflows. Demon launched on Hacker News with 15 points and the project page at daydreamlive.github.io/DEMON describes a fully local, GPU-accelerated approach to music generation. What makes it notable is the combination of real-time performance with diffusion models — a technical achievement that opens up live music creation use cases that were previously impossible with slower batch-generation approaches.

View details
Udio favicon
Udio
No ratings yet

Meet Udio, your AI-powered music creation companion. With Udio, you can effortlessly create and share music using cutting-edge AI technology. This free platform offers a range of tools to produce and refine audio content, from generating diverse music genres to creating vocals and instrumentals in seconds. Whether you're a music enthusiast, content creator, or just looking to add a unique touch to your projects, Udio's AI audio tools have you covered. From crafting melodies to experimenting with text-to-speech capabilities, Udio empowers you to explore the endless possibilities of AI-generated music and audio content. Unleash your creativity and dive into the world of AI music with Udio today!

View details
Dograh favicon
Dograh
No ratings yet

Dograh is an open-source, self-hostable voice-agent platform and a forkable alternative to closed vendors like Vapi and Retell. It gives teams a visual workflow builder for production voice agents, native MCP control so coding assistants can spin up and deploy agents from the IDE, and hybrid speech-to-speech that mixes real recorded clips with TTS for low-latency, natural conversations. Organizations can run it on-prem, in their own VPC, or fully air-gapped, keeping recordings, transcripts, and model inference inside their perimeter—critical for fintech, healthtech, and defense. Telephony integrations (Twilio, Vonage, Telnyx, and more), human handoff, and bring-your-own LLM/STT/TTS make it deployment-flexible. Licensed BSD 2-Clause, it removes vendor lock-in and compliance-review friction while offering managed-cloud and private-cloud options.

View details