Text to speech
This is a companion project that utilizes Sesame AI's Conversational Speech Model (CSM) in an attempt to recreate the original magic of the companions on release. This is a stack, built and used locally.
Voice phone call for Claude Code CLI + official Telegram channel plugin (Chinese STT/TTS, Qwen3-ASR + Azure Neural, PC & phone PWA)
Emotional voice calls for AI companions — tone-aware listening, proactive dialing, streamed speech, and grounded call memories.
Chalk turns one educational prompt into a narrated, hand-drawn explainer video.
VLLM Port of the Chatterbox TTS model, adapter to make audiobooks
Claude Code / Codex skill for authoring supervoice Domain Dictionaries — STT vocabulary, TTS pronunciation, and latency fillers for voice agents.
Turn any Shopify product into an AI promo video with a consistent virtual character — fully automated
Point your existing ElevenLabs, OpenAI Audio or Deepgram app at Sarvam AI by changing one line. Drop-in compatibility gateway for Indic TTS & STT on Bulbul and Saaras: script-aware language detection, grapheme-safe Hindi/Tamil chunking, CER-based voice mapping, caching + request coalescing. TypeScript, Docker, 168 tests, MIT.
Custom local-first bilingual realtime voice stack and ESP32-S3 firmware for Stack-chan
Read any selected text aloud with a local TTS voice
用 Remotion + Edge TTS 制作上方逐笔写字、下方趣味拆字变形的 9:16 汉字短视频
Opensource Audio ASR, TTS, Voice Clone in https://recut.video
Open-source opencode skill: free narrated 9:16 product/tech explainer videos from a long-form write-up. Pillow slides, edge-tts narration, ffmpeg compose.
Clean-room implementation of the diphone concatenative synthesis techniques used by CMU's Flite and Festival
Claude Code plugin that creates demo videos for your projects. Claude is given tools to drive the browser/cli, make recordings, and uses FAL.ai, ElevenLabs, and remotion to build a demo video with scene transitions.