Text to speech
Paced, streaming Piper TTS for real-time voice pipelines - Asterisk AudioSocket, RTP or WebSocket. Optional ElevenLabs with local fallback.
Device-native Qwen3-TTS inference for AMD RX 7900 XTX and NVIDIA RTX 4090 with explicit HIP/CUDA kernels, streaming and Resident execution.
Self-hosted AI video dubbing — translate and re-dub a video's speech, timed to match the original, for a fraction of commercial dubbing tool cost.
Codex skill that announces task progress with the macOS say command.
面向中文家庭的儿童英语单词故事视频生成工具|A Codex skill for children’s word-story videos.
This is a companion project that utilizes Sesame AI's Conversational Speech Model (CSM) in an attempt to recreate the original magic of the companions on release. This is a stack, built and used locally.
Voice phone call for Claude Code CLI + official Telegram channel plugin (Chinese STT/TTS, Qwen3-ASR + Azure Neural, PC & phone PWA)
Official code for "GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech Model"
Emotional voice calls for AI companions — tone-aware listening, proactive dialing, streamed speech, and grounded call memories.
A floating desktop companion that reads Claude Code aloud. Tails your session transcripts and speaks Claude's replies in real time — ElevenLabs, Mistral, offline Piper neural, or Windows TTS — with spoken alerts when Claude is waiting on you. Built with Tauri.
Reusable skill for presenter-led digital human product videos
Chalk turns one educational prompt into a narrated, hand-drawn explainer video.
VLLM Port of the Chatterbox TTS model, adapter to make audiobooks
Claude Code / Codex skill for authoring supervoice Domain Dictionaries — STT vocabulary, TTS pronunciation, and latency fillers for voice agents.
Turn any Shopify product into an AI promo video with a consistent virtual character — fully automated
Point your existing ElevenLabs, OpenAI Audio or Deepgram app at Sarvam AI by changing one line. Drop-in compatibility gateway for Indic TTS & STT on Bulbul and Saaras: script-aware language detection, grapheme-safe Hindi/Tamil chunking, CER-based voice mapping, caching + request coalescing. TypeScript, Docker, 168 tests, MIT.
Custom local-first bilingual realtime voice stack and ESP32-S3 firmware for Stack-chan
Read any selected text aloud with a local TTS voice
用 Remotion + Edge TTS 制作上方逐笔写字、下方趣味拆字变形的 9:16 汉字短视频