Local speech-to-text (transcribe/translate) and text-to-speech (voice generation) using Whisper.cpp and Sesame CSM-1B. Use when (1) transcribing or translating audio/voice messages, (2) generating voice messages or audio from text, (3) user asks to "say something", "read this aloud", or "send a voice message", (4) processing inbound voice notes, (5) voice cloning from reference audio. Fully local on DGX Spark — no API keys, no cloud. For setup issues or background, see references/SETUP_GUIDE.md.
Local speech-to-text (transcribe/translate) and text-to-speech (voice generation) using Whisper.cpp and Sesame CSM-1B. Use when (1) transcribing or translating audio/voice messages, (2) generating voice messages or audio from text, (3) user asks to "say something", "read this aloud", or "send a voice message", (4) processing inbound voice notes, (5) voice cloning from reference audio. Fully local on DGX Spark — no API keys, no cloud. For setup issues or background, see references/SETUP_GUIDE.md.
使用 Whisper.cpp 和 Sesame CSM-1B 的本地语音转文本(转录/翻译)和文本转语音(语音生成)方案。适用于:(1) 转录或翻译音频/语音消息,(2) 从文本生成语音消息或音频,(3) 用户请求“说点什么”、“大声朗读这个”或“发送语音消息”时,(4) 处理收到的语音便笺,(5) 从参考音频进行语音克隆。在 DGX Spark 上完全本地运行——无需 API 密钥,无需云。有关安装问题或背景信息,请参见 references/SETUP_GUIDE.md。
Category: media-generate (媒体生成) · Author: Patvscode · Version: @main
Generate or edit images via Gemini 3 Pro Image (Nano Banana Pro).
Batch-generate images via OpenAI Images API. Random prompt sampler + `index.html` gallery.
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
Extract frames or short clips from videos using ffmpeg.
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
Category:media-generate