Text-to-speech and speech-to-text using fal.ai audio models
Text-to-speech and speech-to-text using fal.ai audio models
Use this skill when you need to work with text-to-speech and speech-to-text using fal.ai audio models.
This skill provides guidance and patterns for text-to-speech and speech-to-text using fal.ai audio models.
For more information, see the source repository.
Generate or edit images via Gemini 3 Pro Image (Nano Banana Pro).
Batch-generate images via OpenAI Images API. Random prompt sampler + `index.html` gallery.
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
Extract frames or short clips from videos using ffmpeg.
Search GIF providers with CLI/TUI, download results, and extract stills/sheets.
Category:media-generate