songsee

246.8k
openclawopenclaw

Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.

191 days ago

openai-whisper-api

246.8k
openclawopenclaw

Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

191 days ago

youtube-downloader

39.6k
ComposioHQComposioHQ

Download YouTube videos with customizable quality and format options. Use this skill when the user asks to download, save, or grab YouTube videos. Supports various quality settings (best, 1080p, 720p, 480p, 360p), multiple formats (mp4, webm, mkv), and audio-only downloads as MP3.

191 days ago

game-audio

21.8k
davila7davila7

Game audio principles. Sound design, music integration, adaptive audio systems.

191 days ago

nemo-curator

21.8k
davila7davila7

GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with RAPIDS. Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora.

Data ProcessingNeMo CuratorData Curation+8
191 days ago

transformers

21.8k
davila7davila7

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

191 days ago

remotion

21.8k
davila7davila7

Best practices and comprehensive guide for Remotion - programmatic video creation in React with animations, compositions, and media handling

VideoReactAnimation+9
191 days ago

markitdown

21.8k
davila7davila7

Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.

191 days ago

audiocraft-audio-generation

21.8k
davila7davila7

PyTorch library for audio generation including text-to-music (MusicGen) and text-to-sound (AudioGen). Use when you need to generate music from text descriptions, create sound effects, or perform melody-conditioned music generation.

MultimodalAudio GenerationText-to-Music+2
191 days ago

whisper

21.8k
davila7davila7

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M params). Use for speech-to-text, podcast transcription, or multilingual audio processing. Best for robust, multilingual ASR.

WhisperSpeech RecognitionASR+7
191 days ago

fal-audio

18.0k
sickn33sickn33

Text-to-speech and speech-to-text using fal.ai audio models

191 days ago

open-notebook

11.1k
K-Dense-AIK-Dense-AI

Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. Use when organizing research materials into notebooks, ingesting diverse content sources (PDFs, videos, audio, web pages, Office documents), generating AI-powered notes and summaries, creating multi-speaker podcasts from research, chatting with documents using context-aware AI, searching across materials with full-text and vector search, or running custom content transformations. Supports 16+ AI providers including OpenAI, Anthropic, Google, Ollama, Groq, and Mistral with complete data privacy through self-hosting.

191 days ago

transformers

10.8k
K-Dense-AIK-Dense-AI

This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.

191 days ago

markitdown

10.8k
K-Dense-AIK-Dense-AI

Convert files and office documents to Markdown. Supports PDF, DOCX, PPTX, XLSX, images (with OCR), audio (with transcription), HTML, CSV, JSON, XML, ZIP, YouTube URLs, EPubs and more.

191 days ago

transcribe

10.4k
openaiopenai

Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.

191 days ago

speech

10.4k
openaiopenai

Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI (`scripts/text_to_speech.py`) with built-in voices and require `OPENAI_API_KEY` for live calls. Custom voice creation is out of scope.

191 days ago

muapi-media-generation

2.8k
SamurAIGPTSamurAIGPT

Generate AI images, videos, music, and audio from the terminal via muapi.ai — supports 100+ models including Flux, Midjourney v7, Kling 3.0, Veo3, and Suno V5

191 days ago

youtube-downloader

2.5k
davepoondavepoon

Download YouTube videos with customizable quality and format options. Use this skill when the user asks to download, save, or grab YouTube videos. Supports various quality settings (best, 1080p, 720p, 480p, 360p), multiple formats (mp4, webm, mkv), and audio-only downloads as MP3.

191 days ago

Developing with Prism

2.3k
prism-phpprism-php

Guide for developing with the Prism PHP package — a Laravel package for integrating LLMs. Enable or use this when working with Prism features such as text generation, structured output, embeddings, image generation, audio processing, streaming, tools/function calling, or any LLM provider integration (OpenAI, Anthropic, Gemini, Mistral, Groq, XAI, DeepSeek, OpenRouter, Ollama, VoyageAI, ElevenLabs). Use for any Prism-related development tasks.

191 days ago

video-transcript-downloader

2.1k
steipetesteipete

Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site. Use when asked to “download this video”, “save this clip”, “rip audio”, “get subtitles”, “get transcript”, or to troubleshoot yt-dlp/ffmpeg and formats/playlists.

191 days ago

markdown-converter

2.1k
steipetesteipete

Convert documents and files to Markdown using markitdown. Use when converting PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx, .xls), HTML, CSV, JSON, XML, images (with EXIF/OCR), audio (with transcription), ZIP archives, YouTube URLs, or EPubs to Markdown format for LLM processing or text analysis.

191 days ago

nlm-skill

1.9k
jacob-bdjacob-bd

Expert guide for the NotebookLM CLI (`nlm`) and MCP server - interfaces for Google NotebookLM. Use this skill when users want to interact with NotebookLM programmatically, including: creating/managing notebooks, adding sources (URLs, YouTube, text, Google Drive), generating content (podcasts, reports, quizzes, flashcards, mind maps, slides, infographics, videos, data tables), conducting research, chatting with sources, or automating NotebookLM workflows. Triggers on mentions of "nlm", "notebooklm", "notebook lm", "podcast generation", "audio overview", or any NotebookLM-related automation task.

191 days ago

media-processing

1.8k
mrgooniemrgoonie

Process multimedia files with FFmpeg (video/audio encoding, conversion, streaming, filtering, hardware acceleration) and ImageMagick (image manipulation, format conversion, batch processing, effects, composition). Use when converting media formats, encoding videos with specific codecs (H.264, H.265, VP9), resizing/cropping images, extracting audio from video, applying filters and effects, optimizing file sizes, creating streaming manifests (HLS/DASH), generating thumbnails, batch processing images, creating composite images, or implementing media processing pipelines. Supports 100+ formats, hardware acceleration (NVENC, QSV), and complex filtergraphs.

191 days ago

ai-multimodal

1.8k
mrgooniemrgoonie

Process and generate multimedia content using Google Gemini API. Capabilities include analyzing audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understanding images (captioning, object detection, OCR, visual Q&A, segmentation), processing videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extracting from documents (PDF tables, forms, charts, diagrams, multi-page), generating images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.

191 days ago

pollinations

1.8k
openclawopenclaw

Pollinations.ai API for AI generation and analysis - text, images, videos, audio, vision, and transcription. Use when user requests AI-powered content (text completion, image generation/editing, video generation, audio/TTS, image/video analysis, audio transcription) or mentions Pollinations. Supports 25+ models with OpenAI-compatible endpoints.

191 days ago

voice-note-to-midi

1.8k
openclawopenclaw

Convert voice notes, humming, and melodic audio recordings to quantized MIDI files using ML-based pitch detection and intelligent post-processing

audiomidimusic+2
191 days ago

assemblyai-transcribe

1.8k
openclawopenclaw

Transcribe audio/video with AssemblyAI (local upload or URL), plus subtitles + paragraph/sentence exports.

191 days ago

tts-whatsapp

1.8k
openclawopenclaw

Send high-quality text-to-speech voice messages on WhatsApp in 40+ languages with automatic delivery

whatsappttsvoice+3
191 days ago

webchat-audio-notifications

1.8k
openclawopenclaw

Add browser audio notifications to Moltbot/Clawdbot webchat with 5 intensity levels - from whisper to impossible-to-miss (only when tab is backgrounded).

webchatnotificationsaudio+3
191 days ago

audio-reply

1.8k
openclawopenclaw

Generate audio replies using TTS. Trigger with "read it to me [public URL]" to fetch and read content aloud, or "talk to me [topic]" to generate a spoken response. Also responds to "speak", "say it", "voice reply".

191 days ago

clipit

1.8k
openclawopenclaw

The master tool for all advanced audio/video processing. Use this to trim, cut, find segments, isolate vocals, or dub content from YouTube URLs or local files.

191 days ago

phone-agent

1.8k
openclawopenclaw

Run a real-time AI phone agent using Twilio, Deepgram, and ElevenLabs. Handles incoming calls, transcribes audio, generates responses via LLM, and speaks back via streaming TTS. Use when user wants to: (1) Test voice AI capabilities, (2) Handle phone calls programmatically, (3) Build a conversational voice bot.

191 days ago

addis-assistant

1.8k
openclawopenclaw

Provides Speech-to-Text (STT) and text Translation using the Addis Assistant API (api.addisassistant.com). Use when the user needs to convert an audio file to text (specifically Amharic), or translate text between languages (e.g., Amharic to English). Requires 'x-api-key'.

191 days ago

airfoil

1.8k
openclawopenclaw

Control AirPlay speakers via Airfoil from the command line. Connect, disconnect, set volume, and manage multi-room audio with simple CLI commands.

191 days ago

duby

1.8k
openclawopenclaw

Convert text to speech using Duby.so API. Supports various voices and emotions.

ttsaudiovoice+1
191 days ago

elevenlabs-stt

1.8k
openclawopenclaw

Transcribe audio files using ElevenLabs Speech-to-Text (Scribe v2).

191 days ago

transcribe

1.8k
openclawopenclaw

Transcribe audio files to text using local Whisper (Docker). Use when receiving voice messages, audio files (.mp3, .m4a, .ogg, .wav, .webm), or when asked to transcribe audio content.

191 days ago

elevenlabs-voices

1.8k
openclawopenclaw

High-quality voice synthesis with 18 personas, 32 languages, sound effects, batch processing, and voice design using ElevenLabs API.

ttsvoicespeech+5
191 days ago

gettr-transcribe-summarize

1.8k
openclawopenclaw

Download audio from a GETTR post (via HTML og:video), transcribe it locally with MLX Whisper on Apple Silicon (with timestamps via VTT), and summarize the transcript into bullet points and/or a timestamped outline. Use when given a GETTR post URL and asked to produce a transcript or summary.

191 days ago

tts

1.8k
openclawopenclaw

Convert text to speech using Hume AI (or OpenAI) API. Use when the user asks for an audio message, a voice reply, or to hear something "of vive voix".

191 days ago

callmac

1.8k
openclawopenclaw

Remote voice control for a Mac from mobile devices using commands like /callmac or /voice. Broadcast announcements, play alarms, tell stories, wake up kids — all triggered by Telegram or WhatsApp messages. Uses edge-tts for robust mixed Chinese/English TTS, plays audio locally on the Mac, and supports looping and volume control for flexible playback.

191 days ago

moodcast

1.8k
openclawopenclaw

Transform any text into emotionally expressive audio with ambient soundscapes using ElevenLabs v3 audio tags and Sound Effects API

191 days ago

youtube-video-downloader

1.8k
openclawopenclaw

Download YouTube videos in various formats and qualities. Use when you need to save videos for offline viewing, extract audio, download playlists, or get specific video formats.

191 days ago

video-subtitles

1.8k
openclawopenclaw

Generate SRT subtitles from video/audio with translation support. Transcribes Hebrew (ivrit.ai) and English (whisper), translates between languages, burns subtitles into video. Use for creating captions, transcripts, or hardcoded subtitles for WhatsApp/social media.

191 days ago

elevenlabs-music

1.8k
openclawopenclaw

Generate music from text prompts using ElevenLabs Eleven Music API. Use when creating songs, soundtracks, jingles, lullabies, or any audio music from descriptions. Supports vocals with AI-generated lyrics, instrumental tracks, and multiple genres/styles. Requires paid ElevenLabs plan.

191 days ago

yt-video-downloader

1.8k
openclawopenclaw

Download YouTube videos in various formats and qualities. Use when you need to save videos for offline viewing, extract audio, download playlists, or get specific video formats.

191 days ago

google-gemini-media

1.8k
openclawopenclaw

Use the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".

191 days ago

parakeet-stt

1.8k
openclawopenclaw

Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU). 30x faster than Whisper, 25 languages, auto-detection, OpenAI-compatible API. Use when transcribing audio files, converting speech to text, or processing voice recordings locally without cloud APIs.

191 days ago