Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
视觉语言预训练框架,连接冻结的图像编码器和大型语言模型。需要图像描述、视觉问答、图像-文本检索或多模态聊天时使用,具备先进的零-shot性能。
Category: developer (开发工具) · Author: davila7 · Version: @main · License: MIT
Tags: Multimodal, Vision-Language, Image Captioning, VQA, Zero-Shot
该 Skill 暂无文档文件。
npx skills add davila7/blip-2-vision-language下载完整 Skill 目录,包含 SKILL.md 及所有相关文件
Search for places (restaurants, cafes, etc.) via Google Places API proxy on localhost.
Interact with GitHub using the `gh` CLI. Use `gh issue`, `gh pr`, `gh run`, and `gh api` for issues, PRs, CI runs, and advanced queries.
Create or update AgentSkills. Use when designing, structuring, or packaging skills with scripts, references, and assets.
Start voice calls via the OpenClaw voice-call plugin.
Notion API for creating and managing pages, databases, and blocks.
Gemini CLI for one-shot Q&A, summaries, and generation.
Category:developer
Tags:Multimodal, Vision-Language, Image Captioning, VQA, Zero-Shot