jimliu/doubao-multimodal-skill · Archived
doubao-multimodal
Async audio/video understanding with Doubao-Seed (火山方舟 Ark). Handles ASR (plain / per-character timestamps / multispeaker), AST translation, speaker…
Installation
npx skills add https://github.com/jimliu/doubao-multimodal-skill
Stronger alternatives
This repository is archived — consider an actively maintained alternative.
Generate Mandarin and multilingual narration with Volcengine Doubao Speech 2.0. Use when creati…
7 installsGenerate images and text via Doubao Web using browser cookies and Playwright UI automation. Use…
276 installsRemove the visible Doubao (豆包) AI watermark from images. Use when asked to remove Doubao wate…
227 installsPDF 解析(Doubao):将 PDF/扫描件转成结构化 Markdown/文本,支持表格与多栏版式。当用户要从 PDF …
160 installsSimilar popular skills
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
基于多模态AI的图片识别与分析。当用户想分析、描述、从图片URL中提取信息、image recognition, image…
1.2K installsUse mmx to generate text, images, video, speech, and music via the MiniMax AI platform. Use whe…
638 installsVision and multimodal capabilities for Claude including image analysis, PDF processing, and doc…
553 installsProcess and generate multimedia content using Google Gemini API. Capabilities include analyze a…
515 installs利用多模态AI分析商品主图,提取视觉特征和提示词。当用户提到分析产品图片、从商品图中提取视觉属性…
368 installsAI驱动的图片生成与编辑工?
365 installsMore details
Agent compatibility
Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
Repository health
main
History
- First seen on skills.sh
- First recorded snapshot · 2 installs