Summary
为非多模态模型(如 deepseek-v4-pro、GLM-5.1、mimo-v2.5-pro 等纯文本模型)提供图片识别能力。当主模型无法识别图片、用户发送了截图/设计稿/UI 截图需要分析、或者用户说'看看这张图'、'分析这个截图'、'这张图片有什么问题'时,自动触发此技能。也适用于用户粘贴了图片但当前模型不支持图…
penfick/skills
为非多模态模型(如 deepseek-v4-pro、GLM-5.1、mimo-v2.5-pro 等纯文本模型)提供图片识别能力。当主模型无法识别图片、用户发送了截图/设计稿/UI 截图需要分析、或?
npx skills add penfick/skills --skill vision-support
为非多模态模型(如 deepseek-v4-pro、GLM-5.1、mimo-v2.5-pro 等纯文本模型)提供图片识别能力。当主模型无法识别图片、用户发送了截图/设计稿/UI 截图需要分析、或者用户说'看看这张图'、'分析这个截图'、'这张图片有什么问题'时,自动触发此技能。也适用于用户粘贴了图片但当前模型不支持图…
Related neighbors and high-traction skills in the same topics — useful to compare before installing.
Guidance for distinctive, intentional visual design when building new UI or reshaping an existi…
866.4K installsBrowser automation CLI for AI agents. Use when the user needs to interact with websites, includ…
810.4K installsReview UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "chec…
617.3K installsBuild, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and …
576.5K installsDebug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe …
568.9K installsDeclared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.
master
Parsed from SKILL.md frontmatter.
Files included with this skill beyond the listing page.
SKILL.md
4,900 B
README.md
4,509 B
SUMMARY.md
863 B
铁律:本技能配置的所有模型仅用于图片内容识别,绝不参与主逻辑推理。
这些模型不会代替主模型做任何决策、分析或编码,它们只负责"看"图片然后把看到的内容用文字描述出来。
/vision 或 /skill:vision-support 手动触发node SKILL_DIR/scripts/vision.mjs init
交互式引导,只需三步:
支持的平台覆盖国内外主流:
| 分类 | 平台 |
|---|---|
| 国际 | OpenAI、Google Gemini、Anthropic Claude、DeepSeek、Groq、Mistral、xAI (Grok)、OpenRouter、Fireworks AI |
| 国内 | 通义千问 (Qwen VL)、智谱 GLM (GLM-4V)、Moonshot (Kimi)、阶跃星辰 (Step)、MiniMax、SiliconFlow (硅基流动)、小米 MiMo |
| 本地 | Ollama、LM Studio |
| 自定义 | 任何 OpenAI 兼容的第三方平台(自填 baseUrl) |
node SKILL_DIR/scripts/vision.mjs config add
同样的交互式引导,添加的模型作为 fallback 回退。主模型失败后自动尝试。
# 交互式
node SKILL_DIR/scripts/vision.mjs init # 初始化主模型
node SKILL_DIR/scripts/vision.mjs config add # 添加 fallback
node SKILL_DIR/scripts/vision.mjs config edit [name] # 编辑模型
# 快捷命令
node SKILL_DIR/scripts/vision.mjs config list # 列出所有模型
node SKILL_DIR/scripts/vision.mjs config primary [name] # 设置主模型
node SKILL_DIR/scripts/vision.mjs config remove <name> # 删除模型
node SKILL_DIR/scripts/vision.mjs config set-key <name> <key> # 设置密钥
node SKILL_DIR/scripts/vision.mjs config set-url <name> <url> # 设置 API 地址
node SKILL_DIR/scripts/vision.mjs config test [name] # 测试连通性
node SKILL_DIR/scripts/vision.mjs ./screenshot.png
node SKILL_DIR/scripts/vision.mjs ./ui.png "这个界面的布局有什么问题?"
node SKILL_DIR/scripts/vision.mjs "https://example.com/img.png" "描述这张图片"
node SKILL_DIR/scripts/vision.mjs img1.png img2.png "对比这两张图的差异"
node SKILL_DIR/scripts/vision.mjs ./screenshots/*.png "分析这些界面截图"
node SKILL_DIR/scripts/vision.mjs ./local.png https://example.com/remote.jpg "描述这两张"
如果用户提到图片但没给路径,先搜索:
find . -name "*.png" -o -name "*.jpg" -o -name "*.webp" | head -20
ls -lt *.png *.jpg *.webp 2>/dev/null
脚本成功后 stdout 输出的纯文本就是识别结果(stderr 是日志不影响)。
config list 中排第一位的 ★ 主模型优先调用。失败后自动依次尝试后续模型。所有模型都失败则非零退出码退出。
| 变量 | 说明 |
|---|---|
VISIONCONFIGPATH |
自定义配置文件路径 |
VISIONDEFAULTMODEL |
临时覆盖主模型(按 name 匹配) |
VISIONAPIKEY |
全局密钥回退 |