pexoai/pexo-skills

videoagent-video-studio

Generate short AI videos from text or images — text-to-video, image-to-video, and reference-based generation — with zero API key setup. Use when the user wants to create a video clip, animate an image, or generate video from a description.

All-time #1464 First seen Mar 7, 2026
8-week activity · all time api

Installation

$ npx skills add pexoai/pexo-skills --skill videoagent-video-studio

Summary

  • Generate short AI videos from text or images using 7 backend models with zero API key setup.
  • Supports three generation modes: text-to-video, image-to-video, and reference-based generation for consistent output Seven models available (minimax, kling, veo, hunyuan, grok, seedance, pixverse) with automatic selection or manual override via --model flag Configurable duration (4–12 seconds), aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4), and automatic prompt enhancement for better results Simple Node.js CLI interface with async job status checking and JSON response format returning video URLs

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from pexoai/pexo-skills · top by installs.

npx skills add pexoai/pexo-skills

Browse all from pexoai/pexo-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 777
License LICENSE
Default branch main
Open issues 6
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version2.1.0
Declared agents clawdbot
More metadata
openclaw
{"emoji":"🎬","install":{"0":"id: node","kind":"node","label":"No dependencies needed — all calls go through the hosted proxy"}}

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 6,528 B
  • docs README.md 4,339 B
  • docs SUMMARY.md 274 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 10,439 installs

SKILL.md

🎬 VideoAgent Video Studio

Use when: User asks to generate a video, create a video from text, animate an image, make a short clip, or produce AI video.

Generate short AI videos with 7 backends. This skill picks the right mode (text-to-video or image-to-video), enhances the prompt for best results, and returns the video URL.


Quick Reference

User Intent Mode Typical Duration
"Make a video of..." (no image) text-to-video 4–10 s
"Animate this image" / "Make this move" image-to-video 4–6 s
"Turn this into a video with..." image-to-video 4–6 s
Cinematic, story, ad Prefer text-to-video with detailed prompt 5–10 s

Generation Modes

Mode Description Models
text-to-video Text prompt only → video minimax, kling, veo, hunyuan, grok, seedance
image-to-video Single image + prompt → animated clip minimax, kling, veo, pixverse, grok, seedance
reference-based Reference images/video → consistent output minimax, kling, veo, hunyuan, grok, seedance

Models (use --model <id>)

Model ID T2V I2V Reference Notes
minimax Subject reference image, character consistency
kling Multi-element / character / keyframe (O3)
veo Google Veo 3.1, multiple reference images
hunyuan Video-to-video style transfer
pixverse Stylized image-to-video
grok Video editing via reference video
seedance Seedance 1.5 Pro, synchronized audio, 4–12 s

Full model details and endpoint reference: [references/models.md](references/models.md).


How to Generate a Video

Step 1 — Choose mode and enhance the prompt

  • Text-to-video: Expand with subject, action, camera movement, lighting, and style. Be specific about motion (e.g. "camera slowly zooms in", "character walks left to right").
  • Image-to-video: Describe the motion to apply to the image (e.g. "gentle breeze in the hair", "camera pans across the scene"). See [references/promptguide.md](references/promptguide.md) for patterns.

Step 2 — Run the script

Text-to-video:

node {baseDir}/tools/generate.js \
  --mode text-to-video \
  --prompt "<enhanced prompt>" \
  --duration <seconds> \
  --aspect-ratio <ratio>

Image-to-video:

node {baseDir}/tools/generate.js \
  --mode image-to-video \
  --prompt "<motion description>" \
  --image-url "<public image URL>" \
  --duration <seconds> \
  --aspect-ratio <ratio>

Parameters:

Parameter Default Description
--mode text-to-video text-to-video or image-to-video
--prompt (required) Scene or motion description
--image-url Required for image-to-video; public image URL
--duration 5 Length in seconds (typically 4–10)
--aspect-ratio 16:9 16:9, 9:16, 1:1, 4:3, 3:4
--model auto Model ID (e.g. kling, veo, grok, seedance); auto = proxy picks

Other commands:

Command Description
node tools/generate.js --list-models List available models from the proxy
node tools/generate.js --status --job-id <id> Check async job status

Step 3 — Return the result

The script returns JSON:

{
  "success": true,
  "mode": "text-to-video",
  "videoUrl": "https://...",
  "duration": 5,
  "aspectRatio": "16:9"
}

Send videoUrl to the user.


Example Conversations

User: "Generate a short video of a cat walking in the rain, cinematic."

node {baseDir}/tools/generate.js \
  --mode text-to-video \
  --prompt "A cat walking through rain, wet streets, neon reflections, cinematic lighting, slow motion, 4K" \
  --duration 5 \
  --aspect-ratio 16:9

User: "Animate this photo" (user uploads a landscape)

node {baseDir}/tools/generate.js \
  --mode image-to-video \
  --prompt "Gentle clouds moving across the sky, subtle grass movement, cinematic atmosphere" \
  --image-url "https://..." \
  --duration 5 \
  --aspect-ratio 16:9

User: "Make a 10-second vertical video of a coffee pour, slow motion."

node {baseDir}/tools/generate.js \
  --mode text-to-video \
  --prompt "Close-up of coffee pouring into a white cup, slow motion, steam rising, soft lighting, product shot" \
  --duration 10 \
  --aspect-ratio 9:16

User: "Use Google Veo for a cinematic shot."

node {baseDir}/tools/generate.js \
  --mode text-to-video \
  --model veo \
  --prompt "A dragon flying through cloudy skies, cinematic lighting, 8s" \
  --duration 8 \
  --aspect-ratio 16:9

User: "Animate this portrait."

node {baseDir}/tools/generate.js \
  --mode image-to-video \
  --model grok \
  --prompt "Gentle smile, subtle head turn" \
  --image-url "https://..." \
  --duration 5

Setup

Zero API keys by default. Requests go through a hosted proxy. Set these for a custom proxy or token:

Variable Required Description
VIDEOSTUDIOPROXY_URL No Proxy base URL
VIDEOSTUDIOTOKEN No Auth token if the proxy requires it

Knowledge Base

  • [references/promptguide.md](references/promptguide.md) — Prompt patterns for text-to-video and image-to-video.
  • [references/models.md](references/models.md) — Model list, capabilities, and selection guide.
  • [references/callingguide.md](references/callingguide.md) — Per-model endpoint details, input parameters, and special handling.