maxgent-ai/maxgent-plugin · Archived

media-understand

AI-powered media understanding and analysis for images, videos, and audio. Use when users ask to describe, analyze, summarize, or extract text (OCR) from media files.

First seen Jan 28, 2026

Installation

$ npx skills add maxgent-ai/maxgent-plugin --skill media-understand

Stronger alternatives

This repository is archived — consider an actively maintained alternative.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from maxgent-ai/maxgent-plugin · top by installs.

npx skills add maxgent-ai/maxgent-plugin

Browse all from maxgent-ai/maxgent-plugin

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Declared
Cline Not declared
OpenCode Not declared

Repository health

License MIT
Default branch main
Open issues 0
Status Archived

Skill metadata

Parsed from SKILL.md frontmatter.

Declared agents gemini

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,435 B
  • docs SUMMARY.md 190 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 8 installs

SKILL.md

Media Understanding

Analyze multimedia content via Maxgent FAL API proxy, using the default route.

Supported Formats

Type Formats Max Size
Image jpg, jpeg, png, gif, webp 20MB
Video mp4, mpeg, mov, webm, YouTube URL 100MB
Audio wav, mp3, aiff, aac, ogg, flac, m4a 100MB

Prerequisites

  1. MAXAPIKEY environment variable (auto-injected by Max)
  2. Bun 1.0+ (built into Max)

Routing

  1. default

- Endpoint: openrouter/router/openai/v1/chat/completions - Model: DEFAULTMMMODEL, defaults to google/gemini-2.5-pro (override with --model)

Usage

bun skills/media-understand/media-understand.js \
  --media PATH_OR_URL --prompt "PROMPT" \
  [--language chinese|english] [--model MODEL_ID] \
  [--max-tokens N] [--temperature X]

Parameters:

  • --media: local file path or YouTube URL
  • --prompt: analysis question
  • --language: chinese (default) or english
  • --model: override the default model
  • --max-tokens: max output tokens (default 4096)
  • --temperature: sampling temperature (default 0.2)

Examples

# Image OCR
bun skills/media-understand/media-understand.js --media ./screenshot.png --prompt "extract all text from this image" --language english

# Video summary (YouTube)
bun skills/media-understand/media-understand.js --media "https://youtube.com/watch?v=xxx" --prompt "summarize this video" --language english

# Local audio analysis
bun skills/media-understand/media-understand.js --media ./meeting.m4a --prompt "summarize key points and list action items" --language english

Instructions

  1. Check MAXAPIKEY.
  2. Identify media type and validate size limits.
  3. Analyze using the default route; override the model with --model if needed.
  4. Local images/videos/audio are auto-uploaded via FAL upload proxy before analysis.
  5. On success, return readable text.
  6. On failure:

- HTTP 402 (insufficient credits): Stop immediately. Do NOT retry. Tell the user their API credits are exhausted. - Other errors: retry once with a different model. If it fails again, stop and clearly indicate whether it's an upload / proxy / model parameter issue.