Summary
- CLI for text, image, video, speech, and music generation via MiniMax AI platform.
- Supports six content modalities: text chat, image generation, video generation, text-to-speech, music generation with lyrics or covers, and image understanding via vision models Includes web search, quota management, and async task polling for long-running operations like video generation Provides agent-friendly flags ( --non-interactive , --quiet , --output json , --async , --dry-run ) and clean stdout/stderr separation for reliable piping and chaining Configurable per-modality defaults, OAuth and API key authentication, region auto-detection, and tool schema export for dynamic agent framework integration