calesthio/openmontage

grok-media

xAI Grok image and video generation guide covering authentication, endpoints, prompt structure, image editing, reference-image video, and async polling.

First seen Apr 7, 2026

Installation

$ npx skills add calesthio/openmontage --skill grok-media

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from calesthio/openmontage · top by installs.

npx skills add calesthio/openmontage

Browse all from calesthio/openmontage

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 56.6K
License LICENSE
Default branch main
Open issues 96
Status Active

Skill metadata

Parsed from SKILL.md frontmatter.

Version1.0.0
More metadata
author
OpenMontage
version
1.0.0
tags
xai, grok, image-generation, video-generation, media

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 3,799 B
  • docs SUMMARY.md 167 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 655 installs

SKILL.md

Grok Media

Use this skill when working with xAI media models in OpenMontage.

Models

  • grok-imagine-image for image generation and image editing
  • grok-imagine-video for text-to-video, image-to-video, and reference-image video

Authentication

  • Env var: XAIAPIKEY
  • Base URL: https://api.x.ai/v1
  • Header: Authorization: Bearer $XAIAPIKEY

Image API

Text-to-image

  • Endpoint: POST /images/generations
  • Core fields:

- model - prompt - n - aspect_ratio - resolution

Image edit

  • Endpoint: POST /images/edits
  • Use image for one source image
  • Use images for multi-image compositing
  • Each source image can be:

- a public HTTPS URL - a base64 data URI

Image prompting

  • Grok responds well to direct natural language
  • For edits, describe only the intended change and preserve everything else implicitly
  • For multi-image merges, explicitly name how each source contributes
  • Prefer one strong scene description over long style-stacking

Video API

Generation

  • Endpoint: POST /videos/generations
  • Polling endpoint: GET /videos/{request_id}
  • Success state: status == "done"
  • Failure states to handle explicitly: failed, expired

Modes

  • Text-to-video:

- prompt-only generation

  • Image-to-video:

- use image: {"url": ...} - this anchors the starting frame

  • Reference-to-video:

- use referenceimages: [{"url": ...}, ...] - this influences who/what appears in the video without locking the first frame - prompts can reference inputs with placeholders like <IMAGE1>, <IMAGE_2>

Video constraints

  • Grok video is best treated as short-form generation
  • Current output resolutions are 480p and 720p
  • Reference-image video supports multiple images and is useful for product placement, wardrobe transfer, and identity consistency
  • Download outputs promptly; provider URLs may be temporary

Pricing

  • grok-imagine-image: $0.02 per generated image
  • grok-imagine-image edits/composites: add $0.002 per input image
  • grok-imagine-video:

- 480p: $0.05 per second - 720p: $0.07 per second

  • grok-imagine-video image-conditioned requests: add $0.002 per input image

Grok-Specific Prompt Guidance

Images

  • Start with subject, action, setting
  • Add one style anchor, not five
  • For edits:

- describe the desired modification - keep the rest of the image stable by omission, not by writing a giant preservation list

Video

  • Keep prompts scene-local: one shot, one main motion idea, one emotional beat
  • For reference-conditioned video, explicitly map source images to roles:

- person from <IMAGE1> - jacket from <IMAGE2> - product from <IMAGE_3>

  • Camera and pacing language helps:

- slow push-in - handheld follow - locked-off medium shot - high-energy whip pan transition

Good Fits

  • Image style transfer
  • Image compositing from multiple sources
  • Reference-conditioned short video
  • Product-led motion clips
  • Character-consistent scenes without hard first-frame lock

Weak Fits

  • Long-form clip generation
  • Heavy reliance on deterministic seeds
  • Overloaded prompts with multiple scene changes

Failure Handling

  • If generation submission succeeds but polling expires, surface it as a provider/runtime issue
  • If a request fails, preserve the endpoint, mode, and prompt summary in the error
  • Do not silently substitute a different provider after xAI was selected without user approval