smithery/lawless-m

Vram-GPU-OOM

Use when GPU services (Ollama, Whisper, ComfyUI/Flux, OCR) contend for VRAM on the RTX 3090 and hit CUDA OOM.

Installation

$ npx skills add smithery/lawless-m --skill vram-gpu-oom

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery/lawless-m.

npx skills add smithery/lawless-m

Browse all from smithery/lawless-m

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 1,631 B
  • docs SUMMARY.md 181 B

History

  1. First recorded snapshot · 0 installs

SKILL.md

GPU OOM / VRAM sharing

Multiple services share one RTX 3090 (24GB). They coordinate without a central scheduler: everyone tries to load normally, catches OOM, waits for others to auto-unload, and retries.

Retry convention

On CUDA OOM: torch.cuda.empty_cache(), time.sleep(30), retry — up to 3 attempts, 30s apart. Re-raise non-OOM errors immediately and re-raise after the final attempt. Same idea in shell: loop a GPU command 3 times with a 30s sleep between failures. Alongside retry, configure every service to unload quickly when idle.

Known services and settings

  • Ollama already handles unload — just set quick keep-alive in /etc/systemd/system/ollama.service.d/override.conf:

`` Environment="OLLAMAKEEPALIVE=30s" ``

  • Invoice OCR (Qwen2-VL) at http://10.99.0.3:8765 — auto-unloads after idle (--auto-unload-minutes, default 5), does 3×/30s OOM retry, and exposes POST /request-unload and GET /status.
  • Whisper large-v3 ≈ 6GB VRAM (see the Whisper-Transcription skill).

Signaling protocol

For faster, more predictable starts, a service can call POST /request-unload on the others before loading a big model instead of relying on OOM-retry delays. The endpoint contracts (/request-unload, /status, the auto-unload background task), the coordinator usage pattern, and worked timelines are in reference-gpu-coordination.md. Helper script: requestgpuunload.py in the OneCuriousRabbit repo.