smithery.ai

strix-halo-setup

Set up and diagnose PyTorch ROCm environments on AMD Strix Halo (gfx1151) Linux systems.

First seen Apr 26, 2026

Installation

$ npx skills add https://smithery.ai

Summary

  • Set up and diagnose PyTorch ROCm environments on AMD Strix Halo (gfx1151) Linux systems.
  • Use for selecting a supported or TheRock nightly stack, verifying real GPU kernels and attention backends, inspecting unified-memory limits, or troubleshooting invalid-device-function and out-of-memory failures.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from smithery.ai · top by installs.

npx skills add https://smithery.ai

Browse all from smithery.ai

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Skill metadata

Parsed from SKILL.md frontmatter.

LicenseMIT
Declared agents claude-code
More metadata
hardware
AMD Strix Halo (gfx1151)
supported_rocm
ROCm 7.2.4/PyTorch 2.9.1 supported; TheRock multi-arch nightlies experimental
tested_date
2026-08-10
skill_version
2.1.0

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 7,343 B
  • docs SUMMARY.md 325 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1 installs

SKILL.md

Strix Halo PyTorch Setup

Set up a reproducible gfx1151 environment, then prove the capabilities the workload needs. Do not infer support from GPU enumeration or a large allocation.

Workflow

Commands below assume the skill is installed globally at ~/.claude/skills/strix-halo-setup/. For a per-project copy, substitute ./.claude/skills/strix-halo-setup/.

  1. Run the system verifier:

``bash ~/.claude/skills/strix-halo-setup/scripts/verify_system.sh ``

  1. Choose an installation track with the user:
Track Use when Tradeoff
AMD supported Stability and AMD's validated matrix matter most Older PyTorch and kernel stack
TheRock multi-arch New PyTorch, Triton, AOTriton, or rapid gfx1151 fixes matter most Moving nightly packages; regressions are possible

Default to the AMD supported track unless the user explicitly prioritizes new features or agrees to nightly risk. See [installation details](docs/INSTALLATION.md).

  1. Create a fresh virtual environment. AMD's manylinux wheels ship cp310 through

cp313, so match the interpreter to the wheel tag rather than assuming 3.12. Never install these wheels into the system Python or an existing environment unless the user requests it.

  1. Install one track. Do not combine AMD supported wheels, PyTorch.org wheels,

old per-family TheRock wheels, or multi-arch TheRock packages in one environment.

  1. Run the capability verifier:

``bash python ~/.claude/skills/strix-halo-setup/scripts/verify_pytorch.py ``

  1. For AOTriton flash attention, start a new process with the experimental

switch and force the backend during verification:

``bash TORCHROCMAOTRITONENABLEEXPERIMENTAL=1 \ python ~/.claude/skills/strix-halo-setup/scripts/verifypytorch.py \ --require flashattention ``

  1. Capture the resolved environment after it passes:

``bash python -m torch.utils.collect_env > collect-env.txt python -m pip freeze > requirements-lock.txt ``

Installation Tracks

AMD Supported

AMD's ROCm Ryzen matrix validates gfx1151 with PyTorch 2.9.1 and FP16. Install AMD's exact repo.radeon.com wheels as documented in [INSTALLATION.md](docs/INSTALLATION.md), not similarly named PyTorch.org wheels. AMD ships patch releases faster than this skill is revised, and every one rewrites the git hash in each filename — list repo.radeon.com/rocm/manylinux/ and take the newest rocm-rel-* rather than trusting a pinned URL.

Use this track when reproducibility is more important than the newest compiler or attention work.

TheRock Multi-Arch Nightly

TheRock replaced new per-family releases with a unified multi-architecture index. Follow the exact [TheRock multi-arch installation](docs/INSTALLATION.md#therock-multi-arch-nightly) and select gfx1151 through the device-gfx1151 package extras.

Do not add --pre by default. The index already publishes ROCm development builds behind stable-looking PyTorch versions; --pre may select a newer PyTorch alpha. Always record the resolved versions and retain the environment until its replacement passes the same capability checks.

Runtime Configuration

Start with no Strix-specific environment overrides. Current packages already identify gfx1151, select visible devices, and choose BLAS/allocator defaults.

Set only the switch required by a tested feature:

export TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1

Do not set these globally:

  • HSAOVERRIDEGFX_VERSION: hides an architecture/package mismatch.
  • PYTORCHROCMARCH: build-time/JIT target selection, not normal runtime setup.
  • HSAENABLESDMA=0: disables DMA copies; use only to isolate a reproduced bug.
  • ROCRVISIBLEDEVICES or HIPVISIBLEDEVICES: restrict devices only when asked.
  • ROCBLASUSEHIPBLASLT=1: leave backend selection on automatic unless a

workload-specific comparison proves otherwise.

  • HSACUMASK, HSAXNACK, HSAFORCEFINEGRAIN_PCIE, and heap percentage

overrides: diagnostic controls, not baseline optimizations.

See [performance features](docs/PERFORMANCE_FEATURES.md) for attention, torch.compile, BLAS, convolution, and dtype guidance.

Unified Memory

Strix Halo has unified physical memory. Linux exposes overlapping VRAM and GTT accounting views; never add them together or describe their sum as usable RAM.

Before changing memory configuration:

  1. Read physical RAM, current GTT, and VRAM separately.
  2. Estimate weights, KV cache, activations, allocator overhead, and host needs.
  3. Leave enough physical RAM for the OS and the workload's CPU allocations.
  4. Prefer AMD's amd-ttm helper over hand-written kernel parameters.
  5. Reboot and re-run verification after a change.

Run the read-only advisor:

~/.claude/skills/strix-halo-setup/scripts/configure_gtt.sh

See [GTT memory configuration](docs/GTTMEMORYFIX.md). Do not claim a model is supported from a synthetic allocation; run a representative inference or training step with the intended precision and context length.

Interpreting Verification

  • verify_system.sh checks host prerequisites without modifying the machine.
  • verify_pytorch.py launches bounded real kernels for FP32, FP16, BF16,

matrix multiplication, MIOpen convolution/backward, SDPA, forced flash SDPA, torch.compile, and a touched allocation.

  • Optional feature warnings do not invalidate basic PyTorch compute. A feature

requested through --require must pass or the script exits nonzero.

  • BF16 passing locally is useful evidence, but AMD's ROCm Ryzen matrix

officially lists FP16 validation for gfx1151.

  • A flash-attention pass proves PyTorch dispatched the forced SDPA backend for

the tested shape. It does not prove every model shape uses that backend.

Troubleshooting Order

  1. Save the exact command and complete error.
  2. Run both verifiers in the affected environment.
  3. Confirm the installed wheel contains a gfx1151 device package.
  4. Remove inherited ROCm/HSA overrides and retry in a new process.
  5. Compare against a fresh environment on the other installation track.
  6. Check TheRock's current test status and issues before changing the host.

Use [TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) for symptom-specific fixes.

Boundaries

  • Focus on Linux PyTorch setup and validation, not model-specific scaffolding.
  • Treat TheRock versions and feature status as time-sensitive; query the index

when installing rather than copying a version from this document.

  • Do not add benchmark numbers or generic model-size compatibility tables.
  • Do not modify BIOS, kernel boot parameters, or system memory without explicit

user approval and a rollback plan.