Strix Halo PyTorch Setup
Set up a reproducible gfx1151 environment, then prove the capabilities the workload needs. Do not infer support from GPU enumeration or a large allocation.
Workflow
Commands below assume the skill is installed globally at ~/.claude/skills/strix-halo-setup/. For a per-project copy, substitute ./.claude/skills/strix-halo-setup/.
- Run the system verifier:
``bash ~/.claude/skills/strix-halo-setup/scripts/verify_system.sh ``
- Choose an installation track with the user:
| Track |
Use when |
Tradeoff |
| AMD supported |
Stability and AMD's validated matrix matter most |
Older PyTorch and kernel stack |
| TheRock multi-arch |
New PyTorch, Triton, AOTriton, or rapid gfx1151 fixes matter most |
Moving nightly packages; regressions are possible |
Default to the AMD supported track unless the user explicitly prioritizes new features or agrees to nightly risk. See [installation details](docs/INSTALLATION.md).
- Create a fresh virtual environment. AMD's manylinux wheels ship cp310 through
cp313, so match the interpreter to the wheel tag rather than assuming 3.12. Never install these wheels into the system Python or an existing environment unless the user requests it.
- Install one track. Do not combine AMD supported wheels, PyTorch.org wheels,
old per-family TheRock wheels, or multi-arch TheRock packages in one environment.
- Run the capability verifier:
``bash python ~/.claude/skills/strix-halo-setup/scripts/verify_pytorch.py ``
- For AOTriton flash attention, start a new process with the experimental
switch and force the backend during verification:
``bash TORCHROCMAOTRITONENABLEEXPERIMENTAL=1 \ python ~/.claude/skills/strix-halo-setup/scripts/verifypytorch.py \ --require flashattention ``
- Capture the resolved environment after it passes:
``bash python -m torch.utils.collect_env > collect-env.txt python -m pip freeze > requirements-lock.txt ``
Installation Tracks
AMD Supported
AMD's ROCm Ryzen matrix validates gfx1151 with PyTorch 2.9.1 and FP16. Install AMD's exact repo.radeon.com wheels as documented in [INSTALLATION.md](docs/INSTALLATION.md), not similarly named PyTorch.org wheels. AMD ships patch releases faster than this skill is revised, and every one rewrites the git hash in each filename — list repo.radeon.com/rocm/manylinux/ and take the newest rocm-rel-* rather than trusting a pinned URL.
Use this track when reproducibility is more important than the newest compiler or attention work.
TheRock Multi-Arch Nightly
TheRock replaced new per-family releases with a unified multi-architecture index. Follow the exact [TheRock multi-arch installation](docs/INSTALLATION.md#therock-multi-arch-nightly) and select gfx1151 through the device-gfx1151 package extras.
Do not add --pre by default. The index already publishes ROCm development builds behind stable-looking PyTorch versions; --pre may select a newer PyTorch alpha. Always record the resolved versions and retain the environment until its replacement passes the same capability checks.
Runtime Configuration
Start with no Strix-specific environment overrides. Current packages already identify gfx1151, select visible devices, and choose BLAS/allocator defaults.
Set only the switch required by a tested feature:
export TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1
Do not set these globally:
HSAOVERRIDEGFX_VERSION: hides an architecture/package mismatch.
PYTORCHROCMARCH: build-time/JIT target selection, not normal runtime setup.
HSAENABLESDMA=0: disables DMA copies; use only to isolate a reproduced bug.
ROCRVISIBLEDEVICES or HIPVISIBLEDEVICES: restrict devices only when asked.
ROCBLASUSEHIPBLASLT=1: leave backend selection on automatic unless a
workload-specific comparison proves otherwise.
HSACUMASK, HSAXNACK, HSAFORCEFINEGRAIN_PCIE, and heap percentage
overrides: diagnostic controls, not baseline optimizations.
See [performance features](docs/PERFORMANCE_FEATURES.md) for attention, torch.compile, BLAS, convolution, and dtype guidance.
Unified Memory
Strix Halo has unified physical memory. Linux exposes overlapping VRAM and GTT accounting views; never add them together or describe their sum as usable RAM.
Before changing memory configuration:
- Read physical RAM, current GTT, and VRAM separately.
- Estimate weights, KV cache, activations, allocator overhead, and host needs.
- Leave enough physical RAM for the OS and the workload's CPU allocations.
- Prefer AMD's
amd-ttm helper over hand-written kernel parameters.
- Reboot and re-run verification after a change.
Run the read-only advisor:
~/.claude/skills/strix-halo-setup/scripts/configure_gtt.sh
See [GTT memory configuration](docs/GTTMEMORYFIX.md). Do not claim a model is supported from a synthetic allocation; run a representative inference or training step with the intended precision and context length.
Interpreting Verification
verify_system.sh checks host prerequisites without modifying the machine.
verify_pytorch.py launches bounded real kernels for FP32, FP16, BF16,
matrix multiplication, MIOpen convolution/backward, SDPA, forced flash SDPA, torch.compile, and a touched allocation.
- Optional feature warnings do not invalidate basic PyTorch compute. A feature
requested through --require must pass or the script exits nonzero.
- BF16 passing locally is useful evidence, but AMD's ROCm Ryzen matrix
officially lists FP16 validation for gfx1151.
- A flash-attention pass proves PyTorch dispatched the forced SDPA backend for
the tested shape. It does not prove every model shape uses that backend.
Troubleshooting Order
- Save the exact command and complete error.
- Run both verifiers in the affected environment.
- Confirm the installed wheel contains a gfx1151 device package.
- Remove inherited ROCm/HSA overrides and retry in a new process.
- Compare against a fresh environment on the other installation track.
- Check TheRock's current test status and issues before changing the host.
Use [TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) for symptom-specific fixes.
Boundaries
- Focus on Linux PyTorch setup and validation, not model-specific scaffolding.
- Treat TheRock versions and feature status as time-sensitive; query the index
when installing rather than copying a version from this document.
- Do not add benchmark numbers or generic model-size compatibility tables.
- Do not modify BIOS, kernel boot parameters, or system memory without explicit
user approval and a rollback plan.