Jamesob's guide to running SOTA LLMs locally(github.com)
405 points by livestyle 49 days ago | 182 comments
tl;dr: A detailed hardware guide for running SOTA LLMs locally, ranging from a $2k dual RTX 3090 setup (running Qwen3.6-27B and Whisper STT) to a $40k rig with 4x RTX Pro 6000s (384GB VRAM) capable of running GLM-5.2-594B at near-Claude-Opus quality. The author details their EPYC-based DDR4 build using c-payne PCIe4 switches for direct GPU-to-GPU communication, along with the finicky BIOS tweaks, ACS disabling, and kernel parameters needed to achieve Gen4 line-rate P2P (27.5 GB/s). Includes ready-to-run Docker configs and notes on power-limiting GPUs to run on a 110V circuit.
HN Discussion:
  • The $40k build cost is understated and local models still underperform cloud offerings
  • Cloud subscriptions or Macbooks are more cost-effective than expensive local rigs
  • ~Cheaper or middle-ground alternatives (single 3090, Arc B70, EVO-X2) offer better value than the recommended builds
  • Skepticism about quantized/pruned model quality claims approaching Claude Opus
  • Questions and additions about specific technical details like STT models and isolation