macOS clamps my M3 Ultra's GPU to 338 MHz before the fans even try. Maxing them doubled my LLM throughput.
My LLM benchmarks kept collapsing 4x mid-session. macmon caught a firmware power limiter clamping the GPU to 338 MHz and holding it while the die cooled, fans never past 70%. Pinning them at max with fanpro: 2.57x sustained decode, and a 100k-context job in 259 s instead of 568+, byte-identical.