Ran Qwen3.8-27B-4bit locally on my M1 Pro 32GB
.
10 tok/s. Each sentence hangs for 2-3 seconds before output starts. Long prompt prefill takes over a minute. 128K context, in theory.
27B boots up fine on M1 Pro, but as a daily driver it falls short. Need double the memory bandwidth to hit 20 tok/s for a smooth experience.