Qwen3.8-27B dropped open source and it's taking shots at Opus4.6.
Dense model, hybrid architecture, native VLM — sees images and video out of the box.
The numbers: SWE-bench Pro 61.7 vs Opus4.6 Max's 53.4, LiveCodeBench 90.3 (best at this parameter count), OSWorld 84.3 vs Opus's 72.7 on computer use. Not even close.
Qwen keeps punching above its weight on the small model side. I've been benchmarking small-model agent capabilities and nothing outside the Qwen family can actually survive long tasks.
🔗 huggingface.co/Qwen/Qwen3.8-27B