A 27B Local Model Got 189 of 191 Fields Right — and Missed the Same One as Gemma 4
I benchmarked Qwen3.8-27B, released by Alibaba under Apache 2.0, across all nine themes of my personal benchmark. On structured output it is the best locally runnable model I have measured, but one-shot HTML generation exposes a different weakness.