The trade-off: Capability correlates with refusal. The models at the top write the best code but refuse the most requests. The models at the bottom refuse nothing but write worse code. GLM 5.2 and Qwen 3.7 Plus are the only models with above-average coding AND >50% compliance — the practical balance point. Use this to select models for your use case, not to judge them.