The company also said that it chose Qwen3.6 over the newer Qwen3.8 (released earlier this month) because the latter runs slower on "today's Macs" since it needs reasoning enabled.
How is this easier than using lmstudio or omlx or whatever your favourite runtime is ?
It’s maybe a bit interesting that Jetbrains are making moves to integrate with local models more closely, but I think claiming this is “easier” is incorrect for most users.
I get up to 90 tokens / second with Jundot/Qwen3.6-35B-A3B-oQ6-mtp, on a M3 Max with 64GB. MoE so it is not a dense model, but that means it runs faster (also, mtp helps). It is a 30GB model, but you should be able to load it, otherwise try the 4-bit quant instead of the 6-bit quant, don't bother quanting your KV Cache (don't enable turboquant in oMLX), since that will slow you down.
It seems to me that Qwen3.6-35B-A3B is still the local LLM leader, because 27B in either of the latest releases is just too slow to be usable compared to OpenCode Zen free models, or OpenRouter free models.
Curious why they're using 3.6-27B and not 3.8-27B which is competitive with Opus 4.6 (https://huggingface.co/Qwen/Qwen3.8-27B)
From the article:
The company also said that it chose Qwen3.6 over the newer Qwen3.8 (released earlier this month) because the latter runs slower on "today's Macs" since it needs reasoning enabled.
But that's incorrect, reasoning can be disabled via the template.
I haven’t found that to be true with LM Studio and Qwen3.8. Works fine without reasoning.
They’re also targeting 64GB M5 Pro and up as “today’s Macs” which perform fine with reasoning enabled.
How is this easier than using lmstudio or omlx or whatever your favourite runtime is ?
It’s maybe a bit interesting that Jetbrains are making moves to integrate with local models more closely, but I think claiming this is “easier” is incorrect for most users.
Anything good to run on Mac M5 Max with 48GB? is this even worth trying? so far I found the responses so slow compared to the paid subscriptions...
I get up to 90 tokens / second with Jundot/Qwen3.6-35B-A3B-oQ6-mtp, on a M3 Max with 64GB. MoE so it is not a dense model, but that means it runs faster (also, mtp helps). It is a 30GB model, but you should be able to load it, otherwise try the 4-bit quant instead of the 6-bit quant, don't bother quanting your KV Cache (don't enable turboquant in oMLX), since that will slow you down.
It seems to me that Qwen3.6-35B-A3B is still the local LLM leader, because 27B in either of the latest releases is just too slow to be usable compared to OpenCode Zen free models, or OpenRouter free models.
Sad a Qwen3.8-35B-A3B model wasn't released.
Also, I think this is just an ad.
Uh...
claude> dl & config best vers of Qwen 3.8 for my system