20 comments

  • beltsazar 10 minutes ago

    As a comparison, Qwen3.6 27B scores 38, which was the highest in its small model category (4B–40B).

    Qwen3.8 27B beats all medium models (40B–150B). It has the same score as DeepSeek V4 Flash 0731, which ranks #5 in large model category (> 150B).

    Sources:

    - https://artificialanalysis.ai/models/open-source/small

    - https://artificialanalysis.ai/models/open-source/medium

    - https://artificialanalysis.ai/models/open-source/large

      phsource a minute ago

      Simon Willison's post about this gives a good context on why exactly this is happening. While it doesn't mention this in the Artificial Analysis page, this is likely with Max reasoning, which has extremely long reasoning traces:

      https://simonwillison.net/2026/Aug/16/qwen-38-27b/

      It seems like the token usage is 2.3x GPT Luna Max and almost 2x Kimi K3!

      https://imgur.com/a/dDSyhr2

      I'm curious if they can make up for this with insanely high tokens-per-second especially when served from hosted providers, though, given how tiny it is (37B!)

  • sottol 16 minutes ago

    A lot of the benchmarks seem often near meaningless these days - really bench-maxxed to the hilt. I tend to still look at the Artificial Analysis rankings to get at least an idea on relative performance of models, is that still warranted?

    What or other opinions on how representative the AA rankings are of real-world performance? Any better indicators?

      Iolaum 11 minutes ago

      Yea and we are reaching the point where this benchmaxing is visible in the model's reported overthinking.

        logicchains 10 minutes ago

        It's not overthinking, it's the right amount of thinking necessary for such a small model to get good results. The dumber the model, the more it has to think to be smart. There's no easy way to reduce the thinking without reducing the model quality.

  • anana_ an hour ago

    For more context, this puts it on par with models like GLM 5.2 and GPT 5.6 Luna, which are far larger

      anana_ 11 minutes ago

      And to read the tea leaves a little:

      3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.

      It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.

      nsingh2 10 minutes ago

      Also with Qwen 3.8 being more token hungry than Luna, using around 2.3x tokens. Which hurts for local deployment.

      bertili 17 minutes ago

      And more context:

      Same score as the latest DeepSeek Flash 0731 which has 284B parameters! (13B active)

      Its also the second best Qwen model, much better than Qwen 3.7 Max, but significantly below Qwen 3.8 Max.

      11 minutes ago
      [deleted]
      johnnyApplePRNG 10 minutes ago

      We don't actually know how large they are, actually.

  • bertili 9 minutes ago

    I can't shake this the existential feeling that this compact series of 27G bytes represent something profound and universal.

  • 6 minutes ago
    [deleted]
  • johnnyApplePRNG 9 minutes ago

    Unbelievable. Bravo Qwen team.

  • colingauvin 15 minutes ago

    It's 7th (!!!) overall on the agentic index, above Terra.

      hadlock 13 minutes ago

      Strangely Qwen 3.8 Max isn't on their list, at all.

  • apitman an hour ago

    Very interesting. I was not expecting anything close to this.

  • manunicholasjac 19 minutes ago

    [flagged]

  • kessler9 an hour ago

    [flagged]