Kimi K3 is not cheap

18 points | by ainch an hour ago

16 comments

  • himata4113 43 minutes ago

    it IS cheap (once the weights are released) and it will only become CHEAPER. For around $3700 a month (via loan purchased hardware + energy cost) you can run around 32 concurrent instances of kimi k3 which can generate nearly a 6.9 billion tokens a day.

    This is napkin math since I'm mostly just extrapolating from glm 5.2 by assuming it's twice as heavy to serve in every single measurement, but I believe you can easily achieve 2500tok/s aggregate compared to 4500tok/s and up to 8000tok/s for glm5.2.

    with nvidia r100 you are likely going to be able to push that number even higher while the cost of hardware appears to be relatively the same, so far I am seeing 21% premium from supermicro which is twice as fast and has nearly twice the vram.

  • sroerick 41 minutes ago

    It feels like this "Kimi is a token hog" meme is 100% astroturfed by Anthropic. It's cheap. Believe your own eyes.

      ofjcihen 38 minutes ago

      I mean at this point their very existence depends on it so I’m not sure if I’d be surprised

        sroerick 31 minutes ago

        Not to mention -

        If you switch the view to "coding tasks" on this website:

          Kimi K3: $3.18 per task
          GLM 5.2: $6.51 per task
          GPT 5.6 Sol: $7.02 per task
          Opus 5: 8.23 per task
          Fable: 11.70 per task
        
        So it's pretty dang cheap lol. Nobody is using frontier inference for "office tasks".
          ofjcihen 27 minutes ago

          Right? The availability of this being in the article that’s pushing the opposite narrative is like…what?

            sroerick 20 minutes ago

            I was actually shocked to see that much improvement on GLM 5.2. I am getting pretty great rates in GLM5.2 right now and I'm extremely happy with the output. I have found Kimi to be generally a little slower but noticeably better at architecture and structuring things. I would have thought is for sure currently more expensive than GLM5.2, particularly with subscriptions etc, but I'm excited for this to decrease further.

  • jszymborski 39 minutes ago

    K2.6 is cheaper than GLM5.2 (at least on DeepInfra) and I've found it works as good as Sonnet for my purposes. Both tend to think themselves into circles a bit and aren't super token efficient, but I've found GLM5.2 much worse on this count making K2.6 even cheaper than the per token price would make seem.

      sroerick 18 minutes ago

      I find them to be about comparable, but I use them both for coding tasks and I'm happy with each. I like K3

  • coder543 42 minutes ago

    This article seems premature to post. Right now, the price is arbitrarily set by a single provider. Why wouldn't Moonshot collect extra revenue during this exclusivity period when they knew there would be hype?

    The model weights are supposed to release tomorrow.

    Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.

      ainch 11 minutes ago

      That's a very fair critique.

      I don't mean to imply that Kimi is not at all cheaper than U.S frontier models. I more wrote this because I believe - since Chinese LLMs entered the public consciousness via DeepSeek R1, which was genuinely ~20x cheaper than o1 - there's a bit of a halo effect around Chinese models which causes people to overestimate the scale of the discount. And relative to that price anchor, Kimi is less extraordinarily cheap.

      At the moment Kimi is ~10% cheaper than GPT-5.6 on the AA benchmark, and as you say that could go down to 20-30% cheaper (although I don't know how inference provider discounts play out on real world usage once you account for quantisation etc...). I'm not trying to suggest that that's nothing, but I do think some of the people driving the Chinese AI discourse would have a harder time pitching their conclusions if they were saying "this new Chinese model is 10% cheaper on some tasks, and it might get another 20% cheaper in the future".

  • ronsor an hour ago

    It's cheap because it won't refuse random tasks. You can't get rid of nannying at any price beyond training your own model, and relative to that, K3 is cheap.

  • cloudie78 an hour ago

    It’s cheap.

  • ofjcihen 41 minutes ago

    I mean you have a chart showing that it’s cheaper than the other models and it also does what I want without argument.

    Additionally, I fully expect the frontier labs to continue increasing prices to meet the profit margins they need to to continue existing.

  • CamperBob2 43 minutes ago

    A lot hinges on what happens tomorrow. I'll believe they'll open the weights when I see the files appear on HF (and when somebody with 24 RTX6000s or whatever reports that they are indeed as good as the closed version.)

  • SwellJoe an hour ago

    In my testing, I'm finding it more expensive than Opus 4.8/5 and GPT 5.6 Sol at API rates, because it chews so much. And, their plan (at least the $19 tier) is much less generous than the ChatGPT $20 plan, like an order of magnitude less, it's basically a demo not a useful amount of usage.

      Stagnant 20 minutes ago

      Yeah can't recommend their $19 plan, only took me a day and a half to hit the weekly usage cap. The $39 plan has 5x the limits so I recommend getting that instead. Having now tested it for a few days, it is the first of the chinese models that actually feels comparable to Opus-tier models