Running Kimi K3 on a M1 Max

48 points | by tito an hour ago

31 comments

  • antirez 35 minutes ago

    SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528

    Soon decent speed across two Mac Studios with 512GB of RAM.

      tito 34 minutes ago

      wow cool! I like watching the new models come out and how they end up crammed in to run on local machines. I learned what mxfp4 is thanks to this latest Kimi release - although it sounds like it means that there's less room for compression in the model compared to others.

  • acmnrs 9 minutes ago

    The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.

      tito 7 minutes ago

      Done, Mac -> Max

  • mips_avatar 11 minutes ago

    Would be interesting to see how fast it would be on 4x mac studio 512gb machines.

  • nlessard 22 minutes ago

    Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).

      trollbridge 2 minutes ago

      No problem at all to read data over and over. In fact, LLM weights are a great candidate for low-quality flash that can't handle a lot of write cycles, and you want a large amount of storage cheaply...

      addaon 20 minutes ago

      Reads are not generally life-limiting for flash. (Well, no more so than power-on time in general. You still have aging mechanisms like electromigration, but these are orders of magnitude slower than write-induced damage.)

  • Azantys 36 minutes ago

    0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

      ggm 15 minutes ago

      So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ?

      (note. I am not a believer in AGI)

      "useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

        tito 3 minutes ago

        Are there any well thought through stories about what this would look like? For example, I'm thinking about like nutrient flow, decision making, energy input, gravitational force, things like that seem to govern the value and speed of intelligence.

      SXX 26 minutes ago

      It is fun.

      Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.

      Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.

      SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.

      tito 32 minutes ago

      I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.

        magicalhippo 20 minutes ago

        > one person on reddit suggested using it in an email interface rather than a chat interface

        Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.

        Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.

        embedding-shape 17 minutes ago

        Email would indeed be fitting for K3 running on a M1 Mac, as it'd take days/weeks to receive a response, which matches with my real-world emailing experience pretty well.

        hugopuybareau 26 minutes ago

        Love the email idea

          tito 21 minutes ago

          Having the right type of interface makes a huge difference.

          It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.

          At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)

      winstonp 17 minutes ago

      16 tokens / s is not nothing.

        Azantys 5 minutes ago

        ? Readme says 60-70s per token

  • tjwebbnorfolk 16 minutes ago

    > ~60–76 s/token

    I don't know if I'd call this "running"

      sermah 14 minutes ago

      Had the same feeling when I first saw min/km units in some (human) running context.

      UPD: I know it's not the same at all, just the reversal of units that gets me

      tito 13 minutes ago

      I commented similarly below, but as a terrible programmer, I probably perform about 1 minute per token too (at Kimi 3 level). It puts into context how I think about intelligence

        brokencode 10 minutes ago

        Maybe in terms of code produces, but one token is only a fragment of a thought for an LLM.

        It’d be like thinking as slowly as Ents talk to each other in Lord of the Rings.

  • als0 34 minutes ago

    Says it requires a 2TB disk? Must it be internal NVMe?

      Fergusonb 19 minutes ago

      You can use an external drive if it's mounted as a writable volume. I would make sure it's fast, maybe thunderbolt 3/4/5 enclosure with a fast drive.

      tito 32 minutes ago

      The Github specifically mentions an option to stream it from an online host. It's extremely slow.

  • piterrro 28 minutes ago

    Will it fit on ESP32??

      embedding-shape 18 minutes ago

      Better questions, how many ESP32s would it take to reach 1 tok/s decoding speed with K3?

      tito 27 minutes ago

      1 button Kimi morse code interface

  • lostmsu 41 minutes ago

    under 0.02 tok/s