Ember-1

77 points | by gmays an hour ago

29 comments

  • nico 2 minutes ago

    > The problem: thinking models think too much

    This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

    It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

  • intothemild 21 minutes ago

    So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?

      DonsDiscountGas a minute ago

      It happens. Most open licenses aren't GPL style copyleft.

      netvarun 11 minutes ago

      Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.

      16 minutes ago
      [deleted]
  • jamienk 24 minutes ago

    Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

      andsoitis 13 minutes ago

      > Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

      I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

      So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

        jamienk 7 minutes ago

        Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port

          andsoitis 5 minutes ago

          > then new work is done on top of stuff that "hits" in a way no one anticipated.

          Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.

          Greatness cannot be planned.

      segmondy 13 minutes ago

      No, because close labs/models borrow but don't contribute back.

      intothemild 20 minutes ago

      Yes, absolutely, but only if people keep contributing in the open.

        jack_pp 13 minutes ago

        not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly

  • netvarun 15 minutes ago

    Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15

      drob518 a few seconds ago

      Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.

  • andsoitis 27 minutes ago

    > The problem: thinking models think too much

    Analysis paralysis stifles not just human intelligence, but other intelligences too.

  • dbuxton 3 minutes ago

    Do they mean Opus 5.5 or Opus 5?

  • themgt 2 minutes ago

    The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

    "Pareto": 8 hits

    "Opus 5.5": zero hits

  • tomrod 25 minutes ago

    Well done, and great iteration.

    The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

  • logicallee 6 minutes ago

    This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.

  • tdhz77 22 minutes ago

    Does anybody know if this would be a good model for creative writing?

  • erichocean 26 minutes ago

    Need this done for DeepSeek, ideally one of the Flash models.

      atemerev 22 minutes ago

      If you have the compute, I have the expertise.

  • ls612 24 minutes ago

    On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.

      spijdar 16 minutes ago

      I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.

      That's just vibes, though.

  • esafak 24 minutes ago

    It looks like it would be similar to GLM 5.3 Flash, had they tested it...

  • monkey_monkey 26 minutes ago

    I don't think the article mentions Pareto frontier enough.

    Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?

      AnodicElegy 14 minutes ago

      I guess they figure "best bang for your buck" comes off a little too colloquial.

  • justmeeew 6 minutes ago

    [dead]

  • huflungdung 27 minutes ago

    [dead]