Flux 3

101 points | by ThouYS an hour ago

20 comments

  • jdthedisciple 2 minutes ago

    It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable.

    But then again I heard the downers have always been the first to leave their dung comments here so let's see...

  • user43928 32 minutes ago

    I hope the open-weight versions will be SOTA.

    > Over the next few weeks and months, we will make the following capabilities available

    > Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

    > We will also release more technical details on the underlying approach.

  • doitright99 a few seconds ago

    AI slop trained on copyrighted content.

  • thisisauserid 40 minutes ago

    - Showed close to zero examples of people.

    - Frivolous use of the term World Model.

    - Claims 20 seconds of video, shows only jumpcuts.

    Coming soon!

      sexy_seedbox 6 minutes ago

      Many video examples on /r/stablediffusion

  • luciana1u 11 minutes ago

    imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.

  • zmmmmm 33 minutes ago

    > It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.

    I'm confused, videos contain images and audio ...?

      ibotty 19 minutes ago

      That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.

      PxldLtd 16 minutes ago

      It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.

  • SubiculumCode 28 minutes ago

    Well, unified multimodal intelligence is the only way we will get to The Terminator, which seems to be the goal now of Silicon Valley and every Nation State with a military budget, so have at.

      UberFly 25 minutes ago

      I wish you were being hyperbolic but I know better.

        SubiculumCode 6 minutes ago

        I don't even know anymore, to be honest. I talk to AIs more than I do with humans, these days...So who knows

  • frotaur 34 minutes ago

    Sorry because pointing this is a bit tired by now, but reading already the first two paragraph thete is this unmistakable stench of LLM slop writing. Immediately disengaged.

  • teiferer 38 minutes ago

    Lots of words about multi-modal but then this:

    > our mission to develop real-world visual intelligence

    Visual is mono-modal, isn't it?

      nerdsniper 36 minutes ago

      Is this really the value-add comment you’re going with?

        camillomiller 34 minutes ago

        Says the one who posts this comment?

  • mattmanser 33 minutes ago

    Open-weight plans are near the bottom (Launch section):

        - Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
        - Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
        - Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
        - Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
  • vouaobrasil 26 minutes ago

    The fact that people keep developing this technology shows that the true problem is not that machines are likely to become intelligent, but that people have already become machines - unthinking and without any care to the future whatsoever.