Rendered at 08:17:04 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
rekpero 5 minutes ago [-]
I have a feeling open-weight models ought to be outperforming proprietary ones by now, but that still hasn’t happened. So far, Nano Banana and GPT-2 Image seem to be the best in class, and Flux still isn’t crossing that quality bar.
user43928 1 hours ago [-]
I hope the open-weight versions will be SOTA.
> Over the next few weeks and months, we will make the following capabilities available
> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
> We will also release more technical details on the underlying approach.
3form 27 minutes ago [-]
Open-weight promise seems nice, but 1) I've seen people commenting that these hopes amounted to nothing for some previous releases (no idea which promises were made though); and 2) if the backbone is released as Dev, what will be missing? I can't easily tell from the post.
musebox35 17 minutes ago [-]
dev variants are usually cfg distilled which means that directly finetuning isn’t as effective. In the past, for the flux2 klein models,they released base versions that are not distilled. So it will probably be a while before you can fully take advantage of the open weights.
iLoveOncall 18 minutes ago [-]
I run some of the 2.3 models locally so I'm not sure where you saw that they didn't follow their words?
thisisauserid 1 hours ago [-]
- Showed close to zero examples of people.
- Frivolous use of the term World Model.
- Claims 20 seconds of video, shows only jumpcuts.
Coming soon!
vitorgrs 49 seconds ago [-]
Yeah, very weird launch. I was like... where's the videos?
sexy_seedbox 45 minutes ago [-]
Many video examples on /r/stablediffusion
bobthebob 26 minutes ago [-]
And they are stunning
jdthedisciple 40 minutes ago [-]
It's incredible how negative and dismissive the comments in here are while here I am thinking the model actually looks impressively capable.
But then again I heard the downers have always been the first to leave their dung comments here so let's see...
muppetman 37 minutes ago [-]
Advertisers love these sorts of reactions.
I see nothing here but them TELLING us how great it is. Not showing us.
jdthedisciple 30 minutes ago [-]
> Not showing us
Have you watched the 46s video full screen on a monitor and not marvelled at the incredible 4K detail of the FPV motorcycle racing clip?
ra 37 minutes ago [-]
It would be interesting to see some time-series sentiment analysis of HN. Subjectively it feels like there's much negativity on HN these days, I hope I'm wrong.
drsalt 31 minutes ago [-]
none of this is exciting compared to holodeck in star trek
vouaobrasil 9 minutes ago [-]
It's possible it's because some people care about the consequences of what we're doing as a society beyond the mindless satiation of curiosity. Nothing wrong with curiosity in my books by the way, but I do think isolating it the way we've done in the technical fields is a dangerous and irresponsible attitude.
saejox 16 minutes ago [-]
i don't get why they are investing money on image/video gen. All generations i have see looked blurry, lacking in fine details and missing the artistic touch (lacks meaning? lifeless?)
Gecko4072 34 minutes ago [-]
I thought the clips were real footage until they were dancing in a flooded room.
zmmmmm 1 hours ago [-]
> It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation.
I'm confused, videos contain images and audio ...?
ibotty 57 minutes ago [-]
That's most likely a disagreement on terms. In the media world, video is only the moving images, not audio. This is separate from images, that are meant to be still images.
PxldLtd 55 minutes ago [-]
It's more a comment about the feature detection I think; all image, video and audio input contribute to the same weights/activations that can produce image, video and audio output.
teiferer 1 hours ago [-]
Lots of words about multi-modal but then this:
> our mission to develop real-world visual intelligence
Visual is mono-modal, isn't it?
cpldcpu 31 seconds ago [-]
its doing video, audio, images and motion. I think that counts as multimodal.
nerdsniper 1 hours ago [-]
Is this really the value-add comment you’re going with?
camillomiller 1 hours ago [-]
Says the one who posts this comment?
SubiculumCode 1 hours ago [-]
Well, unified multimodal intelligence is the only way we will get to The Terminator, which seems to be the goal now of Silicon Valley and every Nation State with a military budget, so have at.
UberFly 1 hours ago [-]
I wish you were being hyperbolic but I know better.
SubiculumCode 45 minutes ago [-]
I don't even know anymore, to be honest. I talk to AIs more than I do with humans, these days...So who knows
luciana1u 50 minutes ago [-]
imagine spending nine figures training a model to learn that the sound has to match the impact. my 8-month-old figured that out by dropping a spoon on the floor twice.
frotaur 1 hours ago [-]
Sorry because pointing this is a bit tired by now, but reading already the first two paragraph thete is this unmistakable stench of LLM slop writing. Immediately disengaged.
cobolexpert 55 minutes ago [-]
Thanks for the heads-up
camillomiller 1 hours ago [-]
Correct reaction
mattmanser 1 hours ago [-]
Open-weight plans are near the bottom (Launch section):
- Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
- Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
- Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
- Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
vouaobrasil 1 hours ago [-]
The fact that people keep developing this technology shows that the true problem is not that machines are likely to become intelligent, but that people have already become machines - unthinking and without any care to the future whatsoever.
> Over the next few weeks and months, we will make the following capabilities available
> Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)
> We will also release more technical details on the underlying approach.
- Frivolous use of the term World Model.
- Claims 20 seconds of video, shows only jumpcuts.
Coming soon!
But then again I heard the downers have always been the first to leave their dung comments here so let's see...
I see nothing here but them TELLING us how great it is. Not showing us.
Have you watched the 46s video full screen on a monitor and not marvelled at the incredible 4K detail of the FPV motorcycle racing clip?
I'm confused, videos contain images and audio ...?
> our mission to develop real-world visual intelligence
Visual is mono-modal, isn't it?