Positives
• It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.
• It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.
• It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.
Negatives
• The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.
Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.
If AI labs get to ignore licenses, so do we.
Qwen-Image 2.1 is definitely a pretty big leap over the last open-weight version, Qwen-Image 1.0, released back in August of last year and managed to score 7 out of 15 as opposed to its predecessor which scored 4 out of 15.
Even though it's significantly smaller, 7b vs 20b, it's multimodal (so you don't need a separate image-to-image model like you did with Qwen-Edit), more coherent, and significantly faster even when outputting at higher 2K resolutions. However, in my testing, I found that I had to play with dialing up the CFG depending on the complexity of the prompt.
I've also added a progress dropdown under Model Performance so you can see how cloud vs. local models have been trending since 2024. Spoiler: June of this year released some of the biggest bangers (Krea 2, Ideogram 4, and the kind of slept-on Boogu-Image 0.1).
Downsides:
- It was clearly trained on at least some level of synthetic training data, and it shows in some of the subpar outputs in terms of fidelity. Some of this you might be able to iron out with a refiner model downstream or a custom LoRA but time will tell.
- They've moved away from the permissive Apache license. Commercial usage is only allowed by request.
Comparisons:
https://genai-showdown.specr.net
If you just want to compare local models only:
Exactly this. I've been very confused by the free software advocates that seemed to hate AI until I realized their reasons for releasing software under an open source license were very different than what I assumed they were.
We hid our “we know better” hubris under the term “disruption” because the reality (breaking everything) was a little too unsettling for us.
We ignore laws and regulations where we know better of course. Don’t you love having a homey place to stay while travelling that has just a few weird rules, a small to-do list and stays spotless thanks to that cleaning fee?
And look at all the good we did! A whole new world of slaves (oops gig workers - sorry!)
And of course we’re doing it again with information and content. We should be in charge of monetizing all your work because you’re not responsible enough to do it right. It furthers our need for power and control. Oops we meant to help make the world a better place (We keep doing that, sorry.)
Is a social contract being broken? Yeah, kind of, but the problems that that contract existed to solve are no longer a thing. Creation is now easy and commodified. We finally have computer we can interact with in natural language, something people tried to do for at least 70 years and never made any significant progress on until LLM arrived.
If you want to gatekeep or only create stuff to boost your own ego or portfolio, then AI might be an issue, if you actually want to build stuff, AI is godsend. We are essentially living in the StarTrek future with Holodecks and replicators and people still find reason to complain.
Relieving yourself of thinking and effort, whether through a machine or another human, is not the liberating force you think it is, especially when that effort isn't just grunt work, but the very process that helps you develop and grow.
The submitted title’s framing “alternative to Google Photos and Immich” is dodgy because this is a soft fork of Immich, adding certain features, described in https://opennoodle.de/noodle-gallery-vs-immich/.
I believe I've spent, perhaps, the most time of anyone on earth on digital photo management (hard to quantify, but since 2005 - 2 startups, one acquisition, and an ongoing open source project used by thousands).
I say that because I've refused to settle for most solutions in this space. Even when I adopted using Google Photos, it was as a read only viewer of my canonical photo library (I wasn't about to let Google take that responsibility). I have such high demands of whatever software I use for managing my photos and videos - they're really the only digital files I actually care about.
Immich is absolutely wonderful. It may have some shortcomings ... like partner sharing and sharing facial recognition between users. But man is it remarkable that an open source project can rival something from Google in terms of quality and experience.
And a soft fork like Noodle is precisely the way to handle it. I am perfectly happy with Immich so am not a user of Noodle. But it's open source flexing its strength.
Jev came in, and added that magic of "you dont need to train your classifier or determine the weights" if you dont want to, and just get the classified answer out. I think that's what is making people see this with a glitter in their eyes.
There are already many Jev-like models in there.
Edit: No affiliation. Just found it and thought others might find it interesting.
For emails, I get 95% accuracy with this method, with only 50-100 examples for training
Training the model takes less than 5 minutes on a CPU
The resulting model is <1MB, and inference is sub 100ms
Some other cool things about this approach:
* the model doesn’t train on some “ideal” or general classification, instead it learns your preferences
* the model runs on pretty much any mobile device and can be retrained online on the device
* privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model)
Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
Gambling regulations exist in most places in the world, including the US. With good reason.
[0] https://www.laliga.com/en-GB/news/official-statement-in-rela...
Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.
We're in the phase where the factions are geographically sorting, identifying allies, and arming. Then we fight it out. Each faction thinks their particular Overton window will win and become dominant afterwards, but likely we'll just end up back in the Stone Age and there won't be much of a society left to have a mainstream.
Also, please refer to a reply to a similar comment: https://news.ycombinator.com/item?id=49781990
A lot of what sounded extraordinary in 2013 is now baked into everyday discussion like metadata and mass surveillance. He is also available less for interviews because he is in Russia.
Basically bigger fish to fry
Snowden did ask them not to get into any direct military secrets that were in the leaks, so that could explain some of the restraint.
While I heard many of the headlines the reporting is very in depth, including a bunch of smaller details worth poking into. I'd recommend people who are interested and who haven't done so yet to just read a couple of these.
Through a combination of blackmail, personality clashes, vanity and ego, incompetence, the whole thing just fizzled out completely. A real shame.
At the time, look at the situation of Chelsea Manning, Julian Assange and the many whistle blowers before Snowden? How can u blame him?
Edit: Am unable to reply to comments. HN don't let you comment 2-3 levels in or is it moderation? - Fixed.
Making thoroughly informed decisions and iterating on a decision doc before committing to a direction and plan is better than every alternative I’ve ever observed in my career.
The criticism I’ve read thus far on this thread seems unwarranted. I give the same kind of feedback to my mentees when their work product or process could use improvement.
There are mind blowing bugs in CC that go unaddressed for months.
Something like 15-20% of all Fable messages in CC are invisible to users. You've most likely noticed this when Claude references something it said but it never said it?
It happens frequently when Fable outputs a message above a certain number of tokens just before doing a tool call.
This has been going on for months. If "users can't see messages the agent sends" isn't a critical issue that gets addressed within 24 hours, I don't really care if you admit you're often wrong, we know.
I don’t think he farms out all of his writing and talking to LLMs. I don’t think Claude code was trained to emulate him or anything.
I think his “voice” has been filed LLM smooth by years of agent based interactions. He claims Anthropic engineers use an average of 500+ agents a day. They are human interfaces to token generators more than human to human communication. They are picking up the tendencies of their most frequent communication partner.
Unfortunately, I see Claude being wrong often enough that I view it as a faulty narrator. Often helpful, sometimes totally full of it.
And now I subconsciously apply this filter to anything that sounds like Claude.
Makes me nervous that my voice may be becoming that of a faulty narrators.
A more generous read, or, at least the reading I took: "once we (think we) know what we're doing, we violently execute."
So, "boo" on you. This guy sounds awesome.
> Sometimes I will give feedback to people when they are missing steps in the framework, or are poorly executing some of the steps. I expect the same feedback in return.
I also don't like that urgency is built in as the standard process either, no wonder everyone is burnt out.
> 6. Act with urgency to achieve the goal
If everything is urgent, then nothing is. This guy sounds miserable to work for and with. Assuming this is accurate and not just hyperbole, he is essentially saying he has no prioritization skills because everything is urgent. I think most people who have been around the block have worked with people like this, and unbeknownst to them, their coworkers develop a default snooze button associated with most of their requests and projects.
Sorry I meant it's too dangerous to be released.
- Have a vague understanding of the problem
- Architect an overcomplicated solution thinking of all possible contingencies
- Pitching the overcomplicated solution to someone else
- Ask them to come up with a simple solution. Ask questions to "birth" to the solution.
- Not providing any feedback as that would mean need you to be accountable for the work
- Trying to convince them they should work out the solution because they are the expert and much smarter then you
- Taking credit for solving the problem#67071 is one of them
Or #83281 which you closed as a duplicate of #67051 which was closed as not planned.
There's a bunch more. Just ask Claude to search repo issues for messages not being shown.
Firefox, Brave and Safari do. Chrome and Edge do not.
https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the...
"Nothing is worse to the demise of a society, than people who want to convince you that the cat is out of the bag and will not go back in, while the cat is being violently shook out of the bag at the same time."
> The mechanism is standard adtech. What has no precedent is running it on an AI chat product.
As someone who has been well aware of this mechanism for quite some time, I still feel icky anytime I re-read the details of it.
What a time to be alive.
why not use your own words? If you are gonna ai generate this blog, just post the prompts instead.
I find that reading on an ereader is many times easier for me, primarily because I can make the text gigantic and I don't lose track of my place on the page nearly as often. I'm not dyslexic (I've been tested), but I do think that there's some issue with my brain where if the text is small I really can't follow it. On all my computers the text is gigantic and even then I will click and highlight text as I'm reading so don't lose place. I would love to be paid to read but I don't think I'd love to be paid to read physical books.
I remember my grandmother was wholly unconvinced that my Kindle was a different thing than an iPad or a Game Boy, and that it was not "real" reading, for whatever reason.
The weird thing is, that's not what AI models seem to be doing. The prose is just weird.
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.
I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.
Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...> The arbitrator also rejected Uber's argument that Proposition 22 -- a California ballot measure approved by voters in 2020 that allows companies to classify app-based drivers as independent contractors instead of employees -- prevented the company from being held liable for Tran's conduct.
The dream of every major tech company, making ridiculous profits while taking zero legal responsibility for what you create...
Mmmm... It seems in the world of politics, esp. populist politics, more facts don't bring better policies. They are simply ignored, or accused of being lies pushed by the ennemies.
> "We have all the facts on this we need. We don't need any more facts. In the land of truth, my friend, the man with one fact is the king."
-- Linton Barwick, In The Loop (2009)
This is only because immigrants are willing to do jobs at wages and working conditions that locals do not accept. Employers in the Finnish berry industry were just jailed for human trafficking. Meanwhile the same companies have been loudly complaining that no Finn wants to work for them.
this is insane. the "speciesist" people aside (the best way for all other species to flourish is for humans to eradicate themselves right now, which is completely mad), i have the opposite problem: the _axiom_ seems to be: we should help ME flourish. "Me" as in people who are raking in trillions for their own very special selves right now, at the cost of everyone else's future, while none of them can be trusted to hold my cell phone for a second.
> if an industry commits to the axiom of helping humanity flourish
where does he see such industries, outside of maybe nonprofits?
Unfortunately, mathematics (especially pure mathematics) is by its very nature very, very poorly understood by those who haven’t worked as a mathematician. Even worse, those who don’t understand are seemingly not at all aware of their misunderstanding and are entirely confident in their (very wrong) characterisation of the subject.
It's about many things, but perhaps the most relevant idea here is that no information matters without understanding. We could generate all possible knowledge, but unless someone--a human--can verify and understand it, it doesn't count. The cure for mortality could be written on the moon, but if no one reads it, it hasn't really been discovered.
[1] https://maskofreason.wordpress.com/wp-content/uploads/2011/0...
Yes, it was developed by Google employees, that does not imply it has the full backing of Google, or Deepmind, or GCP. Notably, the website doesn't seem to claim this either.
The same Google that pulls plugs on a whim?
I've been still just like, making VM's with proxmox, then putting my agent in the machine and letting it run free (with my dotfiles setup script making dev env pretty much free, though I could also just make a VM snapshot). What's wrong with that? Is that not the scalable solution for enterprise rn?
That said, I think Google's ADK ecosystem and this new AX platform is promising--I would expect Google to maintain this and other tooling around this for years to come.
To the Googlers out there: is Google using this at any capacity for internal projects?
> We want to make dealing with agentic infrastructure easier so you can focus on your work. AX is designed with an uncompromising focus on ergonomics, rapid iteration, and joyful workflows for both application developers and AI researchers.
On the the other hand, the readme quickstart section says
> You need a Kubernetes cluster, ko (brew install ko), a container registry your cluster can pull from, and a reachable Agent Substrate Control API (in-cluster default: api.ate-system.svc.cluster.local:443).
Call me old-fashioned but I don't find this "easier". Maybe it's easier in the same way that Kubernetes itself is easier than managing VMs and container deployments at massive scale without such a tool. But there's a vast chasm between what this tool is being sold as and what it actually is.
btw its the same google that has already killed its "gemini cli" and re-introduced it in the form of "antigravity cli"
As in ... "the AI brought it back to life, but it came back /wrong/".
It's one thing if your code is really truly "co-authored by" Claude. It's another thing if it isn't even really co-authored by you.
Evidently the majority of people see no problem with LLM prose, but me and many others think it is some of the most annoying crap possible. Please consider speaking in your own voice.
I know this gets repetitive, but it is important.
These comments are even more annoying. Sides that are winning don't have to constantly tell their opponents that they're losing. I have no idea why you and so many others seem to have attached their ego to the use of a tool that you feel the need to attack people who have complaints about the output it's often used to generate.
Re: why, I don’t think most people understand the very basics of the global economy in mechanical terms, and this was my attempt to explain those mechanics. I wanted the various pieces of the system to be motivated by understandable problems, hence the fable-like story.
I now see some folks were triggered by this stylized approach. Which is a pity. I think the simple setup is worth the payoff. My explanation of money creation, for example, matches the Bank of England’s whitepaper, and I am especially pleased with how the idea of a reserve currency both develops naturally and rhymes with earlier ideas lower in the hierarchy.
The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.
The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.
And the worst thing is that is the best possible strategy for them. It's essentially win-win for everyone but consumers.
- Buyer (AI) has crazy money, so will pay whatever
- Seller doesn't have to build anything new, as the buyer is willing to pay whatever
- Generate ridiculous profits from crazy money
- No oversupply risk in case of reversal
- Return from producing HBM to DDR5 in a single quarter if reversal does happen.
Hold on to your existing hardware, people, and be on lookout for your local deals.
"Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway."
https://en.wikipedia.org/wiki/Qwen#List_of_models
Unfortunately, it looks like this model is using a much more restrictive license:
https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE