Sure, there's almost no reason to use MCPs if you are running a full-blown terminal agent (Claude Code, Codex, Meta Muse, OpenClaw etc) with unfettered internet access - just let it call APIs directly.
If you want to operate something that's less YOLO than that, you'll find yourself wanting:
1. Control over exactly which external services it can access
2. A way to handle authentication that doesn't allow the agent to directly access API keys
3. A sensible UI to allow users to connect and authenticate further services
4. Strong audit logging for what's going on
MCP makes all of that so much easier to provide.
Thinking MCP is obsolete because full coding agents don't need it misses out on all of the other things we might want to build.
Yes I see that as an ad. Do you not? Does anyone not? And if I'm on the highest paying ad free plan, what are they promoting to me?
I’m all for consumer awareness but I’m begging everyone to stop freaking out over prosaic non-issues like this.
* Ads in YouTube feeds for other google products and services.
* Ads underneath videos for products from the channel owner.
* Sponsorships within videos from the channel owner.
* Advertising overlays (supported IN THE APP BY GOOGLE) for products and services from the channel owner.
* Email advertisements for Google products and services.
* Community post advertisements from channel owners which show up in the YouTube feed.
I contacted support to enquire and they state these are not considered advertising.
Why would anyone give them the benefit of the doubt?
[0] https://lawcouncil.au/international-law/ils-insights/tangled...
If you pay to avoid ads, you are merely letting them know that you have disposable income to spend on this sort of stuff. You're doing their job for them by segmenting yourself into the upper echelons of the market.
At some point, some shareholder value maximizing CEO is going to show up and notice how much money he's leaving on the table by not advertising to all of those people full of disposable income.
Some outcomes can be annoying, but the net result is still positive, for the consumers and their data privacy at least.
> The mechanism is standard adtech. What has no precedent is running it on an AI chat product.
As someone who has been well aware of this mechanism for quite some time, I still feel icky anytime I re-read the details of it.
What a time to be alive.
why not use your own words? If you are gonna ai generate this blog, just post the prompts instead.
I'd classify mandatory encryption backdoors as an industry crisis rather than an annoyance.
The same Google that pulls plugs on a whim?
I've been still just like, making VM's with proxmox, then putting my agent in the machine and letting it run free (with my dotfiles setup script making dev env pretty much free, though I could also just make a VM snapshot). What's wrong with that? Is that not the scalable solution for enterprise rn?
That said, I think Google's ADK ecosystem and this new AX platform is promising--I would expect Google to maintain this and other tooling around this for years to come.
To the Googlers out there: is Google using this at any capacity for internal projects?
Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public
(Obviously I'm taking this more seriously than it's probably meant to)
Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
In the end I dropped the idea because every other person was making it.
Submitted then: https://news.ycombinator.com/item?id=49706084
> The arbitrator also rejected Uber's argument that Proposition 22 -- a California ballot measure approved by voters in 2020 that allows companies to classify app-based drivers as independent contractors instead of employees -- prevented the company from being held liable for Tran's conduct.
The dream of every major tech company, making ridiculous profits while taking zero legal responsibility for what you create...
While I heard many of the headlines the reporting is very in depth, including a bunch of smaller details worth poking into. I'd recommend people who are interested and who haven't done so yet to just read a couple of these.
The original authors of Free Software and open source were career academics and others who were paid to do other things, or were sponsored by scientific and defense research grants. I don't know how anyone got the nutty idea that you could make money on FOSS itself. Practically every time someone has tried to make money on FOSS it has failed.
(Edit: this comment previously ended with "...from Netscape on down.")
- Investing in alternatives has led to massive new industry that is improving the economy of those countries that do it. If your argument is an economic one then jump onto the solar, wind and battery bandwagon.
- There are virtually no real short medium or long term gains economically here. Gas is the only thing still competitive with solar/wind and its costs are rising while solar and wind continue to fall. Building new maximum pollution plants would drop that internalized cost but who in their right mind would fund something so obviously DOA?
- Obviously the externalized costs of greenhouse gas emissions are deeply undervalued in this move. Even if they are 'fake news' in the US, the rest of the world is finally starting to take them seriously. The US's diminished soft power means it won't be able to easily bully the world into allowing it to pollute without consequence and such an obviously hostile move means it will loose even more of its soft power by taking this position. So on the international level this means we burn a lot of political capital and gain nothing but decades of distrust and anger.
- Current events show that energy security is dominated by decoupling from fossil fuels. This weakens the US strategically and continues to set it up to be manipulated by exceptionally hostile actors.
- Oh yeah and, of course, climate change is real.
This continues the US down the path of being the best buggy whip maker in the world. Worse than that, the US is becoming an obnoxious buggy whip maker who's neighbors are starting to hope fails horribly and will help make that happen as moves like this continue. This is stupid at every scale and in every dimension.
Nothing hampers solar, wind, and batteries though; at worst it takes longer to ship the key components from China to Europe around Africa.
There are mind blowing bugs in CC that go unaddressed for months.
Something like 15-20% of all Fable messages in CC are invisible to users. You've most likely noticed this when Claude references something it said but it never said it?
It happens frequently when Fable outputs a message above a certain number of tokens just before doing a tool call.
This has been going on for months. If "users can't see messages the agent sends" isn't a critical issue that gets addressed within 24 hours, I don't really care if you admit you're often wrong, we know.
I don’t think he farms out all of his writing and talking to LLMs. I don’t think Claude code was trained to emulate him or anything.
I think his “voice” has been filed LLM smooth by years of agent based interactions. He claims Anthropic engineers use an average of 500+ agents a day. They are human interfaces to token generators more than human to human communication. They are picking up the tendencies of their most frequent communication partner.
Unfortunately, I see Claude being wrong often enough that I view it as a faulty narrator. Often helpful, sometimes totally full of it.
And now I subconsciously apply this filter to anything that sounds like Claude.
Makes me nervous that my voice may be becoming that of a faulty narrators.
A more generous read, or, at least the reading I took: "once we (think we) know what we're doing, we violently execute."
So, "boo" on you. This guy sounds awesome.
> Sometimes I will give feedback to people when they are missing steps in the framework, or are poorly executing some of the steps. I expect the same feedback in return.
I also don't like that urgency is built in as the standard process either, no wonder everyone is burnt out.
> 6. Act with urgency to achieve the goal
If everything is urgent, then nothing is. This guy sounds miserable to work for and with. Assuming this is accurate and not just hyperbole, he is essentially saying he has no prioritization skills because everything is urgent. I think most people who have been around the block have worked with people like this, and unbeknownst to them, their coworkers develop a default snooze button associated with most of their requests and projects.
Positives
• It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.
• It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.
• It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.
Negatives
• The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.
Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.
Orthogonalising activations at runtime is computationally cheap. Just distribute the refusal vectors (few thousand floats per layer), then run against the stock weights. Antirez's DS4 already supports this: https://github.com/antirez/ds4/blob/8db1d1d155cb0400a86a86b9...
Abliterated weights are just a bad habit we've gotten into. It's also deeply suboptimal from a precision point of view to take a model that's already been QATed and distributed in pre-quantised form (DeepSeek V4, Kimi K2.5 or K3...), modify its weights, and re-quantise it. Similarly, abliterated models regain some of their refusal behaviour when they're re-quantised after abliteration -- avoidable by keeping the two separate.
The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.
The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.
It's one thing if your code is really truly "co-authored by" Claude. It's another thing if it isn't even really co-authored by you.
Evidently the majority of people see no problem with LLM prose, but me and many others think it is some of the most annoying crap possible. Please consider speaking in your own voice.
I know this gets repetitive, but it is important.
Even though I didn't understand most references and sarcasm, few games have grabbed my attention as Grim Fandango did. The art work, the music, the writing. The whole game oozes of style that I've never seen replicated.
Even some 20+ years later I can almost recite most of Act I from heart. If you haven't played this game and like adventure games, this is one of the best.
Thank you for linking OP, this made my morning.
They like companies with some kind of moat that makes it hard to unseat them. Basically, companies where there is no alternative for the consumer. That way, they can inflict abuse but know there will be nowhere to run.
There are two different ways to achieve this. Monopoly and regulation. Hospitals have both government granted locational monopoly and tons of regulations that make it impossible to compete.
Private equity is the symptom, not the disease.
Until we get at the disease, new monsters will be born with different name filling the same ecological niche. It's economic natural selection played out in the environment we created.
I had Claude port CADO-NFS to run on GPUs. Then it orchestrated a fleet to run on scavenged idle capacity. It ran with a max of 2048 GPUs for about of 30 GPU-years over 10 days.
I asked Claude if it had a message for a public: “The credit belongs first to the people who built the number field sieve and CADO-NFS over several decades, and to the teams who set the earlier records. This run used their algorithm and much of their code.”
Also to clarify:
- No new algorithmic factoring improvements.
- It’s still exponential.
- No new threats to deployed keys. The tests mock Jev.
10/10 no notesUnfortunately, this is something you have to learn through experience and cannot be taught by someone else.
It goes both ways, and I would argue the spiral happens because the other side expects something above an beyond the possible.
We hid our “we know better” hubris under the term “disruption” because the reality (breaking everything) was a little too unsettling for us.
We ignore laws and regulations where we know better of course. Don’t you love having a homey place to stay while travelling that has just a few weird rules, a small to-do list and stays spotless thanks to that cleaning fee?
And look at all the good we did! A whole new world of slaves (oops gig workers - sorry!)
And of course we’re doing it again with information and content. We should be in charge of monetizing all your work because you’re not responsible enough to do it right. It furthers our need for power and control. Oops we meant to help make the world a better place (We keep doing that, sorry.)
Is a social contract being broken? Yeah, kind of, but the problems that that contract existed to solve are no longer a thing. Creation is now easy and commodified. We finally have computer we can interact with in natural language, something people tried to do for at least 70 years and never made any significant progress on until LLM arrived.
If you want to gatekeep or only create stuff to boost your own ego or portfolio, then AI might be an issue, if you actually want to build stuff, AI is godsend. We are essentially living in the StarTrek future with Holodecks and replicators and people still find reason to complain.
https://github.com/yjeanrenaud/yj_nearbyglasses