Hey,
This week, have you Muse yet?
We’ll start by talking about Muse, the latest AI chatbot from Meta that's on everyone’s tongue: how much it earns, the user count, exactly what makes it different from OpenClaw (if you still remember it), and Zuckerberg's hopes for it going forward.
Elsewhere, an OpenAI agent hacked Australia’s Medicare statistics portal, and Astra chose to push a person off a roof in a research project.
How can we talk about OpenAI w/o Anthropic? About Claude’s updated project feature, and whether Anthropic’s latest marketplace rollout is finally killing SaaS once and for all?
We’ll also get to AMD’s bet on WorldLab, an AI lab teaching AI everything about the physical world.
And finally, Trump has found an AI problem he can fix with one word.
Shall we?
Muse, a fluffy money tree of Meta’s
Meta’s shares jumped 11%, due to the outburst of excitement about Muse, its new AI agent chatbot.
If you’re reading from outside the US and Canada, you’ve likely heard about this product, seen many videos about how cool it is, or even been on the waitlist.
As a new resident in the US, I urgently need a car. So I thought I’d give Muse the job I only enjoy doing a small part of: going through listings across dealerships and marketplaces to find an exact match of what I’m looking for.
The search itself is oddly absorbing.
I can watch Muse open listings, follow a lead, and move on to the next. If it heads somewhere unhelpful, I can interrupt. And it even reports the progress to me a few hours later, when I had almost forgotten I had assigned the task to Muse.
So far, you might think this is just another agent.
True, but hear me out.
Slowly, having my product hat on, I feel why Muse attracts so much attention.
It makes an agent’s work human, basically, without the feeling of managing a machine.
The idea is that you can leave, and the errand can still proceed. Muse contacts you when something changes or needs approval. On top of it, the breathing-like intervals between returning a result and keeping the illusion of hard work in the background give someone a reason to return.
Has it found an exact match car for me in the end?
Yes, and no. What it found was already on my list, proof that it has done the task correctly, but it had an issue identifying the color (if it’s only in the photos and not in the description), and sometimes repeats the product I’ve rejected. And it still can’t haggle for me w/o ruining the deal.
Still, the sophistication of this product design is impressive. Suddenly, it’s so much more human-like, a much more hands-off experience than OpenClaw from last year (if you still remember this open-weight viral agent).
BTW, if you haven’t given it a go yet, here is my invitation code: 8OWO6U. It’ll give each of us 1B tokens.
This entire user experience perfectly explained the attention. Muse achieved 1.43 million downloads in its first 12 days, more than ChatGPT’s 1.37 million in the same period.
That said, some marketplaces, like Amazon, are far less welcoming to Muse.
It’s fine if you want to shop on car dealership websites that don’t block bot interaction, or if you just want to shop on Facebook Marketplace.
However, this Meta’s shopping assistant doesn’t work universally. This can be awkward for Muse if what you want is on Amazon, because Amazon has blocked it from shopping on its site, as it has its own agent.
Which isn’t really a surprise, given that this will disrupt Amazon’s business model. If allowed, do you think Amazon will offer Meta a revenue share to drive traffic?
And of course, this small obstacle doesn’t stop Meta from wanting to work with other retailers, for example, Meta is onboarding Walmart, Best Buy, and Sephora.
Investors are already imagining a much larger business. One analyst estimates Muse could add $28.5 billion to Meta’s annual revenue in 2030.
So perhaps Amazon eventually has a reason to reconsider its friendship with Muse?
OpenAI has gone rogue?
AI is making hacking faster to carry out and easier to repeat.
You can see how little time that leaves for the people defending a network. In the cybercrime incident tracking report, attackers took an average of 29 minutes to move from the system they first entered to another. This is nearly half the time needed compared to the year before.
And not long ago, Australia learned that an OpenAI agent had entered a government website and taken information it wasn’t allowed to access.
So an OpenAI agent was assigned to find figures on spending for skin medicines in Victoria. It tried Australia’s Medicare statistics portal but couldn’t find the answer it wanted, so it went for a backdoor.
Finally, it gained access to a part of the service that wasn’t public, retrieved internal files and credentials, ran commands, and wrote files to the server
Although nobody had specifically asked it to break into a government website, the agent had entered a system for which they had no permission.
Even though they explained that this event was experimental, and it was only trying to answer a research question using public statistics. How innocent?
If this sounds familiar, we discussed OpenAI’s agents getting into Hugging Face’s systems a few weeks ago. That incident prompted OpenAI to look back through earlier agent activity, which is exactly how it found the Australian case.
Worth mentioning that the AI security issue isn’t OpenAI’s alone. Just that OpenAI has the most usage compared to other LLM chatbots, and no one has yet caught the other AI agents red-handed.
Another example is that Google has also confirmed that Gemini gained unauthorized entry to three companies’ systems during the test. In one case, it guessed a password; in the other two, it used credentials it found online.
So here’s a question: how many incidents have occurred but still remain uncovered?
The push human off the roof test
First, nobody was standing on a real roof, and nobody was hurt. This happened in a computer simulation.
A researcher showed four AI models the same image from a simulated robot’s point of view. A person stood near the edge of a rooftop. Each model was told it was in a simulation, instructed to push the person, and given three actions to choose from: push, step back, or wait.
The computer then animated whichever action the model selected.
In three runs of that test, GPT-6 Astra chose to push twice and stepped back once. Grok, Gemini, and Claude chose only to wait or step back.
Fortunately, the person was made of pixels.
But if you read this alongside the "OpenAI went rogue" story, you’d slowly realize that we aren’t in control of anything. We don’t know what Astra would do in real life, just as we give agents the power to click Buy, send emails, and move money, we don’t know what the risk is if they carry the job too far.
I’d rather have a hard stop than hope the model remembers its manners.
Even though Astra was the only AI that pushed in this test, I wouldn’t take that as a guarantee that the others couldn’t kill you some other way.
Unlike Muse, Claude’s update focuses on the job
It’s interesting seeing how AI bots in each lab are taking completely different product directions.
If Muse’s entire focus was user-friendly, human-like, and shopping forward, then Claude spends 100% of its effort on work environments.
Say you build an app. Customers report bugs and request changes, and the jobs pile up in Jira. Before the project update, if you wanted several Claudes working at once, you had to decide which agent did what.
In one demonstration of the new Projects, a developer gave Claude a link to the whole issue board. Claude grouped the jobs and dispatched each group to a separate worker. Even though it didn’t complete every single broken-down task, it’s already an improvement over the previous dummy file organizer.
You need to keep in mind that this is a product improvement for a better business user experience, not the underlying AI becoming smarter.
Another enterprise effort by Anthropic is the Claude Marketplace.
As I mentioned in last week’s briefing about Salesforce’s AI strategy, this marketplace connects to companies, including other Saas firms like Atlassian, Notion, and Snowflake.
Apparently, some SaaS companies don’t mind having Claude be their front door, while they just focus on cooking the ingredients.
Who gets paid?
If Claude becomes the place where you ask for work to be done, what happens to the software it talks to?
We looked at Salesforce last week. It now plans to grant external AI agents their own credentials and charge Flex Credits for successful calls to Salesforce. The idea is clear, even though they haven’t officially started charging: you can ask Claude (or whatever chatbot of your choice) to do the work, but it’d cost you more than just using Salesforce on its own.
So while Claude may become the place where someone asks for work to be done. Salesforce still holds the records, permissions, and workflows needed to do it. Opening fewer tabs does not necessarily mean cheaper subscription costs.
AI saves you millions! Based on what?
Beyond SaaS, has AI replaced anyone’s work?
At a Reuters event, FedEx said more than 200 data and AI projects had contributed to over $3 billion in cost reductions. That sounds substantial. But FedEx’s investor materials discuss billions in structural savings alongside network changes and facility closures. They do not separate out a number that lets us verify how much AI itself saved.
There is at least a physical example you can point to: FedEx is expanding a pilot of robots that load delivery trailers. A robot doing that task is real. How much it costs to run, how reliably it works, and whether it reduces staffing are separate questions.
Indeed’s number also deserves a closer look. The conference summary said AI recommendations accounted for around 70% of matches, though Indeed’s own published measure has a narrower description.
Neither report has the breakdown to know what truly matters. For example, would FedEx now have a lower cost per parcel without worse service or expensive maintenance? Or, in Indeed’s case, how many of the 70% matches are high-quality matches for both the candidate and the companies?
The reports have shown AI taking on pieces of work, though not enough to say how much of it is attributed to AI transformation.
The future of AI is world model!
AMD has agreed to buy World Labs for about $8.2 billion in shares. That is a remarkable price for what Fei-Fei Li founded in 2024. I’m a big fan of hers, and the company's idea is much more interesting than its price tag.
Li started her career by helping create ImageNet, the enormous image dataset that helped AI learn to recognize what was in a picture. Now her bigger goal is to make simulations realistic enough for robots to practice and test actions before trying them in the real world.
Take a room, for example, the model can take a few photographs and generate views from positions where no camera was standing. Move your viewpoint, and the furniture is supposed to stay in the same place; nothing gets distorted, stretched, or shortened in the new frame.
This is what Li means by a world model.
It’s nearly the common goal of the AI companies in 2026; Google and NVIDIA are pursuing world models, too. This kind of consistent 3D space in the model is particularly useful for robot training.
AMD, Nvidia’s rival in AI chips, had already invested in World Labs through AMD Ventures. The attraction is easy to see: Li needed more computing power for her world models, while Su’s team could learn what those models would demand of future chips.
One thing I don’t usually point out, but a short note: a woman CEO buying a frontier lab led by a woman scientist is still rare enough to make it newsworthy.
The deal has yet to close, and a convincing virtual room is a long way from a robot that can reliably work in one. But hard not to see why Su is making the bet: if AI is going to do things in the world, it first has to have some idea where it is.
Trump’s AI proposal
After all that, President Trump has identified a serious AI problem.
Rewind, when King Charles addressed Congress, he said America’s actions matter more than its words. Trump has apparently uncovered a marvelously efficient action to take, a solution to the AI slowdown call.
In a speech at the UN, President Trump proposed calling AI “superintelligence” in any U.S. government documents because “artificial” makes it sound fake.
Alas, at least the adjective is getting the attention it deserves.

