Update, September 20, 10:43 p.m. PT: Muse's browser is back. I reported the outage through the app's own Report an issue form, and a few hours later it was fixed. The details are below.
Hello,
I want to tell you about a week I spent with Meta's new AI agent, because the lesson in it is not really about Meta. It is about how we decide what to depend on, and I think most of us are making that decision backward.
I wrote the long version of this on our site, with the leaderboard screenshots and the full benchmark tables: the complete Meta Muse review.
First, what happened
Meta shipped Muse on September 8. A personal AI agent with its own browser, built to do things on your behalf rather than talk to you about them.
I gave it the admin that eats my week. Signing into accounts. Editing profiles. Setting up recurring posts. Real work, not demos.
It was extraordinary. Faster than anything I have used, running several browser sessions in parallel without losing the thread on any of them, going much further on its own before it needed me. There was a moment early on where I sat and watched it work and thought, this changes my Tuesdays.
Then it started stopping. Tasks would begin, get most of the way, and just halt. Not fail, halt. I would come back to a session sitting idle and have to nudge it along. Then nudge it again.
Then, a few hours in, every browser it had went down at once and stayed down until that night. I tried from different machines. Same dead panel every time.
The error reporting was accurate, which is rarer than it sounds
When I asked what was happening, the explanation it returned turned out to be accurate.
It said its browser backend had been erroring since that morning, that this was infrastructure rather than my account, that chat and files were fine while browser automation was down. Asked whether it was affecting everyone, the reply was that no outage reports could be found and that it could not determine whether the problem was one session or something wider.
When I pushed, the explanation it returned got specific in a way I found oddly reassuring.
Muse does not browse from its own computer. It hands the job to Meta's system, which starts a real Chrome browser on one of their servers and runs the task there. What had broken was that handoff. Its requests were returning "failed before dispatch," which means they never reached a browser at all. The analogy it produced: calling a taxi dispatcher who never picks up. There is no car to inspect.
Failed before dispatch is not a browser crashing. It is requests never getting a browser assigned in the first place, which is what a system under more load than it planned for looks like.
It did not hallucinate an explanation. An LLM asked to account for its own failure will usually generate something plausible and wrong, because generating plausible text is what it does. Getting an accurate diagnosis instead means the system is surfacing real state into the model's context rather than leaving it to guess, and that is an engineering decision worth noticing.
There is a second thing worth telling you, because almost nobody is writing about it. Muse never sees your passwords. They sit in a separate credentials store, and a second system called Sentinel injects them at the network boundary after an action is approved. The model can propose an action. It cannot authorize one.
Which means if it is ever prompt-injected into attempting something it should not, there is nothing for it to leak. It never held the secret.
Early on it surfaced a notice that a stored sign-in had expired and a fresh one was needed, rather than failing quietly or prompting me to paste a password into the chat. In the moment that felt like friction. It is the architecture declining to route around an expired credential, and it is the best security design I have seen in this category.
And then it got fixed
Muse has a Report an issue form built into the app, so I used it. My entire report: "My browser is down, please help!"
A few hours later, that night, it was back. I asked whether it was working. Muse retried with a fresh browser, loaded Google, and replied: "Yeah, it's working. Google loaded clean and the browser's ready to go."
Whether my report or Meta's own monitoring got there first, I cannot say. But a problem reported inside the product was fixed within hours, and the product confirmed the fix itself. That counts, and I want it on the record as clearly as the outage.
Now the part that actually matters to you
Here is the thing I keep chewing on.
Meta solved the hard problem. Getting software to plan a multi-step task, drive a real browser, run several jobs at once, recognize a form, and halt for permission at the right points is genuinely difficult engineering, and they have it working at consumer scale in week one.
What broke was the easy part. Keeping servers up. And that got fixed the same day.
That is a much better position to be in than the reverse, and it is why I am not writing this off. But it exposed something about how I had been thinking, and maybe how you have been too.
When we evaluate a new tool, we watch the demo. The demo is the good day. Every tool is impressive on its good day, because that is what a demo is for.
The question that actually determines whether a tool is worth building a business on is: what happens on the bad day? Not whether it fails, everything fails. Whether it fails loudly, whether you find out immediately, and whether the work you handed over is recoverable.
Muse reported its failure accurately, which is worth a great deal. But it failed in the middle of work I had already delegated and stopped thinking about, which is the expensive kind of failure. The cost is not the outage. The cost is the hours you spent believing something was handled.
What I am actually doing about it
Using it, with a hand on the wheel.
My verdict: it is great. When it is running, Muse is the fastest thing on my desk, and I would use it for real client work, including outreach and social media management. The parallel execution is a real edge.
Use it at your own risk, though. In my experience it needs a lot of oversight: check in on long tasks, nudge the ones that stall, and review its work before it reaches a client.
For what it is worth, I did not leave this on impressions. I went and looked at two independent scoreboards directly, Artificial Analysis and Arena, both on September 20.
Arena keeps a separate leaderboard just for agents. Muse is not in the top ten of it. Claude Fable 5.1 leads at 13.71 percent, and seven of the ten slots are Claude models. That is a blind human preference test, not a vendor's own scorecard, and it matched my week almost exactly.
Muse is not a weak model. It sits fourth on Arena's text board and top ten on web development and vision. On Artificial Analysis it scores 48 on intelligence against 51 for Claude Opus 5 and 53 for Claude Fable 5.1 and GPT-6 Astra. Close enough.
Then you see the other two columns, and the whole product makes sense. Muse runs at 224 to 250 tokens per second where the Claude and GPT frontier models run 54 to 70. It costs $1.60 per task where Claude Opus 5 costs $5.86 and Claude Fable 5.1 costs $7.63.
Five points of intelligence, traded for four times the speed at a quarter of the price.
That is not a company failing to build a frontier model. That is a company deciding most people would rather have fast and cheap and safe than the last five points of clever, and honestly, for a lot of work, they are right.
One more thing, because the timing is too neat to ignore
Muse launched on September 8. By September 18 it was the number one app in the US App Store, above ChatGPT, Gemini, Claude and Instagram, with more than 730,000 downloads in ten days.
My browsers went down two days after that.
Think about what Muse hands every single user: not an API call, but a dedicated virtual machine running a live browser that keeps going after you close the app. That is a whole computer per person. A chat product that goes viral needs more inference capacity. An agent product that goes viral needs more actual machines.
I cannot prove that is what happened, and Meta has not said it. But it fits, and if it is right then what I hit was not bad engineering. It was a company being caught out by its own success at something genuinely hard.
It came back the same day, which fits that reading. Capacity is something you can add in hours. A broken design is not.
Which is the summary of the whole week. The idea is there, and when Muse is running, so is the execution.
The test I would run
If you are weighing Muse, or honestly any agent, do not run the exciting test. Run the boring one.
Give it a multi-step task. Walk away for an hour. Look at the state you come back to. Do that on three separate days.
Those three hours will tell you more than any feature list, any review, or any newsletter, including this one.
I am running it again in a month. Based on what Muse does when it is actually running, I expect to write you an even better letter then.
Until next time,
Jessica
The full review, with every screenshot and the benchmark detail, is here: Meta Muse review on miningwells.com.
If you have been using Muse and your week went differently, hit reply. I read everything, and a second data point is worth more to me than another feature announcement.
Sources: Meta newsroom · ai.meta.com/muse · Meta Help Center · Meta AI Research · Artificial Analysis LLM Leaderboard · Arena Leaderboard · TechCrunch, Muse reaches No. 2 in the US · Muse hits No. 1 with 730,000 US downloads
Benchmark figures read live from Artificial Analysis and Arena on September 20, 2026. Both leaderboards update frequently.


