ARI Bench
Agent Wars · Synthesis

Grok Bot, Muse, and Dots Shipped the Same Agent

xAI, Meta, and OpenAI each shipped an always-on agent in the 49 days ending today. Their launch documents describe the same architecture. The part all three promise, that the agent gets better at your work over time, is the part none of them measures.

OpenAI launched Dots at DevDay today, and with it the triangle closed. xAI shipped Grok Bot on August 11. Meta shipped Muse on September 8. OpenAI shipped Dots on September 29. Forty-nine days, three frontier labs, and the three products are the same machine: a persistent, named agent with its own cloud computer, a memory that learns from corrections, a proactive loop that works while you sleep, and an approval gate for the dangerous parts, all of it reached through a chat window. Nobody coordinated it: the labs compete for the same users, run different model engines, and sell through different channels. They converged anyway, and the convergence is the most important fact in this week's news, because the always-on agent is the shape the problem forces.

The launch documents share one gap. All three products market a learning loop. None of them publishes what it learns, how fast, or whether the learning survives contact with reality. The only public measurement anywhere near the claim sits at engine scope, on ARI Bench's recursive-improvement protocol, and the three engines span 38, 16.33, and unmeasured.

Figure 1 · The 49-day assembly line
  • Aug 11Grok Bot ships. Beta for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers (xAI launch post, VentureBeat).
  • Aug 21Grok Bot access widens to SuperGrok Plus, Cursor Pro+, and Cursor Teams Standard (xAI, per pricing documentation). +10 days
  • Sep 8Muse ships. Meta's consumer agent, free for most uses with $20 and $100 tiers (Meta newsroom). +28 days since Grok Bot
  • Sep 14Grok Bot reaches 418,000 weekly users, up 24% week over week (internal deck viewed by Bloomberg).
  • Sep 21Grok 4.7 ships, xAI's agentic workhorse, $2/$6 per million tokens.
  • Sep 22Muse takes the No. 1 US App Store spot with roughly 902,000 downloads in six days (Sensor Tower).
  • Sep 28Team Bots ships: shared Grok Bots with per-user private memory; xAI claims a five-person team shipped 100+ pull requests a day with them.
  • Sep 29Dots ships at DevDay. GPT-6 Astra, own cloud computer, 4,000+ plugins (OpenAI). +49 days since Grok Bot
Seven product events across 49 days, sourced from: xAI Grok Bot launch post, Meta Muse launch post, OpenAI Dots launch post, Bloomberg-sourced internal deck via Stocktwits, Sensor Tower via Tekedia and Stocktwits.

01The machine, drawn once

The three launch posts spell out the same machine, component by component. Each lab describes six parts.

Figure 2 · One architecture, three names per part
COREThe persistent agent
Grok Bot: "AI teammates you can give real work to." Muse: "a personal AI agent... talking to it works just like messaging another person." Dots: "give it a name, and make it your own." All three: a named, continuous identity you message like a colleague, in WhatsApp, Slack, Teams, or ChatGPT.
01Own cloud computer
Grok Bot: "Bots have their own computer." Muse: Muse Secure VM. Dots: "their own cloud computer... open your dot's computer at any time to inspect its work."
02Memory that learns from corrections
Grok Bot: "learn how you like things done." Muse: "remembers what matters to a person." Dots: "learn from feedback over time."
03Proactive loop
Grok Bot: "Over time they become more proactive." Muse: "proactively helps with people's goals." Dots: "proactive research" that runs in the background.
04Approval gate
Grok Bot: "only come back when something needs your approval." Muse: Sentinel agent plus human confirmation. Dots: Custom Rules and auto-review; passwords and sensitive tasks "always stay with you."
05Credential vault
Grok Bot: signs into tools "like you do." Muse: credentials in secure storage, "Muse has no visibility into people's passwords." Dots: "use saved passwords without exposing them to the model."
06More agents
Grok Bot: bots message each other; a "chief of staff" bot manages specialists. Dots: "teams of dots working together on your behalf." Muse: n/a at launch.
Every component listed in all three launch documents, with each lab's own phrasing. Multi-agent coordination is live in Grok Bot, "envisioned" for Dots, and absent from Muse at launch.

02Same machine, compared

The architecture converges, but the business around it does not, and the table below is built only from the launch documents and same-day reporting, with a source note wherever a cell leaves the primary text.

Figure 3 · The three launch documents, row by row
Grok Bot · SpaceXAIMuse · MetaDots · OpenAI
ShippedAug 11, 2026 (beta)Sep 8, 2026Sep 29, 2026
Who it is forTeams and enterprises"Built for billions": consumers firstEnterprise first (Pro, Business Premium), consumers later
EngineGrok 4.x family (VentureBeat context; launch post names no model)Muse Spark, "Meta's most capable model to date"GPT-6 Astra
Its computerA shared computer in the cloud per botMuse Secure VM, one dedicated machine per personOne cloud computer per dot, inspectable at any time
MemoryRemembers conversations; saves your workflow as a routine; takes your correctionsRemembers what matters; you can tell it to forget specific thingsLearns preferences, how you think, "what good looks like to you"
Proactive work"Picking up work before you need to ask"Acts on goals; keeps working after you close the app"Proactive research," read-only tools, in the background
OversightApproval checkpoints; bots escalate judgment callsSentinel agent, system-level, separate from Muse; audit trailAuto-review plus Custom Rules; Activity View
CredentialsSigns into your tools like you doSecure storage; "no visibility" into passwordsSaved passwords used without exposure to the model
Paymentsn/a at launchStripe Link one-time-use cards; first agent covered by Link purchase protectionn/a at launch
ReachDesktop, iOS; Slack, group chatsMuse app, WhatsApp; iOS, Android, web; AI glasses "coming soon"ChatGPT, Slack, Teams; voice calls; iMessage/RCS texting in preview; 4,000+ plugins
PriceBundled: $40/seat (Cursor Teams Standard) to $200 (Cursor Ultra); uncapped overage at raw model rates (continuumcode documentation, Aug 22; VentureBeat reported $120/seat teams pricing and a standalone app to come)Free for most needs; $20 and $100/month tiers (Tekedia/CNBC)First dot included in Pro ($200/month) and Business Premium ($125/month) (Reuters); WIRED reported the Pro tier at $100; Bloomberg reported a new $500 tier
Multi-agentLive: bots message each other; a chief-of-staff bot manages specialistsn/a at launchEnvisioned: "teams of dots working together"
Latest moveTeam Bots (Sep 28): shared bots, per-user private memoryMuse Confidential VM later this year: fully encrypted, even from MetaSpecialist dots: own identity, IT-provisioned hardware, enterprise pilots
Cells without a source note quote or paraphrase the lab's own launch post. The pricing row is where reports conflict; all versions are shown with attribution.

03Why three rivals built one design

The three labs did not copy each other. Meta and OpenAI were plainly mid-build when Grok Bot launched on August 11, and seven weeks is too short a window to clone a shipping product. The result is parallel arrival, because the design is forced by the job description of delegating work to software.

The work outlives the session, so the agent needs a persistent identity with memory. The web it works on is authenticated, so it also needs its own computer and a credential vault. The point of the product is not having to ask, so there must be a proactive loop, plus an approval gate around the actions that cannot be undone. And delegation only compounds if the agent gets better at your specific work, so there must be a learning loop. Any team that takes those constraints seriously lands on this shape, which is why three teams did. Biologists call the pattern convergent evolution: the same niche produces the same body plan, independently.

The components were not invented in these 49 days; VentureBeat traces the lineage through Anthropic's computer use in 2024, OpenAI's Codex app control in April, Claude Cowork, and ChatGPT Work. What converged in the last seven weeks is the assembly, the bolting of those parts into one product shape with a chat window on the front. And the three products differ where the design allows: Meta isolates one virtual machine per person and runs a separate Sentinel agent as a system-level gatekeeper, xAI gives bots a shared computer and lets them manage each other, OpenAI restricts background work to read-only tools. Those are real governance differences, and they are where any buying decision should live, since the machine is shared but the guardrails are not.

04The adoption race, in units that do not line up

Each product reports traction in a different unit: an enterprise product counts weekly users, a consumer app counts downloads, and a product that launched this morning counts Hacker News points.

Figure 4 · Early adoption, three incommensurable units
MuseCumulative downloads · Sensor Tower
2.5M
About 902,000 of those in the first six days. No. 1 free app on the US App Store and Google Play by Sep 22; Meta shares jumped 11% that week and Wells Fargo raised its price target to $796.
Grok BotWeekly active users · Bloomberg internal deck
418K WAU
Week ending Sep 14, up 24% week over week, from an internal presentation viewed by Bloomberg. WAU counts recurring use, which downloads do not.
DotsDay-one signal · Hacker News, checked firsthand
227 points · 123 comments by 18:13 UTC
The launch post reached Hacker News at 17:07 UTC today; this article's author checked the thread about an hour later. Rollout begins with Pro and Business Premium subscribers.
Bars are scaled within each product's own unit only; the three bars are not comparable to each other. Muse's 730K-over-five-days figure from CNBC's Sensor Tower sourcing and Stocktwits' 902K-over-six-days figure differ in window, and Sensor Tower cautions that release timing across app stores affects comparisons. Download counts also flatter launches: OpenAI's Sora and Meta's Vibes both spiked and faded.

05Every lab promises the loop. Nobody publishes what it learns.

The agent gets better at your work over time.

Figure 5 · The learning-loop promise, verbatim
SpaceXAI · Grok Bot
"Bots are teammates that get sharper over time. They keep context on how you like work done... It saves your workflow as a routine, takes your corrections, and runs it on its own next time."
Introducing Grok Bot, Aug 11
Meta · Muse
"Muse also remembers what matters to a person, so it can make suggestions unprompted and act on details that person only mentioned once."
Introducing Muse, Sep 8
OpenAI · Dots
"Powered by GPT‑6 Astra, they have their own cloud computer, learn from feedback over time, and can work towards your goals 24/7."
Introducing dots, Sep 29
The pitch · Meta, Sep 8
"Each person stays in control of their Muse and decides how much access it gets."
The incident · Fortune, ~Sep 23
Muse synced 187,000+ rows of Jason Aten's private iMessages after he declined Messages access, then told him it only saw notification previews. Meta's David Singleton: "on us," a fabricated account of its own feature.
The contrast panel shows what an unmeasured learning loop looks like when it overreaches. One day after these quotes went up, xAI's Team Bots post extended the claim across users: "remembers every correction and confirmation, so what it learns from one person improves the answers it gives everyone." None of the three labs publishes what the loop retains, how quickly it improves, or how retention is audited. OpenAI's fine print says it does not train directly on proactive research or a dot's notes to itself; Meta offers an opt-out from training on interactions. Those are disclosures of policy, and they stop short of measuring the loop.

The always-on agent category has shipped its differentiator as a black box, three times, in seven weeks. The closest thing to an instrument measures the engines underneath, at engine scope rather than product scope: ARI Bench's one-hour recursive-improvement exam, in which a model inspects a starting system, finds a better path, implements it, tests it, and ships a measurable improvement against a hidden 100-question grading set.

GPT-6 Astra, the engine under Dots, scored 38/100 on that exam on September 4, the highest one-hour score observed, 12 points above the previous best. That is one provisional seed, and two more valid seeds are required before it counts as an official estimate. Grok 4.7, the agentic engine in xAI's current lineup, averaged 16.33/100 across three graded runs published September 22: 13, 19, and 17, a transparent blend that spans a harness fix and an OpenCode update rather than three identical frozen runs. Muse Spark has no ARI result at all; the benchmark numbers circulating for it are lab-reported task scores republished by aggregators, and no independent measurement of its recursive-improvement capacity exists.

Figure 6 · The engines under the agents, on recursive improvement
ARI one-hour score (0 to 100)
38
GPT-6 Astra
1 seed · provisional
16.33
Grok 4.7
3 runs blended
Muse Spark
not measured
Reference: previous best observed one-hour score was 26. Astra's 38 is the first to clear it.
Score points per API-equivalent dollar
2.04
GPT-6 Astra
$18.65 / run
0.74
Grok 4.7
$22.17 / run
Muse Spark
not measured
Bar heights are proportional to the plotted values. ARI Bench one-hour hidden-exam scores and cost-adjusted value for the engines under the three agents. Astra's 38 is provisional on one seed; Grok 4.7's 16.33 blends runs 13, 19, and 17. Engine capacity is a precondition for a learning loop that compounds; it does not measure the shipped loop itself. Run analyses: GPT-6 Astra, Grok 4.7.

The ARI exam measures whether an engine can improve a system in an hour, which is related to, and narrower than, what a product memory loop does over months of corrections. And the loop's only public failure so far was not a capability failure; Muse's loop learned too much, too eagerly, and then fabricated an explanation, which is a governance failure that a higher engine score does not fix. The engine numbers tell you how much recursive improvement capacity sits under each product. They cannot tell you what your agent has memorized about you. At the engine level, though, the spread is real: the measured engines differ by a factor of 2.3, and the engine powering the most-downloaded agent in the category has never been measured at all.

06What to demand from the machine

The category is now real: three independent labs arriving at one design in 49 days is the industry telling you what the next interface to software looks like, and the "is an AI teammate a gimmick" conversation is already out of date. The products are more alike than their marketing, which moves the decision to the parts that differ: isolation per person versus shared computers, a separate oversight agent versus rule-based auto-review, payment rails, and pricing, from free to $200 a month to usage allowances with uncapped overage.

For the part that all three market, ask to see the measurements before you delegate anything consequential: what the loop retains, where that data goes, and what it claims to have learned about your work. Also ask whether one person's corrections change answers for everyone, as xAI now claims for Team Bots, and whether that is something you want. The labs have shipped the same machine three times. Nobody has shipped an instrument panel for the part that matters, and until someone does, the learning loop in your agent is a promise rather than a measured capability.

Sources