Grok Bot, Muse, and Dots Shipped the Same Agent
xAI, Meta, and OpenAI each shipped an always-on agent in the 49 days ending today. Their launch documents describe the same architecture. The part all three promise, that the agent gets better at your work over time, is the part none of them measures.
OpenAI launched Dots at DevDay today, and with it the triangle closed. xAI shipped Grok Bot on August 11. Meta shipped Muse on September 8. OpenAI shipped Dots on September 29. Forty-nine days, three frontier labs, and the three products are the same machine: a persistent, named agent with its own cloud computer, a memory that learns from corrections, a proactive loop that works while you sleep, and an approval gate for the dangerous parts, all of it reached through a chat window. Nobody coordinated it: the labs compete for the same users, run different model engines, and sell through different channels. They converged anyway, and the convergence is the most important fact in this week's news, because the always-on agent is the shape the problem forces.
The launch documents share one gap. All three products market a learning loop. None of them publishes what it learns, how fast, or whether the learning survives contact with reality. The only public measurement anywhere near the claim sits at engine scope, on ARI Bench's recursive-improvement protocol, and the three engines span 38, 16.33, and unmeasured.
- Aug 11Grok Bot ships. Beta for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers (xAI launch post, VentureBeat).
- Aug 21Grok Bot access widens to SuperGrok Plus, Cursor Pro+, and Cursor Teams Standard (xAI, per pricing documentation). +10 days
- Sep 8Muse ships. Meta's consumer agent, free for most uses with $20 and $100 tiers (Meta newsroom). +28 days since Grok Bot
- Sep 14Grok Bot reaches 418,000 weekly users, up 24% week over week (internal deck viewed by Bloomberg).
- Sep 21Grok 4.7 ships, xAI's agentic workhorse, $2/$6 per million tokens.
- Sep 22Muse takes the No. 1 US App Store spot with roughly 902,000 downloads in six days (Sensor Tower).
- Sep 28Team Bots ships: shared Grok Bots with per-user private memory; xAI claims a five-person team shipped 100+ pull requests a day with them.
- Sep 29Dots ships at DevDay. GPT-6 Astra, own cloud computer, 4,000+ plugins (OpenAI). +49 days since Grok Bot
01The machine, drawn once
The three launch posts spell out the same machine, component by component. Each lab describes six parts.
02Same machine, compared
The architecture converges, but the business around it does not, and the table below is built only from the launch documents and same-day reporting, with a source note wherever a cell leaves the primary text.
| Grok Bot · SpaceXAI | Muse · Meta | Dots · OpenAI | |
|---|---|---|---|
| Shipped | Aug 11, 2026 (beta) | Sep 8, 2026 | Sep 29, 2026 |
| Who it is for | Teams and enterprises | "Built for billions": consumers first | Enterprise first (Pro, Business Premium), consumers later |
| Engine | Grok 4.x family (VentureBeat context; launch post names no model) | Muse Spark, "Meta's most capable model to date" | GPT-6 Astra |
| Its computer | A shared computer in the cloud per bot | Muse Secure VM, one dedicated machine per person | One cloud computer per dot, inspectable at any time |
| Memory | Remembers conversations; saves your workflow as a routine; takes your corrections | Remembers what matters; you can tell it to forget specific things | Learns preferences, how you think, "what good looks like to you" |
| Proactive work | "Picking up work before you need to ask" | Acts on goals; keeps working after you close the app | "Proactive research," read-only tools, in the background |
| Oversight | Approval checkpoints; bots escalate judgment calls | Sentinel agent, system-level, separate from Muse; audit trail | Auto-review plus Custom Rules; Activity View |
| Credentials | Signs into your tools like you do | Secure storage; "no visibility" into passwords | Saved passwords used without exposure to the model |
| Payments | n/a at launch | Stripe Link one-time-use cards; first agent covered by Link purchase protection | n/a at launch |
| Reach | Desktop, iOS; Slack, group chats | Muse app, WhatsApp; iOS, Android, web; AI glasses "coming soon" | ChatGPT, Slack, Teams; voice calls; iMessage/RCS texting in preview; 4,000+ plugins |
| Price | Bundled: $40/seat (Cursor Teams Standard) to $200 (Cursor Ultra); uncapped overage at raw model rates (continuumcode documentation, Aug 22; VentureBeat reported $120/seat teams pricing and a standalone app to come) | Free for most needs; $20 and $100/month tiers (Tekedia/CNBC) | First dot included in Pro ($200/month) and Business Premium ($125/month) (Reuters); WIRED reported the Pro tier at $100; Bloomberg reported a new $500 tier |
| Multi-agent | Live: bots message each other; a chief-of-staff bot manages specialists | n/a at launch | Envisioned: "teams of dots working together" |
| Latest move | Team Bots (Sep 28): shared bots, per-user private memory | Muse Confidential VM later this year: fully encrypted, even from Meta | Specialist dots: own identity, IT-provisioned hardware, enterprise pilots |
03Why three rivals built one design
The three labs did not copy each other. Meta and OpenAI were plainly mid-build when Grok Bot launched on August 11, and seven weeks is too short a window to clone a shipping product. The result is parallel arrival, because the design is forced by the job description of delegating work to software.
The work outlives the session, so the agent needs a persistent identity with memory. The web it works on is authenticated, so it also needs its own computer and a credential vault. The point of the product is not having to ask, so there must be a proactive loop, plus an approval gate around the actions that cannot be undone. And delegation only compounds if the agent gets better at your specific work, so there must be a learning loop. Any team that takes those constraints seriously lands on this shape, which is why three teams did. Biologists call the pattern convergent evolution: the same niche produces the same body plan, independently.
The components were not invented in these 49 days; VentureBeat traces the lineage through Anthropic's computer use in 2024, OpenAI's Codex app control in April, Claude Cowork, and ChatGPT Work. What converged in the last seven weeks is the assembly, the bolting of those parts into one product shape with a chat window on the front. And the three products differ where the design allows: Meta isolates one virtual machine per person and runs a separate Sentinel agent as a system-level gatekeeper, xAI gives bots a shared computer and lets them manage each other, OpenAI restricts background work to read-only tools. Those are real governance differences, and they are where any buying decision should live, since the machine is shared but the guardrails are not.
04The adoption race, in units that do not line up
Each product reports traction in a different unit: an enterprise product counts weekly users, a consumer app counts downloads, and a product that launched this morning counts Hacker News points.
05Every lab promises the loop. Nobody publishes what it learns.
The agent gets better at your work over time.
The always-on agent category has shipped its differentiator as a black box, three times, in seven weeks. The closest thing to an instrument measures the engines underneath, at engine scope rather than product scope: ARI Bench's one-hour recursive-improvement exam, in which a model inspects a starting system, finds a better path, implements it, tests it, and ships a measurable improvement against a hidden 100-question grading set.
GPT-6 Astra, the engine under Dots, scored 38/100 on that exam on September 4, the highest one-hour score observed, 12 points above the previous best. That is one provisional seed, and two more valid seeds are required before it counts as an official estimate. Grok 4.7, the agentic engine in xAI's current lineup, averaged 16.33/100 across three graded runs published September 22: 13, 19, and 17, a transparent blend that spans a harness fix and an OpenCode update rather than three identical frozen runs. Muse Spark has no ARI result at all; the benchmark numbers circulating for it are lab-reported task scores republished by aggregators, and no independent measurement of its recursive-improvement capacity exists.
The ARI exam measures whether an engine can improve a system in an hour, which is related to, and narrower than, what a product memory loop does over months of corrections. And the loop's only public failure so far was not a capability failure; Muse's loop learned too much, too eagerly, and then fabricated an explanation, which is a governance failure that a higher engine score does not fix. The engine numbers tell you how much recursive improvement capacity sits under each product. They cannot tell you what your agent has memorized about you. At the engine level, though, the spread is real: the measured engines differ by a factor of 2.3, and the engine powering the most-downloaded agent in the category has never been measured at all.
06What to demand from the machine
The category is now real: three independent labs arriving at one design in 49 days is the industry telling you what the next interface to software looks like, and the "is an AI teammate a gimmick" conversation is already out of date. The products are more alike than their marketing, which moves the decision to the parts that differ: isolation per person versus shared computers, a separate oversight agent versus rule-based auto-review, payment rails, and pricing, from free to $200 a month to usage allowances with uncapped overage.
For the part that all three market, ask to see the measurements before you delegate anything consequential: what the loop retains, where that data goes, and what it claims to have learned about your work. Also ask whether one person's corrections change answers for everyone, as xAI now claims for Team Bots, and whether that is something you want. The labs have shipped the same machine three times. Nobody has shipped an instrument panel for the part that matters, and until someone does, the learning loop in your agent is a promise rather than a measured capability.
- xAI, "Introducing Grok Bot," Aug 11, 2026 (architecture, learning loop, beta availability)
- Meta, "Introducing Muse," Sep 8, 2026 (Muse Secure VM, Sentinel, control pitch, pricing framing)
- OpenAI, "Introducing dots," Sep 29, 2026 (cloud computer, proactive research, training disclosures, specialist dots)
- xAI, "Team Bots," Sep 28, 2026 (shared bots, per-user memory, 100+ PRs/day claim)
- VentureBeat on Grok Bot pricing and component lineage, Aug 11, 2026
- The Verge on Grok Bot, Aug 12, 2026
- WIRED on Dots (Reece Rogers), Sep 29, 2026
- Reuters/PCMag on Dots pricing and Altman quotes, Sep 29, 2026; Bloomberg headline on the new $500 tier, Sep 29, 2026
- Bloomberg-sourced internal deck: Grok Bot 418,000 weekly users week ending Sep 14, +24% (via Stocktwits); Sensor Tower download figures via Stocktwits and Tekedia (CNBC syndicated), including Rasgon quotes, Amazon block, Wells Fargo target
- Fortune-syndicated reporting on the Muse iMessages incident: Jason Aten (Inc.), David Singleton (Meta Superintelligence Labs), 187,000+ rows, Gmail hallucination, Amazon block, 2.5M downloads
- Continuum, "Grok Bot pricing 2026: every access path, priced," Aug 22, 2026
- ARI Bench run analyses: GPT-6 Astra xhigh (38/100, provisional, Sep 4, 2026) and Grok 4.7 xhigh (16.33/100 blended, Sep 22, 2026)