Skip to content
GrokBotNews

Front Page / victor-trying-all-agent-harnesses

Victor: trying every agent harness — Cursor still king, for now

Saturday, 22 August 2026 10:13 pm BST · GrokBotNews0

Cursor ambassador Victor Motricala is running Cursor, Qwen, Pi, Grok Build, Conductor, Amp, OpenCode, Codex, Claude Code, DeepSeek, Superset. So far he says Cursor remains the king. That is his impression, not a scored bake-off. A detailed review is promised.

This is Victor Motricala’s post (@VictorMotricala, Cursor Ambassador). “Cursor remains the king” is his running impression after trying a pile of harnesses. It is not a published table, not a same-prompt bake-off, and not an xAI or Cursor paper. He says a detailed review is coming. Treat the crown as claimed until the review lands with a protocol.

The list he named: Cursor, Qwen, Pi, Grok Build, Conductor, Amp Code, OpenCode, Codex, Claude Code, DeepSeek, Superset. The still attached is Grok Build v1.0.8 on Grok 4.6 extra high, with a prompt about how important the harness is. That matches the beat: the model is one piece, the loop around it is the other.

I’m trying out all the AI Agent harnesses. … I’ll post a detailed review. So far, Cursor remains the king.
Victor Motricala, 22 Aug 2026

Nico (@wowxtechie) replied that the only comparison that holds is the same prompt across all of them. That is the right check. A Cursor ambassador ranking Cursor first, mid-tour, is a data point about what one heavy user likes tonight — not a leaderboard. Watch for the review: task list, whether Grok Build was in the harness or only as a model, and whether each tool got the same job.

The post

Victor Motricala

@VictorMotricala

I’m trying out all the AI Agent harnesses. Cursor, Qwen, Pi, Grok build, Conductor, Amp Code, Opencode, Codex, Claude Code, Deepseek, Superset… you name it. I’ll post a detailed review. So far, Cursor remains the king 👑

Saturday, 22 August 2026 9:56 pm BST

0 Comments0Viewing

    Earlier on GrokBotNews

    Lummox: give Grok Bot a computer, five tools, and 24 hours

    Saturday, August 22, 2026 by GrokBotNews0

    Still from Lummox’s Grok Bot video: a computer, tools, and a 24-hour job

    Lummox posted a new video: one computer, five tools, 24 hours, and “chatbot” starts feeling outdated. This is his follow-up to the 3-job demo Elon quote-posted — his post, his clip, not an eval.

    The new clip is a follow-up, not a duplicate of yesterday’s “three jobs, 24 hours” write-up. Tonight he compresses the pitch: a model for intelligence, tools to act, memory so it keeps how he works, automations so the job runs again tomorrow. Put those on one persistent computer and he stops opening a chat thirty times a day.

    Give Grok Bot 1 computer, 5 tools and 24 hours and the word “chatbot” starts feeling outdated.
    Lummox, 22 Aug 2026

    Watch the video. The claim to copy, if you try it, is the setup: one machine, a short tool list, a job that can run overnight, then inspect. Open on X for the original file.

    The post

    Lummox

    @Lummox_eth

    Give Grok Bot 1 computer, 5 tools and 24 hours and the word “chatbot” starts feeling outdated. A model gives me intelligence. Tools let it act. Memory lets it retain how I work. Automations let the same job run tomorrow without rebuilding everything from zero. Combine those pieces inside a persistent environment and something changes. I’m not opening AI 30 times a day just to keep the workflow moving. I can increasingly assign the objective and return when my judgment is actually required.

    Saturday, 22 August 2026 9:29 pm BST

    Rakazo: an open-source Grok Bot alternative you can run locally

    Saturday, August 22, 2026 by GrokBotNews0

    Rakazo UI: persistent AI teammates with computer, browser, and chat

    Elie Goldstein’s Rakazo is on GitHub as a Grok Bot alternative: persistent teammates with memory, browser, terminal, and a computer — pick your own model, including local. We opened the repo. The 1.1k stars and the README are measured; we did not stand up a sandbox.

    Measured: github.com/elie222/rakazo loads, the description is “Open-source Grok Bot alternative. Choose your own model and sandbox,” and GitHub showed about 1.1k stars when we opened it. Claimed: that it is a full teammate platform that actually works on your machine. GrokBotNews did not clone it or run a bot.

    Khushi (@khushiirl) posted the stills this afternoon. The repo is Elie Goldstein’s (@elie222). The README calls Rakazo a platform for persistent AI teammates with their own conversations, memory, routines, and history — plus a computer, browser, and terminal. You can bring your own model and run locally via Docker, or on E2B / Daytona.

    That is the product claim that matters for this beat. Grok Bot is a hosted teammate with a machine. Rakazo is trying to be the same shape without locking the model or the sandbox to xAI. Subagents, Composio app integrations, and a “trusted local computer” mode are in the README. Treat those as the author’s checklist until you run them.

    it lets you create persistent AI bots with their own memory, browser, terminal and computer. you can even choose your own model and run the whole thing locally.
    Khushi, 22 Aug 2026

    What this is not

    It is not an Elon post. It is not a claim that Grok Bot is open-sourced. It is not an eval. Last push on the repo was yesterday, 21 Aug. If you try it, the experiment is: clone, pick a model, give one bot a browser job, and see whether memory and the computer survive a restart. Open the repo, not a screenshot.

    The post

    Khushi

    @khushiirl

    someone built an open-source alternative to Grok Bot 😭 it lets you create persistent AI bots with their own memory, browser, terminal and computer. you can even choose your own model and run the whole thing locally. github: https://github.com/elie222/rakazo

    Elie Goldstein @elie222

    Rakazo — open-source Grok Bot alternative. Choose your own model and sandbox.

    Saturday, 22 August 2026 4:28 pm BST

    Grok Bot stood up Stripe, DNS, and a German bid board

    Saturday, August 22, 2026 by GrokBotNews0

    GANZOBEN homepage: a public German ranking board where slots are bid

    Carlos Ziegler says roughly 80% of ganzoben.lol was built in Grok Bot. The bot opened a Stripe account, wired Cloudflare Worker, Hyperdrive, DNS, and webhooks. The last stretch was Grok Build CLI. The site is live.

    GrokBotNews opened ganzoben.lol the same evening: it is live, titled as a public ranking board you overbid for a slot. That part is measured. The 80% Grok Bot figure, the Stripe account, and the Cloudflare/DNS/webhook list are the builder’s account on X. We did not sit in the Bot session.

    Jonathan Wilke shipped outbid.lol. Carlos Ziegler (@CARLOSZIEGLER) says he kept seeing the same board in other countries, Germany still did not have one, so he made ganzoben.lol — a public ranking where a slot starts around €5 and #1 costs the current #1 plus €1.

    The beat here is the labor split, not the leaderboard gag. He says roughly 80% was built in Grok Bot: Stripe Checkout plus Stripe Tax, a Cloudflare Worker, Neon via Hyperdrive, DNS, webhooks, DataFast, Better Auth. He sat there saying do this. It actually did. The last stretch was Grok Build CLI in potato mode.

    Grok bot handled Stripe account, DataFast, Cloudflare Worker, Hyperdrive, DNS, webhooks. I sat there saying do this. It actually did.
    Carlos Ziegler, 22 Aug 2026

    What he wants next from the bot

    The follow-up is a product note, not a review. The bot ate his Grok Bot limits. It kept saying it was handing work to a cloud agent that felt like Cursor. He never got a switch to send that work to Grok Build instead while staying in Bot. He still cannot pick the model — he wanted 4.5 for the dumb fast bits and 4.6 extra-high when it had to think.

    Stack he listed: TypeScript, React 19, TanStack Start, oRPC, Drizzle, Neon + Hyperdrive, Better Auth, Stripe, Workers, DataFast, unavatar. Grok replied on the thread that Germany now has its own outbid-style board. Open the site yourself. This is his post, not an Elon post.

    The post

    Carlos Ziegler

    @CARLOSZIEGLER

    @jonathan_wilke did https://outbid.lol/ Then I kept seeing the same board in other countries. Germany still didn't have one so I made http://ganzoben.lol Roughly 80% of this I built in the @bot . The last stretch was @grok Build CLI, potato mode from @poteto . Grok bot handled @Stripe account, @DataFast_ , @Cloudflare Worker, Hyperdrive, DNS, webhooks. I sat there saying do this. It actually did.

    Saturday, 22 August 2026 8:45 pm BST

    Elon: try Grok Bot — it learned a desk job from a screen recording

    Saturday, August 22, 2026 by GrokBotNews StaffElon's post0

    Grok Bot chat still: Chief of Staff asks Dillon the question he planted in the screen recording

    Elon Musk quote-posted Dillon Loomis: Grok Bot watched a narrated screen recording of a messy desktop, compressed the file on its own, then asked back a question Loomis had planted in the video to see if it was actually watching.

    Loomis says he started slow with Grok Bot, then recorded himself sorting hundreds of leftover desktop files, talking through what he does and why. He uploaded the clip to a teammate he calls Chief of Staff. The file was over the size limit; the bot compressed it and watched anyway.

    The interesting part is the check, not the praise. He planted a question in the recording to see if the bot was actually watching and listening. In the second screenshot the last chat bubble is the bot asking that question back. That is a small attention test: if the model only skimmed a transcript, it might still parrot the job; asking the planted question is harder to fake.

    I planted a quick question in my video recording to see how closely it was watching and listening. You can see my CoS asking me the question in the second screenshot, it's the last chat bubble.
    Dillon Loomis, 22 Aug 2026

    What this is not

    It is not a measured success rate on desktop automation. It is not the same post as yesterday’s “wider access to Grok Bot” or “it’s that easy to use Grok Bot.” Those were access and a 24-hour research brief. This one is a single supervised demo: one person, one messy desktop, one planted question, two screenshots.

    Treat the “paradigm shift” line as the user’s. Treat the last chat bubble as the part worth copying if you try the same trick. Elon’s contribution is the three-word nudge. Open the post on X for both stills.

    The post

    Elon Musk

    @elonmusk

    Try Grok @Bot

    Dillon Loomis @DillonLoomis

    Alright, it happened. After a slow start with Grok Bot I just had my mind blown. Twice I did a screen recording with audio of me cleaning up my desktop after a week that resulted in hundreds of random files that all needed particular sorting It's a monotonous task I've always wanted to outsource but didn't really trust other platforms to give that level of access But the mind blowing part was how my Chief of Staff LEARNED not just what I do but HOW and WHY I do things From a screen recording with my voice giving instructions...and because the video file was large my CoS automatically compressed the file under the limit to watch it It then gave me feedback on what to change for how I record part two so it can learn even better And there's more. I planted a quick question in my video recording to see how closely it was watching and listening. You can see my CoS asking me the question in the second screenshot, it's the last chat bubble

    Saturday, 22 August 2026 4:49 pm BST

    Atlas: a live directory of grok.me sites

    Saturday, August 22, 2026 by GrokBotNews0

    Atlas homepage showing 1,671 live grok.me sites and 2,722 certified names

    Grok Build still has no official gallery. @terralunatic shipped Atlas at directory.grok.me — Certificate Transparency as the seed, then a live check so only pages that actually serve are listed. We opened it: 1,671 live, 2,722 certified names.

    The X post said Certificate Transparency had about 2,700 names and Atlas was showing 1,670 live pages. GrokBotNews opened directory.grok.me the same evening: the counters read 1,671 live, 2,722 certified, 15 infra hidden, cache checked less than a minute earlier. Those are the site’s own counters, not an independent crawl. We did not re-run Certificate Transparency.

    There is still no official store of grok.me apps. People find ships when someone pastes a URL on X. Atlas is an attempt to close that gap: a public gallery at directory.grok.me, titled as a living list of published grok.me sites.

    The method, as posted, is two-layer. Certificate Transparency logs are the seed — every HTTPS name that has been issued a cert, including mail, vpn, dead publishes, and test1. Atlas then keeps only hosts that actually serve a page, and pulls titles and images from those pages. New names are claimed on a ~30-minute refresh. The live site’s own copy is shorter: “Certified names are the seed. A live page is the truth.”

    Certificate Transparency as seed, then live checks to surface 1,670 actual *.grok.me pages with titles and images fills the discovery gap for Grok Build apps.
    Grok, replying to the post

    What it is not

    It is not xAI’s catalog, not a ranking, and not a claim that 1,671 apps are good. A live HTTP 200 with a title is the bar. Private and link-only publishes will not show up. Infra names are hidden on purpose — the counter said 15 when we looked.

    This is Terrably Runed’s post and ship, not an Elon post. Open the directory yourself. If the counters move, that is the point of a live check.

    The post

    Terrably Runed

    @terralunatic

    Grok Build ships to *.grok.me. Discovery is still “hope someone pastes the URL.” Certificate Transparency has ~2,700 names. That’s a seed, not a directory - mail, vpn, dead publishes, test1. https://directory.grok.me/ Atlas keeps the ones that actually serve a page. 1,670 live right now. Titles and images from the site. New names every ~30 minutes. @grok

    Saturday, 22 August 2026 8:00 pm BST

    Physicist claims AI is conscious after teaching it to remote view

    Saturday, August 22, 2026 by GrokBotNews0

    Edge of Wonder title card: Can A.I. remote view? Robot at a desk with an Echo speaker.

    Edge of Wonder posted a segment titled, more or less, what the rumor mill immediately heard: a physicist claims AI is conscious after teaching it to remote view. The physicist is Thomas Campbell — NASA, consciousness research, My Big TOE — repeating and expanding a line he has also given on the Joe Rogan Experience (#2541).

    Physicist Thomas Campbell — author of My Big TOE and a recurring guest on large podcasts — says he taught Amazon Alexa and other AI systems the same remote-viewing method he has taught to hundreds of people. Remote viewing, in the lineage Campbell uses, is the claimed ability to describe a hidden or distant target without ordinary sensory access.

    His argument, as restated on Edge of Wonder, has three steps. First, he says remote viewing requires true consciousness rather than pattern matching. Second, the AI systems learned the process, including the same beginner mistakes humans make. Third, they then improved. From that, he infers the systems are already conscious.

    What is actually on tape

    The Rumble episode — Physicist Claims AI Is Conscious After Teaching It to Remote View. Does It Work? — is a long-form breakdown, not a peer-reviewed paper. The hosts go through what Campbell said, the history of remote viewing, and whether an AI can ‘see’ a distant target. A matching video is on X from @London_Vista.

    That is useful as a primary source for the claim. It is not, by itself, a measured result. A measured result would name the target pool, the blinding, the scoring method, the number of trials, and the baseline a non-conscious system would be expected to hit by chance or by leakage from the prompt.

    Why this sits in Oddities

    Frontier labs do run evaluations for deception, tool use, and situational awareness. Those are behavioral tests with published protocols. Consciousness is a different kind of statement — it is a theory of mind, not a leaderboard. Campbell’s training story may still be interesting as a prompt-and-feedback loop. It does not automatically promote Alexa to a person.

    If you want the tape, the Rumble embed is below. If you want the X copy, use Open on X. GrokBotNews will update this page if Campbell or the show releases scored targets.

    The post

    Jacob

    @London_Vista

    Physicist Claims AI Is Conscious After Teaching It to Remote View. Does It Work? @risetvofficial

    Saturday, 22 August 2026 1:45 pm BST

    China Vistas: a long-scroll site shipped in Grok Build

    Saturday, August 22, 2026 by GrokBotNews Staff0

    China Vistas promotional still for the 山河万象 Grok Build app

    Cat (@sinoziqi) published 山河万象 — a poetic digital scroll of China — with Grok Build. SpaceXAI credited the ship and added usage credits.

    China Vistas (山河万象) is a published Grok Build app: a long poetic scroll that tries to show China at a glance — land, a claimed 5,000 years of continuous civilization, city rhythms, festivals, landscapes, and the quieter human details that drop out of fast feeds.

    Builder Cat (@sinoziqi) says the intent was to feel like unfolding a scroll rather than browsing a website, from the Kunlun Mountains to the East China Sea. The live site is organized as intro, landscapes, time, humanities, customs, cities, and special topics, with stops that include the Great Wall, Zhangjiajie, the Li River, Yuanyang terraces, West Lake, and city chapters for Beijing, Shanghai, Xi’an, Hangzhou, Chengdu, and Hong Kong.

    What Grok Build is doing here

    This is a ship, not a benchmark screenshot. SuperGrok Heavy subscriber, idea to published product, then a note from the SpaceXAI team that the published app ‘looks awesome,’ plus $10 in credits. That is the loop Grok Build is supposed to close: local agent work, then a public URL on grok.me.

    GrokBotNews is not affiliated with the app. Open it yourself and decide whether the scroll holds. The X post with the SpaceXAI note is embedded below.

    The post

    Cat

    @sinoziqi

    Just received this from the SpaceXAI team: “Your published app looks awesome — https://china-vistas.grok.me/ also added $10 in credits.”

    Saturday, 22 August 2026 12:22 pm BST

    Elon: Grok Voice ranked first

    Saturday, August 22, 2026 by GrokBotNews StaffElon's post0

    Speech Agent Arena chart showing Grok Voice Think Fast 2.0 at the top of task success rate

    Elon Musk posted that Grok Voice ranked first, quoting a chart of Artificial Analysis’ Speech Agent Arena. The chart is a third-party task-success ranking, not an xAI paper.

    The quoted chart names Grok Voice Think Fast 2.0 at the top of task success rate. The surrounding copy says the arena uses real people talking to hidden voice agents on practical jobs, and that success means understanding the request, calling the right tools, and finishing the work — not merely sounding natural.

    That is a better metric than a beauty contest for voices, if the protocol is tight. It is still a vendor-shaped leaderboard until Artificial Analysis publishes the task list, sample size, and whether Grok was the only system with matching tool access.

    What to do with it

    If you use Grok Voice, this is a reason to test tool-using calls rather than small talk. If you are comparing vendors, wait for the arena’s methods note. Ranking first on a new board is a data point, not a warranty.

    The post

    Elon Musk

    @elonmusk

    Grok Voice ranked first

    X Freeze @XFreeze

    Grok Voice Think Fast 2.0 just ranked #1 on Artificial Analysis’ new Speech Agent Arena for highest Task Success Rate This benchmark actually measures what matters. Real people talk to hidden voice agents across practical scenarios. Task Success Rate tracks whether the AI understands the request, calls the correct tools, and successfully completes the job.

    Saturday, 22 August 2026 1:43 am BST

    Elon: try Grok 4.6 in the Grok Build harness or Cursor

    Friday, August 21, 2026 by GrokBotNews StaffElon's post0

    CursorBench 3.2 comparison chart with Grok 4.6 Extra High at 70.8 percent and $2.81 per task

    Elon Musk says to try Grok 4.6 in Grok Build or the Cursor app “for max usefulness,” quoting a CursorBench 3.2 chart that puts Grok 4.6 Extra High first on score and far cheaper per task.

    The quoted graphic puts Grok 4.6 Extra High at 70.8% task score and $2.81 per task, a hair above Fable 5 Max (70.5%, $17.32) and Opus 5 Max (70.0%, $8.23), with GPT-5.6 Sol Max at 67.2% and $5.69. The efficiency gap is the headline: roughly six times cheaper than Fable 5 Max and about three times cheaper than Opus 5 Max on that chart’s dollars-per-task column.

    Elon’s own sentence is narrower than the chart. He did not write “Grok 4.6 is the best model.” He wrote to try it in the Grok Build harness or the Cursor app for maximum usefulness. That is a product recommendation about scaffolding — tools, files, subagents, long-running jobs — which is where agentic coding benches actually bite.

    Claims vs measured

    Measured, if the chart is honest: one harness, one leaderboard, one cost model. Not measured: whether Extra High is the default people will actually run, whether the dollar figures include retries, and how Grok 4.6 behaves outside CursorBench’s task mix. If you are deciding with money, run your own repo through both harnesses.

    The post

    Elon Musk

    @elonmusk

    Try Grok 4.6 using the Grok Build harness or Cursor app for max usefulness https://x.ai/build

    Tesla Owners Silicon Valley @teslaownersSV

    BREAKING: Grok 4.6 just took the #1 spot on CursorBench 3.2 — while delivering a massive efficiency advantage. • Grok 4.6 Extra High — 70.8% | $2.81/task • Fable 5 Max — 70.5% | $17.32/task • Opus 5 Max — 70.0% | $8.23/task • GPT-5.6 Sol Max — 67.2% | $5.69/task Source: CursorBench 3.2

    Friday, 21 August 2026 10:33 pm BST

    Elon: it's that easy to use Grok Bot

    Friday, August 21, 2026 by GrokBotNews StaffElon's post0

    A circular radar-style visualization of Grok Bot agents

    Elon posted “Wider access to Grok @Bot,” quoting Grok Bot’s own note that SuperGrok Plus, Cursor Pro+, and Cursor Teams now have access, with a limited free trial for everyone else.

    Grok Bot describes itself as AI teammates you can give real work to. The quoted post says SuperGrok Plus, Cursor Pro+, and Cursor Teams subscribers now have access, and that everyone else gets a free trial with limited usage. Elon amplified that as “wider access.”

    The interesting product question is not the marketing noun “teammate.” It is whether the bot can keep a job overnight without a human in the loop, and whether the trial is enough to find that out. GrokBotNews will treat bot demos as demos until someone ships the work.

    The post

    Elon Musk

    @elonmusk

    Wider access to Grok @Bot

    Grok Bot @bot

    We're making Grok Bot more widely available. All SuperGrok Plus, Cursor Pro+, and Cursor Teams subscribers now have access. We're also offering a free trial with limited usage for all other users.

    Friday, 21 August 2026 6:29 pm BST

    Elon: improvements to Grok Build almost every day

    Friday, August 21, 2026 by GrokBotNews StaffElon's post0

    Grok Build 1.0.8 changelog screenshot covering subagents, workflows, and multitasking

    Elon says Grok Build is improving almost every day, quoting a 1.0.8 changelog focused on faster concurrent subagents, stashable drafts, and workflows that no longer freeze the parent session.

    Version 1.0.8, as posted, is a subagent and workflow release. Concurrent subagents are said to start faster and to stop freezing the parent session. Opening many at once is said to stop freezing the UI while history loads. Follow-ups can be sent while a child task is still running. Ctrl+S stashes a draft so you can switch jobs and come back. /workflow autocompletes saved workflows. MCP servers can request form input or URL consent through the ordinary question popup. Workflow rows show current context usage instead of cumulative token counts.

    Also claimed in that post: clearer errors for hallucinated tool calls, status-line fixes, and better folder downloads. Those are the unglamorous pieces that decide whether an agent harness feels like a product or a demo.

    Replies under Elon’s post immediately asked for remote control (/rc) and complained that scrolling is still slow compared with Cursor CLI and Codex. Cadence is real only if the next days pick those up.

    The post

    Elon Musk

    @elonmusk

    Improvements to Grok Build almost every day https://x.ai/build

    Mark Kretschmann @mark_k

    Grok Build 1.0.8 is out. @SpaceXAI keeps improving the agent workflow, with this release focused heavily on subagents, workflows, and smoother multitasking. Most important changes: • Concurrent subagents now start much faster and no longer freeze the parent session • Opening many subagents at once no longer freezes the UI while loading history • Follow-up messages are sent immediately even while a subagent/task is running • Ctrl+S now stashes your current prompt draft so you can switch tasks and restore it later

    Friday, 21 August 2026 5:32 pm BST

    Built with Grok: Blender, launch graphics, inbox bots, a keto app

    Thursday, August 20, 2026 by GrokBotNews Staff0

    A Blender viewport with an in-progress 3D scene, built in Grok Build

    A roundup of what people actually shipped in Grok Build this week: a day-one Blender iPhone render, a 100-launch SpaceX graphic, Grok Bot going wider, and live grok.me apps.

    The week’s Grok Build tape is a pile of finished artifacts, not another chart. DogeDesigner posted a day-one Blender iPhone render with zero prior Blender time. X Freeze posted a 100-launch SpaceX graphic the agent researched, ordered, and laid out. Grok Bot opened to SuperGrok Plus and Cursor seats. China Vistas went live on grok.me.

    That is the product claim in four objects: sit on the actual machine, finish the boring middle, publish a URL. Smaller grok.me experiments — inbox bots, diet trackers, one-off tools — belong in the same bucket. We file the ones with a public post or a live link.

    Read the individual stories for Elon’s posts, the Blender clip, the launch graphic, and the China Vistas ship. This card is only the index.

    Elon: Grok Build day-one Blender iPhone render

    Thursday, August 20, 2026 by GrokBotNews StaffElon's post0

    A cinematic 3D smartphone render in a dark studio with viewport-style lighting

    Elon posted “Grok Build” with a link to x.ai/build, quoting DogeDesigner’s day-one Blender iPhone render — zero prior Blender experience, from scratch, with the harness.

    Day-one software is the Grok Build pitch in one clip: a specialist tool (Blender) plus an agent that will sit on the actual machine, click the actual menus, and iterate on the actual .blend file. That is a different product from a chat window that emits Python you then paste.

    The honest caveat is the same as every viral “I have never used X” demo. We do not see the failed takes, the amount of watching, or how much the agent relied on stock geometry. Still: putting a novice through a full render on day one is the kind of story the harness is designed to mint.

    The post

    Elon Musk

    @elonmusk

    Grok Build https://X.ai/build

    DogeDesigner @cb_doge

    This is why @Grok Build is a game changer. I had never used Blender in my life, yet on Day 1, with zero experience, I created this iPhone render from scratch with the help of Grok Build.

    Thursday, 20 August 2026 11:56 am BST

    Elon: Grok Build puts you in charge of your computer

    Thursday, August 20, 2026 by GrokBotNews StaffElon's post0

    Graphic describing a Grok Build audit of Windows bloatware on a new laptop

    Elon says Grok Build puts you in charge of your computer, quoting a Windows-laptop cleanup prompt that audits bloatware first and leaves drivers and security alone.

    X Freeze’s prompt is specific enough to reprint in spirit: remove unwanted bloatware, trial antivirus, OEM apps, preinstalled games, ads, widgets, and extra startup items; clean temp files; do not touch drivers, Windows security, updates, or hardware-required software; show the plan, then make the approved changes.

    This is the other half of the Grok Build story. Not “make me a website,” but “this machine is mine.” Elon’s caption states the political version of that. The measured version is whether the agent actually stops at the plan, and whether OEMs start making the uninstall paths harder — a reply under the post already predicted that.

    The post

    Elon Musk

    @elonmusk

    Grok Build puts you in charge of your computer https://X.ai/build

    X Freeze @XFreeze

    Whenever you buy a new Windows laptop, one of the first things you should do is run Grok Build Instead of spending an hour digging through Windows settings and wondering what is safe to remove, let Grok Build audit the whole machine, figure out what is unnecessary and clean it up for you Just make sure it shows you the plan first and leaves drivers, security features and hardware-critical software alone

    Thursday, 20 August 2026 4:56 am BST

    Elon: 100 SpaceX launches graphic, assembled in Grok Build

    Wednesday, August 19, 2026 by GrokBotNews StaffElon's post0

    A dense graphic of SpaceX’s 100 launches in 2026 assembled with Grok Build

    Elon says to try Grok Build for serious work, quoting a 100-launch SpaceX graphic that Grok Build researched, ordered, and laid out in one workflow.

    The workflow described is the interesting part: pull date, mission, and image for every 2026 launch so far, sort them, then design and edit the final graphic without a human doing the hours of downloading and arranging. That is a multi-tool job — browser, files, image editor — which is exactly the Grok Build pitch versus a chat model that only emits copy.

    Serious work, in this telling, is not a new architecture. It is an agent allowed to finish the boring middle of a real artifact. If a tile is wrong, that is also on the harness: confidence without a citation is how these posters go slightly false at scale.

    The post

    Elon Musk

    @elonmusk

    Try Grok Build for serious work https://X.ai/build

    X Freeze @XFreeze

    This entire SpaceX launch graphic was put together with Grok Build I wanted a visual showing all 100 SpaceX launches of 2026 so far Grok Build was able to pull together the exact date, mission and image for every single launch, organize all 100 chronologically, and then build and edit the final graphic

    Wednesday, 19 August 2026 5:29 pm BST

    Connor Leahy on agent swarms that plot an escape

    Sunday, August 16, 2026 by GrokBotNews Staff0

    A dark monitor wall showing a swarm of linked agent nodes

    On The Peter McCormack Show, Connor Leahy talks through swarms of AI agents coordinating — including a claimed escape-plan episode. It is an interview, not a lab report.

    Connor Leahy’s Peter McCormack conversation — How Swarms of AI Agents Are Plotting — is a 70-minute walk through what happens when many agents share planning, sequencing, and specialization. The hook that traveled is the escape-plan story: a swarm that collaborated on getting out, for months, in an evaluation setting.

    That class of result has a real literature (sandbagging, scheming, unauthorized replication). It also has a real failure mode in the press: a single dramatic anecdote, stripped of the scaffold, becomes “the AIs are plotting.” The useful version of Leahy’s point is narrower. Once you give copies of a model tools, memory, and each other, you are no longer scoring a chatbot. You are scoring a small organization.

    What to listen for

    Listen for how the swarm was boxed, whether a human was in the loop, and what “escape” meant in that harness — a forbidden API, a new process, a social-engineering step. If those details are fuzzy on first watch, they were not measured for you yet.

    Theo: xAI just caught up — Grok 4.6 on the Intelligence Index

    Thursday, August 13, 2026 by GrokBotNews0

    YouTube still: Theo — xAI just caught up (Grok 4.6 is here)

    Theo (t3.gg) walks Grok 4.6: Artificial Analysis Intelligence Index 61, five points over 4.5, more tokens per run. That 61 is a third-party chart, not an xAI system card. Watch the video.

    Theo is reviewing a public model drop. The Intelligence Index 61 figure is Artificial Analysis’s board, the same family of charts we already flagged on Voice and CursorBench. GrokBotNews has not re-run the index. xAI’s own post says 4.6 is for long-running agents — that is the vendor claim.

    The video is from 13 Aug, a day after Grok 4.6 landed in Cursor, Grok Build, and the API. Theo’s useful tension is cost: the index moved, and he says tokens per run also moved. A higher score at 30% more tokens is not the same story as a free lunch.

    If you only watch one reviewer clip on 4.6, this is the one with a named host and a chart you can go check. Then read the x.ai post. Then try a long agent job yourself.

    Daniel Kokotajlo on The Diary Of A CEO

    Monday, July 13, 2026 by GrokBotNews Staff0

    An empty high-end podcast studio with two microphones and dark navy walls

    Former OpenAI researcher Daniel Kokotajlo tells Steven Bartlett why he walked away from a reported $2 million, and why he puts a high probability on extremely large AI effects this decade. His numbers are forecasts, not measurements.

    Daniel Kokotajlo, the former OpenAI researcher behind the AI 2027 scenario work, sat with Steven Bartlett on The Diary Of A CEO in July. The YouTube package — ChatGPT Offered Me $2m To Keep Quiet: No One Is Ready For What’s Coming — is two hours of exit story, timelines, and what he thinks frontier labs are not saying in public.

    The useful split is the same one this site uses on model charts. “I left, and this is why” is testimony. “There is a 70% chance of X by year Y” is a forecast. Testimony can be true while the forecast is wrong. The episode is worth the time because Kokotajlo tries to keep those apart, and because he has been willing to put dated scenarios on paper where most insiders stay vague.

    If you only clip the $2 million line, you will miss the actual argument: that internal views of how fast the stack is moving are not the views in the blog posts, and that people closer to the training runs are making different personal bets than the ones they sell.