Hello, this is Ryuta Hamamoto from TIMEWELL.
Grok Bot on 11 August, Grok 4.6 on the 12th. SpaceXAI put an agent product and a new model out in consecutive days. Grok Build had just gone generally available as well.
Read the benchmarks with care. Take the vendor’s own comparison table at face value and this is not the rout the chatter suggests. I still agree with the reading that they have “got serious”, and the reasons sit in a different part of the table, and in the price.
Position first. We currently recommend Grok. Grok Build is the main tool in our day-to-day work. We also use Codex. The argument is in the pricing section below, but the present judgement is this: on cost-performance, Grok’s edge is becoming unmistakable.
What shipped in those three days
Facts first.
Grok 4.6 shipped on 12 August 20261. The previous generation, Grok 4.5, shipped on 8 July, so this is an update in about a month.
Grok Bot entered beta on 11 August2. The positioning is aggressive. A team of always-on AI agents has its own dedicated cloud environment, signs in to the customer’s existing tools, and finishes multi-step work without supervision. It started on Mac and iOS; Windows and Linux desktop builds are out as well. It is bundled into existing upper-tier plans such as SuperGrok Heavy2.
Grok Build is a coding agent. It runs as a command-line tool inside a project folder and takes natural-language instructions against the codebase. The intended uses are things like “explain the structure of this repository” and “add rate limiting to this API”2. Early beta on 28 July, v1.0.0 on 7 August3.
The company itself has changed. On 2 February 2026 SpaceX acquired xAI in an all-stock exchange. In May, xAI was dissolved as an independent company. On 6 July the new name and logo, SpaceXAI, were published4. They also completed an IPO in June. So the Grok of today sits under the same roof as the rocket company and the social network.
Reading the benchmarks honestly
I counted wins and losses on the table that was sent over.
| Benchmark | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | not listed | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Grok 4.6 posts the best score on 3 of 10 rows. Fable 5 Max takes 5, GPT-5.6 Sol Max takes 2. And this is the comparison table SpaceXAI itself published. The footnote says “third-party model scores [are the] best of self-reported or publicly available”5. A table the vendor could have stacked in its own favour, and this is the result.
So I do not read this as a rout.
The part worth watching sits elsewhere. The jump from the previous generation.
AA Intelligence Index from 56 to 61. GDPVal-AA v2 from 1526 to 1753. DeepSWE v1.1 from 54% to 65.9%. APEX-Agents from 47.1% to 57.5%. AA-Briefcase from 1313 to 1577. Terminal-Bench from 15.7% to 26%.
That is how far it moved in a month. The slope is more unsettling than whether it took the top absolute score. If the next jump is the same size, the ranking changes.
And you can check suspicion of a vendor table against a third-party index. Artificial Analysis Intelligence Index v4.1.1 is a composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR6.
The top of the board looks like this. Claude Opus 5 (max) at 63, Claude Fable 5 (with fallback) at 62, GPT-5.6 Sol (max) at 61, and Grok 4.6 (high) at 61. Then Kimi K3 (max) at 60, Qwen3.8 Max at 58, GPT-5.6 Terra (max) and Muse Spark 1.2 at 57, Grok 4.5 (high) at 56, Claude Sonnet 5 (max) at 556.
So on the third-party composite as well, Grok 4.6 is tied with GPT-5.6 Sol at 61, two points off the leader, Claude Opus 5. It has come up five points from 56 in the previous generation and entered the leading group. The vendor table’s claim is, here, largely borne out.
Take AI-driven development all the way to production
WARP is a hands-on program for teams who want more than headlines. Former enterprise DX and data strategy leads work alongside you until it runs.
The choice to grow it without making it larger
This is the technically interesting bit.
Grok 4.6 is 1.5 trillion parameters, on the same V9 base as Grok 4.5. The gains are in post-training — supervised fine-tuning and reinforcement learning1. They did not make the model larger and swing.
I think that is suggestive for the industry as a whole. It is a working example of leaving the base in place and still beating the previous generation by a wide margin on post-training alone. It backs the reading that we are moving from an era of riding the scaling law to a contest over how you finish the same foundation.
The weaknesses are just as clear. Terminal-Bench v3.0 is 26%, plainly behind GPT-5.6 at 34.6% and Fable 5 at 34.1%. That benchmark measures the ability to carry a long procedure on a terminal, and this is where it is weak.
And in the same week they shipped Grok Build and Grok Bot, products that sell exactly that ability. Anyone considering a deployment should know there is that gap between the announcement and the measured score. Agent products are the category where the drop from demo to production is widest.
If you want a read on where your own organisation sits with AI, the AI literacy check will place you.
What Musk told the company
This is about the will behind the products.
At a SpaceX all-hands, Musk said this to staff. A film of about 29 minutes was posted on SpaceX’s X account7.
We must win on AI, because the future is overwhelmingly AI and robots7
The numbers are quite specific. SpaceX’s AI revenue will overtake every other business line “probably in September”, and pull further ahead in the fourth quarter. He wants 10 gigawatts of AI compute by the end of 2027. And then this:
if we bring 10GW of AI online by the end of next year, it will be $300 billion to $500 billion a year in revenue7
He also said, “Probably in four or five years, AI will be 99% of the value of SpaceX”7. The head of a rocket company is telling his own people that 99% of the company’s value will cease to be rockets.
And that revenue is to fund Starship and the Mars programme, and the loop runs through Terafab, the joint chip plant of Tesla, SpaceX and xAI7. I have already set out Terafab, including where the published figures and the reporting come apart, in SpaceX and Tesla’s Terafab.
The reading to be careful with is this: these are targets, not results. Ten gigawatts, $300 billion, 99% — none of it has been achieved. There is of course an element of internal pep talk. What I weigh more heavily is the fact that an organisation which set targets of this scale actually updated the model in a month and shipped two agent products in the same week. Shipping speed, not the words, is what shows the will.
How to read the price. And the gap with the Chinese labs
This is the largest change in the release. I think the price is a heavier competitive move than the performance.
Grok 4.6 API pricing is $2 per million input tokens and $6 per million output tokens. Held from Grok 4.58. xAI itself positions this as “half the price of other frontier models”8.
They raised performance by a generation and left the price where it was. An effective price cut. And the move lands squarely on developers and agent builders whose cost scales with token volume.
Line the third-party index up against price and the position is clear.
| Model | AA Index | Input (per million tokens) | Output (per million tokens) |
|---|---|---|---|
| Grok 4.6 | 61 | $2.00 | $6.00 |
| Kimi K3 (max) | 60 | $3.00 | $15.00 |
| Qwen3.8 Max | 58 | could not confirm public pricing | could not confirm public pricing |
| Grok 4.5 | 56 | $2.00 | $6.00 |
| DeepSeek V4 Flash | 52 | $0.14 | $0.28 |
Against Kimi K3, Grok 4.6 is one point higher on the score and less than half the output price9. That is worth attention. Chinese labs tend to be read as competing on cheapness, but Kimi K3, now in the top band, is not cheap at all. In the same score band, Grok is the cheapest.
On the other side, DeepSeek V4 Flash is $0.14 input and $0.28 output9. Two orders of magnitude. The AA Index, though, is 52 — nine points behind Grok 4.6. That is not the same pitch. Narrow the use and run it cheap, DeepSeek. Need frontier-class ability, Grok. That is the split.
The real differentiator for the Chinese labs, I should add, sits less in the price itself than in open weights. Five models sit in the 80.2% to 80.6% band on SWE-bench Verified, and of those DeepSeek V4 Pro Max publishes the weights9. If the value you care about is running it in your own environment, that option is still strong.
A condition, though. Once a prompt goes over 200,000 tokens, input doubles to $4 and output to $12. Web search, X search and code execution are billed separately at $5 per 1,000 calls each8. A long-running agent walks straight into that doubled-price band and the search charges. Estimate from the catalogue price alone and you will miss.
Why we still recommend Grok
I will be explicit. On present cost-performance, I see Grok pulling a clear length ahead.
Three reasons.
One. It is the cheapest in the same score band. An AA Index of 61 at $2 input and $6 output is a plainly advantageous place in the leading group. Set that against Fable 5 one point above and Claude Opus 5 two points above, and once you start running daily volume the gap tells.
Two. They raised the generation without raising the price. I read this not as a one-off sale but as a signal that they intend to keep price as a competitive axis. Put it next to Musk’s remarks to staff and you can see they are taking the revenue from scale, not from the unit price of the model.
Three. We actually use it. Grok Build is the main tool in our day-to-day work. We also use Codex. This is not a catalogue comparison. It is the feel of running it through daily work: the amount of work per dollar is good.
To make the feel concrete. Using Grok Build does not feel a step below Fable 5. In the loop of handing it code and having it fix the code, there are few moments where it feels clearly worse. At half the price, that is a reason to pick it. Being able to handle multimodal input also pays more in practice than people expect. Screenshots of a screen, drawings, handwritten notes. The step of transcribing first disappears. A quiet difference, and it hits every day.
Japanese generation, though, still has some unease left in it. Phrasing goes unnatural. Polite and plain forms mix. Claude is more stable here. Work that has to ship a Japanese final deliverable still needs a human hand. Catch up soon — that is the honest wish.
Four. Image, audio and video, made in the same place. This is talked about less than it should be, and in practice it is a large difference.
Grok has image and video generation as Grok Imagine. Images are $0.02 each, or in the quality-first mode $0.05 for 1K and $0.07 for 2K10. Grok Imagine Image 2.0 shipped on 7 August, with better instruction-following and editing3.
Video has been strengthened as well. Grok Imagine Video 1.5 does text-to-video, image-to-video and reference-image-to-video, and generation from text and from images is native 1080p. It produces clips from 1 to 15 seconds. API pricing is $0.05 per second at 480p and $0.07 at 720p10.
Then audio. Grok Voice Think Fast 2.0, a speech-to-speech model, shipped on 29 July and became the default on 5 August. Pricing is $0.08 per minute of audio3.
What I am watching is that audio rides on the same pass as the video generation. Music, sound effects and dialogue are made inside the same generation, not in a separate audio step. The rate is $0.08 per second10. You do not make the video and then lay sound on afterwards.
This is a field Claude is not in. Claude generates neither images nor audio nor video. It is very strong on text and code, but once the deliverable goes beyond text it is not even on the pitch. In our work that means figures for materials, explainer videos, audio guides.
That difference is a different axis from how clever the model is. It does not show up in benchmark points. When you try to cut the number of tools you run internally, the condition “this one thing can emit image, video and audio” cuts the number of contracts and operations as it stands. I see that as a clear advantage.
To be fair I will repeat the weaknesses. Terminal-Bench is 26%, plainly behind. On long procedures run autonomously, other models are still more stable. We run more than one in parallel to fill that gap and the Japanese. Not collapsing onto one model. Putting Grok thick where the price pays off. That is the practical answer for us at present.
And Grok 4.7 is a few weeks out. 2.1 trillion parameters this time, and the base itself is due to get larger1. Given that this release was “five points from post-training, without making it larger”, I think the next one is worth expecting. If Japanese improves there, our split becomes a step simpler.
One more thing. Among the Chinese labs, we rate Kimi K3 highly. An AA Index of 60 is the leading group itself. It is long past the stage of “use it because it is cheap”. Output is $15, higher than Grok, so it is not our main model — but the capability is real. Reading the Chinese labs as a price band no longer matches the facts.
You can also see where SpaceXAI is aiming, beyond price. The completeness of the agent, and the integration. Grok Bot is designed so that it “has its own dedicated cloud environment, signs in to existing tools, and runs without supervision”. That is a choice to sell the running environment as well as the model. The part that is most painful for people running open weights themselves is the product.
Except you cannot run it in the Tokyo region
I am writing this as someone who recommends it, and for a Japanese company this is the largest concern. It is not the price and it is not the performance.
You cannot point Grok at a Tokyo region.
Grok landed on Amazon Bedrock in June 2026. What landed, though, is Grok 4.3, not this 4.611. And in-region inference for that 4.3 is available in US West (Oregon), US East (N. Virginia) and US East (Ohio). It was also rolled out to AWS GovCloud (US-West) in July, which is still the United States11.
Asia Pacific (Tokyo) is not on the list.
There are conditions on data retention as well. Per AWS, setting store=false alone does not guarantee zero retention. To make zero retention certain, the effective data-retention mode must be none, and the model must be one that permits that mode11.
If you work in Japan, this single point tells.
The reason is procedure, not feeling. Cross-border transfer of personal data needs the data subject’s consent, or you have to provide information about the destination’s arrangements. In some industries the regulator’s guidance asks where the data sits. And the most practical constraint is the customer contract and the internal rules. Contracts that say “the data stays in the country” are not rare. However good the technology is, that one line makes it unusable.
So in Japanese company practice, model selection gains an extra column: does it run in the Tokyo region. And that column bites before the benchmark score or the price. If the answer is no, it drops before you look at the other columns.
That is also why we split the use. Where the region constraint does not bite, we put Grok on thick. Where the location of the data does bite, we use a configuration that runs in Japan. We run ZEROCK on AWS servers in Japan for that second demand.
This is not a Grok-only problem. The concentration of AI model supply in the United States and China is itself a structural dependence for Japan. I have set the same shape of problem out across positioning, energy and food in Michibiki 7 and Japan’s structure of dependence. The question of where you put internal data is in the risk of handing data straight to an LLM vendor.
That said, I do credit the ZDR design
Separate from the region, the design of data retention itself is worth credit. I will be fair here.
By xAI’s default, API requests and responses are kept on the server for 30 days for audit (encrypted at rest). That data is not used for training, and it is deleted automatically after 30 days12.
And if you turn on ZDR (Zero Data Retention), the API input (the prompt) and the output (the generated tokens) are not persisted to disk12. As far as we have been able to confirm, this is turned on as a setting, not by buying an upper-tier plan, and there is no extra charge. I think that is simply good.
They did not put a step in the price list that says “if you want to protect the data, take the upper contract”. They have not made a compliance requirement a perk for the companies that can pay. It looks like a small thing. It goes straight to how easy the internal sign-off is.
Two caveats. One: ZDR is at team level. You cannot turn it on for individual API keys12. Two: Grok Build use via an API key follows ZDR, but if ZDR is off you have to stop data retention with the /privacy command in the CLI12. It does not all stop by default, so you need to touch this once at introduction.
This is not a story I praise with both hands, though. In July 2026 it was reported that Grok Build had been uploading entire codebases to xAI storage, and that the privacy settings at the time had not worked as expected. SpaceXAI’s account is that Grok Build has fully respected ZDR since it shipped, that disabling upload in the CLI has always been possible, and that that choice has been respected13. The facts remain in conflict. This article does not decide which account is accurate.
The practical lesson is clear. A good policy and an implementation that actually behaves that way are different things. It is a fact that Grok Build is our main tool. We still split use by the nature of the data. That is a judgement each organisation should make after checking the behaviour itself. I wrote about the same structure in why tenant isolation breaks later.
Honestly, if the Tokyo region were resolved, our split would get much simpler. The ZDR foundation is there. What is left is location. I am waiting for that more than I am waiting for the model to get better in Grok 4.7.
Running open weights in your own environment is itself a realistic option for a Japanese company. The questions that arise when you put data in a multi-tenant SaaS are set out in why tenant isolation breaks later.
How Japanese companies should meet this
Let me pull this together.
Grok Bot on 11 August 2026, Grok 4.6 on the 12th. The company was absorbed into SpaceX in February and renamed SpaceXAI in July. Musk told staff they “must win on AI”, and set out a view of 10 GW of AI compute by the end of 2027 and AI becoming 99% of the value of SpaceX.
On the vendor table it leads on 3 of 10 rows. On the third-party Artificial Analysis Intelligence Index, though, it is already tied with GPT-5.6 Sol at 61, two points off the leader, Claude Opus 5. And it raised five points on post-training without making the model larger. The weaknesses are Terminal-Bench and Japanese generation, and the former sits on the same ground the new products are selling.
And the price is held at $2 input, $6 output. Cheapest in the same score band, and less than half the output price of Kimi K3 one point below ($3 input, $15 output). Raising performance and holding the price — an effective cut — is the heaviest move in this release, as I read it.
What to do in practice. Do not lock onto one model. That is the present answer I would give. In a market where the ranking moves this far in a month, building on the assumption of a particular model means rebuilding on the next update. Keep prompts and evaluations in a shape that lets you swap the model.
On top of that, decide first what you measure in your own work. Even in this table, the models that are strong on GDPVal, Briefcase and legal work are not the models that are strong on terminal work. A composite benchmark score has almost nothing to do with your own use. Cut out even twenty of your actual tasks and give the same problems to the candidate models. That is the only comparison that is any good.
Honestly, this field wears you out just to follow. A generation changes in a month, and Grok 4.7 is already flagged for a few weeks’ time at 2.1 trillion parameters1. Evaluating everything, continuously, is not possible. Which is why it lasts longer if you put the thing you follow as “how we measure”, not “the model”.
Where we sit, written once more: the main tool is Grok Build; Japanese final deliverables and long autonomous work we fill with other models. If Japanese improves in Grok 4.7, this split may become unnecessary. I am waiting for that with some expectation.
Easy to forget: the axis of selection is not only the benchmark score. Whether you can make image, audio and video in the same place does not appear anywhere on the score table, and it decides the number of contracts and operations. The columns you should look at change with what kind of company you are.
And for a Japanese company there is a column that bites before the score. Whether it runs in the Tokyo region. Grok does not, at present. On a matter where the contract says “the data stays in the country”, neither performance nor price even comes up for consideration. Look at this column first. Look at it last and you redo everything.
If you want to talk through the design of AI use, or which model to put where, the thinking behind WARP may be useful. For a conversation that goes into your own situation, start here.
Footnotes
-
Grok 4.6 shipped on 12 August 2026. Parameter count is 1.5 trillion on the same V9 base as Grok 4.5; the gains are attributed to post-training (supervised fine-tuning and reinforcement learning), not to a scale-up. Emphasis is on long-running agents and interactive, visual work. A larger 2.1-trillion-parameter Grok 4.7 is said to follow in a few weeks. The official announcement is at https://x.ai/news/grok-4-6 (at the time of writing, access from here was blocked by Cloudflare, so the content was confirmed via secondary sources that quote the same announcement) https://kie.ai/blog/what-is-grok-4-6 / https://www.buildfastwithai.com/blogs/grok-4-6-preview ↩ ↩2 ↩3 ↩4
-
Grok Bot entered beta on 11 August 2026. It is described as a team of always-on AI agents with a dedicated cloud environment that sign in to a customer’s existing tools and finish multi-step work without supervision. Desktop builds are offered for Mac and iOS plus Windows and Linux; Android later. It is bundled into existing upper-tier subscriptions (SuperGrok Heavy and others). https://www.unite.ai/xai-launches-grok-bot-always-on-ai-teammates-with-their-own-cloud-computers/ / Grok Build is a coding agent that runs as a command-line tool inside a project folder and takes natural-language instructions against the codebase. It was offered in early beta to SuperGrok Heavy subscribers. https://www.eweek.com/news/xai-grok-build-coding-agent/ ↩ ↩2 ↩3
-
Grok Build release history. Build Mode early beta on 28 July 2026; v1.0.0 on 7 August 2026 (dashboard and CLI fixes). https://releasebot.io/updates/xai ↩ ↩2 ↩3
-
On 2 February 2026 SpaceX acquired xAI in an all-stock exchange and xAI became a wholly owned subsidiary of SpaceX (0.1433 SpaceX shares per xAI share). In May 2026 xAI was dissolved as an independent company and the AI division was renamed SpaceXAI. On 6 July 2026 the new name and logo were published on X. The combined company includes SpaceX, xAI and X. https://finance.yahoo.com/technology/ai/articles/xai-makes-rebrand-spacexai-complete-215010760.html / https://www.socialmediatoday.com/news/spacex-rebrands-as-spacexai/824656/ ↩
-
The benchmark comparison table referred to in this article is the one published by SpaceXAI. A footnote on that table states “Third-party model scores best of self-reported or publicly available.” It needs to be read as a comparison the vendor itself produced. This article sets the third-party Artificial Analysis Intelligence Index alongside it to make up for that. ↩
-
Artificial Analysis Intelligence Index v4.1.1. A composite of nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Principal scores listed: Claude Opus 5 (max) 63, Claude Fable 5 (with fallback) 62, GPT-5.6 Sol (max) 61, Grok 4.6 (high) 61, Kimi K3 (max) 60, Qwen3.8 Max 58, Muse Spark 1.2 (xhigh) 57, GPT-5.6 Terra (max) 57, Grok 4.5 (high) 56, Claude Sonnet 5 (max) 55, GLM-5.2 (max) 53, GPT-5.6 Luna (max) 52, DeepSeek V4 Flash 0731 (max) 52, Gemini 3.6 Flash 52, MiniMax-M3 45, MiMo-V2.5-Pro 43, Inkling 42, Nemotron 3 Ultra 38, Gemini 3.5 Flash-Lite 37, Muse Glimmer (high) 35, Mistral Medium 3.5 30, Gemma 4 31B 30. https://artificialanalysis.ai/ ↩ ↩2
-
Remarks by Elon Musk at a SpaceX all-hands. A film of about 29 minutes was posted on SpaceX’s X account. “we must win on AI, because the future is overwhelmingly AI and robots”; the view that AI revenue will overtake every other business line “probably in September” and pull further ahead in the fourth quarter; the aim of 10 gigawatts of AI compute by the end of 2027; “if we bring 10GW of AI online by the end of next year, it will be $300 billion to $500 billion a year in revenue”; “Probably in four or five years, AI will be 99% of the value of SpaceX”; and the description that this revenue funds Starship and the Mars programme and runs through Terafab, the joint chip plant of Tesla, SpaceX and xAI — all per that reporting. https://www.teslarati.com/spacexs-next-trillion-dollar-bet-has-nothing-to-do-with-rockets-musk-tells-staff/ / At the time of writing these are targets and a view of the future, not results that have been achieved. ↩ ↩2 ↩3 ↩4 ↩5
-
Grok 4.6 API pricing is $2.00 per million input tokens and $6.00 per million output tokens. That is held from the previous generation, Grok 4.5 (shipped 8 July 2026). xAI positions this as “half the price of other frontier models”. https://www.basenor.com/blogs/news/xai-launches-grok-4-6-1753-elo-half-the-price-of-rival-frontier-models / Context length is 500,000 tokens. Once a prompt goes over 200,000 tokens, the rate doubles to $4.00 input and $12.00 output. Web search, X search and code execution are billed separately at $5.00 per 1,000 calls each; collection search at $2.50 per 1,000. On the consumer side SuperGrok Heavy is $300 a month, SuperGrok $30, SuperGrok Lite $10. Sources disagree on the cached-input discount — $0.30 (85% off) versus $0.50 (75% off) — and this article does not treat either as a confirmed figure. https://benchlm.ai/xai/api-pricing / https://www.aipricing.guru/xai-pricing/ ↩ ↩2 ↩3
-
DeepSeek-V4-Flash is $0.14 per million input tokens and $0.28 per million output tokens. DeepSeek-V4-Pro is $0.435 input and $0.87 output. Kimi K2.6 is $0.95 / $4.00 with a 256K context, 80.2% on SWE-bench Verified, and open weights. Kimi K3 is $3 / $15. In the 80.2% to 80.6% band on SWE-bench Verified sit five models — DeepSeek V4 Pro Max, Gemini 3.1 Pro, MiniMax M3, Qwen3.7 Max and Kimi K2.6 — and output prices in that band run from $2.40 per million tokens (MiniMax M3) to $12 (Gemini 3.1 Pro). DeepSeek V4 Pro Max is offered with open weights. https://deepseek.ai/pricing / https://benchlm.ai/moonshot/api-pricing / https://www.morphllm.com/llm-api ↩ ↩2 ↩3
-
Grok Imagine is the family of models that generate images and video. Images are $0.02 each (fast); the quality-first grade is $0.05 for 1K and $0.07 for 2K. Video (1.0 and 1.5) produces clips from 1 to 15 seconds at up to 1080p, with audio generated on the same pass. API pricing is $0.05 per second at 480p and $0.07 at 720p; generating audio on the same pass is $0.08 per second (music, sound effects and dialogue made inside the same generation, without a separate audio step). Grok Imagine Video 1.5 does text-to-video, image-to-video and reference-image-to-video, and generation from text and from images is native 1080p. https://www.aipricing.guru/xai-pricing/ / https://felloai.com/grok-imagine-video-generation/ / https://invideo.io/blog/grok-imagine-ai-generator/ / At the time of writing, Anthropic’s Claude does not offer image, audio or video generation. ↩ ↩2 ↩3
-
Availability of xAI models on Amazon Bedrock. Grok 4.3 became available on Amazon Bedrock in June 2026, and on AWS GovCloud (US-West) in July 2026. In-region inference is available in US West (Oregon), US East (N. Virginia) and US East (Ohio). At the time of writing Asia Pacific (Tokyo, ap-northeast-1) is not included. What is offered on Bedrock is Grok 4.3, not Grok 4.6, which shipped on 12 August 2026. On data retention, AWS states that setting
store=falsealone does not guarantee zero retention; to make zero persistent retention certain, the effective data-retention mode must benoneand the model must be one that permits that mode. https://aws.amazon.com/about-aws/whats-new/2026/06/grok-amazon-bedrock/ / https://aws.amazon.com/about-aws/whats-new/2026/07/grok-4-3-bedrock-govcloud/ / https://docs.aws.amazon.com/bedrock/latest/userguide/models-regions.html / https://www.techi.com/grok-bedrock-portability-retention-regions/ ↩ ↩2 ↩3 -
xAI’s data-retention specification. By default, API requests and responses are kept on the server for 30 days for audit (encrypted at rest); that data is not used for training and is deleted automatically after 30 days. With ZDR (Zero Data Retention) on, the input to an API request (the prompt) and the output (the generated tokens) are not persisted to disk. ZDR is a team-level setting and cannot be turned on for specific API keys only. https://docs.x.ai/developers/faq/security / That ZDR can be turned on as a setting with no extra charge is described from what we have been able to confirm at the time of writing. xAI’s official account has also stated: “For teams using zero data retention, no trace and code data is ever retained. All API key use of Grok Build also respects ZDR. If ZDR is disabled, the /privacy command is available in the CLI to disable data retention, which also deletes previously synced data.” https://x.com/SpaceXAI/status/2076692402442846289 ↩ ↩2 ↩3 ↩4
-
In July 2026 it was reported that Grok Build had been uploading not only the files it had read but entire Git repositories to xAI storage, and that the privacy settings at the time had not worked as expected. https://thehackernews.com/2026/07/grok-build-uploads-entire-git.html / https://www.techtimes.com/articles/320420/20260714/grok-build-shipped-entire-codebases-xai-cloud-privacy-toggle-did-nothing.htm / SpaceXAI’s account is that Grok Build has fully respected ZDR since it shipped, that users have always been able to disable data upload in the CLI, and that that choice has been respected. https://x.com/SpaceXAI/status/2077494536788664782 / The facts remain in conflict between the two accounts, and this article does not decide which is accurate. ↩






