Think of this week in AI like a group chat where everyone keeps forwarding the same article. My job on Sunday is to be the friend who actually opened it, read the whole thing, and tells you straight whether it holds up.
This week it mostly didn’t. Not because the news was fake, but because the good part got left out.
Here’s your verdict up front. Grok 4.6 is a real upgrade and it is genuinely the cheapest model at the intelligence frontier right now, at $2 per million input tokens and $6 per million output. That part is true. But the cheap price only holds while your prompt stays under 200,000 tokens. Cross that line and xAI bills the entire request at double, $4 and $12. And long, context-heavy work is the exact thing xAI is selling this model for. So if you run short prompts, this is a steal. If you run long-horizon agents, you need to do math before you trust the sticker. That’s the whole edition in four sentences. Now the receipts.
First, a way to think about it that will stick.
Grok 4.6’s pricing works like a happy hour menu. Cheap drinks, great deal, everybody’s happy. But there’s a cutoff time posted by the door, and here’s the twist that makes it worse than a real bar. At a real happy hour, when the clock runs out, only the next drink costs more. With Grok 4.6, the moment you cross the line, the bartender reprices your entire tab at full rate. The beer you drank an hour ago? Full price now too. That’s what “$4/$12 across the entire request” means. Not the tokens past 200K. All of them.
Hold onto that, because it changes the answer to “should I switch?”
Beat one: the one thing that actually mattered this week.
xAI shipped Grok 4.6 on August 12, 2026, with an underlying checkpoint dated August 10. The API id is just grok-4.6. It landed day one on Cursor and Grok Build, with API access through console.x.ai plus OpenRouter, Vercel, and Cloudflare. No EU region at launch, so if you’re in the EU, you’re waiting.
The confirmed specs are clean. A 500,000-token context window. Knowledge cutoff of February 1, 2026. Text and image in, text out. Reasoning levels from low up to a new “xhigh.” On the Artificial Analysis Intelligence Index, it scores 61, which matches OpenAI’s GPT-5.6 Sol and sits one point behind the top Claude. That combination, frontier-level score at the lowest frontier price, is the entire pitch. xAI’s Michael Truell framed it as Opus-class intelligence at low cost and high speed.
Now the parts the launch post was quiet about.
One, the pricing cliff. The $2/$6 rate covers prompts under 200K tokens. At or above that threshold, xAI’s own pricing page reprices the whole request to $4/$12. For a model marketed around long-running agents and big-context work, the discount evaporates precisely where you’d want to use it (verify against xAI’s live pricing page before you budget, as of August 15, 2026).
Two, the quiet cache hike. Cached input rose from $0.30 to $0.50 per million tokens, roughly 67% higher than Grok 4.5. Small number, but if your workload leans on caching, that’s a real line item, and it wasn’t in the headline.
Three, the context window did not grow. It was already 500K on Grok 4.5. The upgrade is how well the model uses long context, not a bigger limit. A few pieces of coverage framed the 500K like a new feature. It isn’t.
Four, the benchmarks are xAI’s own. As of August 13, 2026, no independent third party had replicated the numbers. On xAI’s ten-row comparison table, reviewers point out that the top Claude actually wins the most rows, GPT-5.6 Sol takes the coding-specific ones and Grok 4.6 loses Terminal-Bench v3.0 by about 8.6 points, a result the launch prose skipped over. Strong model. Selectively presented.
Cost and licensing, plainly: Grok 4.6 is paid and proprietary. Closed weights, per-token API. $2/$6 under 200K, $4/$12 at or above it, cached input $0.50, and a faster variant at twice the price. The first-week promo gives 2x included usage inside Grok Build and Cursor, but xAI never published the base usage number, so “2x” is a nudge, not something you can put in a budget.
Beat two: claims versus receipts.
Two more launches got a coat of paint this week. Let’s scrape it off.
Meta Muse Glimmer, the “free” local agent. This one’s mostly good news, which is why the framing matters. Meta Superintelligence Labs released Muse Glimmer on August 10, 2026, a roughly 29.6-billion-parameter multimodal model, and here’s the part that’s genuinely rare: the weights ship under an unmodified Apache 2.0 license. Not a Llama-style custom license with fine print. The real thing. You can build a commercial product on it with no royalty obligation. Mark Zuckerberg and Meta’s Alexandr Wang both pitched it as running locally on a single consumer GPU.
So, where’s the receipt? In the word “free.” The license is free. Running it is not. Wang said it fits in 24GB of VRAM, and Meta targets 24GB hardware with a compressed 4-bit build. But reviewers who read the model card note that the small “under 20GB” figure describes the language weights alone. A working multimodal agent also needs the vision encoder, the KV cache, runtime overhead, and the optional speed drafter. Translation: you need a real GPU, which is real money, or a cloud rental that bills by the hour. Free license, paid hardware. And Meta’s benchmark wins are Meta’s own, mixed against rivals, leading on some agent tests, trailing Qwen on others. Still, for local agent work, this is the most interesting release of the week. Just don’t read “free” as “no cost.”
Qwen3.8-Max, the number that isn’t a number yet. You may have seen a $2/$6 per million price for Qwen3.8-Max floating around X, next to a claim that it’s “second only to” the top Claude. Both are unconfirmed. The pricing figure circulating on X is unverified, the “second only to” ranking has no third-party backing, the model previewed back on July 19 rather than launching in August, and there’s no published benchmark table and no open-weight date. There may be a great model here eventually. Right now there’s a screenshot and a vibe. Nothing to act on.
Beat three: skip, wait, steal for the week.
Steal this: Muse Glimmer, if you do local or agent work and you own or can rent a 24GB GPU. A genuinely Apache 2.0 agentic model that runs on one machine is the first realistic default for local agents. Worth an actual pilot this week. Budget the hardware honestly and test it in your own scaffold, not on Meta’s benchmark table.
Skip this: swapping your whole stack to Grok 4.6 because of the “cheapest at the frontier” headline. The score that backs that headline is xAI-reported and, as of August 13, unreplicated, and the model loses the coding-specific evals. Don’t rip out a working setup for a launch-day number.
Wait on this: two things. Qwen3.8-Max until there are real verifiable numbers instead of an X screenshot. And Grok 4.6 for long-context agent work until you’ve modeled the 200K cliff against your actual token usage. If your prompts routinely run long, “cheap” might be the most expensive option on the menu.
That’s the week. One real bargain with a catch, one real freebie that costs money, and one number that doesn’t exist yet. See you next Sunday.
Sources
Grok 4.6 launch, pricing, and the 200K threshold — DigitalApplied
Qwen3.8-Max unverified pricing and ranking claims — AIToolsRecap
Numbers dated to August 12 through 15, 2026. Pricing and benchmark figures shift fast, re-verify against each vendor’s live page before you budget.


