Beat 1: The Deep Pick
Everybody in my group chat has a friend who took a free puppy. Nobody in my group chat has a friend who kept a free puppy for free.
That is Muse Glimmer.
Meta released it on August 10, 2026: a 30 billion parameter model, weights on Hugging Face, Apache 2.0 license, built to run agents on a single consumer GPU. Every headline said “free” and “open source.” Both words are true. Neither one tells you what you are about to pay.
My verdict, so you can stop reading if you need to: if you already own a 24 GB GPU or a 32 GB Mac, download it this week and put it on a real tool-calling job. If you do not own that hardware, do not buy it for this model, and do not rent it hosted at current prices. Details below.
What “free” means here, exactly.
Free license. Apache 2.0 lets you use it commercially, modify it, redistribute it, and ship it inside a product without paying Meta. That is a real change. Every prior Meta open release used a custom Llama license with strings attached, including that 700 million monthly user cutoff people liked to joke about. Apache 2.0 has no such clause. Lawyers relax. Engineers download.
Not a free service. Meta ships no hosted endpoint for Glimmer. There is no per-token fee to Meta because there is no Meta API for it. You either run it yourself or pay a third party like Together AI or Fireworks to run it for you.
The bill driver: memory.
The full precision weights are roughly 60 GB. Meta’s 4-bit quantized build shrinks the language model to under 20 GB, and the model card says that leaves room for the KV cache, the vision encoder, and the speed-boosting “drafter” inside a 24 GB or 32 GB envelope. Independent teardown of the actual files puts the main 17 GB quant at 16.76 GB, plus a 1.4 GB vision projector, plus the drafter. Call it just under 20 GB before you type a single prompt.
So, the puppy needs a yard. An RTX 5090 class card, or an M4 Max or M5 Max Mac with 32 GB or more. If you have that, the marginal cost is electricity and a weekend. If you do not, you are looking at a four-figure hardware purchase, and at that point “free” has stopped meaning anything.
Speed, per Meta’s own numbers on the model card: 74.9 tokens per second on an RTX 5090 without the drafter, 233.4 with it. On an M4 Max, 23.7 without, 37.8 with. Those are vendor figures. Note the gap: the drafter roughly triples speed on a discrete GPU and adds about half again on a Mac.
The hosted path, if you skip the hardware.
Artificial Analysis lists Glimmer at $0.32 per million input tokens and $1.35 per million output tokens across providers as of August 11, 2026, and flags both as expensive against the median for open-weights models of similar size ($0.05 in, $0.15 out). Read that twice. The “free” model, rented, costs more per token than the open models it is competing with. Prices move fast in launch weeks. Re-verify before you commit budget.
Beat 2: Claims vs. Receipts
Meta’s claim (model card, August 10, 2026): Glimmer “performs strongly for its size class” against Gemma 4 31B and Qwen3.6 27B. Meta’s own comparison table bolds Glimmer as the best of the three on roughly half the rows. Standouts: MCP Atlas 75.5 versus Qwen’s 62.5, DeepSearch QA 74.6 versus 71.1, SWE-Bench Pro 51.2 versus 50.2.
What the same table admits: Qwen3.6 27B beats Glimmer on SWE-Bench Verified (77.2 vs 76.0), OSWorld-Verified (75.6 vs 65.9), and TerminalBench 2.1 (60.7 vs 51.7). On GDPval-AA v2, the real-world work test, Glimmer scores 953 to Qwen’s 1141. Meta printed those losses. Credit where due.
A methodology note reviewers flagged: analysts who read Meta’s evaluation report say Meta used whichever was more favorable to the competitor, the competitor’s self-reported score or Meta’s own reproduction, except where Artificial Analysis covered all three models. That is a defensible choice. It is also not the same thing as an independent, like-for-like run.
The independent receipt (Artificial Analysis, August 10 to 11, 2026): Glimmer scores 35 on the Artificial Analysis Intelligence Index. Qwen3.6 27B scores 38. Gemma 4 31B scores 30. Kimi K2.5, a one trillion parameter model, scores 36. Llama 4 Maverick, Meta’s last open release, scored 14.
So, the independent read: Glimmer is a big jump for Meta, a clear win over Gemma at the same size, and a few points behind Qwen. Not the sweep the launch chatter implied. A solid second place in its weight class, with the best license in the room.
Still unverified as of August 23, 2026: Meta’s claim that the 17 GB quant loses only 1.0 percent accuracy versus full precision. I found no independent replication. Treat it as a vendor number until someone outside Meta runs it.
Beat 3: The Call for This Week
Do this: If the hardware is already on your desk, pull the K-Quant-17GB build and run it against a tool-calling workflow you actually use. Judge it on your tasks, not Meta’s table. Apache 2.0 means anything you build on it is yours to ship.
Skip this: Buying a GPU for Glimmer. Renting Glimmer hosted at current listed prices when Qwen3.6 27B scores higher independently and typically costs less per token.
Wait on this: Muse Spark 1.2 open weights. Zuckerberg said “soon” on August 10. No date, no license named. If that lands under Apache 2.0, it is a bigger story than Glimmer. Also wait on the quantization degradation claim until an outside lab reproduces it.
Steal this line: “Free is the sticker on the box. The bill is the box.”
See you next Sunday.


