Kimi K3, GLM-5.3-Flash, and Qwen3.8-Flash-Next versus Claude Fable 5, plus the flagship that dropped Friday. Every figure dated August 31, 2026.
You want Claude Fable 5 behavior without the Claude Fable 5 invoice. By the end of this you’ll know which open-weight model belongs in the expensive seat, which one belongs in the cheap seats, and which launch-week claim should not survive contact with a calendar.
Verdict first.
No single open-weight model is Fable. The closest single model is Kimi K3, which scores 60 on the Artificial Analysis Intelligence Index v4.1.1 to Fable’s 62, as of August 31. Two points. That’s the good news. (Z.ai’s GLM-5.3 flagship went open-weight on Friday and also scores 60. More on that below.)
The closest system is not one model at all. It’s GLM-5.3-Flash doing the work, Kimi K3 doing the planning and the final review, and Qwen3.8-Flash-Next brought in only when the agent has to look at a screen.
The rest of this is the receipts, and one number that made me put my coffee down.
What “Fable-like” even means
Anthropic sells Fable 5 as a model that runs long, ambiguous, multi-step jobs with few check-ins. A 1M-token context, up to 128K output tokens, $10 per million input and $50 per million output, June 9 launch.
So “Fable-like” is not a score. It’s two things.
Score is the leaderboard number.
Stamina is whether the model is still making sense at turn 80.
Twenty years in intelligence taught me analysts were never graded on how sharp they sounded in the morning brief. They were graded on whether the assessment held up six months later. Same test here. ## First, a word about the word “open-source”
I asked this question as “which open-source model gets me closest.” That framing was wrong and I’m fixing it in public.
None of these three are open source in the strict sense. They’re open weight, and the licenses differ.
GLM-5.3-Flash ships under MIT. That’s the cleanest of the three.
Kimi K3 ships under the Kimi K3 License. MIT-shaped, with conditions: run a model-as-a-service business past $20 million in revenue over any 12 months and you need a separate agreement with Moonshot. Cross 100 million monthly users or $20 million in monthly revenue and “Kimi K3” has to appear on your product’s interface.
Qwen3.8-Flash-Next ships under the Qwen Community License 1.0. Not Apache. Similar shape to Kimi’s terms.
Since Friday there’s a fourth example. GLM-5.3, Z.ai’s 753B flagship, went open-weight on August 28 under a new GLM-5.3 License. It reads like MIT until the clause that says companies with over $10 billion in revenue across any 12 months need a Z.ai security review before hosting it commercially. Z.ai’s own model cards use both “open-weights” and “open source” for it. Only one of those is accurate.
And open weights are not free compute. Kimi K3 is 2.8 trillion parameters. Third-party reporting puts the download near 594 GB, and Northflank describes a 64-accelerator deployment. GLM-5.3-Flash needs roughly 306 GiB of FP8 weights on Hopper-class hardware or newer. Qwen3.8-Flash-Next sits at about 180B parameters on disk once you count its 51B n-gram table.
GLM-5.3’s full-precision download is about 756 GB.
Translation: you and I are renting these through an API. Which is why the pricing section matters more than the parameter count.
Kimi K3: closest on paper, and the number that hurt
Moonshot released K3 on July 16 and posted the weights July 27. It’s 2.8T total, 104B active, native vision, 1M context.
API price on August 31: $3 per million input, $15 per million output, $0.30 cached. Next to Fable’s $10 and $50, that reads like 70% off.
Now the ugly number.
Artificial Analysis runs a private benchmark called AA-Briefcase, realistic knowledge-work tasks with real file inputs. Kimi K3 scored an Elo of 1543 there, second only to Fable at 1574. Great result. Then AA published the bill: $10.57 per task, 83 turns per task versus 67 for Fable, and an average of 56.4 minutes per task, roughly 2.5 times Fable’s time.
Cheaper per token. Not cheaper per task. Because K3 takes more turns and writes more output to get there.
Moonshot’s own model card, July 2026, shows where it wins. Terminal-Bench 2.1: 88.3 to Fable’s 88.0. SWE-Marathon: 42.0 to 35.0, with Moonshot noting Fable hit safety fallbacks on 35% of those tasks. BrowseComp: 91.2 to 88.0, though that 91.2 uses context compaction at 300K tokens.
And where it trails. Toolathlon-Verified: 76.5 to 77.9. JobBench: 54.3 to 57.4. GDPval-AA v2: 1686 Elo to 1747.
Two operational quirks from Moonshot’s own docs: K3 requires you to pass its full reasoning content back on every turn, and Moonshot describes it as “proactive by nature,” which is vendor-speak for “it will make decisions you didn’t ask for unless you fence it in.”
Score: near. Stamina: it finishes, but you pay for the scenic route.
GLM-5.3-Flash: two days old, and the cheap seats
Z.ai released GLM-5.3-Flash on August 26, after a week of running it anonymously on OpenRouter as “Ox Alpha.” It’s 320B total, 18B active, natively multimodal, 1M context, MIT weights.
List price: $0.15 per million input, $0.50 per million output, $0.03 cached. A launch promotion halves that through September 9, 2026, per multiple pricing trackers; Z.ai’s own page cites the discounted per-task figure.
Run the ratio against Kimi. Twenty times cheaper on input. Thirty times cheaper on output.
Artificial Analysis scores it 57 on the Intelligence Index at $0.09 per index task, and reports the entire index cost $138.02 to run. Same score as GPT-5.6 Terra, per AA, at about 5.7 times lower cost per task. Z.ai’s launch numbers: Terminal-Bench 2.1 at 84.3, DeepSWE 1.1 at 63.4, AutomationBench at 48.8, and Z.ai describes the whole package as “approaching Claude Opus 4.8.” Vendor figures, dated August 26.
Two catches. AA now measures 45 tokens per second on Z.ai’s API (down from 50.2 on Friday) and flags it as notably slow. And thinking cannot be disabled, so you pay for reasoning tokens on every call, even the dumb ones.
Z.ai also says the entire preview week ran on Chinese AI chips.
GLM-5.3: the Friday arrival
Not in my original brief, because it wasn’t downloadable when I wrote it. Z.ai announced GLM-5.3 on August 14, held the weights two weeks for a safety review over its cybersecurity results, and posted them August 28.
It’s text-only, 753B total with roughly 40B active, 1M context, 128K output. API price: $1.40 input, $4.40 output, $0.26 cached, unchanged from GLM-5.2. AA scores it 60, the same as Kimi K3, at $0.68 per index task, and notes it generated 170M tokens running the index, well above the median. Verbose is the polite word.
So the planner seat has two candidates at 60. Kimi K3 has vision, a Fable comparison in its own model card, and a measured AA-Briefcase bill. GLM-5.3 has output tokens at roughly a third of Kimi’s price and no AA-Briefcase number yet, so I can’t tell you what a real task costs. If your agent never looks at a screenshot, test it in that seat and measure.
Qwen3.8-Flash-Next: the one that can see
Alibaba’s weights went live August 24, formal release August 26. Only 6B active parameters. AA scores it 56 and measures 73.4 tokens per second, the fastest of the three.
Its superpower is the screen. Vendor numbers: AndroidWorld 84.5, OSWorld 2.0 at 52.3 partial (19.4 binary), Toolathlon-Verified 73.5.
Pricing has a small conflict. Qwen’s launch announcement lists the production Qwen3.8-Flash at $0.16 input and $0.47 output. Artificial Analysis lists $0.15 and $0.47. I’m showing both.
And the context caveat: the open weights run 262K natively, extendable to 1M with YaRN. The hosted Qwen3.8-Flash gets 1M by default.
Now look at who Qwen’s launch table compares against: Claude Opus 4.6 Max, a February 2026 model. Not Fable. Not even Opus 4.8.
The build
Anthropic already published the pattern. Their advisor strategy runs Sonnet 5 as the executor and calls Fable 5 only for guidance. Per Anthropic’s own results, that pair hit about 92% of Fable’s standalone SWE-bench Pro performance at roughly 63% of the cost.
Flip it onto open weights.
GLM-5.3-Flash does the typing: research, tool calls, code, document work. Kimi K3 writes the plan, gets consulted on hard decisions, and reviews the final output. Qwen3.8-Flash-Next joins the moment a screenshot enters the loop.
Cost drivers, in order: the planner’s output tokens and turn count (Kimi at $15 per million, GLM-5.3 at $4.40; cap the advisor’s output either way), Flash’s mandatory reasoning tokens, the September 9 promo cliff, and the harness engineering nobody budgets for.
Before anyone asks: Claude Opus 5, released July 24, scores 63 on the same index, one point above Fable. I’m holding Fable as the target because it’s the model Anthropic built for long-horizon autonomy and the one in its advisor pattern. Stamina over score.
Claims vs. Receipts
Claim: “Kimi K3 scores 57 on Artificial Analysis.” Most July coverage said this. Receipt: 60, as of August 31. AA moved to index v4.1.1, and Fable moved too, from 59.9 to 62. Both models shifted. Always check the version number on the leaderboard.
Claim: “Kimi K3 is 70% cheaper than Fable.” Receipt: Per token, yes. Per AA-Briefcase task, $10.57 and 56 minutes, among the most expensive models on that benchmark.
Claim: “Qwen3.8-Flash-Next beats Claude.” Receipt: Beats Claude Opus 4.6 Max on several vendor-run benchmarks. Opus 4.6 is a February 2026 model. Of the three labs, only Moonshot puts Fable 5 in its comparison table.
Claim: “GLM-5.3 is open source.” Z.ai’s own materials. Receipt: Open-weight under a custom license with a revenue-gated security review. The Flash sibling is MIT. The flagship is not.
Correction to my own brief: I called these open-source models. They are open-weight, and three of the four now carry commercial conditions.
So what
Do this: Build with GLM-5.3-Flash as the worker and Kimi K3 as advisor and reviewer. If your tasks are text-only, run GLM-5.3 in the advisor seat as a second test. Cap the advisor’s output. Measure cost per finished task, never per token.
Skip this: Kimi K3 as your only model on high-volume work. The turn count will eat the discount.
Wait on this: Locking in GLM-5.3-Flash pricing in your budget before September 9. Budget the list price, not the promo.
Steal this line: “Per token is the sticker. Per task is the bill.”
The gap between open weights and Fable is now small enough that the interesting question has moved. It’s no longer who has the smartest model. It’s who builds the best harness around the models they can afford.
Sources
Anthropic: Introducing Claude Fable 5 and Claude Mythos 5 (Claude Platform docs)
Anthropic webinar: Claude Fable 5 and model orchestration patterns
The Decoder: Anthropic’s advisor pattern, Sonnet 5 plus Fable 5 results
Hugging Face: moonshotai/Kimi-K3 model card and benchmark table
Artificial Analysis: Kimi K3 on AA-Briefcase, cost and time per task
LLM Stats: GLM-5.3-Flash launch, list price, and promo end date
MarkTechPost: GLM-5.3-Flash release and self-hosting requirements
Hugging Face: Qwen/Qwen3.8-Flash-Next model card and benchmark tables
The Decoder: Alibaba releases Qwen3.8-Flash-Next, launch pricing
CellCog: Qwen3.8-Flash-Next license terms and production API status
Hugging Face: zai-org/GLM-5.3 model card, released August 28
Artificial Analysis: GLM-5.3 (max) model page, score and pricing
Artificial Analysis on X: GLM-5.3-Flash at $0.09 per index task
BenchLM: Artificial Analysis Intelligence Index mirror, updated August 29


