Sunday AI Signal is the weekly wrap where I take the loudest AI claims of the week, hold them against dated receipts, and tell you what to do, skip, or wait on.
You know that feeling when a download bar says “3 days remaining,” and you just close the laptop lid and go live your life? Tencent released a 770-billion-parameter model this week, and the download is 1.56 terabytes. That is not a download. That is a relationship.
Read this, and you’ll know whether the biggest open model of the year deserves a dollar of your budget, which two chip-and-benchmark headlines to stop repeating at work, and one line to steal for your next meeting.
The verdict up front: Hy4 preview is worth testing through the API and not worth buying hardware for. OpenAI’s chip numbers are real measurements with a self-graded report card. And “OpenAI cuts off Cursor” is smaller than the headline, at least by Cursor’s own count. Every number this week came with a question attached: who ran the test?
THE DEEP PICK: TENCENT HY4 PREVIEW
Picture a hospital with 256 specialists on staff. You walk in with a problem, and only eight of them come into the exam room, plus one general practitioner who sees everybody. That is a mixture-of-experts model. Hy4 preview has 770 billion parameters total, but only 49 billion wake up for any given token. You get the knowledge of the whole hospital and the bill of eight doctors. Now stretch the waiting room to a million tokens of context, which is roughly a full codebase in one conversation, and you have the pitch.
The specs, dated August 28, 2026, from the Hugging Face model card: 770B total, 49B active, 1M-token context, and a built-in speculative-decoding layer so it drafts its own guesses and checks them, which is where the speed comes from. Weights ship in BF16 and FP8, with vLLM and SGLang serving recipes on day one.
Free, paid, freemium, or open source? All four, depending on which door you use.
Open source: the weights are under a standard, unmodified Apache 2.0 license. No monthly-active-user cap, no field-of-use clause, commercial use allowed. Earlier Hunyuan releases carried custom terms; this one does not. Freemium: Tencent’s WorkBuddy and CodeBuddy apps are free for two weeks from August 28. Paid: the API on Tencent Cloud TokenHub and OpenRouter runs $0.834 per million input tokens, $2.501 per million output tokens, and $0.042 per million for cached input (Tencent-listed, August 28; re-verify before it goes in a budget).
What it costs beyond the sticker. Free weights are free the way a puppy is free.
The model card’s serving recipe runs FP8 across eight GPUs, and the FP8 weights alone land near 780 GB, so you need eight cards in the 141 GB class or better. If you don’t own that box, you rent it by the hour, and that hourly rate is your real cost. Self-hosting pays off at sustained volume or when data can’t leave the building. Below that, the API wins.
Per-task math, because per-token prices lie by omission. Take a long agentic coding session: 3 million input tokens (mostly the same context re-sent each turn) and 150,000 output tokens. Uncached, that session runs about $2.88. At an 80 percent cache-hit rate, it drops to about $0.98. The cost driver is how much context you re-send and whether it hits cache. One 1M-token “read my whole repo” call is about 83 cents of input before the model says a word.
CLAIMS VS RECEIPTS
Claim 1: Hy4 preview “beats GLM 5.3 and Kimi K3.” (Tencent, August 28) Receipt: the win comes from Tencent’s internal blind evaluation, where 163 employees rated 203 engineering tasks and Hy4 scored 2.99 out of 4 against 2.92 and 2.94. It lost roughly 40 percent of head-to-head matchups in that same test. On public GPQA Diamond, coverage of Tencent’s own benchmark appendix puts Hy4 at 92.3 and Kimi K3 at 93.5. Artificial Analysis had no independent Intelligence Index score for Hy4 preview as of August 30.
Claim 2: OpenAI’s Jalapeño chip delivers “1.5 to 1.9x more work per watt than Nvidia.” (OpenAI at Hot Chips, August 25) Receipt: OpenAI ran the tests. SemiAnalysis engineers watched in OpenAI’s lab but did not execute the suite themselves, per Forbes and Tom’s Hardware. The comparison targets were Nvidia’s GB200 and GB300, which carry prior-generation memory, not the Rubin parts that would be the fair match. Tom’s Hardware reports an appendix comparison using all-in utility power narrows the gap, and against a GB300 running multi-token prediction the lead shrinks to roughly 1.5x. The part tested was early-stepping silicon, and volume deployment is a 2027 story. The measurements are real. The framing belongs to the vendor.
Claim 3: “OpenAI cuts off Cursor.” (OpenAI statement, August 28) Receipt: OpenAI says it will wind down the contract by November 12, 2026, citing what it describes as a history of Musk-owned companies violating contracts. The notice came two weeks after SpaceX closed its roughly $60 billion purchase of Cursor’s parent. Cursor CEO Michael Truell replied August 29 that OpenAI models serve about 5 percent of Cursor’s user traffic. That figure is company-provided and unaudited, but it moves the story from “Cursor loses its brain” to “Cursor loses a tab.” Existing models run until the cutoff. If your team standardized on GPT inside Cursor, you have about ten weeks to pick a model, an IDE, or a direct API key.
ALSO ON THE TAPE
A federal judge ruled Thursday night, August 27, that the Pentagon’s supply-chain-risk label on Anthropic was unlawful retaliation for the company’s criticism; the government is expected to appeal (Fortune/AP, August 28). Sony Music Publishing and Warner Chappell sued Anthropic on August 28, alleging lyrics were used in training; that is a complaint, not a finding (Music Business Worldwide). Nvidia paused its $36 billion AI Compute Partnership under two months after launch over antitrust concerns raised by its own employees, per the Wall Street Journal (August 29). And a YouGov survey of 1,250 U.S. workers found 3 percent said they lost a job to AI since 2023 (Fortune, August 29).
THE SO WHAT
Do this: Route Hy4 preview through OpenRouter into whatever evaluation set you already run this week, while the free window is open. Log your cache-hit rate. That one number decides whether it is cheap for you.
Skip this: Buying an eight-GPU box because the weights are “free.” And repeating the 1.9x Jalapeño figure without saying who measured it.
Wait on this: Production traffic on Hy4 until Artificial Analysis posts an independent score, and any Nvidia conclusion from Jalapeño until someone independent runs it against Rubin.
Steal this line: “Every benchmark comes with a question attached: who ran the test?”
Every number here is dated and will move. Re-verify pricing, the Artificial Analysis listing, and the Cursor shutoff date before you build on them.
SOURCES
Tencent Hy4-preview-FP8 model card, Hugging Face (Aug 28, 2026)
Tencent press release: Tencent Releases and Open-Sources Tencent Hy4 preview (Aug 28, 2026)
DataLearner: Hy4 preview benchmark and pricing summary (Aug 28, 2026)
Data Science in Your Pocket: Tencent Hy4 Preview vs GLM 5.3, Qwen 3.8, Kimi K3 (Aug 28, 2026)
Artificial Analysis Intelligence Index v4.1.1 (checked Aug 30, 2026)
Tom’s Hardware: OpenAI’s 700W Jalapeño ASIC outpaces Nvidia flagship GPU (Aug 26, 2026)
Forbes: OpenAI Publishes First Jalapeño Benchmarks Against Nvidia Blackwell (Aug 27, 2026)
OpenAI: Our decision on Cursor following its acquisition by SpaceX (Aug 28, 2026)
CNBC: OpenAI to end model access to Cursor after acquisition by SpaceX (Aug 29, 2026)
The Decoder: OpenAI cuts off Cursor after SpaceX acquisition (Aug 29, 2026)
Fortune/AP: Judge: Pentagon punished Anthropic for ‘arrogance,’ and that’s illegal (Aug 28, 2026)
Music Business Worldwide: Sony Music Publishing and Warner Chappell sue Anthropic (Aug 28, 2026)
Yahoo Finance/WSJ: Nvidia pauses $36B AI cloud financing program (Aug 29, 2026)
Fortune: Only 3% of US workers lost a job to AI since 2023, YouGov survey (Aug 29, 2026)
AI Weekly: AI News Today, August 30, 2026 (aggregator used for story discovery only)


