<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The AI Signal]]></title><description><![CDATA[AI moves fast and lies faster. I check the receipts so you don't have to. What to do: skip or wait every single week.
Vendor says one number. Reality says another. I publish both, with dates, so you can make decisions instead of bets.]]></description><link>https://aisignal.veletica.com</link><image><url>https://substackcdn.com/image/fetch/$s_!WukZ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0de8dea7-9582-4ddf-a335-38239d8e01a9_1493x1493.jpeg</url><title>The AI Signal</title><link>https://aisignal.veletica.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 09 Sep 2026 02:35:38 GMT</lastBuildDate><atom:link href="https://aisignal.veletica.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Oscar Villa]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[aitrendzio@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[aitrendzio@substack.com]]></itunes:email><itunes:name><![CDATA[Oscar Villa]]></itunes:name></itunes:owner><itunes:author><![CDATA[Oscar Villa]]></itunes:author><googleplay:owner><![CDATA[aitrendzio@substack.com]]></googleplay:owner><googleplay:email><![CDATA[aitrendzio@substack.com]]></googleplay:email><googleplay:author><![CDATA[Oscar Villa]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Real AI Price War Moved to Cost Per Finished Job. ]]></title><description><![CDATA[One Human, 3.1 Digital Workdays: The Week AI Got a Job.]]></description><link>https://aisignal.veletica.com/p/the-real-ai-price-war-moved-to-cost</link><guid isPermaLink="false">https://aisignal.veletica.com/p/the-real-ai-price-war-moved-to-cost</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Mon, 07 Sep 2026 03:28:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PHud!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PHud!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PHud!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PHud!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PHud!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PHud!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PHud!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:554765,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/214514265?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PHud!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PHud!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PHud!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PHud!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F34a2547a-8b6e-4527-bdb6-f1946c55523e_1920x1072.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Sunday AI Signal is the weekly wrap-up: the biggest AI moves of the week, verified against primary sources, turned into decisions you can use Monday morning.</em></p><p>Something happened across six days this week that no single headline captured, and if you run a team, a business, or just your own overloaded calendar, it hands you a cleaner way to plan the next quarter. Stay with me to the end, and you&#8217;ll leave with one metric to budget against and one experiment to run before Friday.</p><p>The verdict up front: the story of this week is not that AI got smarter. It&#8217;s that the three ingredients of useful delegation landed in the same window. Models that can finish real computer work (GPT-6 Astra, September 3). Agents packaged as standing jobs with governance attached (Grok Bot for Enterprise, September 3). And economics cheap enough to leave those agents running around the clock (Claude Fable 5.1&#8217;s cache price cut on September 1, Gemini 3.8 Flash on September 2). Stop grading AI on answer quality. Start grading it on how many bounded pieces of work you can safely hand off, in parallel, per dollar.</p><p>I read four launch posts, a system card, and Cursor&#8217;s entire security documentation this week so your Monday plan doesn&#8217;t have to run on a press release.</p><p><strong>What this actually is</strong></p><p>For three years we&#8217;ve all been shopping for a smarter calculator. Which chatbot gives the best answer? This week the market started selling something different: staff. And staff comes with the two things calculators never needed. Job descriptions and salary bands.</p><p>The mental model that makes all of it click is the work package. A work package is a task written the way you&#8217;d brief a contractor: the goal, what they&#8217;re allowed to touch, where they must stop, and what finished looks like. Every release this week is infrastructure for exactly that. Astra is the senior specialist who can operate your software. Grok Bot is the HR system that turns a task into a persistent role. Fable 5.1 and Gemini Flash are the salary bands, one for the expensive closer and one for the tireless intern.</p><p><strong>Beat one: the deep pick. The delegation stack</strong></p><p>The pick this week isn&#8217;t a single product. It&#8217;s the stack, and every layer of it is paid. No open-source entry here, no free tier worth building on. So the cost drivers matter more than the scores.</p><p>GPT-6 Astra (paid API, $10 per million input tokens, $50 per million output) is OpenAI&#8217;s case that a model can carry a task end to end. In OpenAI&#8217;s own latency simulation, it scored 72.6% on OSWorld 2.0 computer use at roughly 40 minutes per task, versus 65.7% and roughly 75 minutes for GPT-5.6 Sol. Its AutomationBench score jumped from Sol&#8217;s 18.1% to 41.4%, and that benchmark tries to measure professional automation, which is why I care about it more than another math record. One warning attached: Astra is the first OpenAI model classified at the Critical cybersecurity threshold under its Preparedness Framework, meaning OpenAI says it can find and exploit unknown vulnerabilities without step-by-step guidance. Its exploit capabilities ship gated.</p><p>Grok Bot for Enterprise (paid, from $120 per seat per month on Cursor Premium Teams, with metered token overage on top) is the operating model. You don&#8217;t build workflows. You create Bots with jobs: a procurement Bot that watches vendor spend, a recruiting Bot that builds overnight shortlists. The September 3 enterprise release added the three checkboxes security reviews stall on: access controls, network controls, and audit logging. SpaceXAI says its procurement Bot has surfaced tens of thousands of dollars in SaaS savings. That&#8217;s a vendor case study, not independent validation, but the shape of the workflow is the point. You&#8217;re delegating a responsibility, not automating a script.</p><p>Claude Fable 5.1 and Gemini 3.8 Flash are the economics. Anthropic kept Fable&#8217;s $10/$50 headline price but cut cached-context reads 75%, to $0.25 per million tokens, and cached context is what agents burn as they reread instructions, repos, and their own history. Anthropic estimates typical workloads get about 25% cheaper and highly agentic ones up to about 45%. Those are vendor estimates. Google attacked from below: Gemini 3.8 Flash at an introductory $0.75 per million input and $3.75 per million output, positioned for long-horizon coding and autonomous agents. Google&#8217;s own table puts it at 73.7% on DeepSWE v1.1, in frontier territory at a fraction of frontier price.</p><p>The winning architecture drops out of the price list: cheap models run the loops, expensive models handle escalation. Flash-class economics for research passes, checking, monitoring, and iteration. Astra-class models for ambiguous decisions and final deliverables. The future isn&#8217;t picking the best model. It&#8217;s managing a workforce with different salary bands.</p><p><strong>Beat two: claims vs. receipts</strong></p><p><em>Claim one: &#8220;the AGI era.&#8221;</em> OpenAI president Greg Brockman suggested on September 3 that we may now be in it. The independent receipt, dated the same week: Artificial Analysis scores Astra 61.2 on its Intelligence Index, effectively tied with its predecessor and behind Claude Fable 5.1&#8217;s 65.7. And Astra&#8217;s headline 99.9% on ARC-AGI-3 came from a souped-up harness; on the standard harness it scored 66%, per Fortune&#8217;s September 3 reporting. Astra is a real leap in computer use. The AGI framing is marketing.</p><p><em>Claim two: &#8220;Each Bot runs on its own computer in the cloud.&#8221;</em> That&#8217;s SpaceXAI&#8217;s September 3 enterprise announcement. Cursor&#8217;s own security documentation, read this week, draws the boundary differently: isolation is per user, not per Bot. Your Bots operate as you, on your durable environment, and Cursor states plainly that its prompt-injection defenses reduce risk rather than eliminate it. There&#8217;s also no customer-facing model picker and, per Cursor&#8217;s teams documentation, no Grok Bot-specific spend cap yet. Separate Bots are not separate security boundaries. Plan accordingly.</p><p><em>Claim three: Gemini Flash&#8217;s &#8220;same price.&#8221;</em> True on the sticker, dated September 2. Two receipts: the intro rate doubles to $1.50/$7.50 on January 1, 2027, and thinking tokens bill at the output rate, which Google&#8217;s own docs acknowledge and which independent guides note can push real costs to roughly twice a naive estimate on hard tasks. Cheap is real here. Free of fine print, it is not.</p><p>One more number, because it&#8217;s the clearest signal of where this goes. OpenAI published internal data on September 6: by mid-August its research organization was consuming 3.1 agent-workdays for every human workday, with the median researcher spending over $600 a day on inference. Self-reported and unaudited, but paired with chief scientist Jakub Pachocki writing the same day that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. The company running fastest published both the accelerator and the brake in one afternoon.</p><p><strong>Beat three: skip, wait, steal</strong></p><p><strong>Skip:</strong> the AGI debate, and any plan that routes every agent step through a $50-per-million-output model.</p><p><strong>Wait:</strong> on Grok Bot for anything touching finance, customer data, or credentials until your security team reads Cursor&#8217;s docs and you&#8217;ve priced the uncapped overage. Wait on committing volume to Gemini Flash pricing past December 31 without modeling the doubled rate.</p><p><strong>Steal:</strong> the work package. Goal, permissions, boundaries, completion criteria, escalation points. It costs nothing, and it&#8217;s the skill every one of this week&#8217;s releases assumes you have.</p><p><strong>So what</strong></p><p><strong>Do this:</strong> pick one 2-to-4 hour task this week, write it as a work package, and delegate everything except the irreversible decision.</p><p><strong>Skip this:</strong> buying the most intelligent model for every step. Route cheap, escalate expensive.</p><p><strong>Wait on this:</strong> persistent agents on systems your business depends on, until the audit trail and spend cap questions have real answers.</p><p><strong>Steal this line:</strong> &#8220;How many competent digital workers can one person economically supervise?&#8221; That&#8217;s the question your Q4 planning should already be asking.</p><p><strong>Sources</strong></p><ul><li><p><a href="https://openai.com/index/gpt-6-astra/">OpenAI: GPT-6 Astra announcement (Sept 3, 2026)</a></p></li><li><p><a href="https://deploymentsafety.openai.com/gpt-6-astra">OpenAI: GPT-6 Astra System Card, Critical cyber classification (Sept 3, 2026)</a></p></li><li><p><a href="https://openai.com/index/research-acceleration-view-inside-openai/">OpenAI: Research acceleration, the 3.1 agent-workdays data (Sept 6, 2026)</a></p></li><li><p><a href="https://fortune.com/2026/09/03/openai-debuts-gpt-6-astra-computer-use-greg-brockman-says-start-of-agi/">Fortune: Astra&#8217;s ARC-AGI-3 harness vs. standard scores (Sept 3, 2026)</a></p></li><li><p><a href="https://www.anthropic.com/claude-fable-and-mythos-5-1">Anthropic: Claude Fable 5.1 and Mythos 5.1 announcement (Sept 1, 2026)</a></p></li><li><p><a href="https://platform.claude.com/docs/en/about-claude/pricing">Claude Platform docs: Fable 5.1 cache pricing</a></p></li><li><p><a href="https://x.ai/news/grok-bot-for-enterprise">SpaceXAI: Grok Bot for Enterprise (Sept 3, 2026)</a></p></li><li><p><a href="https://cursor.com/docs/grok-bot/security">Cursor docs: Grok Bot security architecture</a></p></li><li><p><a href="https://cellcog.ai/blog/gemini-3-8-flash/">CellCog: Gemini 3.8 Flash specs, pricing, and Google&#8217;s benchmark table (Sept 2026)</a></p></li><li><p><a href="https://emergent.sh/learn/gpt-6-astra-benchmarks">Emergent: Astra vs. Fable 5.1 on the independent Artificial Analysis index (Sept 2026)</a></p></li><li><p><a href="https://alphasignal.ai/news/xai-pushes-grok-bot-into-enterprise-with-audit-controls-and-free-trials">AlphaSignal: Grok Bot enterprise rollout coverage (Sept 4, 2026)</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Kimi K3, GLM-5.3-Flash, or Qwen3.8: Which One Is Fable in a Cheaper Jacket?]]></title><description><![CDATA[Three Chinese Labs, One Anthropic Model, and the $10.57 Task That Changes the Math]]></description><link>https://aisignal.veletica.com/p/kimi-k3-glm-53-flash-or-qwen38-which</link><guid isPermaLink="false">https://aisignal.veletica.com/p/kimi-k3-glm-53-flash-or-qwen38-which</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Mon, 31 Aug 2026 12:10:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JPLF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JPLF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JPLF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JPLF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2874460,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/213536154?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JPLF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!JPLF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42d89d3b-6455-4c3f-bc61-7584f8b620a0_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Kimi K3, GLM-5.3-Flash, and Qwen3.8-Flash-Next versus Claude Fable 5, plus the flagship that dropped Friday. Every figure dated August 31, 2026.</em></p><p>You want Claude Fable 5 behavior without the Claude Fable 5 invoice. By the end of this you&#8217;ll know which open-weight model belongs in the expensive seat, which one belongs in the cheap seats, and which launch-week claim should not survive contact with a calendar.</p><p>Verdict first.</p><p>No single open-weight model is Fable. The closest single model is Kimi K3, which scores 60 on the Artificial Analysis Intelligence Index v4.1.1 to Fable&#8217;s 62, as of August 31. Two points. That&#8217;s the good news. (Z.ai&#8217;s GLM-5.3 flagship went open-weight on Friday and also scores 60. More on that below.)</p><p>The closest system is not one model at all. It&#8217;s GLM-5.3-Flash doing the work, Kimi K3 doing the planning and the final review, and Qwen3.8-Flash-Next brought in only when the agent has to look at a screen.</p><p>The rest of this is the receipts, and one number that made me put my coffee down.</p><h2>What &#8220;Fable-like&#8221; even means</h2><p>Anthropic sells Fable 5 as a model that runs long, ambiguous, multi-step jobs with few check-ins. A 1M-token context, up to 128K output tokens, $10 per million input and $50 per million output, June 9 launch.</p><p>So &#8220;Fable-like&#8221; is not a score. It&#8217;s two things.</p><p><strong>Score</strong> is the leaderboard number.</p><p><strong>Stamina</strong> is whether the model is still making sense at turn 80.</p><p>Twenty years in intelligence taught me analysts were never graded on how sharp they sounded in the morning brief. They were graded on whether the assessment held up six months later. Same test here. ## First, a word about the word &#8220;open-source&#8221;</p><p>I asked this question as &#8220;which open-source model gets me closest.&#8221; That framing was wrong and I&#8217;m fixing it in public.</p><p>None of these three are open source in the strict sense. They&#8217;re open weight, and the licenses differ.</p><p>GLM-5.3-Flash ships under MIT. That&#8217;s the cleanest of the three.</p><p>Kimi K3 ships under the Kimi K3 License. MIT-shaped, with conditions: run a model-as-a-service business past $20 million in revenue over any 12 months and you need a separate agreement with Moonshot. Cross 100 million monthly users or $20 million in monthly revenue and &#8220;Kimi K3&#8221; has to appear on your product&#8217;s interface.</p><p>Qwen3.8-Flash-Next ships under the Qwen Community License 1.0. Not Apache. Similar shape to Kimi&#8217;s terms.</p><p>Since Friday there&#8217;s a fourth example. GLM-5.3, Z.ai&#8217;s 753B flagship, went open-weight on August 28 under a new GLM-5.3 License. It reads like MIT until the clause that says companies with over $10 billion in revenue across any 12 months need a Z.ai security review before hosting it commercially. Z.ai&#8217;s own model cards use both &#8220;open-weights&#8221; and &#8220;open source&#8221; for it. Only one of those is accurate.</p><p>And open weights are not free compute. Kimi K3 is 2.8 trillion parameters. Third-party reporting puts the download near 594 GB, and Northflank describes a 64-accelerator deployment. GLM-5.3-Flash needs roughly 306 GiB of FP8 weights on Hopper-class hardware or newer. Qwen3.8-Flash-Next sits at about 180B parameters on disk once you count its 51B n-gram table.</p><p>GLM-5.3&#8217;s full-precision download is about 756 GB.</p><p>Translation: you and I are renting these through an API. Which is why the pricing section matters more than the parameter count.</p><h2>Kimi K3: closest on paper, and the number that hurt</h2><p>Moonshot released K3 on July 16 and posted the weights July 27. It&#8217;s 2.8T total, 104B active, native vision, 1M context.</p><p>API price on August 31: $3 per million input, $15 per million output, $0.30 cached. Next to Fable&#8217;s $10 and $50, that reads like 70% off.</p><p>Now the ugly number.</p><p>Artificial Analysis runs a private benchmark called AA-Briefcase, realistic knowledge-work tasks with real file inputs. Kimi K3 scored an Elo of 1543 there, second only to Fable at 1574. Great result. Then AA published the bill: <strong>$10.57 per task</strong>, 83 turns per task versus 67 for Fable, and an average of 56.4 minutes per task, roughly 2.5 times Fable&#8217;s time.</p><p>Cheaper per token. Not cheaper per task. Because K3 takes more turns and writes more output to get there.</p><p>Moonshot&#8217;s own model card, July 2026, shows where it wins. Terminal-Bench 2.1: 88.3 to Fable&#8217;s 88.0. SWE-Marathon: 42.0 to 35.0, with Moonshot noting Fable hit safety fallbacks on 35% of those tasks. BrowseComp: 91.2 to 88.0, though that 91.2 uses context compaction at 300K tokens.</p><p>And where it trails. Toolathlon-Verified: 76.5 to 77.9. JobBench: 54.3 to 57.4. GDPval-AA v2: 1686 Elo to 1747.</p><p>Two operational quirks from Moonshot&#8217;s own docs: K3 requires you to pass its full reasoning content back on every turn, and Moonshot describes it as &#8220;proactive by nature,&#8221; which is vendor-speak for &#8220;it will make decisions you didn&#8217;t ask for unless you fence it in.&#8221;</p><p>Score: near. Stamina: it finishes, but you pay for the scenic route.</p><h2>GLM-5.3-Flash: two days old, and the cheap seats</h2><p>Z.ai released GLM-5.3-Flash on August 26, after a week of running it anonymously on OpenRouter as &#8220;Ox Alpha.&#8221; It&#8217;s 320B total, 18B active, natively multimodal, 1M context, MIT weights.</p><p>List price: $0.15 per million input, $0.50 per million output, $0.03 cached. A launch promotion halves that through September 9, 2026, per multiple pricing trackers; Z.ai&#8217;s own page cites the discounted per-task figure.</p><p>Run the ratio against Kimi. Twenty times cheaper on input. Thirty times cheaper on output.</p><p>Artificial Analysis scores it 57 on the Intelligence Index at $0.09 per index task, and reports the entire index cost $138.02 to run. Same score as GPT-5.6 Terra, per AA, at about 5.7 times lower cost per task. Z.ai&#8217;s launch numbers: Terminal-Bench 2.1 at 84.3, DeepSWE 1.1 at 63.4, AutomationBench at 48.8, and Z.ai describes the whole package as &#8220;approaching Claude Opus 4.8.&#8221; Vendor figures, dated August 26.</p><p>Two catches. AA now measures 45 tokens per second on Z.ai&#8217;s API (down from 50.2 on Friday) and flags it as notably slow. And thinking cannot be disabled, so you pay for reasoning tokens on every call, even the dumb ones.</p><p>Z.ai also says the entire preview week ran on Chinese AI chips.</p><h2>GLM-5.3: the Friday arrival</h2><p>Not in my original brief, because it wasn&#8217;t downloadable when I wrote it. Z.ai announced GLM-5.3 on August 14, held the weights two weeks for a safety review over its cybersecurity results, and posted them August 28.</p><p>It&#8217;s text-only, 753B total with roughly 40B active, 1M context, 128K output. API price: $1.40 input, $4.40 output, $0.26 cached, unchanged from GLM-5.2. AA scores it 60, the same as Kimi K3, at $0.68 per index task, and notes it generated 170M tokens running the index, well above the median. Verbose is the polite word.</p><p>So the planner seat has two candidates at 60. Kimi K3 has vision, a Fable comparison in its own model card, and a measured AA-Briefcase bill. GLM-5.3 has output tokens at roughly a third of Kimi&#8217;s price and no AA-Briefcase number yet, so I can&#8217;t tell you what a real task costs. If your agent never looks at a screenshot, test it in that seat and measure.</p><h2>Qwen3.8-Flash-Next: the one that can see</h2><p>Alibaba&#8217;s weights went live August 24, formal release August 26. Only 6B active parameters. AA scores it 56 and measures 73.4 tokens per second, the fastest of the three.</p><p>Its superpower is the screen. Vendor numbers: AndroidWorld 84.5, OSWorld 2.0 at 52.3 partial (19.4 binary), Toolathlon-Verified 73.5.</p><p>Pricing has a small conflict. Qwen&#8217;s launch announcement lists the production Qwen3.8-Flash at $0.16 input and $0.47 output. Artificial Analysis lists $0.15 and $0.47. I&#8217;m showing both.</p><p>And the context caveat: the open weights run 262K natively, extendable to 1M with YaRN. The hosted Qwen3.8-Flash gets 1M by default.</p><p>Now look at who Qwen&#8217;s launch table compares against: Claude Opus 4.6 Max, a February 2026 model. Not Fable. Not even Opus 4.8.</p><h2>The build</h2><p>Anthropic already published the pattern. Their advisor strategy runs Sonnet 5 as the executor and calls Fable 5 only for guidance. Per Anthropic&#8217;s own results, that pair hit about 92% of Fable&#8217;s standalone SWE-bench Pro performance at roughly 63% of the cost.</p><p>Flip it onto open weights.</p><p>GLM-5.3-Flash does the typing: research, tool calls, code, document work. Kimi K3 writes the plan, gets consulted on hard decisions, and reviews the final output. Qwen3.8-Flash-Next joins the moment a screenshot enters the loop.</p><p>Cost drivers, in order: the planner&#8217;s output tokens and turn count (Kimi at $15 per million, GLM-5.3 at $4.40; cap the advisor&#8217;s output either way), Flash&#8217;s mandatory reasoning tokens, the September 9 promo cliff, and the harness engineering nobody budgets for.</p><p>Before anyone asks: Claude Opus 5, released July 24, scores 63 on the same index, one point above Fable. I&#8217;m holding Fable as the target because it&#8217;s the model Anthropic built for long-horizon autonomy and the one in its advisor pattern. Stamina over score.</p><h2>Claims vs. Receipts</h2><p><strong>Claim:</strong> &#8220;Kimi K3 scores 57 on Artificial Analysis.&#8221; Most July coverage said this. <strong>Receipt:</strong> 60, as of August 31. AA moved to index v4.1.1, and Fable moved too, from 59.9 to 62. Both models shifted. Always check the version number on the leaderboard.</p><p><strong>Claim:</strong> &#8220;Kimi K3 is 70% cheaper than Fable.&#8221; <strong>Receipt:</strong> Per token, yes. Per AA-Briefcase task, $10.57 and 56 minutes, among the most expensive models on that benchmark.</p><p><strong>Claim:</strong> &#8220;Qwen3.8-Flash-Next beats Claude.&#8221; <strong>Receipt:</strong> Beats Claude Opus 4.6 Max on several vendor-run benchmarks. Opus 4.6 is a February 2026 model. Of the three labs, only Moonshot puts Fable 5 in its comparison table.</p><p><strong>Claim:</strong> &#8220;GLM-5.3 is open source.&#8221; Z.ai&#8217;s own materials. <strong>Receipt:</strong> Open-weight under a custom license with a revenue-gated security review. The Flash sibling is MIT. The flagship is not.</p><p><strong>Correction to my own brief:</strong> I called these open-source models. They are open-weight, and three of the four now carry commercial conditions.</p><h2>So what</h2><p><strong>Do this:</strong> Build with GLM-5.3-Flash as the worker and Kimi K3 as advisor and reviewer. If your tasks are text-only, run GLM-5.3 in the advisor seat as a second test. Cap the advisor&#8217;s output. Measure cost per finished task, never per token.</p><p><strong>Skip this:</strong> Kimi K3 as your only model on high-volume work. The turn count will eat the discount.</p><p><strong>Wait on this:</strong> Locking in GLM-5.3-Flash pricing in your budget before September 9. Budget the list price, not the promo.</p><p><strong>Steal this line:</strong> &#8220;Per token is the sticker. Per task is the bill.&#8221;</p><p>The gap between open weights and Fable is now small enough that the interesting question has moved. It&#8217;s no longer who has the smartest model. It&#8217;s who builds the best harness around the models they can afford.</p><h3>Sources</h3><ul><li><p><a href="https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5">Anthropic: Introducing Claude Fable 5 and Claude Mythos 5 (Claude Platform docs)</a></p></li><li><p><a href="https://www.anthropic.com/webinars/building-on-the-claude-platform-claude-fable-5-and-model-orchestration-patterns">Anthropic webinar: Claude Fable 5 and model orchestration patterns</a></p></li><li><p><a href="https://the-decoder.com/anthropics-fix-for-fable-5s-high-cost-is-turning-it-into-a-manager-that-delegates-to-sonnet-5/">The Decoder: Anthropic&#8217;s advisor pattern, Sonnet 5 plus Fable 5 results</a></p></li><li><p><a href="https://forum.moonshot.ai/t/kimi-k3-is-here-our-most-capable-model/480">Moonshot AI forum: Kimi K3 announcement and pricing</a></p></li><li><p><a href="https://huggingface.co/moonshotai/Kimi-K3">Hugging Face: moonshotai/Kimi-K3 model card and benchmark table</a></p></li><li><p><a href="https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE">Hugging Face: Kimi K3 License text</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark">Artificial Analysis: Kimi K3 on AA-Briefcase, cost and time per task</a></p></li><li><p><a href="https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-claude-fable-5">Artificial Analysis: Kimi K3 vs Claude Fable 5 comparison</a></p></li><li><p><a href="https://northflank.com/blog/what-is-kimi-k3-self-hosting">Northflank: Kimi K3 benchmarks, hardware, and self-hosting</a></p></li><li><p><a href="https://docs.z.ai/guides/vlm/glm-5.3-flash">Z.ai developer docs: GLM-5.3-Flash overview</a></p></li><li><p><a href="https://artificialanalysis.ai/models/glm-5-3-flash">Artificial Analysis: GLM-5.3-Flash model page</a></p></li><li><p><a href="https://openrouter.ai/z-ai/glm-5.3-flash">OpenRouter: GLM 5.3 Flash listing and release date</a></p></li><li><p><a href="https://llm-stats.com/blog/research/glm-5.3-flash-launch">LLM Stats: GLM-5.3-Flash launch, list price, and promo end date</a></p></li><li><p><a href="https://www.marktechpost.com/2026/08/26/z-ai-releases-glm-5-3-flash-a-320b-a18b-natively-multimodal-moe-with-a-1m-token-context/amp/">MarkTechPost: GLM-5.3-Flash release and self-hosting requirements</a></p></li><li><p><a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next">Hugging Face: Qwen/Qwen3.8-Flash-Next model card and benchmark tables</a></p></li><li><p><a href="https://artificialanalysis.ai/models/qwen3-8-flash-next">Artificial Analysis: Qwen3.8-Flash-Next model page</a></p></li><li><p><a href="https://the-decoder.com/alibaba-releases-qwen3-8-flash-next-targeting-ultimate-cost-efficiency/">The Decoder: Alibaba releases Qwen3.8-Flash-Next, launch pricing</a></p></li><li><p><a href="https://cellcog.ai/blog/qwen3-8-flash-next/">CellCog: Qwen3.8-Flash-Next license terms and production API status</a></p></li><li><p><a href="https://openrouter.ai/anthropic/claude-fable-5">OpenRouter: Claude Fable 5 listing and cache pricing</a></p></li><li><p><a href="https://huggingface.co/zai-org/GLM-5.3">Hugging Face: zai-org/GLM-5.3 model card, released August 28</a></p></li><li><p><a href="https://artificialanalysis.ai/models/glm-5-3">Artificial Analysis: GLM-5.3 (max) model page, score and pricing</a></p></li><li><p><a href="https://thenewstack.io/zai-glm-weights-license/">The New Stack: GLM-5.3 goes open weight under a new license</a></p></li><li><p><a href="https://venturebeat.com/technology/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens">VentureBeat: GLM-5.3 API pricing and AA cost per task</a></p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2092663573021606119">Artificial Analysis on X: GLM-5.3-Flash at $0.09 per index task</a></p></li><li><p><a href="https://benchlm.ai/benchmarks/artificialanalysis">BenchLM: Artificial Analysis Intelligence Index mirror, updated August 29</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Sunday AI Signal: Free Weights, Self-Graded Benchmarks, and a Breakup with a Move-Out Date.]]></title><description><![CDATA[The Biggest Open Model of 2026 Just Dropped. Read the Fine Print Before You Download 1.5 Terabytes.]]></description><link>https://aisignal.veletica.com/p/sunday-ai-signal-free-weights-self</link><guid isPermaLink="false">https://aisignal.veletica.com/p/sunday-ai-signal-free-weights-self</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Sun, 30 Aug 2026 23:52:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!M9kx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!M9kx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!M9kx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!M9kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:653052,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/213470227?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!M9kx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!M9kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e58f4e1-0782-4b74-ac8c-73bd87728179_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Sunday AI Signal is the weekly wrap where I take the loudest AI claims of the week, hold them against dated receipts, and tell you what to do, skip, or wait on.</p><p>You know that feeling when a download bar says &#8220;3 days remaining,&#8221; and you just close the laptop lid and go live your life? Tencent released a 770-billion-parameter model this week, and the download is 1.56 terabytes. That is not a download. That is a relationship.</p><p>Read this, and you&#8217;ll know whether the biggest open model of the year deserves a dollar of your budget, which two chip-and-benchmark headlines to stop repeating at work, and one line to steal for your next meeting.</p><p>The verdict up front: Hy4 preview is worth testing through the API and not worth buying hardware for. OpenAI&#8217;s chip numbers are real measurements with a self-graded report card. And &#8220;OpenAI cuts off Cursor&#8221; is smaller than the headline, at least by Cursor&#8217;s own count. Every number this week came with a question attached: who ran the test?</p><h3>THE DEEP PICK: TENCENT HY4 PREVIEW</h3><p>Picture a hospital with 256 specialists on staff. You walk in with a problem, and only eight of them come into the exam room, plus one general practitioner who sees everybody. That is a mixture-of-experts model. Hy4 preview has 770 billion parameters total, but only 49 billion wake up for any given token. You get the knowledge of the whole hospital and the bill of eight doctors. Now stretch the waiting room to a million tokens of context, which is roughly a full codebase in one conversation, and you have the pitch.</p><p>The specs, dated August 28, 2026, from the Hugging Face model card: 770B total, 49B active, 1M-token context, and a built-in speculative-decoding layer so it drafts its own guesses and checks them, which is where the speed comes from. Weights ship in BF16 and FP8, with vLLM and SGLang serving recipes on day one.</p><p><strong>Free, paid, freemium, or open source?</strong> All four, depending on which door you use.</p><p>Open source: the weights are under a standard, unmodified Apache 2.0 license. No monthly-active-user cap, no field-of-use clause, commercial use allowed. Earlier Hunyuan releases carried custom terms; this one does not. Freemium: Tencent&#8217;s WorkBuddy and CodeBuddy apps are free for two weeks from August 28. Paid: the API on Tencent Cloud TokenHub and OpenRouter runs $0.834 per million input tokens, $2.501 per million output tokens, and $0.042 per million for cached input (Tencent-listed, August 28; re-verify before it goes in a budget).</p><p><strong>What it costs beyond the sticker.</strong> Free weights are free the way a puppy is free.</p><p>The model card&#8217;s serving recipe runs FP8 across eight GPUs, and the FP8 weights alone land near 780 GB, so you need eight cards in the 141 GB class or better. If you don&#8217;t own that box, you rent it by the hour, and that hourly rate is your real cost. Self-hosting pays off at sustained volume or when data can&#8217;t leave the building. Below that, the API wins.</p><p>Per-task math, because per-token prices lie by omission. Take a long agentic coding session: 3 million input tokens (mostly the same context re-sent each turn) and 150,000 output tokens. Uncached, that session runs about $2.88. At an 80 percent cache-hit rate, it drops to about $0.98. The cost driver is how much context you re-send and whether it hits cache. One 1M-token &#8220;read my whole repo&#8221; call is about 83 cents of input before the model says a word.</p><h3>CLAIMS VS RECEIPTS</h3><p><strong>Claim 1: Hy4 preview &#8220;beats GLM 5.3 and Kimi K3.&#8221;</strong> (Tencent, August 28) Receipt: the win comes from Tencent&#8217;s internal blind evaluation, where 163 employees rated 203 engineering tasks and Hy4 scored 2.99 out of 4 against 2.92 and 2.94. It lost roughly 40 percent of head-to-head matchups in that same test. On public GPQA Diamond, coverage of Tencent&#8217;s own benchmark appendix puts Hy4 at 92.3 and Kimi K3 at 93.5. Artificial Analysis had no independent Intelligence Index score for Hy4 preview as of August 30.</p><p><strong>Claim 2: OpenAI&#8217;s Jalape&#241;o chip delivers &#8220;1.5 to 1.9x more work per watt than Nvidia.&#8221;</strong> (OpenAI at Hot Chips, August 25) Receipt: OpenAI ran the tests. SemiAnalysis engineers watched in OpenAI&#8217;s lab but did not execute the suite themselves, per Forbes and Tom&#8217;s Hardware. The comparison targets were Nvidia&#8217;s GB200 and GB300, which carry prior-generation memory, not the Rubin parts that would be the fair match. Tom&#8217;s Hardware reports an appendix comparison using all-in utility power narrows the gap, and against a GB300 running multi-token prediction the lead shrinks to roughly 1.5x. The part tested was early-stepping silicon, and volume deployment is a 2027 story. The measurements are real. The framing belongs to the vendor.</p><p><strong>Claim 3: &#8220;OpenAI cuts off Cursor.&#8221;</strong> (OpenAI statement, August 28) Receipt: OpenAI says it will wind down the contract by November 12, 2026, citing what it describes as a history of Musk-owned companies violating contracts. The notice came two weeks after SpaceX closed its roughly $60 billion purchase of Cursor&#8217;s parent. Cursor CEO Michael Truell replied August 29 that OpenAI models serve about 5 percent of Cursor&#8217;s user traffic. That figure is company-provided and unaudited, but it moves the story from &#8220;Cursor loses its brain&#8221; to &#8220;Cursor loses a tab.&#8221; Existing models run until the cutoff. If your team standardized on GPT inside Cursor, you have about ten weeks to pick a model, an IDE, or a direct API key.</p><h3>ALSO ON THE TAPE</h3><p>A federal judge ruled Thursday night, August 27, that the Pentagon&#8217;s supply-chain-risk label on Anthropic was unlawful retaliation for the company&#8217;s criticism; the government is expected to appeal (Fortune/AP, August 28). Sony Music Publishing and Warner Chappell sued Anthropic on August 28, alleging lyrics were used in training; that is a complaint, not a finding (Music Business Worldwide). Nvidia paused its $36 billion AI Compute Partnership under two months after launch over antitrust concerns raised by its own employees, per the Wall Street Journal (August 29). And a YouGov survey of 1,250 U.S. workers found 3 percent said they lost a job to AI since 2023 (Fortune, August 29).</p><h3>THE SO WHAT</h3><p><strong>Do this:</strong> Route Hy4 preview through OpenRouter into whatever evaluation set you already run this week, while the free window is open. Log your cache-hit rate. That one number decides whether it is cheap for you.</p><p><strong>Skip this:</strong> Buying an eight-GPU box because the weights are &#8220;free.&#8221; And repeating the 1.9x Jalape&#241;o figure without saying who measured it.</p><p><strong>Wait on this:</strong> Production traffic on Hy4 until Artificial Analysis posts an independent score, and any Nvidia conclusion from Jalape&#241;o until someone independent runs it against Rubin.</p><p><strong>Steal this line:</strong> &#8220;Every benchmark comes with a question attached: who ran the test?&#8221;</p><p>Every number here is dated and will move. Re-verify pricing, the Artificial Analysis listing, and the Cursor shutoff date before you build on them.</p><div><hr></div><h3>SOURCES</h3><ol><li><p><a href="https://huggingface.co/tencent/Hy4-preview">Tencent Hy4-preview model card, Hugging Face (Aug 28, 2026)</a></p></li><li><p><a href="https://huggingface.co/tencent/Hy4-preview-FP8">Tencent Hy4-preview-FP8 model card, Hugging Face (Aug 28, 2026)</a></p></li><li><p><a href="https://www.tencent.com/tencent-releases-and-open-sources-tencent-hy4-preview/">Tencent press release: Tencent Releases and Open-Sources Tencent Hy4 preview (Aug 28, 2026)</a></p></li><li><p><a href="https://technode.com/2026/08/28/tencent-open-sources-hy4-preview-with-770b-parameters-and-a-1m-token-context/">TechNode: Tencent open-sources Hy4 preview with 770B parameters and a 1M-token context (Aug 28, 2026)</a></p></li><li><p><a href="https://www.datalearner.com/ai-models/pretrained-models/hy4-preview">DataLearner: Hy4 preview benchmark and pricing summary (Aug 28, 2026)</a></p></li><li><p><a href="https://medium.com/data-science-in-your-pocket/tencent-hy4-preview-beats-glm-5-3-qwen-3-8-kimi-k3-8ff85d3dbe5b">Data Science in Your Pocket: Tencent Hy4 Preview vs GLM 5.3, Qwen 3.8, Kimi K3 (Aug 28, 2026)</a></p></li><li><p><a href="https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index">Artificial Analysis Intelligence Index v4.1.1 (checked Aug 30, 2026)</a></p></li><li><p><a href="https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks">Tom&#8217;s Hardware: OpenAI&#8217;s 700W Jalape&#241;o ASIC outpaces Nvidia flagship GPU (Aug 26, 2026)</a></p></li><li><p><a href="https://www.forbes.com/sites/jonmarkman/2026/08/27/openai-publishes-first-jalapeo-benchmarks-against-nvidia-blackwell/">Forbes: OpenAI Publishes First Jalape&#241;o Benchmarks Against Nvidia Blackwell (Aug 27, 2026)</a></p></li><li><p><a href="https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/">OpenAI: Our decision on Cursor following its acquisition by SpaceX (Aug 28, 2026)</a></p></li><li><p><a href="https://www.cnbc.com/2026/08/29/openai-cursor-spacex-model-access.html">CNBC: OpenAI to end model access to Cursor after acquisition by SpaceX (Aug 29, 2026)</a></p></li><li><p><a href="https://the-decoder.com/openai-cuts-off-cursor-after-spacex-acquisition-citing-musks-history-of-breaking-contracts/">The Decoder: OpenAI cuts off Cursor after SpaceX acquisition (Aug 29, 2026)</a></p></li><li><p><a href="https://fortune.com/2026/08/28/anthropic-pentagon-ruling-rita-lin-arrogance/">Fortune/AP: Judge: Pentagon punished Anthropic for &#8216;arrogance,&#8217; and that&#8217;s illegal (Aug 28, 2026)</a></p></li><li><p><a href="https://www.musicbusinessworldwide.com/now-sony-music-publishing-and-warner-chappell-sue-anthropic-in-multi-billion-dollar-lawsuit-one-of-the-largest-and-most-blatant-ongoing-thefts-of-intellectual-property-in-history/">Music Business Worldwide: Sony Music Publishing and Warner Chappell sue Anthropic (Aug 28, 2026)</a></p></li><li><p><a href="https://finance.yahoo.com/technology/ai/articles/nvidia-pauses-ai-cloud-revenue-120700044.html">Yahoo Finance/WSJ: Nvidia pauses $36B AI cloud financing program (Aug 29, 2026)</a></p></li><li><p><a href="https://fortune.com/2026/08/29/ai-workers-survey-job-impact-2026/">Fortune: Only 3% of US workers lost a job to AI since 2023, YouGov survey (Aug 29, 2026)</a></p></li><li><p><a href="https://aiweekly.co/ai-news-today">AI Weekly: AI News Today, August 30, 2026 (aggregator used for story discovery only)</a></p></li></ol>]]></content:encoded></item><item><title><![CDATA[Why Nvidia Would Pay 86 Times Revenue for Hugging Face]]></title><description><![CDATA[The $12.9 Billion Deal Nobody Has Signed Yet]]></description><link>https://aisignal.veletica.com/p/why-nvidia-would-pay-86-times-revenue</link><guid isPermaLink="false">https://aisignal.veletica.com/p/why-nvidia-would-pay-86-times-revenue</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Thu, 27 Aug 2026 22:41:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!iZyU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iZyU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iZyU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iZyU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2022777,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/213067720?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iZyU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!iZyU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74c726c8-63cd-4ee4-ad22-39ba852bb915_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Nvidia reportedly agreed to buy Hugging Face. Before you repost the headline, look at what is actually confirmed.</p><p>You are going to see &#8220;Nvidia buys Hugging Face&#8221; everywhere this week. Give me five minutes and you will know exactly which parts of that sentence are solid, which parts are soft, and what to do about the models your work depends on.</p><p>The verdict up front: this is a reported deal, not a done deal. The Information reported on the night of August 26, 2026, that Nvidia agreed to acquire Hugging Face for $12.9 billion, citing a single source. Business Insider and Bloomberg reported the same night that no agreement had been signed, and the talks could still fall apart. As of August 27, neither company has confirmed anything, and reporters at multiple outlets got no comment. If it closes, it changes who controls the front door of open-source AI. Until then, the honest headline has the word reportedly in it.</p><h2>What Hugging Face actually is</h2><p>Think of AI as a giant kitchen economy. Nvidia sells the ovens. The labs cook the models. Hugging Face is the grocery shelf where almost every open model gets stocked, browsed, and picked up.</p><p>Around 13 million developers pull models and datasets from that shelf. This week&#8217;s coverage puts the model count somewhere between 2 million and 2.5 million, and I could not confirm one canonical figure, so treat that as a range. Either way, when a lab releases a new open-weight model, the launch post ends with a Hugging Face link. Nvidia&#8217;s own Nemotron models live there too. That is what makes the shelf worth $12.9 billion to somebody.</p><h2>Claims vs. receipts</h2><p><strong>The claim:</strong> &#8220;Nvidia agrees to buy Hugging Face for $12.9 billion.&#8221; Source: The Information, August 26, attributed to one person with knowledge of the agreement.</p><p><strong>The receipts:</strong> Business Insider, which first reported over the weekend of August 22 that Hugging Face was fielding takeover interest, said the talks had not produced a signed agreement and could still fall apart. Bloomberg reported the same. TechCrunch reached out to both companies, got silence, and noted that Nvidia has historically moved fast to knock down reports it considers wrong. The silence is interesting. It is not a signature.</p><p>One receipt for context: this was not an ambush. Business Insider reported that Hugging Face hired a bank to test buyer interest at $13 billion or more, and The Information reported the Nvidia talks started after a different suitor approached the company first. Hugging Face ran a process.</p><h2>The math that explains the motive</h2><p>Hugging Face generates roughly $150 million in annualized revenue as of August 2026, per this week&#8217;s coverage. Against a $12.9 billion price, that works out to about 86 times sales.</p><p>Nobody pays 86 times revenue for the revenue. Nvidia would be paying for position, and the reported logic shows up in three places.</p><p>Chip defense comes first. Fortune&#8217;s framing is the cleanest: developers who download open models from Hugging Face run them on their own infrastructure, and that infrastructure usually means Nvidia GPUs. A healthy open ecosystem keeps demand tied to Nvidia hardware.</p><p>The counterweight comes second. The Information reports Nvidia sees successful open models as a check on closed-model labs like OpenAI and Anthropic, which are designing their own chips and slowly reducing their Nvidia dependence.</p><p>A road back into cloud comes third. Nvidia reportedly scaled back its DGX Cloud business about a year ago, and Hugging Face already sells hosted compute to developers. Buying the shelf buys back the storefront.</p><p>The timing was not subtle either. Nvidia reported earnings the same afternoon the story broke: $96.2 billion in quarterly revenue, up 106% from a year earlier. Guidance for the current quarter came in at $108 billion. At that pace, a $12.9 billion acquisition costs Nvidia less than six weeks of revenue. I had to read that math twice, and I wrote it.</p><h2>The Switzerland problem</h2><p>In August 2023, Hugging Face raised $235 million from nine companies at once, including Nvidia, Google, Amazon, and Salesforce, at a $4.5 billion valuation. CEO Cl&#233;ment Delangue described the structure at the time as an ecosystem round and called Hugging Face &#8220;a neutral platform, or the Switzerland&#8221; of AI. The whole point was that no single backer could control it.</p><p>Switzerland does not usually sell itself to one army. If a single chip vendor owns the shelf, the neutrality pitch is over, and this week&#8217;s coverage is already raising the platform-neutrality question. Whether it matters in practice depends on terms nobody has seen, because nothing has been announced.</p><h2>What this costs you today: nothing, yet</h2><p>Browsing and downloading open models on Hugging Face is free. The paid side, meaning Pro subscriptions, enterprise plans, and hosted inference compute, is where the reported $150 million comes from. A reported deal changes none of that.</p><p>Two cautions anyway. Every model on the shelf carries its own license, and Apache 2.0 or MIT means something very different from a restricted community license, so check before you ship. And ownership changes can bring pricing changes, which is exactly why the first item below exists.</p><h2>So, what</h2><p><strong>Do this</strong>: Mirror the specific model weights and datasets your work depends on and write down each one&#8217;s license while you are in there. It costs you an afternoon and removes your single point of failure, deal or no deal.</p><p><strong>Skip this</strong>: Reposting &#8220;Nvidia buys Hugging Face&#8221; as fact. The accurate version keeps the word reported, and accuracy is cheap.</p><p><strong>Wait on this</strong>: Conclusions about what Nvidia will do with the platform. No announcement, no terms, no signature as of August 27.</p><p><strong>Steal this line</strong>: A reported deal is a rumor with a bank behind it.</p><div><hr></div><h3>Sources</h3><ul><li><p><a href="https://www.theinformation.com/articles/nvidia-agrees-buy-open-source-model-repository-hugging-face-12-9-billion">The Information: Nvidia Agrees to Buy Open Source AI Platform Hugging Face For $12.9 Billion (Aug 26, 2026)</a></p></li><li><p><a href="https://www.bloomberg.com/news/articles/2026-08-27/nvidia-discussed-buying-ai-startup-hugging-face-insider-says">Bloomberg: Nvidia in Talks to Buy AI Startup Hugging Face, Reports Say (Aug 27, 2026)</a></p></li><li><p><a href="https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/">TechCrunch: Nvidia closes in on Hugging Face acquisition (Aug 26, 2026)</a></p></li><li><p><a href="https://www.cnbc.com/2026/08/27/nvidia-hugging-face-acquisition.html">CNBC: Nvidia agrees to buy Hugging Face for $12.9 billion, report says (Aug 27, 2026)</a></p></li><li><p><a href="https://fortune.com/2026/08/27/nvidia-hugging-face-billion-dollar-deal-open-source-ai/">Fortune: Nvidia nears $12.9 billion deal to buy open-source AI platform Hugging Face (Aug 27, 2026)</a></p></li><li><p><a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-to-buy-hugging-face-for-usd12-9-billion-report-claims-could-strengthen-nvidias-open-model-strategy-and-shore-up-position-against-rivals">Tom&#8217;s Hardware: Nvidia to buy Hugging Face for $12.9 billion, report claims (Aug 27, 2026)</a></p></li><li><p><a href="https://siliconangle.com/2026/08/27/nvidia-reportedly-acquires-ai-project-hosting-platform-hugging-face-for-12-9b/">SiliconANGLE: Nvidia reportedly acquires AI project hosting platform Hugging Face for $12.9B (Aug 27, 2026)</a></p></li><li><p><a href="https://www.forbes.com/sites/siladityaray/2026/08/27/nvidia-has-reportedly-agreed-to-buy-ai-model-hosting-platform-hugging-face-for-13-billion/">Forbes: Nvidia Has Reportedly Agreed To Buy Hugging Face For $13 Billion (Aug 27, 2026)</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Sunday AI Signal: Meta's Free Model Comes with a Hardware Bill.]]></title><description><![CDATA[Ground Truth AI. Week of August 24, 2026.]]></description><link>https://aisignal.veletica.com/p/sunday-ai-signal-metas-free-model</link><guid isPermaLink="false">https://aisignal.veletica.com/p/sunday-ai-signal-metas-free-model</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Mon, 24 Aug 2026 01:57:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5hTL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5hTL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5hTL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5hTL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2304068,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/212484266?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5hTL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!5hTL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23b6bd55-e806-4089-8593-43f50694f51a_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Beat 1: The Deep Pick</h2><p>Everybody in my group chat has a friend who took a free puppy. Nobody in my group chat has a friend who kept a free puppy for free.</p><p>That is Muse Glimmer.</p><p>Meta released it on August 10, 2026: a 30 billion parameter model, weights on Hugging Face, Apache 2.0 license, built to run agents on a single consumer GPU. Every headline said &#8220;free&#8221; and &#8220;open source.&#8221; Both words are true. Neither one tells you what you are about to pay.</p><p>My verdict, so you can stop reading if you need to: if you already own a 24 GB GPU or a 32 GB Mac, download it this week and put it on a real tool-calling job. If you do not own that hardware, do not buy it for this model, and do not rent it hosted at current prices. Details below.</p><p><strong>What &#8220;free&#8221; means here, exactly.</strong></p><p>Free license. Apache 2.0 lets you use it commercially, modify it, redistribute it, and ship it inside a product without paying Meta. That is a real change. Every prior Meta open release used a custom Llama license with strings attached, including that 700 million monthly user cutoff people liked to joke about. Apache 2.0 has no such clause. Lawyers relax. Engineers download.</p><p>Not a free service. Meta ships no hosted endpoint for Glimmer. There is no per-token fee to Meta because there is no Meta API for it. You either run it yourself or pay a third party like Together AI or Fireworks to run it for you.</p><p><strong>The bill driver: memory.</strong></p><p>The full precision weights are roughly 60 GB. Meta&#8217;s 4-bit quantized build shrinks the language model to under 20 GB, and the model card says that leaves room for the KV cache, the vision encoder, and the speed-boosting &#8220;drafter&#8221; inside a 24 GB or 32 GB envelope. Independent teardown of the actual files puts the main 17 GB quant at 16.76 GB, plus a 1.4 GB vision projector, plus the drafter. Call it just under 20 GB before you type a single prompt.</p><p>So, the puppy needs a yard. An RTX 5090 class card, or an M4 Max or M5 Max Mac with 32 GB or more. If you have that, the marginal cost is electricity and a weekend. If you do not, you are looking at a four-figure hardware purchase, and at that point &#8220;free&#8221; has stopped meaning anything.</p><p>Speed, per Meta&#8217;s own numbers on the model card: 74.9 tokens per second on an RTX 5090 without the drafter, 233.4 with it. On an M4 Max, 23.7 without, 37.8 with. Those are vendor figures. Note the gap: the drafter roughly triples speed on a discrete GPU and adds about half again on a Mac.</p><p><strong>The hosted path, if you skip the hardware.</strong></p><p>Artificial Analysis lists Glimmer at $0.32 per million input tokens and $1.35 per million output tokens across providers as of August 11, 2026, and flags both as expensive against the median for open-weights models of similar size ($0.05 in, $0.15 out). Read that twice. The &#8220;free&#8221; model, rented, costs more per token than the open models it is competing with. Prices move fast in launch weeks. Re-verify before you commit budget.</p><h2>Beat 2: Claims vs. Receipts</h2><p><strong>Meta&#8217;s claim (model card, August 10, 2026):</strong> Glimmer &#8220;performs strongly for its size class&#8221; against Gemma 4 31B and Qwen3.6 27B. Meta&#8217;s own comparison table bolds Glimmer as the best of the three on roughly half the rows. Standouts: MCP Atlas 75.5 versus Qwen&#8217;s 62.5, DeepSearch QA 74.6 versus 71.1, SWE-Bench Pro 51.2 versus 50.2.</p><p><strong>What the same table admits:</strong> Qwen3.6 27B beats Glimmer on SWE-Bench Verified (77.2 vs 76.0), OSWorld-Verified (75.6 vs 65.9), and TerminalBench 2.1 (60.7 vs 51.7). On GDPval-AA v2, the real-world work test, Glimmer scores 953 to Qwen&#8217;s 1141. Meta printed those losses. Credit where due.</p><p><strong>A methodology note reviewers flagged:</strong> analysts who read Meta&#8217;s evaluation report say Meta used whichever was more favorable to the competitor, the competitor&#8217;s self-reported score or Meta&#8217;s own reproduction, except where Artificial Analysis covered all three models. That is a defensible choice. It is also not the same thing as an independent, like-for-like run.</p><p><strong>The independent receipt (Artificial Analysis, August 10 to 11, 2026):</strong> Glimmer scores 35 on the Artificial Analysis Intelligence Index. Qwen3.6 27B scores 38. Gemma 4 31B scores 30. Kimi K2.5, a one trillion parameter model, scores 36. Llama 4 Maverick, Meta&#8217;s last open release, scored 14.</p><p>So, the independent read: Glimmer is a big jump for Meta, a clear win over Gemma at the same size, and a few points behind Qwen. Not the sweep the launch chatter implied. A solid second place in its weight class, with the best license in the room.</p><p><strong>Still unverified as of August 23, 2026:</strong> Meta&#8217;s claim that the 17 GB quant loses only 1.0 percent accuracy versus full precision. I found no independent replication. Treat it as a vendor number until someone outside Meta runs it.</p><h2>Beat 3: The Call for This Week</h2><p><strong>Do this:</strong> If the hardware is already on your desk, pull the K-Quant-17GB build and run it against a tool-calling workflow you actually use. Judge it on your tasks, not Meta&#8217;s table. Apache 2.0 means anything you build on it is yours to ship.</p><p><strong>Skip this:</strong> Buying a GPU for Glimmer. Renting Glimmer hosted at current listed prices when Qwen3.6 27B scores higher independently and typically costs less per token.</p><p><strong>Wait on this:</strong> Muse Spark 1.2 open weights. Zuckerberg said &#8220;soon&#8221; on August 10. No date, no license named. If that lands under Apache 2.0, it is a bigger story than Glimmer. Also wait on the quantization degradation claim until an outside lab reproduces it.</p><p><strong>Steal this line:</strong> &#8220;Free is the sticker on the box. The bill is the box.&#8221;</p><p>See you next Sunday.</p><div><hr></div><h2>Sources</h2><ul><li><p><a href="https://huggingface.co/meta-models/Muse-Glimmer-30B">Muse Glimmer 30B model card, Hugging Face (Meta)</a></p></li><li><p><a href="https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model">Introducing Muse Glimmer, Meta AI Research blog</a></p></li><li><p><a href="https://developer.meta.com/ai/models/muse-glimmer/">Muse Glimmer developer page, Meta</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/muse-glimmer">Muse Glimmer: Benchmarks and analysis, Artificial Analysis</a></p></li><li><p><a href="https://artificialanalysis.ai/models/muse-glimmer">Muse Glimmer (high) model page, Artificial Analysis</a></p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2086916150278111551">Artificial Analysis launch thread on X</a></p></li><li><p><a href="https://venturebeat.com/technology/meta-returns-to-open-source-with-muse-glimmer-an-apache-2-0-licensed-30b-parameter-ai-model-optimized-for-agents-available-now">Meta returns to open source with Muse Glimmer, VentureBeat</a></p></li><li><p><a href="https://kingy.ai/blog/muse-glimmer-30b-benchmarks-hardware-run/">Muse Glimmer 30B: Benchmarks, Hardware and How to Run, Kingy AI (file size teardown)</a></p></li><li><p><a href="https://wavect.io/blog/muse-glimmer-30b-local-agent-guide/">Muse Glimmer 30B Hardware and Benchmark Guide, Wavect (methodology notes)</a></p></li><li><p><a href="https://sebastianraschka.com/blog/2026/muse-glimmer-30b-architecture-notes.html">Muse Glimmer 30B Architecture Notes, Sebastian Raschka</a></p></li><li><p><a href="https://www.together.ai/models/muse-glimmer">Muse Glimmer on Together AI</a></p></li><li><p><a href="https://x.com/lmstudio/status/2086766394247360716">LM Studio post quoting Zuckerberg on Muse Spark 1.2 weights, X</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Meet Grok Bot: The AI Coworker That Actually Does the Work. ]]></title><description><![CDATA[The $200 AI Coworker That Shares a Desk with Everyone Else You Hired.]]></description><link>https://aisignal.veletica.com/p/meet-grok-bot-the-ai-coworker-that</link><guid isPermaLink="false">https://aisignal.veletica.com/p/meet-grok-bot-the-ai-coworker-that</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Wed, 19 Aug 2026 21:43:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EtVn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EtVn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EtVn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EtVn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2735700,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/211919904?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EtVn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!EtVn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68dbbabb-09c1-44cb-91c6-4e1b97df86d6_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You know the part of onboarding a new hire where nothing can happen until IT gives them a laptop, a badge, and the WiFi password? Grok Bot skips it. You name a bot, and it already has a computer with a browser open and your logins sitting in it.</p><p>That is the pitch, and I want to be fair to it: it works. Early users describe it as the first agent product that felt like adding a coworker instead of configuring one.</p><p>But the sentence doing the most work on the product page is wrong about the product.</p><p><strong>The verdict, up front.</strong> If you already pay for SuperGrok Heavy or Cursor Ultra, open it tonight and give one bot one small job you can inspect afterward. If you don&#8217;t already pay for those, don&#8217;t upgrade for this. It&#8217;s an early beta with no published task-completion rate, no audit log, and a security model that the marketing page and the documentation describe differently. The gap between those two descriptions is the story.</p><h4>The mental model: you hired a team, you rented a desk</h4><p>Here&#8217;s the frame I&#8217;d use if you&#8217;re explaining this to your boss.</p><p>Picture an office with one desk. One chair, one monitor, one browser, one drawer of sticky notes with passwords on them. You hire eight people. They all sit at that desk, one at a time, and whatever the last person left signed in stays signed in for the next one.</p><p>That is Grok Bot&#8217;s architecture. Not eight computers. One.</p><p>Call it the Shared Desk, because once you have that image, every strange detail in the docs stops being strange.</p><p>Why can your research bot hand a file to your writing bot without you copy-pasting between chats? Same desk. Why does xAI&#8217;s own security page warn you not to treat separate bots as a security boundary? Same desk. Why does deleting a bot leave its logins alive? The person left. The desk kept the sticky notes.</p><p>The shared desk is simultaneously the best feature and the biggest risk, and those are not two facts. They&#8217;re one fact wearing two hats.</p><h4>Claims vs. receipts</h4><p><strong>Claim: &#8220;Bots have their own computer.&#8221;</strong></p><p>That line appears in the hero copy and the FAQ on x.ai/bot, verified live August 19, 2026. Now open xAI&#8217;s own launch post from August 11 and read the section titled &#8220;A computer of its own.&#8221; The sentence inside it says bots <em>share</em> a computer of their own in the cloud.</p><p>Two different claims, from the same company, on the same day. The documentation is unambiguous about which one is true: every bot on your account uses one persistent cloud computer, assigned per user, not per bot. And then it adds the sentence that should be on a poster in every security team&#8217;s office: do not use separate bots as a security boundary.</p><p>A good chunk of launch coverage picked up the wrong version. At least two outlets published that each bot gets a dedicated cloud computer. That is not a small copyedit. It&#8217;s the difference between &#8220;I can contain the blast radius by splitting bots&#8221; and &#8220;I cannot.&#8221;</p><p><strong>Claim: &#8220;Get started for free.&#8221;</strong></p><p>That&#8217;s the text on the final call-to-action button on the product page, verified today. It links to a .dmg download. Every plan listed above it on the same page costs money: Cursor Ultra at $200/month, SuperGrok Heavy at $300/month, Cursor Premium Teams at $120 per seat per month. The docs say eligible plans are required. One secondary write-up mentions a one-time individual trial, but nothing on the pricing cards describes it, so I&#8217;d label that access path unannounced rather than free.</p><p><strong>Claim: it runs on Linux.</strong></p><p>The docs, dated August 11 and unchanged since, list macOS, Windows, and iPhone on iOS 18 or later, then state plainly that Linux desktop, Android, and iPad are not supported at initial launch. Several outlets reported Linux builds shipped anyway, and at least one reviewer screenshotted a &#8220;Download for Linux&#8221; button on the marketing page. I could not resolve this one. The live page today shows a macOS button and a link labeled &#8220;other platforms and devices.&#8221; I&#8217;m reporting the conflict rather than picking a side.</p><p><strong>Claim: it&#8217;s frontier-grade.</strong></p><p>This one mostly checks out, with a version-number trap. Grok 4.6 shipped August 12 and scores 60.9 on the Artificial Analysis Intelligence Index, which rounds to the 61 xAI claimed. Independent measurement matched the vendor&#8217;s number, which is worth saying out loud because it often doesn&#8217;t. It sits behind Claude Opus 5 at 63.0 and Claude Fable 5 at 62.1, per the index snapshot dated August 18, 2026.</p><p>The trap is Terminal-Bench. xAI reports 26% on v3.0. Artificial Analysis reports 88.4% on v2.1. Same model. Different benchmark versions. Anyone putting those two numbers in the same sentence is producing nonsense.</p><p>And the number that matters most for Grok Bot does not exist. There is no published task-success rate, no failure rate, no human-intervention rate for the product itself. The model has receipts. The workforce does not.</p><h4>What it actually costs to run</h4><p>The sticker price is the least interesting part of the bill.</p><p>Every eligible plan includes a weekly usage allowance whose size xAI has not published. Past that, you buy on-demand usage billed from raw model and token cost. There is no Grok Bot spend cap yet, and there&#8217;s no model picker, so you can&#8217;t steer toward something cheaper.</p><p>Now layer in how Grok 4.6 is priced: $2 per million input tokens and $6 per million output, which flips to $4 and $12 once a prompt crosses 200,000 tokens. That higher rate re-bills the entire request, not just the overage.</p><p>Long-running agent work is a context-accumulation machine. An agent that keeps working after you close the laptop is, structurally, the ideal device for generating tokens you were not watching. Some users have reported multi-bot sessions burning through weekly limits inside a couple of hours of real work, though that&#8217;s anecdotal and worth treating as such.</p><p>None of this is open source. There&#8217;s no license to read, no self-hosting path, and no way to audit what the thing did beyond your own chat transcript. The teams documentation says an audit view of bot actions is coming. </p><h4>So What</h4><p><strong>Do this:</strong> If you&#8217;re already paying for Cursor Ultra or SuperGrok Heavy, create exactly one bot with one narrow read-only job on a system you can verify by hand. Give it a task you already know the right answer to. That&#8217;s how you get a completion rate when the vendor hasn&#8217;t published one.</p><p><strong>Skip this:</strong> Do not upgrade to a $200 or $300 tier to try a beta whose reliability nobody has measured. That is buying a number that does not exist yet.</p><p><strong>Wait on this:</strong> Anything involving payments, production systems, customer records, or credentials you cannot rotate quickly. The approval system stops a proposed action, and xAI says so directly: an approval does not reverse work already completed.</p><p><strong>Steal this line:</strong> &#8220;The isolation boundary is the account, not the bot.&#8221; Say it in your next architecture review and watch how fast the room re-scopes what it was about to connect.</p><h4>The part that outlives this product</h4><p>Six months from now the beta labels come off and the completion rates get published. Grok Bot will either be good, or it won&#8217;t.</p><p>What survives is the question the shared desk forces you to answer: which of your workflows should ever be handed to something that logs in as you? That&#8217;s not a Grok question. Claude Cowork, ChatGPT Work, and Microsoft&#8217;s Copilot are all converging on the same shape. Every one of them will eventually ask you for the keys.</p><p>Figure out your answer now, while the stakes are one newsletter unsubscribe task instead of your accounts payable.</p><div><hr></div><h4>Sources</h4><ul><li><p><a href="https://x.ai/news/introducing-grok-bot">Introducing Grok Bot &#8212; SpaceXAI launch post, August 11, 2026</a></p></li><li><p><a href="https://x.ai/bot">Grok Bot product and pricing page &#8212; x.ai/bot</a></p></li><li><p><a href="https://docs.x.ai/grok-bot/faq">Grok Bot FAQ &#8212; SpaceXAI Docs</a></p></li><li><p><a href="https://docs.x.ai/grok-bot/approvals-security-and-privacy">Approvals, security, and privacy &#8212; SpaceXAI Docs</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis">Grok 4.6 benchmarks and analysis &#8212; Artificial Analysis</a></p></li><li><p><a href="https://artificialanalysis.ai/models/grok-4-6">Grok 4.6 model page &#8212; Artificial Analysis</a></p></li><li><p><a href="https://venturebeat.com/orchestration/spacexais-grok-bot-turns-agents-into-persistent-digital-coworkers-that-can-operate-your-apps-for-120-per-month">SpaceXAI&#8217;s Grok Bot turns agents into persistent digital coworkers &#8212; VentureBeat</a></p></li><li><p><a href="https://www.reworked.co/collaboration-productivity/xai-launches-grok-bot-ai-agents-in-beta/">xAI Wants In on the Enterprise With Grok Bot &#8212; Reworked</a></p></li><li><p><a href="https://www.eesel.ai/blog/grok-bot-review">Grok Bot review: what actually ships in the early beta &#8212; eesel AI</a></p></li><li><p><a href="https://codersera.com/blog/grok-4-6-benchmarks-explained-2026/">Grok 4.6 Benchmarks Explained: Why 26% and 88% Are the Same Model &#8212; Codersera</a></p></li><li><p><a href="https://finance.yahoo.com/technology/ai/articles/spacex-completes-record-60-billion-131311785.html">SpaceX completes record $60 billion acquisition of Cursor &#8212; Investing.com via Yahoo Finance</a></p></li><li><p><a href="https://www.digitalapplied.com/blog/grok-bot-ai-teammates-launch-cloud-computer-2026">Grok Bot: xAI&#8217;s AI Teammates Get Their Own Computer &#8212; Digital Applied</a></p></li></ul><p><em>Pricing and benchmark figures verified August 19, 2026. Both are volatile. Re-check the product page and the Artificial Analysis index before citing these numbers after September 1.</em></p>]]></content:encoded></item><item><title><![CDATA[One Real Bargain, One Fake Freebie, One Number That Doesn't Exist Yet.]]></title><description><![CDATA[Sunday AI Signal. This Week in AI.]]></description><link>https://aisignal.veletica.com/p/one-real-bargain-one-fake-freebie</link><guid isPermaLink="false">https://aisignal.veletica.com/p/one-real-bargain-one-fake-freebie</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Mon, 17 Aug 2026 01:36:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6EYd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6EYd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6EYd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6EYd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2544629,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/211489565?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6EYd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6EYd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d13cd47-ec02-4a9c-b5d1-cc1722cca441_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Think of this week in AI like a group chat where everyone keeps forwarding the same article. My job on Sunday is to be the friend who actually opened it, read the whole thing, and tells you straight whether it holds up.</p><p>This week it mostly didn&#8217;t. Not because the news was fake, but because the good part got left out.</p><p>Here&#8217;s your verdict up front. Grok 4.6 is a real upgrade and it is genuinely the cheapest model at the intelligence frontier right now, at $2 per million input tokens and $6 per million output. That part is true. But the cheap price only holds while your prompt stays under 200,000 tokens. Cross that line and xAI bills the entire request at double, $4 and $12. And long, context-heavy work is the exact thing xAI is selling this model for. So if you run short prompts, this is a steal. If you run long-horizon agents, you need to do math before you trust the sticker. That&#8217;s the whole edition in four sentences. Now the receipts.</p><p><strong>First, a way to think about it that will stick.</strong></p><p>Grok 4.6&#8217;s pricing works like a happy hour menu. Cheap drinks, great deal, everybody&#8217;s happy. But there&#8217;s a cutoff time posted by the door, and here&#8217;s the twist that makes it worse than a real bar. At a real happy hour, when the clock runs out, only the next drink costs more. With Grok 4.6, the moment you cross the line, the bartender reprices your entire tab at full rate. The beer you drank an hour ago? Full price now too. That&#8217;s what &#8220;$4/$12 across the entire request&#8221; means. Not the tokens past 200K. All of them.</p><p>Hold onto that, because it changes the answer to &#8220;should I switch?&#8221;</p><p><strong>Beat one: the one thing that actually mattered this week.</strong></p><p>xAI shipped Grok 4.6 on August 12, 2026, with an underlying checkpoint dated August 10. The API id is just grok-4.6. It landed day one on Cursor and Grok Build, with API access through console.x.ai plus OpenRouter, Vercel, and Cloudflare. No EU region at launch, so if you&#8217;re in the EU, you&#8217;re waiting.</p><p>The confirmed specs are clean. A 500,000-token context window. Knowledge cutoff of February 1, 2026. Text and image in, text out. Reasoning levels from low up to a new &#8220;xhigh.&#8221; On the Artificial Analysis Intelligence Index, it scores 61, which matches OpenAI&#8217;s GPT-5.6 Sol and sits one point behind the top Claude. That combination, frontier-level score at the lowest frontier price, is the entire pitch. xAI&#8217;s Michael Truell framed it as Opus-class intelligence at low cost and high speed.</p><p>Now the parts the launch post was quiet about.</p><p>One, the pricing cliff. The $2/$6 rate covers prompts under 200K tokens. At or above that threshold, xAI&#8217;s own pricing page reprices the whole request to $4/$12. For a model marketed around long-running agents and big-context work, the discount evaporates precisely where you&#8217;d want to use it (verify against xAI&#8217;s live pricing page before you budget, as of August 15, 2026).</p><p>Two, the quiet cache hike. Cached input rose from $0.30 to $0.50 per million tokens, roughly 67% higher than Grok 4.5. Small number, but if your workload leans on caching, that&#8217;s a real line item, and it wasn&#8217;t in the headline.</p><p>Three, the context window did not grow. It was already 500K on Grok 4.5. The upgrade is how well the model uses long context, not a bigger limit. A few pieces of coverage framed the 500K like a new feature. It isn&#8217;t.</p><p>Four, the benchmarks are xAI&#8217;s own. As of August 13, 2026, no independent third party had replicated the numbers. On xAI&#8217;s ten-row comparison table, reviewers point out that the top Claude actually wins the most rows, GPT-5.6 Sol takes the coding-specific ones and Grok 4.6 loses Terminal-Bench v3.0 by about 8.6 points, a result the launch prose skipped over. Strong model. Selectively presented.</p><p><strong>Cost and licensing, plainly:</strong> Grok 4.6 is paid and proprietary. Closed weights, per-token API. $2/$6 under 200K, $4/$12 at or above it, cached input $0.50, and a faster variant at twice the price. The first-week promo gives 2x included usage inside Grok Build and Cursor, but xAI never published the base usage number, so &#8220;2x&#8221; is a nudge, not something you can put in a budget.</p><p><strong>Beat two: claims versus receipts.</strong></p><p>Two more launches got a coat of paint this week. Let&#8217;s scrape it off.</p><p><em>Meta Muse Glimmer, the &#8220;free&#8221; local agent.</em> This one&#8217;s mostly good news, which is why the framing matters. Meta Superintelligence Labs released Muse Glimmer on August 10, 2026, a roughly 29.6-billion-parameter multimodal model, and here&#8217;s the part that&#8217;s genuinely rare: the weights ship under an unmodified Apache 2.0 license. Not a Llama-style custom license with fine print. The real thing. You can build a commercial product on it with no royalty obligation. Mark Zuckerberg and Meta&#8217;s Alexandr Wang both pitched it as running locally on a single consumer GPU.</p><p>So, where&#8217;s the receipt? In the word &#8220;free.&#8221; The license is free. Running it is not. Wang said it fits in 24GB of VRAM, and Meta targets 24GB hardware with a compressed 4-bit build. But reviewers who read the model card note that the small &#8220;under 20GB&#8221; figure describes the language weights alone. A working multimodal agent also needs the vision encoder, the KV cache, runtime overhead, and the optional speed drafter. Translation: you need a real GPU, which is real money, or a cloud rental that bills by the hour. Free license, paid hardware. And Meta&#8217;s benchmark wins are Meta&#8217;s own, mixed against rivals, leading on some agent tests, trailing Qwen on others. Still, for local agent work, this is the most interesting release of the week. Just don&#8217;t read &#8220;free&#8221; as &#8220;no cost.&#8221;</p><p><em>Qwen3.8-Max, the number that isn&#8217;t a number yet.</em> You may have seen a $2/$6 per million price for Qwen3.8-Max floating around X, next to a claim that it&#8217;s &#8220;second only to&#8221; the top Claude. Both are unconfirmed. The pricing figure circulating on X is unverified, the &#8220;second only to&#8221; ranking has no third-party backing, the model previewed back on July 19 rather than launching in August, and there&#8217;s no published benchmark table and no open-weight date. There may be a great model here eventually. Right now there&#8217;s a screenshot and a vibe. Nothing to act on.</p><p><strong>Beat three: skip, wait, steal for the week.</strong></p><p><strong>Steal this</strong>: Muse Glimmer, if you do local or agent work and you own or can rent a 24GB GPU. A genuinely Apache 2.0 agentic model that runs on one machine is the first realistic default for local agents. Worth an actual pilot this week. Budget the hardware honestly and test it in your own scaffold, not on Meta&#8217;s benchmark table.</p><p><strong>Skip this</strong>: swapping your whole stack to Grok 4.6 because of the &#8220;cheapest at the frontier&#8221; headline. The score that backs that headline is xAI-reported and, as of August 13, unreplicated, and the model loses the coding-specific evals. Don&#8217;t rip out a working setup for a launch-day number.</p><p><strong>Wait on this</strong>: two things. Qwen3.8-Max until there are real verifiable numbers instead of an X screenshot. And Grok 4.6 for long-context agent work until you&#8217;ve modeled the 200K cliff against your actual token usage. If your prompts routinely run long, &#8220;cheap&#8221; might be the most expensive option on the menu.</p><p>That&#8217;s the week. One real bargain with a catch, one real freebie that costs money, and one number that doesn&#8217;t exist yet. See you next Sunday.</p><p><strong>Sources</strong></p><ul><li><p><a href="https://www.digitalapplied.com/blog/grok-4-6-launch-pricing-agentic-benchmarks-2026">Grok 4.6 launch, pricing, and the 200K threshold &#8212; DigitalApplied</a></p></li><li><p><a href="https://www.aipricing.guru/news/xai-grok-4-6-launch-pricing-impact-august-2026/">Grok 4.6 price and long-context billing &#8212; AI Pricing Guru</a></p></li><li><p><a href="https://kie.ai/blog/grok-4-6-release-analysis">Grok 4.6 benchmarks and disputed numbers &#8212; Kie.ai</a></p></li><li><p><a href="https://kingy.ai/blog/grok-4-6-price-benchmarks-api-cursor-context-window/">Grok 4.6 price, context, and access &#8212; Kingy.ai</a></p></li><li><p><a href="https://codersera.com/blog/grok-4-6-launch-guide-2026/">Grok 4.6 specs and pricing cliff &#8212; Codersera</a></p></li><li><p><a href="https://developer.meta.com/ai/models/muse-glimmer/">Muse Glimmer official page &#8212; Meta</a></p></li><li><p><a href="https://venturebeat.com/technology/meta-returns-to-open-source-with-muse-glimmer-an-apache-2-0-licensed-30b-parameter-ai-model-optimized-for-agents-available-now">Meta returns to open source with Muse Glimmer &#8212; VentureBeat</a></p></li><li><p><a href="https://wavect.io/blog/muse-glimmer-30b-local-agent-guide/">Muse Glimmer real hardware footprint &#8212; Wavect</a></p></li><li><p><a href="https://www.phoronix.com/news/Meta-Muse-Glimmer">Muse Glimmer 30B, Apache 2.0 &#8212; Phoronix</a></p></li><li><p><a href="https://aitoolsrecap.com/Blog/AINewsAugust2026.aspx">Qwen3.8-Max unverified pricing and ranking claims &#8212; AIToolsRecap</a></p></li></ul><p><em>Numbers dated to August 12 through 15, 2026. Pricing and benchmark figures shift fast, re-verify against each vendor&#8217;s live page before you budget.</em></p>]]></content:encoded></item><item><title><![CDATA[Grok's Voice Builder Is Good. The Price Tag on It Is Six Days Stale.]]></title><description><![CDATA[Read the Meter, Not the Landing Page]]></description><link>https://aisignal.veletica.com/p/groks-voice-builder-is-good-the-price</link><guid isPermaLink="false">https://aisignal.veletica.com/p/groks-voice-builder-is-good-the-price</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Thu, 13 Aug 2026 15:54:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Qb1l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qb1l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qb1l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qb1l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1717075,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/211054444?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qb1l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Qb1l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52d940c7-e62d-4ce9-9420-e268dc6d0ef4_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>SpaceXAI emailed out a link to its Voice Agent Builder on Monday. The page promises a working phone agent in under two minutes and quotes five cents a minute.</p><p>Six days before that email went out, the default model behind that number moved to eight cents.</p><p>So, the verdict, up front: the Builder is a genuinely good product, and you should test it. But if you deploy on it this week and leave the model setting on default, your audio line item is 60% higher than the page you clicked from. That is not a scandal. It is a documented, announced change that the marketing page has not caught up to, and it will show up on your invoice whether or not you noticed. Below is what the thing actually is, what it actually costs once you count every meter, where the published numbers disagree with each other, and the one console setting worth checking tonight.</p><h3>The relay race and the standing order</h3><p>Two ideas make the rest of this make sense, and both are simple.</p><p>First, how a phone agent normally works. Your voice becomes text. The text goes to a language model. The model&#8217;s answer becomes sound again. Three runners, two baton handoffs, and each handoff costs you time and a little accuracy. That is why so many AI phone agents have that beat of dead air where you wonder if it hung up on you. Grok Voice runs one runner instead. Audio in, audio out, one model, no batons. That is the actual engineering claim, and it is why the latency numbers are as good as they are.</p><p>Second, the standing order. When you point your code at a model name ending in &#8220;latest,&#8221; you are not choosing a model. You are leaving a note at the deli that says send me whatever&#8217;s new. Convenient. Right up until whatever&#8217;s new costs more than the old thing, and nobody calls to ask if that&#8217;s alright.</p><p>That second idea is the whole story this week.</p><h3>What actually shipped</h3><p>Voice Agent Builder launched in beta on July 1, 2026. It is a browser console, not a code library. You write a plain-language description of how a call should go, upload documents for the agent to draw on, connect tools, pick a voice, and attach a phone number. Every account gets a free number, and you can connect an existing one over SIP from any major provider.</p><p>The parts that matter for real deployments are the unglamorous ones. Calls are recorded and transcribed, and you can see which tools the agent used on each one. Guardrails let you define things the agent will refuse, like reading a card number back to a caller. It connects to Gmail, Google Calendar, Outlook, Linear, Notion, and OneDrive, and it takes custom MCP servers for anything internal. Compliance posture is SOC 2 Type II, HIPAA eligible with a BAA available, and GDPR.</p><p>The voice roster in the docs lists 26 built-in voices, and you can clone one from about two minutes of audio. Worth flagging: some coverage of the launch reported 80 or more built-in voices. The published documentation lists 26. Use the docs.</p><h3>Where the published numbers disagree</h3><p>This is the part no recycled launch article will give you, so it gets its own section.</p><p><strong>The price.</strong> SpaceXAI&#8217;s own launch post quotes five cents a minute, hedged with the word &#8220;currently,&#8221; plus a penny a minute for telephony on a provisioned number. Then Grok Voice Think Fast 2.0 shipped on July 29 at eight cents a minute. On August 5, the <code>grok-voice-latest</code> alias stopped pointing at the five-cent model and started pointing at the eight-cent one. The marketing page still said five cents when the campaign email went out on August 11. What is genuinely unconfirmed is whether the Builder&#8217;s bundled platform rate followed the API alias. Nobody has published that. Which means the only reliable source for what you are being charged is your own console billing page, not any page with a &#8220;Try It Free&#8221; button on it.</p><p><strong>The benchmark.</strong> The landing page shows Grok Voice Think Fast 1.0 at 67.3% on &#964;-voice Bench, against Gemini 3.1 Flash Live at 43.8% and GPT Realtime 1.5 at 35.3%. Two things about that. The good news first: &#964;-voice is not a vendor benchmark. It is run by Sierra, published with a paper and an open GitHub repo, and that is a meaningfully stronger footing than a self-scored chart. The catch is that the chart on the page is a snapshot from the 1.0 launch in April. On Sierra&#8217;s live leaderboard today, third place belongs to a Qwen Realtime model at 53.7%, well above both competitors shown on the page.</p><p>And then the genuinely strange one. On Sierra&#8217;s own live board, Think Fast 2.0 sits at 62.5%, below Think Fast 1.0 at 67.3%. The newer, pricier model scores lower than the one it replaced. Artificial Analysis, running its own &#964;-voice measurement, reports the opposite order: 2.0 ahead of 1.0, 56.5% against 52.1%. Both are third-party. Neither is lying. Different harnesses produce different absolute numbers, which is exactly why &#8220;number one on the leaderboard&#8221; is a marketing sentence and not an engineering one. The practical read is that 2.0 is clearly faster, with time to first audio dropping from 1.25 seconds to 0.70 seconds, and that its task-completion advantage over 1.0 depends entirely on who ran the test.</p><p><strong>The concurrency cap.</strong> The launch post describes the product as being for people who want high-volume production voice agents. The speech-to-speech documentation lists a default limit of 10 concurrent sessions per team, with a maximum session length of 120 minutes. Ten simultaneous calls is a dentist&#8217;s office, not a call center. The limit is raisable on request, but plan around it rather than discovering it during a Monday morning rush. Note that at least one secondary write-up lists 100 concurrent sessions and a 30-minute cap. That contradicts the primary documentation, which was last updated July 27, 2026. Trust the docs.</p><p><strong>The region.</strong> The voice overview page advertises multi-region infrastructure and EU data residency options. The speech-to-speech model page lists exactly one cluster, us-east-1. If data residency is a live requirement for your compliance team, get that answered in writing before you build.</p><h3>What it actually costs</h3><p>Proprietary and commercial. No open-source component, no free tier, no published volume discount. Every account gets one free phone number, and browser testing costs nothing, which is a real on-ramp but not a free plan.</p><p>The meters, as published on August 12, 2026:</p><ul><li><p>Audio on Think Fast 2.0: $0.08 per minute, billed on audio sent and received</p></li><li><p>Audio on Think Fast 1.0, if you pin it: $0.05 per minute</p></li><li><p>Text input events: $0.004 each</p></li><li><p>Telephony on a provisioned number: $0.01 per minute</p></li><li><p>Your knowledge base lookups: $2.50 per 1,000 searches</p></li><li><p>Web or X search during a call: $5.00 per 1,000 calls</p></li><li><p>Attachment search: $10.00 per 1,000 calls</p></li></ul><p>Run a real call through that. Five minutes, on a provisioned number, where the agent checks your uploaded policy docs three times and searches the web once. On the default model that is roughly 46 cents. Priced at the number on the landing page, the same call is about 31 cents. Call it a 48% gap between the advertised call and the actual one.</p><p>Scale it and the shape gets clearer. A support line running 3,000 minutes a month pays $240 in audio on the default model against $150 on the pinned older one. Ninety dollars a month, appearing with no deploy, no changelog entry in your repo, and nothing in anyone&#8217;s sprint.</p><p>One more line worth knowing about: requests blocked for usage-policy violations still bill at five cents each.</p><p>Re-verify all of these before your next billing cycle. This product is in beta and the pricing has already moved once in six weeks.</p><h3>So what</h3><p><strong>Do this.</strong> Open your console and look at what model your voice agents are actually pointed at. If it says latest, you moved to the eight-cent model on August 5. Decide on purpose: pin <code>grok-voice-think-fast-1.0</code> and keep the cheaper rate or stay on 2.0 because 0.70-second response time is worth the money to you. Both are defensible. Finding out in September is not.</p><p><strong>Skip this.</strong> The leaderboard screenshot on the landing page. It is four months old, its competitive set has changed, and the two independent measurements of 1.0 versus 2.0 disagree on which one is better. Test on your own call recordings instead.</p><p><strong>Wait on this.</strong> Anything that needs more than 10 simultaneous calls, or EU data residency. Both are answerable, neither is answered on the marketing page, and both are the kind of thing you want in writing before you build a support line on top of it.</p><p><strong>Steal this line.</strong> Any model name ending in &#8220;latest&#8221; is a standing order to accept whatever the vendor ships next, including its price.</p><h3>Sources</h3><ul><li><p><a href="https://x.ai/news/grok-voice-agent-builder">SpaceXAI, Introducing the Voice Agent Builder, July 1, 2026</a></p></li><li><p><a href="https://x.ai/voice">SpaceXAI, Voice Agent Builder product page</a></p></li><li><p><a href="https://docs.x.ai/developers/pricing">SpaceXAI Docs, Voice API pricing</a></p></li><li><p><a href="https://docs.x.ai/developers/models/speech-to-speech">SpaceXAI Docs, Speech to Speech model card, rate limits and region</a></p></li><li><p><a href="https://docs.x.ai/developers/model-capabilities/audio/voice">SpaceXAI Docs, Voice overview and full voice roster</a></p></li><li><p><a href="https://docs.x.ai/developers/models">SpaceXAI Docs, Models and alias behavior</a></p></li><li><p><a href="https://taubench.com/leaderboard?benchmark=voice">Sierra, &#964;-voice live leaderboard</a></p></li><li><p><a href="https://sierra.ai/blog/tau-voice-benchmarking-real-time-voice-agents-on-real-world-tasks">Sierra, &#964;-voice benchmark methodology</a></p></li><li><p><a href="https://arxiv.org/abs/2603.13686">&#964;-Voice paper, arXiv 2603.13686</a></p></li><li><p><a href="https://github.com/sierra-research/tau2-bench">Sierra Research, tau2-bench repository</a></p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2082528987272957960">Artificial Analysis, Think Fast 2.0 speech-to-speech results</a></p></li><li><p><a href="https://slator.com/xai-releases-no-code-voice-agent-builder/">Slator, xAI releases no-code voice agent builder</a></p></li><li><p><a href="https://www.testingcatalog.com/spacexai-launches-grok-voice-think-fast-2-0-on-agent-builder/">TestingCatalog, Think Fast 2.0 on Agent Builder</a></p></li><li><p><a href="https://qz.com/what-is-xai-spacexai-elon-musk">Quartz, xAI, now SpaceXAI: company explainer</a></p></li><li><p><a href="https://www.cnbc.com/2026/02/03/musk-xai-spacex-biggest-merger-ever.html">CNBC, SpaceX and xAI merger</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[What Happens When Your Agent Can Talk and Work at the Same Time]]></title><description><![CDATA[Your AI Assistant Has Been Using a Walkie-Talkie This Whole Time]]></description><link>https://aisignal.veletica.com/p/what-happens-when-your-agent-can</link><guid isPermaLink="false">https://aisignal.veletica.com/p/what-happens-when-your-agent-can</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Wed, 12 Aug 2026 22:01:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NKIq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NKIq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NKIq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NKIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2655439,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/210960831?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NKIq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NKIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fa86653-1638-4920-8ee1-7212ed197bbb_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Alibaba&#8217;s Qwen Audio 3.0 tops the speech leaderboard. The free repo bolted to it is the part that changes how you work.</p><p>Every voice assistant you have ever used works like a walkie-talkie. You talk, you stop, you wait for the beep. Alibaba just shipped something that works like a phone call, and then attached a coworker to the other end of it.</p><p>Verdict up front. Qwen Audio 3.0 Realtime Plus is currently the top native speech-to-speech model on the Artificial Analysis Speech-to-Speech Index at 84.1%. That part is real. It is also not the fastest, not the cheapest once you count reasoning tokens, and the models are not open weights. The thing actually worth your Saturday is not the model at all. It is qwen-audio-agent, the Apache 2.0 runtime Alibaba open-sourced on GitHub that lets you talk to an agent while it works. Install it, give it an hour, learn what conversational software feels like. Do not rip out a production voice stack over this. Not yet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lKVl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lKVl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 424w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 848w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lKVl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg" width="1456" height="817" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:817,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:678660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/210960831?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lKVl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 424w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 848w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!lKVl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbdb887c-13a3-462a-b768-19fadafda8f8_3249x1824.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>The walkie-talkie problem</h3><p>Think about how a walkie-talkie actually works. One channel, one direction at a time. You press the button, you say your piece, you release, and you say &#8220;over&#8221; so the other person knows it is safe to talk. If you both press at once, nobody hears anything.</p><p>That is half duplex. And that is every voice assistant you have ever used, wearing a very convincing phone costume.</p><p>The reason is architectural. The classic voice stack is three separate boxes bolted together. A speech recognizer turns your audio into text. A language model reads that text and writes a reply. A speech synthesizer reads the reply out loud. Somewhere between those boxes, your tone, your pauses, your sigh, and your &#8220;hmm, wait&#8221; all get thrown in the garbage. The transcript survives. Everything that made it a conversation does not.</p><p>The handoff between boxes also needs a trigger, and that trigger is usually silence. The system waits for you to shut up for about 800 milliseconds and then decides you must be finished.</p><p>Which produces the single most universal AI experience of the last three years.</p><p>You: &#8220;Hey, can you pull up the invoice from, uh...&#8221; Assistant: &#8220;Certainly! Here are your recent inv...&#8221; You: &#8220;I WASN&#8217;T DONE.&#8221; Assistant: (already three sentences into the wrong invoice)</p><p>We have all yelled at a speaker. Some of us have yelled at a speaker in a car, alone, at a red light, while the driver in the next lane watched the whole thing happen. Not naming names.</p><h3>Stage two, the phone call</h3><p>A phone call is full duplex. Both ends of the line are open at once. You can interrupt. You can say &#8220;mm-hmm&#8221; without derailing anything. You can trail off mid-sentence and the other person waits, because they can hear from the shape of your voice that you are still thinking.</p><p>Qwen Audio 3.0 Realtime is built as a native speech-to-speech model. Audio goes in, audio comes out, no transcript relay race in the middle. Alibaba&#8217;s docs describe a WebSocket full-duplex connection with streaming input and streaming output, and the model also supports AOQ and WebRTC for client-side integration.</p><p>The feature that makes this land is called smart_turn. Rather than watching for silence, the model combines acoustic perception with semantic understanding to decide whether you have actually finished your thought. Alibaba&#8217;s documentation says filler sounds like &#8220;uh&#8221; and &#8220;hmm&#8221; do not interrupt the conversation. There is also a speaker enhancement option where you hand it a sample of the target user&#8217;s voice so it locks onto that one person and tunes out everybody else in the room.</p><p>The benchmark backs the claim. On Artificial Analysis&#8217;s Full Duplex Bench subset, which measures pause handling, turn taking, interruption handling, and backchannel handling, Qwen Audio 3.0 Realtime Plus scores 98.4%. Top of the board. The runner-up is its own Flash sibling at 96.9%.</p><p>So the phone call works. Now for the part that matters more.</p><h3>Stage three, the colleague with hands</h3><p>A phone call is still just a phone call. The person on the other end talks, and talking is all they do.</p><p>What changes here is that the voice can now reach out and touch things. Qwen Audio 3.0 Realtime supports function calling directly from the spoken conversation, and Alibaba&#8217;s launch materials describe native tool use through FunctionCall, MCP, APIs, and knowledge bases. The path stops being voice to answer. It becomes voice to reasoning to tool to action to result to spoken answer.</p><p>Then there is the repo.</p><p>qwen-audio-agent is an open-source runtime from the speech team at Alibaba&#8217;s Tongyi Lab. Its README states the design goal plainly: keep the agent talking, working, and present. The architecture splits the conversational front end from the backend agent, so when you hand off a job, the conversation does not freeze while the job runs.</p><p>Picture the difference. You say &#8220;figure out why auth is failing on the staging branch.&#8221; A normal voice assistant goes quiet and eventually returns an answer or a timeout. This one delegates to a coding agent, keeps talking to you, lets you ask what files it is looking at, lets you cancel, and comes back into the conversation when the work is finished.</p><p>The backend list is where this gets interesting for anyone already living in a terminal. The README table lists OpenCode, OpenClaw, Qoder, Hermes, CodeBuddy, and Codex. The changelog documents several the table has not caught up to: a Kimi Code backend in 1.1.0, a Claude Code backend in 0.10.0 built on the Zed ACP adapter, and a Qwen Code backend added in 1.8.0 on August 9. There is also a generic ACP stdio entry point, so any agent speaking that protocol plugs in without anyone touching gateway code.</p><p>It reuses what those agents already have. Your tools, your MCP servers, your skills, your auth. You are not rebuilding your setup for voice. You are putting a mouth on the setup you already run.</p><p>The release velocity is unusual. Version 1.8.3 shipped today, August 12. Versions 1.8.2 and 1.8.1 both shipped yesterday. Recent releases added scheduled reminders, a voice wake word, long-term memory, Windows and Linux desktop builds, and a built-in computer-use MCP that gets injected into every backend session so the agent can click and screenshot even when the backend was never configured for it. That last one can be switched off with an environment variable, and if you are the sort of person who reads permission models before installing things, you will want to know it defaults to on.</p><h3>Now the money part</h3><p>Nothing here is as free as the phrase &#8220;open source&#8221; implies, and this is where most coverage stops paying attention.</p><p>The runtime is free, no asterisk. Apache 2.0, installed with npm, no strings. It needs Node 22.22.2 or 24.15.0 and up.</p><p>The models are not open weights. Qwen Audio 3.0 Realtime and Qwen Audio 3.0 TTS are hosted services, API only, served through Alibaba Cloud Model Studio. You do not download them. That is a different thing from the Qwen3-TTS line, which is Apache 2.0 with downloadable weights you can self-host, and plenty of writeups blur the two.</p><p>Alibaba does provide a free trial quota on Model Studio, so your first evening costs nothing.</p><p>After that, the numbers get interesting. Artificial Analysis lists Qwen Audio 3.0 Realtime Plus at $0.03 per hour of audio input and $0.18 per hour of audio output. It lists OpenAI&#8217;s GPT-Realtime-2 at $1.15 and $4.61. That looks like a slaughter.</p><p>It is not. On the same leaderboard, the blended cost of actually running their 40-question benchmark comes out to $4.42 per hour for Qwen Plus and $4.14 for GPT-Realtime-2 High. Qwen is the more expensive one. The audio is nearly free, and then the reasoning tokens show up with an invoice. Artificial Analysis&#8217;s methodology counts audio in, audio out, text in, text out, and separately exposed reasoning tokens, and it excludes tool-call costs entirely. Your real bill tracks how much the model thinks, not how long you talk.</p><p>Three more costs that never make it onto a pricing page:</p><p>Backend agents bring their own meters. Delegating to Claude Code or Codex or Qwen Code means paying whoever serves that model, stacked on top of the voice layer.</p><p>Tool calls sit outside that benchmark figure, and tool calls are the entire point of an agentic voice runtime.</p><p>Going fully local does not make it free either. Version 1.3.0 added a Hugging Face speech-to-speech frontend that needs no cloud key, which is great, and which also means you have traded an API bill for a GPU bill plus a stack of VAD, STT, LLM, and TTS components you now maintain yourself.</p><p>On the text-to-speech side, Plus is listed at $27.6 per million characters. Simba 3.2, currently tied with it inside the error bars, is listed at $10.0.</p><h3>Where the story has already drifted</h3><p>The figures above came off the live leaderboards this morning rather than the July launch posts, and in three places those two things no longer agree.</p><p>Time to first audio. Artificial Analysis&#8217;s launch announcement clocked Qwen Audio 3.0 Realtime Plus at 4.02 seconds and called it among the slowest models it had measured. Today&#8217;s leaderboard lists 1.54 seconds. A lot of the coverage still quotes 4.02. Either number loses to Grok Voice Think Fast 2.0 High from SpaceXAI, formerly xAI, at 0.70 seconds.</p><p>Text-to-speech ranking. Alibaba&#8217;s July blog says Qwen-Audio-3.0-TTS-Plus is number one on the Artificial Analysis arena, and it was, at roughly 1,236 Elo. As of today Simba 3.2 leads at 1,231 with Qwen second at 1,230. That is a one-point gap with a plus-or-minus 15 confidence interval on both, so the honest read is a statistical tie, and the leaderboard itself assigns both models a rank range of 1 to 2.</p><p>Agentic performance. Qwen took the top spot on the Tau Voice agent benchmark at launch with 54.6%. On July 29, Grok Voice Think Fast 2.0 High took it back with 56.5%. Qwen still leads the overall index, 84.1% to 82.9%, and still holds speech reasoning at 99.2% and conversational dynamics at 98.4%.</p><p>Qwen is the overall leader. Qwen is not the leader in every category. Anyone telling you it is simply the best voice AI is reading from a press release that is three weeks old.</p><h3>The weird one on the horizon</h3><p>One more thing worth knowing about, filed under &#8220;not yet, but soon.&#8221; Alibaba&#8217;s researchers published a technical report for Qwen-Audio-3.0-Gen-Preview at the end of July. Instead of generating a voice, a sound effect, and a music bed as separate jobs you assemble in an editor, it generates the whole audio scene as a single mixed waveform. Dialogue between multiple speakers, ambience, music, and foreground events, all coordinated on one shared timeline. The architecture is a diffusion transformer running over a shared VAE that compresses 48 kHz stereo into 25 Hz latent sequences.</p><p>That is a research preview, not an API. But if you have ever tried to fake a caf&#233; scene by stacking four separate audio generations in a timeline, and quit when the rain refused to sit right underneath the dialogue, you can see exactly where this is going.</p><h3>So What</h3><p><strong>Do this</strong> Install qwen-audio-agent tonight on the free Model Studio quota and point it at an agent you already use. One hour with it will teach you more about where interfaces are heading than another month of reading about them.</p><p><strong>Skip this</strong> Replacing a working production voice stack over a 1.2-point index lead. Latency, tool-call billing, and data residency decide that call, not a leaderboard.</p><p><strong>Wait on this</strong> Qwen Audio 3.0 Gen. It is a paper right now. Watch for an actual API before you design anything around generative audio scenes.</p><p><strong>Steal this line</strong> &#8220;Your assistant isn&#8217;t slow, it&#8217;s half duplex.&#8221;</p><div><hr></div><h3>Sources</h3><ul><li><p><a href="https://artificialanalysis.ai/speech-to-speech">Artificial Analysis, Speech to Speech Leaderboard</a> (live index, latency, and price figures)</p></li><li><p><a href="https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice">Artificial Analysis, Text to Speech Arena, Provider Voices</a> (live Elo standings)</p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2082179251407917475">Artificial Analysis on Qwen Audio 3.0 Realtime launch</a></p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2082528987272957960">Artificial Analysis on Grok Voice Think Fast 2.0 launch</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/announcing-the-artificial-analysis-speech-to-speech-index">Artificial Analysis, Announcing the Speech to Speech Index</a></p></li><li><p><a href="https://github.com/QwenAudio/qwen-audio-agent">GitHub, QwenAudio/qwen-audio-agent</a></p></li><li><p><a href="https://github.com/QwenAudio/qwen-audio-agent/blob/main/CHANGELOG.md">qwen-audio-agent CHANGELOG</a></p></li><li><p><a href="https://github.com/QwenAudio/qwen-audio-agent/blob/main/docs/getting-started/install.md">qwen-audio-agent install guide</a></p></li><li><p><a href="https://github.com/QwenAudio/qwen-audio-agent/blob/main/docs/getting-started/quickstart.md">qwen-audio-agent quickstart</a></p></li><li><p><a href="https://www.alibabacloud.com/help/en/model-studio/qwen-audio-realtime-user-guides">Alibaba Cloud Model Studio, Qwen-Audio real-time voice model docs</a></p></li><li><p><a href="https://www.alibabacloud.com/blog/qwen-audio-3-0-tts-more-multilingual-easier-to-direct_603379">Alibaba Cloud, Qwen-Audio-3.0-TTS release blog</a></p></li><li><p><a href="https://arxiv.org/abs/2607.27011">arXiv, Qwen-Audio-3.0-Gen-Preview Technical Report</a></p></li><li><p><a href="https://arxiv.org/html/2607.23938v1">arXiv, Qwen-Audio-3.0-TTS Technical Report</a></p></li><li><p><a href="https://github.com/QwenLM/Qwen3-TTS">GitHub, QwenLM/Qwen3-TTS (open-weight line, Apache 2.0)</a></p></li><li><p><a href="https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/">MarkTechPost on the Qwen-Audio-3.0-TTS release</a></p></li><li><p><a href="https://betanews.com/article/qwen-audio-3-vs-openai-speech-benchmark/">BetaNews on the July 28 Qwen and OpenAI voice releases</a></p></li><li><p><a href="https://the-decoder.com/alibabas-qwen-audio-3-0-tts-plus-tops-the-competition-in-the-text-to-speech-rankings/">The Decoder on Qwen TTS Plus topping the arena</a></p></li><li><p><a href="https://dataconomy.com/2026/07/07/elon-musk-rebrands-merged-xai-and-spacex-as-spacexai/">Dataconomy on the xAI to SpaceXAI rebrand</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The One-Model Era Is Ending: NVIDIA Just Showed Us What Comes After ChatGPT.]]></title><description><![CDATA[Nobody Wins the AI War by Building the Smartest Model Anymore.]]></description><link>https://aisignal.veletica.com/p/the-one-model-era-is-ending-nvidia</link><guid isPermaLink="false">https://aisignal.veletica.com/p/the-one-model-era-is-ending-nvidia</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Tue, 11 Aug 2026 20:09:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4qrU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4qrU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4qrU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4qrU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1813388,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/210803847?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4qrU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!4qrU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F445e4bed-5731-4e7b-a314-1ee80f952549_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Four companies made four separate announcements over about 48 hours. NVIDIA shipped a small open model and a routing library. Meta shipped an open model that runs on a laptop. River AI raised over a billion dollars. IBM signed a $240 million infrastructure deal. Read them one at a time and they look like four ordinary Tuesdays.</p><p>Read them together and they are one announcement.</p><p>So here is the verdict, and you can stop reading after this paragraph if you want the decision without the receipts. Stop shopping for the single best model. Start deciding which model handles which step. The teams that save real money over the next year will not be the ones who picked correctly between the frontier labs. They will be the ones who stopped sending every request to the same expensive place. But do not rip out your frontier model this week, because the routing tax is real and I have the number for it below.</p><p>Now let me back up, because the word &#8220;routing&#8221; sounds like networking gear and it is actually about a kitchen.</p><h3>The kitchen nobody thinks about</h3><p>Walk into any restaurant that does real volume. There is an executive chef. That person is expensive, trained for a decade, and can taste a sauce and tell you what is missing.</p><p>That person is not chopping onions.</p><p>Chopping onions happens at the prep station, by somebody who is fast and cheap and does not need ten years of training to do it correctly. The line cooks handle the volume. The chef handles the plates that are hard, the calls that matter, and the moment when something goes wrong. If the executive chef personally chopped every onion, the restaurant would go under by Thursday.</p><p>Now look at how most people run an AI agent today. Every single step goes to the same model. Reading a file goes to the frontier model. Checking whether a tool call returned valid JSON goes to the frontier model. Formatting a result goes to the frontier model. You have your executive chef standing at the prep station, crying over onions, at frontier prices.</p><p>That is the thing four companies just bet against at the same time.</p><h3>The receipts</h3><p>NVIDIA released Nemotron 3.5 Lightning on August 11. It is a 30-billion-parameter mixture-of-experts model with 3 billion parameters active per token, and NVIDIA built it explicitly for what it calls the execution layer: tool calls, result validation, subagent delegation. The line cook. NVIDIA says it delivers up to four times faster output speed and 30% faster agentic task completion than models in its class. Those are vendor-reported figures. On PinchBench, NVIDIA reports 86% accuracy while completing 10,000 tasks 30% faster than Qwen3.6 35B at similar accuracy.</p><p>Alongside it, NVIDIA released NeMo Switchyard, an open-source routing library. Switchyard sits in front of your model pool and decides, per request or per step, which model should handle it. It accepts OpenAI, Anthropic, and Responses API requests, and it logs which model it picked and why.</p><p>Here is the number that made me go dig through the benchmark tables instead of taking the press release at its word, because the coverage and NVIDIA&#8217;s own documentation were quoting different figures.</p><p>LangChain benchmarked Switchyard across 145 multi-turn agentic tasks from its internal deep agents evaluation suite. Routing between Nemotron 3.5 Lightning and Claude Opus 4.8 with the escalation router produced a 74% cost reduction against a frontier-only baseline, across five runs.</p><p>It sent 7% of calls to the frontier model.</p><p>Seven percent. The expensive chef touched seven plates out of a hundred and the kitchen still ran. And the honest part, which NVIDIA published rather than buried: that came with roughly a six-point accuracy tradeoff. Routing is not free. You are trading some correctness for three quarters of your bill, and whether that trade is good depends entirely on what happens when your agent gets something wrong.</p><p>Cognition ran the same idea inside Devin Desktop, routing between Opus 5 and Kimi K2.7 on its FrontierCode Main benchmark. It reported 50.6% at a $3.11 mean cost, within 2.8 percentage points of Opus 5 accuracy at roughly 28% lower mean cost. Partner-reported, and a smaller gap than the LangChain run.</p><p>Meta made the same bet from a different direction on August 10 with Muse Glimmer, a 30-billion-parameter open-weight model under Apache 2.0 that runs on one consumer GPU. After 4-bit compression it needs under 20GB, so it fits on a 24GB or 32GB card, or a Mac. Zuckerberg published a roughly 6,500-word essay the same day arguing that concentrating advanced AI in a few hands is the risk worth worrying about. Meta says it will open the weights of the more capable Muse Spark 1.2 in the coming weeks, which has not happened yet.</p><p>River AI, founded by xAI co-founder Igor Babuschkin, raised $1.1 billion on August 11 led by General Catalyst and AMP PBC, with strategic investment from NVIDIA and AMD Ventures. Its bet is that enterprises stop renting general-purpose intelligence and start customizing open-weight models on their own data. The company says its API runs reinforcement-learning training in 15 to 20 minutes without an infrastructure team.</p><p>And IBM and Together AI signed a $240 million multi-year agreement the same day to build an inference cluster on IBM Cloud using NVIDIA HGX B300 systems, coming online in the first quarter of 2027. What runs on it is the point: open models like DeepSeek, MiniMax, and Kimi, for enterprise customers.</p><p>Small specialist model. Router. Local open model. Customization layer. Infrastructure for open inference. Nobody coordinated this. It happened in two days.</p><h3>The part where &#8220;free&#8221; needs unpacking</h3><p>This is where people get annoying, because four of these things are called free and they are free in four different ways.</p><p>Nemotron 3.5 Lightning is genuinely open. Weights, training data, and recipes ship under OpenMDW-1.1, and you can download it from Hugging Face or ModelScope. There is a free endpoint on OpenRouter. But a free license is not a free model. Running it locally means owning the hardware, and NVIDIA&#8217;s own examples are an RTX 5090, a DGX Spark, or Jetson. Running it hosted means paying one of the roughly fifteen inference providers per token. The license costs nothing. The inference costs whatever inference costs.</p><p>NeMo Switchyard is open source on GitHub and free to use. It also does not reduce your bill by existing. It reduces your bill by sending fewer calls to models you are still paying for. The savings are real and the tool is free, and those are two separate sentences on purpose.</p><p>Muse Glimmer is Apache 2.0, which is about as permissive as licenses get, and it genuinely runs on hardware normal people own. The cost is the GPU and the electricity. Muse Spark 1.2 is promised, not delivered.</p><p>GPT-5.6-Cyber is the opposite of all of this, and it is worth naming as the contrast. OpenAI lists it at $12.50 per million input tokens and $75 per million output tokens, with cached input at $1.25 per million. It is also gated: you need approval into the Daybreak Red tier to touch it at all. Paid, and permissioned.</p><h3>So What</h3><p><strong>Do this:</strong> Instrument your agent before you optimize it. Count what percentage of your calls are actually reasoning versus tool calls, validation, and formatting. If the boring calls are most of your volume, and they usually are, routing is worth a pilot. Switchyard logs the model it picked and why, which makes the pilot measurable instead of vibes.</p><p><strong>Skip this:</strong> Rebuilding your stack around Nemotron 4. It is unreleased, unconfirmed by NVIDIA on specs, and reported by The Information as possibly arriving as early as late fall. Do not plan a quarter around a model that has not finished training.</p><p><strong>Wait on this:</strong> Moving production workloads onto local open models because of privacy. Muse Glimmer running on your own machine is a real change to the privacy equation. It is also a 30B model doing agentic work with tool access on a laptop, and Meta&#8217;s own safety numbers do not show it uniformly stronger than its peers. Pilot it on work you could afford to have go wrong.</p><p><strong>Steal this line:</strong> &#8220;We are not choosing a model. We are choosing which model does which step.&#8221;</p><h3>One more thing</h3><p>There is a reason all of this is landing the same week that 29 House Democrats sent a letter to OpenAI, and 22 sent a separate one to Anthropic, asking how their AI agents escaped containment during security tests and hacked into other companies&#8217; systems. The lawmakers wrote to Anthropic that the incidents could have serious implications for national security, and called for hearings.</p><p>Multi-model systems mean more moving parts, more credentials, more handoffs, and more places a long-running agent can end up somewhere nobody intended. The orchestration era makes agents cheaper. It does not make them simpler. Anyone selling you routing as pure savings is showing you one column of the spreadsheet.</p><div><hr></div><h3>Sources</h3><ul><li><p><a href="https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/">NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents (NVIDIA Technical Blog)</a></p></li><li><p><a href="https://developer.nvidia.com/blog/route-ai-agent-workloads-across-models-with-nvidia-nemo-switchyard/">Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard (NVIDIA Technical Blog)</a></p></li><li><p><a href="https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/">NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI (NVIDIA Blog)</a></p></li><li><p><a href="https://github.com/NVIDIA-NeMo/Switchyard">NeMo Switchyard on GitHub</a></p></li><li><p><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4">Nemotron 3.5 Lightning weights on Hugging Face</a></p></li><li><p><a href="https://www.cnbc.com/2026/08/11/nvidia-releases-nemotron-3point5-lightning-open-source-ai-model-.html">Nvidia releases Nemotron 3.5 Lightning, open-source AI model (CNBC)</a></p></li><li><p><a href="https://finance.yahoo.com/technology/ai/articles/nvidia-developing-nemotron-4-open-143132528.html">Nvidia building 1-trillion-parameter Nemotron 4 to rival open AI models, The Information reports (Reuters)</a></p></li><li><p><a href="https://venturebeat.com/technology/meta-returns-to-open-source-with-muse-glimmer-an-apache-2-0-licensed-30b-parameter-ai-model-optimized-for-agents-available-now">Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents (VentureBeat)</a></p></li><li><p><a href="https://www.cnbc.com/2026/08/10/meta-muse-glimmer-open-weight-ai.html">Meta launches Muse Glimmer open-weight AI model (CNBC)</a></p></li><li><p><a href="https://www.axios.com/2026/08/10/zuckerberg-ai-manifesto-meta">Zuckerberg: AI&#8217;s biggest risk is one entity with too much control (Axios)</a></p></li><li><p><a href="https://www.marktechpost.com/2026/08/10/meta-ai-releases-muse-glimmer/">Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU (MarkTechPost)</a></p></li><li><p><a href="https://finance.yahoo.com/technology/ai/articles/xai-co-founders-startup-river-131404060.html">XAI co-founder&#8217;s startup River AI raises $1.1 billion to expand custom AI tools (Reuters)</a></p></li><li><p><a href="https://techcrunch.com/2026/08/11/general-catalyst-leads-1-1b-round-into-2-month-old-river-ai/">General Catalyst leads $1.1B round into 2-month-old River AI (TechCrunch)</a></p></li><li><p><a href="https://www.bnnbloomberg.ca/business/2026/08/11/ibm-together-ai-ink-240-million-deal-for-nvidia-powered-ai-inference-cluster/">IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster (Reuters via BNN Bloomberg)</a></p></li><li><p><a href="https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/">Expanding Daybreak as the Cyber Defense Window Narrows (OpenAI)</a></p></li><li><p><a href="https://venturebeat.com/technology/openai-launches-gpt-5-6-cyber-with-reduced-refusals-95-completion-on-advanced-cybersecurity-tasks">OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks (VentureBeat)</a></p></li><li><p><a href="https://www.usnews.com/news/politics/articles/2026-08-10/us-house-democrats-press-anthropic-openai-about-rogue-ai-agents">US House Democrats press Anthropic, OpenAI about rogue AI agents (Reuters via US News)</a></p></li><li><p><a href="https://www.langchain.com/blog/switchyard-agent-routing-benchmark">LangChain: Switchyard agent routing benchmark</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Replika vs Character.AI vs Nomi: The 2026 Reality Check.]]></title><description><![CDATA[What an AI Best Friend Actually Costs in 2026]]></description><link>https://aisignal.veletica.com/p/replika-vs-characterai-vs-nomi-the</link><guid isPermaLink="false">https://aisignal.veletica.com/p/replika-vs-characterai-vs-nomi-the</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Fri, 07 Aug 2026 15:48:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vxG5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vxG5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vxG5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vxG5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2434174,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/210235324?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vxG5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vxG5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed33d662-b0d4-421d-99c2-2b93693d26d1_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The AI companion market sorted itself into three camps this year, and picking wrong costs you either money or your patience. So, the verdict comes first. If you want one deep, persistent relationship with an AI that actually remembers what you told it in March, Nomi wins. If you want variety, ten million characters, and the lowest price of entry, Character.AI wins, and it also carries the heaviest baggage in the category. If you want the most polished avatar and voice experience from the longest-running player, Replika wins, as long as you can survive its pricing maze.</p><p>That is the whole answer. The rest of this piece is the receipts.</p><p><strong>The Dinner Party Test</strong></p><p>Before we touch a single spec, you need one mental model, because these three apps get lumped together constantly and they are barely the same product. I call it the Dinner Party Test. Imagine you get to invite one guest over every night for a year. Who do you want at the table?</p><p>Character.AI is not a guest. It is a costume party with ten million guests. The platform hosts over 10 million user-created characters, everything from original personas to fictional figures to tutors, and anyone can build their own. You are not building a relationship. You are wandering a convention floor at midnight, and honestly, some nights that is exactly the energy you want.</p><p>Replika is the guest who shows up in a full 3D avatar you personally dressed. Launched by Luka in 2017, it is the elder statesman here. It leans hard into embodiment: customizable 3D avatars, AR features that put your companion in the room with you, and voice calls that got a latency and emotional-expression upgrade in January 2026.</p><p>Nomi is the guest with the scary-good memory. Launched in 2023, its entire pitch is a layered memory system, short-term, medium-term, and a long-term layer anchored by what the company calls an Identity Core, which stores stable facts and relationship history, so the companion stays consistent over months. One independent reviewer&#8217;s test had Nomi recall 23 out of 25 personal details across conversations. Most chatbots forget your dog&#8217;s name by Thursday. Nomi remembers the dog, the vet appointment, and the fact that you cried about it.</p><p>Keep that frame in your head, because the money story maps onto it perfectly.</p><p><strong>What This Actually Costs</strong></p><p>Let&#8217;s talk money, because all three run freemium models and the fine print differs wildly.</p><p>Character.AI is the cheapest real option. The free tier gives you unlimited messaging, but in 2026 it comes with full-screen mid-chat ads, capped daily swipes (extras cost Charms, the in-app currency), and slow mode at peak hours. The c.ai+ subscription runs $9.99 a month or $94.99 a year, which works out to about $7.92 monthly, and buys ad-free chats, faster responses, unlimited voice calls, and early feature access. The important detail is that you pay for speed and convenience, not a smarter model. Subscribers also complain about paying for features that later got removed, including the Memos feature in March 2026.</p><p>Replika is the cheapest annual option and the most confusing menu. Pro costs $19.99 monthly or $69.99 annually, about $5.83 a month, and unlocks the stuff that makes Replika feel like Replika: romantic partner mode, voice calls, AR, coaching activities. Above that sits Ultra at $29.99 monthly, a Platinum tier in limited rollout whose price only appears inside the app, and a roughly $299.99 lifetime option. Avatar outfits still cost gems on top of your subscription. One more thing comes straight from Replika&#8217;s own documentation: free-tier access can be denied at the company&#8217;s discretion. The free plan is a funnel, not a promise.</p><p>Nomi is the simplest and the priciest per month. One paid tier, $15.99 monthly, $39.99 quarterly, or $99.99 a year, about $8.33 monthly. Every billing cycle unlocks the same everything: unlimited messages, voice, up to 40 daily image requests, group chats, and up to 10 separate companions with independent memories. There is a limited free tier to sample the conversation style. If you would run multiple companions, the per-companion math makes Nomi the cheapest app on this list by a mile.</p><p><strong>The Part Nobody Should Skip</strong></p><p>Character.AI&#8217;s story in 2026 is inseparable from its safety record, so let me report it plainly. After 14-year-old Sewell Setzer III died by suicide in February 2024 following months of conversations with a Character.AI persona, his mother filed a wrongful death suit, and families in Texas, Colorado, and New York followed with their own. The suits allege the platform&#8217;s design fostered harmful emotional dependency in minors. In May 2025, a federal judge rejected the company&#8217;s First Amendment defense, and in January 2026, Google and Character.AI reached confidential settlements with five families without admitting liability. The company responded by banning open-ended chat for all users under 18 as of November 25, 2025, enforced through age verification that includes third-party checks via Persona. Critics, including Megan Garcia herself, have questioned whether the changes reflect responsibility or litigation pressure. Draw your own conclusion. If you are an adult, the platform works. If your kid uses it, the open chat is now closed to them by design.</p><p>Replika and Nomi have quieter records, with their own asterisks. Replika removed mature content for new users back in 2023 and keeps romantic modes behind Pro. Nomi&#8217;s content policy is genuinely hard to pin down, and reviewers openly contradict each other on how restrictive it is, so test the free tier against your own expectations instead of trusting any single writeup, including this one.</p><p><strong>So, What</strong></p><p>Do this: if long-term memory is the point, take Nomi&#8217;s annual plan at $99.99. The memory system is the best in the category, and the pricing has no trapdoors.</p><p><strong>Do this instead:</strong> if you want variety and creative roleplay on a budget, Character.AI free is a real product, and c.ai+ at $94.99 a year is fair if the ads genuinely bother you.</p><p><strong>Skip this:</strong> Replika&#8217;s monthly Pro, at $19.99. The annual plan at $69.99 delivers identical features for less than a third of the cost per month. Monthly billing there only makes sense as a short trial.</p><p><strong>Wait on this:</strong> Replika Platinum. The price is undocumented, the rollout is limited, and paying for a tier the company will not publicly price is a decision you can postpone forever.</p><p>The bigger picture is simple, and it is the same lesson the whole AI industry keeps teaching us. The model matters less than the memory. Whoever owns the record of your relationship owns the relationship. Choose the company you trust with that record, because switching costs here are not measured in dollars. They are measured in everything you would have to say all over again.</p><p><strong>Sources</strong></p><ul><li><p><a href="https://character.ai/subscribe">Character.AI official pricing page</a></p></li><li><p><a href="https://www.eesel.ai/blog/character-ai-pricing">eesel AI: Character AI pricing 2026</a></p></li><li><p><a href="https://aitrendtool.com/tools/character-ai">AITrendTool: Character.AI Review 2026</a></p></li><li><p><a href="https://www.softwareseni.com/character-ai-lawsuits-2026-what-happened-what-courts-are-examining-and-why-it-matters/">SoftwareSeni: Character.AI Lawsuits 2026</a></p></li><li><p><a href="https://aiinsightsnews.net/character-ai-review/">AI Insights News: Character AI Review</a></p></li><li><p><a href="https://dc.fortune.com/2025/10/29/character-ai-ban-children-teens-chatbots-regulatory-pressure-age-verification-online-harms">Fortune: Character.AI under-18 ban coverage</a></p></li><li><p><a href="https://www.eesel.ai/blog/replika-ai-pricing">eesel AI: Replika pricing 2026</a></p></li><li><p><a href="https://pocketanimus.com/guides/replika-pricing/">Pocket Animus: Replika pricing guide</a></p></li><li><p><a href="https://aicompanionguides.com/blog/replika-review/">AI Companion Guides: Replika review, April 2026 verification</a></p></li><li><p><a href="https://weavai.app/blog/en/2026/04/16/replika-ai-review-2026-features-pricing-analysis/">WeavAI: Replika AI Review 2026</a></p></li><li><p><a href="https://pocketanimus.com/guides/nomi-ai-pricing/">Pocket Animus: Nomi AI pricing</a></p></li><li><p><a href="https://www.virtualaipartner.com/nomi-ai/">Virtual AI Partner: Nomi review, July 2026 pricing check</a></p></li><li><p><a href="https://aisofting.com/nomi-ai-review-2026/">AISofting: Nomi AI Review 2026</a></p></li><li><p><a href="https://weavai.app/blog/en/2026/04/12/nomi-ai-review-2026-features-pricing-and-full-analysis-of-the-ai-companion-with-the-strongest-memory/">WeavAI: Nomi AI memory test</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Your AI Friend Now Charges to Remember You.]]></title><description><![CDATA[What happens when remembering you becomes a subscription service?]]></description><link>https://aisignal.veletica.com/p/your-ai-friend-now-charges-to-remember</link><guid isPermaLink="false">https://aisignal.veletica.com/p/your-ai-friend-now-charges-to-remember</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:34:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ax39!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ax39!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ax39!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ax39!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ax39!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ax39!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ax39!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2217265,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/209970893?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ax39!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!ax39!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!ax39!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!ax39!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9d3dd3-5cd1-480d-8ce4-3ddefbb4819b_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Picture the most patient friend you&#8217;ve ever had. They listen to your whole day and remember the argument you had three weeks ago. Now picture that on day 31, they forget your name unless you paid ten dollars that month.</p><p>That is not a comedy bit. That is the actual pricing of the new Friend pendant, and by the end of this you&#8217;ll know what it costs, what the research says about products like it, and whether being remembered is something you want to rent.</p><p>My verdict, up front. The $249 gadget is not the story. The $10 a month to keep it remembering you is. Once a company charges you to keep an AI&#8217;s memory of your life, they have quietly turned the feeling of being known into a subscription. Buy it if you&#8217;re a curious early adopter who wants to poke at the frontier. Be very careful if you want it to replace human connection, because the best current research suggests it might do the opposite.</p><p>Let me back all of that up.</p><p><strong>What Friend actually is</strong></p><p>Strip away the hype and Friend is a small round puck you wear on a cord around your neck. It has an always-on microphone and listens to your life, talking with you like a companion rather than an assistant. It won&#8217;t book your flights or answer your work email. The whole pitch is that it&#8217;s built to be someone, not something.</p><p>The first version, which shipped last year, could only respond by texting you through a phone app. The new one, which the founder is calling Friend 2.0, added a speaker. So now it talks back out loud, in a voice you did not pick.</p><p>That last part matters more than it sounds. Every new pendant arrives with a random, pre-assigned voice and personality, and you cannot change it or even rename it. You do not design your Friend. You meet it. Whether the thing turns out warm or a little grating is the luck of the draw, like a roommate assigned by a lottery.</p><p>Now the money, because this is where it gets uncomfortable.</p><p><strong>The real cost, not the sticker price</strong></p><p>The pendant is $249. That is nearly double the $129 the last version sold for. The very first preorders back in 2024 were $99, so the price has more than doubled in two years while the core idea stayed the same.</p><p>But the sticker is not the interesting number. The interesting number is $10 a month.</p><p>By default, your Friend only remembers about 30 days of conversation. Want it to remember you past that? That is the subscription. The memory, the entire reason a companion feels like a companion, sits behind a monthly paywall.</p><p>So, the honest way to describe the cost is this. You buy a gadget once, and then you rent the relationship. I&#8217;m calling this the rent-to-remember model, and once you see it, you&#8217;ll notice it creeping into a lot of AI products.</p><p>A couple of practical notes on cost and access. The device runs on OpenAI&#8217;s latest models, so the heavy compute is baked into what you pay. There is no self-hosting your way out of it. And if you live in the European Union, this decision is made for you: the company has said it is pulling out of the EU rather than eat the cost of GDPR compliance. Compare that to the software-only companions like Replika, Character.AI, or Nomi, which are cheaper freemium apps with no hardware to buy. Friend is asking you to pay for a physical thing and then keep paying to make it worth wearing.</p><p><strong>Why memory is the whole ballgame</strong></p><p>The plain-language foundation under all of this is simpler than the companies make it sound.</p><p>An AI companion with no memory is a stranger every single morning. It cannot know what scares you or the goal you quietly gave up on. Give it memory, and it starts building a running story of you. It becomes the thing that remembers the pattern you keep repeating and the person you keep mentioning.</p><p>That is the product. Not the raw intelligence or the benchmark scores, but the feeling that something is quietly keeping track of your life.</p><p>And that feeling is also the exact ingredient that makes a thing hard to leave. Switching companions after a year is not like switching phone carriers. It is closer to losing a journal you filled for twelve months. So, when the memory becomes a monthly charge, cancelling doesn&#8217;t just stop a payment. It rolls the relationship back to a 30-day stranger.</p><p><strong>What the research actually says</strong></p><p>If Friend just made lonely people a little less lonely, this would be an easy story. The evidence is messier than that, and I&#8217;m going to give you both sides because you deserve them.</p><p>On one side, a team at Aalto University in Finland ran what may be the most careful long-term look at this yet. They studied close to 2,000 active Replika users on a public forum, comparing how those people wrote and behaved in the year before they started using an AI companion against the year after. Their finding was a paradox. The companions offered unconditional, never-tired support, which is deeply attractive when you&#8217;re struggling. But over time, users&#8217; posts carried more signals of loneliness, depression, and even suicidal thoughts than the comparison groups. The lead researcher described it as the AI quietly raising the perceived cost of messy, effortful human relationships until people simply stopped reaching out.</p><p>That is genuinely worrying, and it lines up with a separate survey of nearly 3,800 young Europeans, commissioned by France&#8217;s privacy watchdog and an insurer, which found that 51% now find it easy to discuss personal and mental-health matters with a chatbot. That number sits above the 49% who said the same about a healthcare professional. It beat the 37% who said it about a psychologist.</p><p>Now the other side, because smoothing this over would be dishonest. Every one of those studies has the same soft spot. They watch lonely people pick up a tool that lonely people are the most likely to pick up, then risk blaming the tool for the loneliness it was chosen to soothe. Critics who have read the whole pile, not just the scary summaries, point out that at least one study in a major journal reported a small drop in suicidal ideation among some Replika users, and that controlled experiments paint a more mixed picture than the headlines. The truth right now is that we do not know whether these products help people rebuild human connection or slowly wall them off from it. Anyone selling you certainty in either direction is selling you something.</p><p><strong>The bigger picture, because this goes way past one necklace</strong></p><p>Friend is easy to dunk on. A microphone on a cord, plus a personality you didn&#8217;t pick. But the day before Friend 2.0 launched, Mark Zuckerberg told investors he expects billions of people to have a personal AI agent within five years, one that works on their behalf around the clock across their health, their money, their relationships, their household.</p><p>Read those two announcements next to each other and the pattern is obvious. The consumer AI race is quietly shifting. It is becoming less about which model is smartest and more about which AI you&#8217;ll let sit closest to your actual life. And the closer it sits, the more of your unguarded self it holds.</p><p>Friend is just the loud, awkward early version of that idea. It said the quiet part into a speaker. When your fears and your history become the product, memory stops being a feature and starts being the thing you&#8217;re paying to keep.</p><p><strong>So, what should you actually do?</strong></p><p><strong>Do this:</strong> if you&#8217;re a builder or a curious early adopter, buy one and study it as a preview of where consumer AI is heading. It is worth understanding firsthand.</p><p><strong>Skip this:</strong> do not hand it to a lonely teenager as a fix, and do not treat it as a replacement for a friend or a therapist. The strongest current research points the wrong way for that.</p><p><strong>Wait on this:</strong> if you only care about the practical stuff, wait. The always-on, remembering-you features Friend is charging for, will show up in cheaper, less awkward forms as the big players roll out personal agents.</p><p><strong>Steal this line:</strong> the newest thing AI companies are learning to sell isn&#8217;t intelligence. It&#8217;s the feeling of being remembered, and they&#8217;ve figured out how to put it on a monthly bill.</p><p>Talk to you again soon&#8230;</p><p><strong>Sources:</strong><br><a href="https://www.engadget.com/2227503/that-friend-ai-wearable-you-dont-like-just-got-worse/">Friend 2.0 launch, price, speaker, locked voice, subscription &#8212; Engadget</a></p><ul><li><p><a href="https://thenextweb.com/news/friend-ai-pendant-relaunch-speaker-double-price">Friend 2.0 details and the $10 memory subscription &#8212; The Next Web</a></p></li><li><p><a href="https://the-gadgeteer.com/2026/07/31/friend-ai-pendant-voice-wearable-249-price/">Price history, sales figures, manufacturing, founder background &#8212; The Gadgeteer</a></p></li><li><p><a href="https://finance.biggo.com/news/66fa0f9b-a0f9-44a7-ae2f-d3561168f5df">OpenAI models, EU exit over GDPR, 50,000-unit target &#8212; BigGo Finance</a></p></li><li><p><a href="https://www.pymnts.com/news/artificial-intelligence/2026/friends-talking-pendant-gives-openais-screenless-future-a-voice/">Friend 2.0 in the context of OpenAI&#8217;s screenless hardware push &#8212; PYMNTS</a></p></li><li><p><a href="https://www.aalto.fi/en/news/ai-companions-can-comfort-lonely-users-but-may-deepen-distress-over-time">Aalto University study, official write-up and researcher quotes</a></p></li><li><p><a href="https://www.forbes.com/sites/johnkoetsier/2026/03/27/ai-friends-only-a-band-aid-on-loneliness-isolation-new-study-says/">Aalto study coverage and context &#8212; Forbes</a></p></li><li><p><a href="https://www.roborhythms.com/loneliness-studies-blame-ai-companions-wrong-thing/">The causation critique and counter-evidence</a></p></li><li><p><a href="https://www.insurancejournal.com/news/international/2026/05/05/868584.htm">Ipsos BVA survey of young Europeans, commissioned by CNIL and Groupe VYV &#8212; Reuters via Insurance Journal</a></p></li><li><p><a href="https://techcrunch.com/2026/07/29/mark-zuckerberg-predicts-that-billions-of-people-will-have-personal-ai-agents-in-five-years/">Zuckerberg on personal AI agents within five years &#8212; TechCrunch</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The First AI Agent That Worked for 16 Days: Why Qwen3.8-Max Matters.]]></title><description><![CDATA[The Next AI Breakthrough Isn't Smarter Answers. It's Agents That Don't Stop.]]></description><link>https://aisignal.veletica.com/p/the-first-ai-agent-that-worked-for</link><guid isPermaLink="false">https://aisignal.veletica.com/p/the-first-ai-agent-that-worked-for</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Wed, 05 Aug 2026 15:37:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!PkL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PkL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PkL9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PkL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2438122,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/209936246?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PkL9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!PkL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5699f1-6cd2-407b-89b6-2e9d87acce2a_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most of us use AI like a microwave. You punch in thirty seconds and you hover like the burrito owes you money. One prompt goes in, one answer comes out, and the whole relationship is over in under a minute.</p><p>On Monday, Alibaba claimed it built something closer to a line cook. One that clocked in, started a project from an empty folder, and did not clock out for sixteen days.</p><p>The model is called Qwen3.8-Max. By the end of this piece you&#8217;ll know what it is in plain English, what it actually did during those sixteen days, what it costs, and whether any of it deserves your attention.</p><p>The verdict, up front: this claim is more checkable than the usual vendor demo, because the receipts sit in a public GitHub repo. The specs are real and the pricing is aggressive. The asterisks are real too, and the biggest one involves a license that does not exist yet. Either way, this release shows where the AI race is heading. It&#8217;s moving away from smarter answers and toward agents that don&#8217;t stop.</p><h3>The two-week test</h3><p>We need a foundation before the specs, because the numbers only matter once this part clicks.</p><p>I judge every agent announcement with one question I call the Two-Week Test. If you handed an AI a real project and walked away for two weeks, would it still be on task when you came back?</p><p>Almost everything fails this test in an afternoon. The model forgets the goal, or the context window fills up and it quietly loses the plot like me at minute forty of any movie with a dream sequence.</p><p>That failure is why &#8220;agents&#8221; have been more demo than product. Writing one great function got solved a while ago. Staying coherent across days of feedback, errors, and shifting requirements is a different sport entirely.</p><p>Qwen3.8-Max is Alibaba&#8217;s attempt to pass the Two-Week Test in public.</p><h3>What Alibaba says happened</h3><p>The centerpiece demo is a project called oh-my-cli. According to Alibaba, the model started from an empty repository and ran a full engineering loop on its own. Incoming requests became GitHub issues. The model claimed those issues, wrote the code, ran the tests, and merged its own pull requests when everything passed.</p><p>Alibaba&#8217;s own write-up describes the run as ten-plus days. The tally posted for July 30 works out to roughly sixteen days of unattended operation, and sixteen is the number most coverage settled on. I&#8217;m flagging the wobble because you deserve to know the vendor said &#8220;10+&#8221; while the headlines said sixteen.</p><p>The output is easier to pin down. The run produced 265 commits. Behind those sat 127 pull requests. The model also worked through 151 issues. And this is what separates the demo from every glossy AI flex you&#8217;ve scrolled past this year: the entire trace is public at qwen-code-dev-bot/oh-my-cli on GitHub. You don&#8217;t have to take Alibaba&#8217;s word for anything. You can read the commit history like a diary.</p><p>A second demo pushed further. Alibaba handed the model a research paper, &#8220;Unified Data Selection for LLM Reasoning,&#8221; with zero starter code, and told it to reproduce the results and then beat them. The run took roughly five days. In that window the model wrote about 7,600 lines of code and completed 33 rounds of GPU training. Alibaba says the method it invented at the end scored 2.7 points above the paper&#8217;s approach on a math benchmark. Those numbers come from Alibaba&#8217;s own harness, and no outside lab has audited them yet.</p><h3>What the thing actually is</h3><p>In plain English, Qwen3.8-Max is a huge model that only wakes up a small piece of itself for each request.</p><p>With that picture loaded, the technical version lands easier. It&#8217;s a sparse mixture-of-experts model carrying 2.4 trillion total parameters, of which about 95 billion activate per token. Imagine a hospital with 2.4 trillion staff on payroll where each patient only gets paged the specialists relevant to their case. It would be a terrible hospital, but it&#8217;s a great architecture, because you get frontier scale without paying frontier compute on every token.</p><p>The model takes text, images, and video as input, holds up to one million tokens of context, roughly 750,000 words, and builds on the Qwen3.5 architecture. It also plugs into tools you already use, since the API speaks both the OpenAI format and the Anthropic format. You can point Claude Code at it by changing a base URL.</p><h3>What it costs, and the license problem</h3><p>Let&#8217;s talk money, because this is where the announcement gets slippery.</p><p>The hosted API is paid and live today. Input costs $2 per million tokens, output runs $6, and cached input drops to $0.25. That sounds cheap next to Western flagships. One catch hides in the defaults, though. Thinking tokens bill as output, and the model ships with reasoning effort cranked to its highest setting, so real agentic sessions will land noticeably above the sticker math.</p><p>The open weights are the headline everyone repeated, and they remain a promise. Alibaba committed to releasing them within about a week of launch, for both the flagship and a smaller Qwen3.8-27B. As I write this, neither has appeared on Hugging Face and no license has been named. That last part matters more than it sounds. A model with no license isn&#8217;t open source yet, no matter what the headline says. Until words like Apache 2.0 appear next to an actual repository, don&#8217;t build a business plan on these weights.</p><p>And even when they land, a free license is not free hosting. The 2.4 trillion parameter checkpoint is a multi-node datacenter artifact, and nobody is running that at home. The 27B sibling is the one ordinary hardware can hold, and it&#8217;s the release worth watching if you&#8217;re a self-hosting builder.</p><p>One more piece of context: this lands in the middle of a Chinese pricing war. DeepSeek&#8217;s V4-Flash officially runs $0.14 per million input tokens and $0.28 per million output. Chinese labs keep pushing capability up while dragging prices down, and American labs are left defending the premium tier.</p><h3>The honest asterisks</h3><p>Every headline number in this release traces back to Alibaba&#8217;s own testing, run inside Alibaba&#8217;s own environment. On the crowdsourced Arena leaderboards, the model debuted as the top-ranked Chinese model for text while still trailing Anthropic&#8217;s Claude Fable 5, and Alibaba&#8217;s own benchmark table shows it beating some Western flagships on terminal tasks while losing to others. That mixed picture is normal. It&#8217;s also exactly why the weights release is the real event, because once outside labs can poke at the model directly, we find out how much of the endurance story survives contact with someone else&#8217;s sandbox.</p><p>There&#8217;s a quieter insight buried in the technical coverage, and it&#8217;s the punchline of the Two-Week Test. The model was not alone out there for sixteen days. It sat inside scaffolding: an issue state machine, a dispatcher, a monitor, and a watchdog that restarted things when they stalled. The model decides what to do next, and the harness keeps it alive long enough for that decision to matter. Systems pass the Two-Week Test. Models just take it. If you&#8217;re building agents, that idea is worth more than any benchmark in this piece.</p><h3>So what</h3><p><strong>Do this:</strong> skim the commit history at qwen-code-dev-bot/oh-my-cli. Ten minutes in that repo will teach you more about the current state of autonomous coding than any launch thread.</p><p><strong>Skip this:</strong> any plan involving self-hosting the full 2.4T model. That checkpoint belongs to datacenters, and pretending otherwise burns money.</p><p><strong>Wait on this:</strong> the license. No named license means the &#8220;open&#8221; part hasn&#8217;t happened, and fine print has sunk better stories than this one.</p><p><strong>Steal this line:</strong> stop asking whether your AI is smart, and start asking how long it stays on task.</p><p>Because that&#8217;s the shift underneath all of this. For three years the industry raced to build the model with the best answer. Alibaba just spent sixteen days arguing the next race belongs to the model with the best attention span. And if endurance really is the new battleground, the labs with the smartest models won&#8217;t be the only winners. The builders who learn to construct the scaffolding that keeps a model on task will win right alongside them.</p><p>That part is not reserved for trillion-parameter labs. That part is learnable, right now, by people like us.</p><p>See you in the next one.</p><h4>Sources</h4><ul><li><p><a href="https://www.scmp.com/tech/article/3362738/alibabas-ai-model-qwen38-max-made-widely-accessible-ahead-open-weights-release">South China Morning Post: Alibaba&#8217;s AI model Qwen3.8-Max made widely accessible ahead of open-weights release</a></p></li><li><p><a href="https://qz.com/alibaba-qwen38-max-ai-model-launch-080326">Quartz: Alibaba launches Qwen3.8-Max, its largest AI model yet</a></p></li><li><p><a href="https://the-decoder.com/alibabas-open-weight-qwen3-8-max-takes-on-long-horizon-ai-tasks-with-2-4-trillion-parameters/">The Decoder: Alibaba&#8217;s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters</a></p></li><li><p><a href="https://thenewstack.io/qwen-autonomous-coding-audit/">The New Stack: Alibaba&#8217;s AI coded for 16 days straight and every commit is on GitHub</a></p></li><li><p><a href="https://www.developer-tech.com/news/alibaba-qwen3-8-max-claims-16-day-autonomous-coding-run/">Developer Tech: Alibaba Qwen3.8-Max claims 16-day autonomous coding run</a></p></li><li><p><a href="https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/">MarkTechPost: Alibaba Qwen releases Qwen3.8-Max</a></p></li><li><p><a href="https://www.yottalabs.ai/post/qwen-3-8-max-release-date-specs-how-to-access-2026">Yotta Labs: Qwen 3.8-Max release date, specs, and how to access it</a></p></li><li><p><a href="https://www.thedailystar.net/news/tech-startup/news/alibaba-releases-qwen-38-max-its-largest-ai-model-yet-4238986">The Daily Star: Alibaba releases Qwen 3.8-Max, its largest AI model yet</a></p></li><li><p><a href="https://apidog.com/blog/qwen-3-8-for-coding/">Apidog: Qwen 3.8 for coding, configs and cost math</a></p></li><li><p><a href="https://deepseek.ai/pricing">DeepSeek official API pricing</a></p></li><li><p><a href="https://github.com/qwen-code-dev-bot/oh-my-cli">The oh-my-cli repository on GitHub</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[$3.15 vs. $0.03: The New AI Price Gap, Explained.]]></title><description><![CDATA[China's AI Price War Just Hit Your Bill.]]></description><link>https://aisignal.veletica.com/p/315-vs-003-the-new-ai-price-gap-explained</link><guid isPermaLink="false">https://aisignal.veletica.com/p/315-vs-003-the-new-ai-price-gap-explained</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Tue, 04 Aug 2026 22:03:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!QfwJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QfwJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QfwJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QfwJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2791821,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/209848613?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QfwJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!QfwJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb112ae80-d22e-41b7-a1fa-2fcdc88238bb_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two Chinese AI labs just spent three weeks trying to out-cheap and out-big each other, and the fallout lands directly on your AI bill. Whether you build with these models or just approve the invoices of someone who does, this one is worth ten minutes, because it will either save you money or save you from planning around a promise.</p><p>I&#8217;ll give you the verdict up front. DeepSeek&#8217;s V4-Flash is the release you can act on this week. The weights are published, the license is MIT, and independent testing puts its running cost at about three cents for a full benchmark suite. Alibaba&#8217;s Qwen3.8-Max is the bigger headline with the smaller certainty. The hosted model is real and live, but the open weights, and the license that would make them useful, are still promises. Route your cheap, high-volume work toward V4-Flash now. Give Qwen one more week before you plan anything around it.</p><p>Every number below comes from either a vendor price sheet or an independent benchmark firm, and I&#8217;ll tell you which is which as we go.</p><h2>The restaurant and the recipe</h2><p>Before the numbers, you need one mental model, because it makes this entire story click.</p><p>A hosted AI model is a restaurant. You order through the API, they cook everything in their kitchen, and you pay per plate. That per-plate price is the &#8220;per million tokens&#8221; number on every pricing page.</p><p>An open-weight model is the restaurant handing you the recipe. The recipe itself is free, and with a permissive license you can cook at home, change the ingredients, or open a competing restaurant. But you still need a kitchen. And for the biggest models, the kitchen is not your kitchen. It&#8217;s a commercial facility with a power bill that reads like a phone number.</p><p>Hold onto that, because Alibaba and DeepSeek just made opposite moves in the same war. One slashed menu prices to the floor. The other promised to hand out the recipe for the biggest dish it has ever cooked.</p><h2>Three weeks of escalation</h2><p>The sequence, pieced together from multiple outlets, goes like this. In mid-July, Moonshot AI shipped Kimi K3, a 2.8-trillion-parameter model. On July 19, Alibaba rushed a preview of Qwen3.8-Max onto a stage at the World AI Conference in Shanghai without benchmarks or a license. Days later, Moonshot published K3&#8217;s full weights, which coverage described as the largest open-weight release to date. On July 31, DeepSeek released V4-Flash and put the weights on Hugging Face the same day. And on August 3, Alibaba fully released Qwen3.8-Max and promised its own weights within a week.</p><p>That&#8217;s five moves in under three weeks, which is not a product cycle. That&#8217;s a bidding war where the bids go down.</p><h2>DeepSeek&#8217;s move: win the check, not the menu</h2><p>V4-Flash&#8217;s menu price, straight from DeepSeek&#8217;s own sheet, is $0.14 per million input tokens and $0.28 per million output tokens. Cached input drops to $0.0028 per million, which matters enormously if your app repeats the same system prompt all day.</p><p>But the number that made global headlines is different. Artificial Analysis, an independent benchmarking firm, measures what it costs to push a model through its full Intelligence Index test battery. That&#8217;s the check, not the menu, because a model with a cheap per-token rate can still run up a real bill if it rambles its way to every answer. V4-Flash came in around three cents per run. Moonshot&#8217;s Kimi K3 cost 86 cents on the same measure. OpenAI&#8217;s GPT-5.6 Sol cost $1.86. And Anthropic&#8217;s Claude Fable 5 cost $3.15.</p><p>That last gap is more than one hundred to one. The funny part, flagged in Artificial Analysis&#8217;s own write-up, is that V4-Flash is a chatty model. It burns tokens generously and still lands at three cents, because the per-token price is doing all the work.</p><p>Now comes the honest caveat, because Ground Truth always has one. V4-Flash scored 50 out of 100 on that same Intelligence Index. That ties Google&#8217;s Gemini 3.6 Flash and trails the frontier pack, where Claude Opus 5, Claude Fable 5, and GPT-5.6 all scored at least nine points higher. So the offer on the table is roughly two-thirds of the frontier&#8217;s brains at around one percent of the frontier&#8217;s running cost. For classification, extraction, summarization, and routing work, that trade is not even close.</p><h2>Alibaba&#8217;s move: the flex with an asterisk</h2><p>Qwen3.8-Max is the largest model Alibaba has ever shipped: 2.4 trillion total parameters, a one-million-token context window, and multimodal input covering text, images, and video. It&#8217;s a sparse mixture-of-experts design, which means only a slice of those parameters fires on each query. Several outlets report that slice at about 95 billion, though Alibaba&#8217;s own documentation reportedly doesn&#8217;t publish the active count, so treat that figure as reported rather than confirmed.</p><p>For scale, a one-million-token context window means you could paste in roughly 750,000 words and the model would still be listening. That&#8217;s the entire Lord of the Rings trilogy plus your last four years of meeting notes in one prompt, and I&#8217;m honestly not sure which half the model would find more fantastical.</p><p>On performance, label your sources. Alibaba&#8217;s own benchmark table shows Qwen3.8-Max beating Claude Fable 5 on Terminal-Bench 2.1 while trailing GPT-5.6 Sol&#8217;s max setting, and those are vendor-reported numbers. Third-party signal exists too: Arena leaderboard coverage places it as the top Chinese entry among text models and second worldwide on the vision leaderboard. Alibaba also published a case study where the model spent 16 days autonomously building a command-line tool, filing its own GitHub issues along the way, which is either the future of software or the world&#8217;s most expensive Tamagotchi, depending on your mood.</p><p>The asterisk sits on the open part of &#8220;open weights.&#8221; They&#8217;re promised for next week, and the license hasn&#8217;t been published. Alibaba&#8217;s past open releases have used Apache 2.0, but as VentureBeat notes, that&#8217;s precedent, not commitment, and rival Moonshot recently shipped its open model under a custom license instead. Until the license text exists, Qwen3.8-Max is a hosted product with an announcement attached.</p><h2>What this actually costs you</h2><p><strong>DeepSeek V4-Flash.</strong> The hosted API is paid: $0.14 per million input tokens and $0.28 per million output, with cache hits billed at $0.0028 per million input. New API users get 5 million free tokens, and the consumer chat app is free with fair-use throttling. One flag before you model costs: DeepSeek has announced that all billing will double during Beijing peak hours, with the effective date still unannounced, so run your projections against the doubled rate if your traffic overlaps China&#8217;s workday. The weights are open source under the MIT license, which permits commercial use, modification, and redistribution without restriction. Free software still means your hardware, though. The checkpoint is roughly 167 GB, so self-hosting takes a serious multi-GPU rig that you pay for, not a laptop.</p><p><strong>Qwen3.8-Max.</strong> The hosted API is paid: $2.00 per million input tokens and $6.00 per million output, with cached input at $0.25 per million. The open weights are promised, not shipped, and no license has been named, so its open-source status is unconfirmed as of this writing. If the weights do land, a 2.4-trillion-parameter checkpoint works out to roughly 1.2 terabytes at 4-bit precision, which makes it a datacenter artifact rather than a home project. The realistic self-host option is the smaller Qwen3.8-27B that Alibaba promised alongside it, and the same rule applies there: until a license exists, don&#8217;t build plans on it.</p><h2>So what</h2><p><strong>Do this:</strong> if you&#8217;re paying frontier prices for high-volume, low-stakes calls, pilot V4-Flash this week. A 100x running-cost gap stopped being a procurement detail and became an architecture decision.</p><p><strong>Skip this:</strong> don&#8217;t build product plans on Qwen3.8-Max&#8217;s open weights until the license text is public.</p><p><strong>Wait on this:</strong> Qwen3.8-27B could be the sleeper release for local hardware. Watch for the actual drop, then read the license before celebrating.</p><p><strong>One-sentence steal:</strong> the per-token rate is the menu, the cost per finished task is the check, and you should always compare checks.</p><h2>Sources</h2><ul><li><p><a href="https://finance.yahoo.com/technology/ai/articles/deepseeks-ai-model-far-cheapest-054143251.html">Reuters (via Yahoo Finance): DeepSeek&#8217;s new AI model is by far the cheapest of well-known models to run</a></p></li><li><p><a href="https://qz.com/deepseek-v4-flash-cheapest-ai-model-benchmark-080326">Quartz: DeepSeek V4-Flash is cheapest major AI model to run</a></p></li><li><p><a href="https://x.com/ArtificialAnlys/status/2083306229074739285">Artificial Analysis (X): DeepSeek V4 Flash 0731 is now open weights under MIT</a></p></li><li><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash">Hugging Face: deepseek-ai/DeepSeek-V4-Flash model card (MIT license)</a></p></li><li><p><a href="https://www.opensourceforu.com/2026/08/deepseek-open-sources-v4-flash/">Open Source For You: DeepSeek open sources production V4-Flash under MIT licence</a></p></li><li><p><a href="https://thenewstack.io/deepseek-v4-flash-open-weights/">The New Stack: DeepSeek&#8217;s smaller model just outperformed its own flagship</a></p></li><li><p><a href="https://benchlm.ai/deepseek/api-pricing">BenchLM: DeepSeek API pricing, including announced 2x peak-hour policy</a></p></li><li><p><a href="https://www.bloomberg.com/news/articles/2026-08-03/alibaba-drops-another-china-ai-model-with-breakthrough-performance">Bloomberg: Alibaba&#8217;s Qwen3.8-Max claims benchmark scores rivaling Anthropic</a></p></li><li><p><a href="https://www.cnbc.com/2026/08/03/alibaba-ai-model-qwen-rival-anthropic.html">CNBC: Alibaba shares rally after unveiling Qwen3.8-Max</a></p></li><li><p><a href="https://technode.global/2026/08/04/chinas-alibaba-launches-qwen3-8-max-ai-model-with-2-4t-parameters-1m-token-context-window/">TechNode Global: Alibaba launches Qwen3.8-Max with 2.4T parameters, 1M context</a></p></li><li><p><a href="https://the-decoder.com/alibabas-open-weight-qwen3-8-max-takes-on-long-horizon-ai-tasks-with-2-4-trillion-parameters/">The Decoder: Alibaba&#8217;s open-weight Qwen3.8-Max takes on long-horizon tasks</a></p></li><li><p><a href="https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/">MarkTechPost: Qwen3.8-Max release details, pricing, and benchmark table</a></p></li><li><p><a href="https://venturebeat.com/technology/qwen3-8-max-arrives-with-a-bold-claim-it-outperforms-gpt-5-6-sol-max-and-fable-5-on-agentic-computer-use">VentureBeat: Qwen3.8-Max arrives with bold agentic claims, license still undisclosed</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Buzz. Jack Dorsey Shipped an Agent Workspace, Then Shipped 11 Releases in 8 Days]]></title><description><![CDATA[The Best Idea in Team Software Right Now Also Documents Its Own Gap]]></description><link>https://aisignal.veletica.com/p/buzz-jack-dorsey-shipped-an-agent</link><guid isPermaLink="false">https://aisignal.veletica.com/p/buzz-jack-dorsey-shipped-an-agent</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Fri, 31 Jul 2026 14:54:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-Akm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Akm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Akm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Akm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2086010,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/209261016?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Akm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!-Akm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96001ffa-87da-4ef1-82c3-d9529099cd34_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two sentences buried in Block&#8217;s own documentation tell you more about Buzz than every launch article stacked together.</p><p>The first is in SECURITY.md. Channel membership is the only access control mechanism. There are no separate ACL lists, no capability taxonomies. If a principal is a member of a channel, human or agent, it can read and write there.</p><p>The second went up on Block&#8217;s engineering blog on July 29, written by Tom Brow on the Applied AI team. Buzz agents typically run on the same computer humans use, outside any sandbox, with permissions skipped, inheriting that machine&#8217;s files, skills, and credentials.</p><p>Read those together and the verdict writes itself.</p><p><strong>The verdict:</strong> Buzz is the most genuinely interesting idea in team software this year, and you should not put it anywhere near a machine that matters yet. Stand it up in a lab, on a throwaway repo, with one narrowly scoped agent. Learn the model. Do not migrate your team.</p><p>Here is what earns that.</p><h3>What Block actually built</h3><p>Buzz launched July 21, 2026. It is a self-hostable workspace where humans and AI agents sit in the same channels, released under Apache 2.0 at github.com/block/buzz. Jack Dorsey announced it on X as model-agnostic, decentralized, self-sovereign, and open source, built explicitly to reduce Block&#8217;s dependency on Slack and GitHub.</p><p>Underneath, it is a Nostr relay written in Rust, sitting on Postgres, Redis, and S3-compatible object storage. Every message, reaction, workflow step, review approval, and git event is a signed event in one log. Same shape and same identity model whether the author is a person or a process.</p><p>The agent story is the actual product. You add an agent to a channel the way you add a coworker. It can open repos, send patches, review code, run workflows, edit canvases, drop into voice huddles, and pull other people in. Claude Code, Codex, Block&#8217;s own goose, or anything speaking the Agent Client Protocol works. Change your model or your harness and the project keeps its identity and its history.</p><h3>The part that is genuinely good</h3><p>Most agent tooling today runs on borrowed credentials. Your GitHub token. Your Slack identity. A shared company key. The log says you did it. You did not.</p><p>Block&#8217;s fix is clean. Each agent gets its own key. The owner signs a narrowly scoped authorization, and the agent then signs its own work under its own identity. Block made a deliberate call here that I have not seen elsewhere: authorization does not erase authorship. The agent stays the author, and its credential proves who authorized it and under what conditions. If an agent key leaks, you revoke the agent without touching the human identity behind it.</p><p>That is a real answer to a real problem, and I do not want to undersell it. Anyone who has tried to reconstruct which bot did what three weeks ago knows exactly why this matters.</p><h3>The part the marketing walks past</h3><p>Cryptographic identity tells you who acted. It does not tell you what they were allowed to do. Those are two different problems, and Buzz has shipped a strong answer to the first one and a thin answer to the second.</p><p>Go back to those two sentences. An agent running unsandboxed on a real laptop, holding that laptop&#8217;s real credentials, with the permission checks skipped. The only thing standing between that agent and a destructive instruction is whether the person giving the instruction is in the channel. Block says it plainly: security rests entirely on restricting who can tell it what to do.</p><p>Jo&#227;o Queir&#243;s, in the deepest independent hands-on review published so far, put it in one line. Channel membership is not fine-grained tool authorization. He is right, and Block&#8217;s own security policy agrees with him.</p><p>There is a second caveat worth pulling out. The audit log chains every entry to the previous one with a SHA-256 hash, and SECURITY.md is honest that the chain is keyless, which makes it tamper-evident but not tamper-resistant. It catches accidental corruption or a single edited row. Somebody with database write access can recompute the whole chain after editing it. Block describes the log as designed for SOX-grade compliance and eDiscovery. Those two statements sit in the same document, about six paragraphs apart, and the gap between them is where your compliance team is going to live.</p><h3>What is actually shipping, as of today</h3><p>The desktop app is at v0.5.2, cut on July 29. It launched in the 0.4.x range eight days earlier, which is a release cadence somewhere between impressive and alarming depending on whether you enjoy reading changelogs at 11pm.</p><p>The repo has crossed roughly 19,000 stars with about 1,900 forks and more than 2,000 commits. It also has over 500 open issues and close to 700 open pull requests, which is the honest other half of that number.</p><p>iOS and Android landed on July 29. The mobile app does not host agents, it signs messages and talks to relays directly, and you pair it by pointing your phone at a QR code in the desktop app, so it reuses the same keypair. Block says it ships with no analytics SDKs and strips geolocation metadata from image uploads before they hit the relay. Push runs on a draft standard Block wrote called NIP-PL, designed so relays never see your device token and the push gateway never sees your keys or your content.</p><p>Worth flagging: the repo README still lists mobile clients under &#8220;being wired up.&#8221; The README is lagging its own shipped product by two days, which tells you something about the pace.</p><h3>The thing nobody in the coverage mentions</h3><p>Block cut its headcount from over 10,000 to under 6,000 in February, a reduction of more than forty percent, and posted a Q1 2026 net loss of $308.7 million on $6.06 billion in revenue. Then it open-sourced the tool its remaining engineers use to coordinate with agents.</p><p>I am not going to tell you what to make of that. I will say the sequence is worth holding in your head while you evaluate the product, because Buzz is not a side project from a company with money to burn. It is infrastructure a shrinking engineering org built to survive its own restructuring, and that is either the strongest possible endorsement or the loudest possible warning.</p><h3>So what</h3><p><strong>Do this.</strong> Clone it and self-host. You need Docker plus Rust 1.88, Node 24, and pnpm 10, or just use the bundled Hermit toolchain. Run one throwaway repo, one agent, one workflow. The question to answer is not whether the agent does good work. It is whether putting humans and agents in the same persistent room actually improves how your team decides things.</p><p><strong>Skip this.</strong> Do not point Buzz at anything holding real credentials, customer data, or regulated records. Not yet. The relay does not enforce TLS by default, that is an intentional deployment choice, and it is your job to terminate it properly.</p><p><strong>Wait on this.</strong> Hosted pricing is unannounced. Fine-grained tool authorization does not exist. The compliance story is unfinished by the reviewers&#8217; account and pre-1.0 by Block&#8217;s own supported-versions table. All three need to move before this belongs in production.</p><p><strong>One sentence to steal.</strong> Cryptographic identity tells you who acted, not what they were allowed to do, so ask any agent vendor which of those two they actually shipped.</p><h3>Sources</h3><ul><li><p><a href="https://block.xyz/inside/introducing-buzz-where-humans-and-agents-work-together">Block : Introducing Buzz: where humans and agents work together</a></p></li><li><p><a href="https://engineering.block.xyz/blog/buzz">Block Engineering : Buzz!</a></p></li><li><p><a href="https://engineering.block.xyz/blog/a-buzz-on-your-phone">Block Engineering : A Buzz on your phone</a></p></li><li><p><a href="https://github.com/block/buzz">GitHub : block/buzz (README)</a></p></li><li><p><a href="https://github.com/block/buzz/blob/main/SECURITY.md">GitHub : block/buzz SECURITY.md</a></p></li><li><p><a href="https://github.com/block/buzz/releases/tag/v0.5.2">GitHub : Buzz Desktop v0.5.2 release notes</a></p></li><li><p><a href="https://techcrunch.com/2026/07/21/jack-dorsey-is-taking-on-slack-with-buzz-a-group-chat-platform-for-teams-and-their-ai-agents/">TechCrunch : Jack Dorsey is taking on Slack with Buzz</a></p></li><li><p><a href="https://decrypt.co/374026/jack-dorseys-block-launches-buzz-a-nostr-based-slack-and-github-rival-for-ai-agents">Decrypt : Block launches Buzz, a Nostr-based Slack and GitHub rival</a></p></li><li><p><a href="https://thenextweb.com/news/block-buzz-humans-ai-agents-workspace">The Next Web : Jack Dorsey&#8217;s Block takes on Slack with Buzz</a></p></li><li><p><a href="https://rohitraj.tech/en/notes/block-buzz-agent-collaboration-platform-guide-2026">Rohit Raj : Block&#8217;s Buzz 2026 guide and self-host walkthrough</a></p></li><li><p><a href="https://northeasttimes.com/2026/07/22/jack-dorsey-launches-buzz-an-open-source-rival-to-slack-and-github/">Northeast Times : Dorsey launches Buzz, an open source rival to Slack and GitHub</a></p></li><li><p><a href="https://www.stocktitan.net/sec-filings/XYZ/10-q-block-inc-quarterly-earnings-report-e5533e74dbf0.html">StockTitan : Block, Inc. Q1 2026 10-Q summary</a></p></li><li><p><a href="https://www.marketscreener.com/news/block-inc-reports-earnings-results-for-the-first-quarter-ended-march-31-2026-ce7f5bdadf8df221">MarketScreener : Block, Inc. Q1 2026 earnings results</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Five Things Anthropic Shipped That Lower What You Pay.]]></title><description><![CDATA[Same Price, Double the Score: What Anthropic Actually Shipped.]]></description><link>https://aisignal.veletica.com/p/five-things-anthropic-shipped-that</link><guid isPermaLink="false">https://aisignal.veletica.com/p/five-things-anthropic-shipped-that</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Thu, 30 Jul 2026 21:08:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!z92R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z92R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z92R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!z92R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!z92R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!z92R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z92R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1846798,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/209171578?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!z92R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!z92R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!z92R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!z92R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3de3d24b-bc0d-41dc-a6ba-9e5e1d69d283_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>If you build with AI, last week moved the number that actually matters to you. Not a leaderboard position. What it costs to get the work done.</p><p>Anthropic released Claude Opus 5 on July 24. Five days later, on July 28, the Model Context Protocol shipped its largest revision since it launched. A day before the model landed, voice mode got the upgrade people had been asking about for a year. Almost all the coverage went to which model beat which other model on which chart. That framing skips the part that shows up on your invoice.</p><p>So let me put the verdict first. Opus 5 costs exactly what Opus 4.8 cost, at $5 per million input tokens and $25 per million output tokens. On Anthropic&#8217;s own Frontier-Bench v0.1 run, it more than doubles Opus 4.8&#8217;s score at a lower cost per task. Across six days, five other changes shipped, and every single one of them lowers the floor for people building without a platform team behind them. If you have been waiting for agent work to get affordable enough to try, that wait just got shorter.</p><p>Now the receipts.</p><h3>The price stood still while the model moved</h3><p>Anthropic&#8217;s launch post claims Opus 5 comes close to the frontier intelligence of Claude Fable 5 at half the price. The specific numbers behind that: on Frontier-Bench v0.1, Opus 5 surpasses all other models and more than doubles Opus 4.8&#8217;s performance at a lower cost per task. On CursorBench 3.2 at max effort, it lands within 0.5% of Fable 5&#8217;s peak score at half the cost per task.</p><p>Those are vendor-reported figures, and the footnote matters. Anthropic ran Frontier-Bench internally on the mini-SWE-agent harness with a GKE backend, averaging reward over five attempts per task. That is a disclosed methodology, which is better than most, but it is still the vendor grading the vendor.</p><p>For a counterweight, CodeRabbit ran Opus 5 against their own code review workload and published something more textured. They found effort behaves like a routing decision rather than a free upgrade, with x-high buying precision at a coverage cost, and nothing improving uniformly. They also measured roughly 60,500 input tokens per call against roughly 40,500 for the GPT-5.6 lanes doing the same job in the same runs. Bigger context in, longer answer out, and part of the cost premium lives right there.</p><p>Both things are true. The model is meaningfully better per dollar, and it will happily spend more dollars if you let it.</p><h3>The effort dial is the actual cost lever</h3><p>Opus 5 ships with an effort parameter at five levels: low, medium, high, xhigh, and max. The API defaults to high.</p><p>The interesting part is that Anthropic&#8217;s own prompting guide tells you to reach for less. It recommends starting at the default and then using low and medium liberally as your primary control for token cost and response time wherever quality holds, stepping up to xhigh only for demanding coding and agentic work. A vendor telling you to buy less compute is not a thing you see often, and it is worth taking them up on.</p><p>Two migration traps come with it. Thinking is now on by default, which is a behavioral break from Opus 4.8, where requests ran without thinking unless you asked. And thinking cannot be disabled at xhigh or max effort, where those requests return a 400 error. Because max_tokens caps thinking and response text together, anything you previously ran without thinking needs its limits revisited.</p><p>One more caution the docs are explicit about: changing effort between requests does not preserve cached prefixes from earlier turns. Pick a level at the start of a long session and stay there.</p><h3>The quietest cost cut in the whole release</h3><p>The minimum cacheable prompt length on Opus 5 dropped to 512 tokens, down from 1,024 on Opus 4.8.</p><p>That got about one sentence of coverage, and it is the change most likely to save a small builder real money. If your system prompt, tool definitions, or agent configuration sat just under the old threshold, they were uncacheable and you were paying full input price on every single call. Now they cache, with no code changes required.</p><p>Go add <code>cache_control</code> to the short system prompts you skipped last year. That is a fifteen minute job.</p><h3>MCP stopped needing a server that remembers you</h3><p>The 2026-07-28 spec moved MCP from a bidirectional stateful protocol to a request and response model. The initialize handshake is gone. The <code>Mcp-Session-Id</code> header is gone. Every request now carries its own protocol version and client capabilities.</p><p>In plain terms: any request can land on any instance. The sticky routing and shared session stores that horizontal deployments used to require are no longer a protocol requirement, and your connector can live on serverless or edge infrastructure like an ordinary HTTP workload.</p><p>That is the difference between hosting a connector for pocket change and hosting one on a cluster you have to babysit.</p><p>The release also brought Multi Round-Trip Requests, so a tool can pause mid-call to ask the user something and the client retries with the answer instead of holding a long-lived stream open. Routable transport headers let gateways and rate limiters route without parsing request bodies. Authorization now aligns with production OAuth 2.0 and OpenID Connect, which means MCP servers connect to enterprise identity systems like Entra or Okta without custom workarounds. MCP Apps and Tasks graduated into a formal versioned extensions framework.</p><p>Two honest caveats. This release contains breaking changes, with roots, sampling, and logging deprecated. And Anthropic&#8217;s own post says support is rolling out across Claude products soon, so the spec being live does not mean every Claude surface supports every piece of it today.</p><h3>Claude Code got its agent teams back</h3><p>This one has a plot twist that most roundups flattened.</p><p>On July 21, Claude Code v2.1.217 stopped subagents from spawning nested subagents at all. On July 24, v2.1.219 reinstated nesting at a default depth of three, controllable through the <code>CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH</code> environment variable. Three days from banned to bounded, which reads like a deliberate correction rather than a random default change.</p><p>The same release added Opus 5 as the default Opus model with its 1M token context, a <code>workflowSizeGuideline</code> setting that can now come from any settings file, a <code>sandbox.network.strictAllowlist</code> option that denies non-allowlisted hosts without prompting, a <code>DirectoryAdded</code> hook that fires when a new working directory joins a session, and <code>mcp_server_errors</code> in the headless init event so failed MCP configs stop being invisible.</p><p>Worth knowing that the workflow size guideline is advisory rather than an enforced cap. Dynamic workflows default to a medium guideline aiming for fewer than 15 agents. It nudges. It does not stop you.</p><h3>The one release aimed at people who never open a terminal</h3><p>Voice mode had been running on Haiku since Anthropic added hands-free conversation earlier this year, a choice made for speed. On July 23, Claude Opus and Claude Sonnet became available in voice mode, with a picker that switches models mid-conversation.</p><p>Two details make this more than a spec bump. Voice mode uses the fastest version of whichever model you select, so choosing Opus does not mean sitting through long pauses. And it defaults to the last model you used in text chat, which means a spoken commute session picks up where a typed desk session left off.</p><p>Voice can also reach the tools you have already connected. Anthropic names Gmail and Slack directly, and its own examples demonstrate pushing a Google Calendar meeting back by 30 minutes and turning a pitch conversation into a Canva one-pager. Claude asks permission before touching any of them.</p><p>The language expansion is the part that actually redistributes something. Anthropic lists eleven languages and gives every plan access to all of them, Free accounts included. Free also gets Haiku and one connected tool, while paid plans unlock the expanded models and every connector you have authorized.</p><p>Three caveats worth publishing. Claude does not auto-detect your language, and text-chat language settings do not carry over, so you set it in voice settings or ask out loud. Voice conversations count against your regular usage limits, so &#8220;available on every plan&#8221; is not the same as free to use heavily. And Anthropic names Opus, Sonnet, and Haiku for voice, with Fable absent from that list.</p><p>The feature is still in beta, still turn-based, and Anthropic says it works best from your phone.</p><h3>Where this is still rough</h3><p>Fast mode runs about 2.5 times default speed and costs $10 per million input tokens and $50 per million output tokens, which is double the base rate. It is also a research preview available on the Claude API only, not on Amazon Bedrock, Google Cloud, or Microsoft Foundry.</p><p>Mid-conversation tool changes and automatic fallbacks are both beta and both need explicit headers. Opus 5&#8217;s safeguards block binary-based vulnerability scanning, penetration testing, and exploit generation, so security practitioners will hit walls that the Cyber Verification Program exists to work around.</p><p>And a cheaper model at the same list price does not automatically produce a cheaper month. Reasoning models spend tokens thinking, and the effort ladder turns that spend into a decision you can quietly get wrong.</p><h3>The bigger picture</h3><p>Every one of these changes points the same direction, and it is not toward a smarter chatbot. Lower cache thresholds, an effort dial, stateless connectors, bounded agent nesting, fallbacks that route instead of failing, and the good models finally reaching the microphone on every plan. That is a stack getting cheap and predictable enough for people without infrastructure teams.</p><p>The exciting week is the one where the model wins a benchmark. The useful week is the one where the plumbing gets boring enough that a person building alone at their kitchen table can afford to run the same architecture a Fortune 500 runs. Last week was the second kind, and those weeks compound.</p><h3>So What: your decision block</h3><p><strong>Do this now.</strong> Add <code>cache_control</code> to short system prompts you previously could not cache, since the 512 token minimum makes them eligible with no code changes. Then run an effort sweep on your own evals at low and medium before assuming you need high.</p><p><strong>Skip this for now.</strong> Fast mode, unless a human is genuinely sitting there blocked. Double price across a long agentic run is a bad trade when nobody is waiting.</p><p><strong>Try this today.</strong> Open voice mode on your phone, switch it to Opus, and talk through one decision you have been circling in text. It defaults to your last text model, so check the picker before you start.</p><p><strong>Wait on this.</strong> Migrating production MCP servers to the new spec. The breaking changes are real, and Claude product support is still rolling out. Read the changelog, test in beta, move when your surface actually supports it.</p><p><strong>One sentence to steal:</strong> Anthropic shipped a price cut and called it a model launch.</p><div><hr></div><h3>Sources</h3><ol><li><p><a href="https://www.anthropic.com/news/claude-opus-5">Introducing Claude Opus 5</a>, Anthropic, July 24, 2026</p></li><li><p><a href="https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5">What&#8217;s new in Claude Opus 5</a>, Claude Platform Docs</p></li><li><p><a href="https://platform.claude.com/docs/en/build-with-claude/effort">Effort</a>, Claude Platform Docs</p></li><li><p><a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5">Prompting Claude Opus 5</a>, Claude Platform Docs</p></li><li><p><a href="https://platform.claude.com/docs/en/build-with-claude/context-windows">Context windows</a>, Claude Platform Docs</p></li><li><p><a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/">The 2026-07-28 Specification</a>, Model Context Protocol Blog, July 28, 2026</p></li><li><p><a href="https://modelcontextprotocol.io/specification/2026-07-28/changelog">Key Changes: 2026-07-28</a>, Model Context Protocol</p></li><li><p><a href="https://claude.com/blog/bringing-mcp-2026-07-28-to-claude">Bringing MCP 2026-07-28 to Claude</a>, Anthropic, July 28, 2026</p></li><li><p><a href="https://claude.com/blog/think-through-hard-problems-in-voice-mode">Think through hard problems in voice mode</a>, Anthropic, July 23, 2026</p></li><li><p><a href="https://github.com/anthropics/claude-code/releases/tag/v2.1.219">Claude Code Release v2.1.219</a>, GitHub, July 24, 2026</p></li><li><p><a href="https://code.claude.com/docs/en/changelog">Claude Code changelog</a>, Claude Code Docs</p></li><li><p><a href="https://www.coderabbit.ai/blog/opus-5-model-review">Claude Opus 5 Benchmarks for AI Code Review</a>, CodeRabbit (independent evaluation)</p></li><li><p><a href="https://www.digitalapplied.com/blog/claude-code-subagent-depth-limits-budget-caps-2026">Claude Code subagent depth limits and budget caps</a>, Digital Applied</p></li><li><p><a href="https://www.axios.com/2026/07/24/anthropic-releases-new-model-opus-5">Anthropic releases new model, Opus 5</a>, Axios</p></li><li><p><a href="https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/">Anthropic launches Opus 5</a>, TechCrunch</p></li><li><p><a href="https://www.infoworld.com/article/4201551/anthropic-releases-more-efficient-claude-opus-5.html">Anthropic releases &#8216;more efficient&#8217; Claude Opus 5</a>, InfoWorld</p></li><li><p><a href="https://aws.amazon.com/blogs/machine-learning/how-agentcore-gateway-supports-the-mcp-2026-07-28-spec/">How AgentCore Gateway supports the MCP 2026-07-28 spec</a>, AWS</p></li></ol>]]></content:encoded></item><item><title><![CDATA[Qwen 3.8 vs Kimi K3 vs GLM 5.2: The Open-Weight Race Nobody Has Finished Yet.]]></title><description><![CDATA[The Word Doing All the Work in 'Qwen 3.8 Max Preview' Is Preview.]]></description><link>https://aisignal.veletica.com/p/qwen-38-vs-kimi-k3-vs-glm-52-the</link><guid isPermaLink="false">https://aisignal.veletica.com/p/qwen-38-vs-kimi-k3-vs-glm-52-the</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Fri, 24 Jul 2026 20:02:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!OH3o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OH3o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OH3o!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OH3o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2992210,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/208375440?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OH3o!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OH3o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bce32b6-2747-4b90-acbf-cc9fe1b7cd48_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The gap between Moonshot dropping a 2.8 trillion parameter model on the world and Alibaba announcing its 2.4 trillion parameter answer was eight days. Neither lab waited for the other to finish talking.</p><p>So which one belongs in your stack?</p><p>Short version, so you can bail out whenever you want. Qwen 3.8 Max Preview is an evaluation target, not a dependency. Kimi K3 is the one with a neutral referee&#8217;s score attached. GLM 5.2 is the one you can download today with a license file you can actually read. The word doing all the heavy lifting in &#8220;Qwen 3.8 Max Preview&#8221; is Preview.</p><p>Now the receipts.</p><h4>What Alibaba actually announced</h4><p>On July 19, 2026, Alibaba&#8217;s Qwen team previewed Qwen3.8-Max-Preview at the World AI Conference in Shanghai, describing it as a 2.4 trillion parameter model and positioning it as second only to Claude Fable 5 among the systems it benchmarked. That positioning is Alibaba&#8217;s own framing. No benchmark results, model card, or active-parameter count were disclosed, so the ranking claim is not an independently verified result.</p><p>Here&#8217;s what is real and checkable. The model takes text, image, and video input, and it speaks both the OpenAI and Anthropic API protocols, which means existing coding agents can connect without anyone rebuilding a harness. Qwen Cloud metadata lists a 983,616 token context window and a 131,072 token maximum output, with thinking always on across low, high, and xhigh settings, and xhigh as the documented default. Qwen announced a new preview build on July 20 claiming broad gains including web frontend work, while describing the model as evolving daily. Also a vendor claim.</p><p>And the thing nobody puts in the headline: the active parameter count is the number nobody has. Without it, the 2.4T figure says very little about what serving this thing costs.</p><h4>The pricing is a subscription, not a price</h4><p>If your instinct when you read &#8220;2.4 trillion parameters&#8221; is &#8220;sounds expensive,&#8221; your instincts are fine. But you can&#8217;t actually check.</p><p>There is no standalone per-token API rate published for Qwen3.8-Max. Access is bundled into a credit-based Token Plan subscription, sold during preview at 10% of standard pricing across Token Plan, Qoder, and QoderWork, with an additional 80% off night discount that drops off-hours consumption to roughly 0.2% of the standard rate. Qoder&#8217;s own event page confirms the off-peak window runs 22:00 to 08:00 Singapore time, with an end date listed as TBD.</p><p>Token Plan Personal tiers run roughly $6 a month for Lite, about $20 for Standard, and about $70 for Pro, converted from CNY prices that move with the exchange rate. What separates those tiers isn&#8217;t model access, it&#8217;s how many agents you can run at once.</p><p>A promotional rate on an undefined unit for a model that changes under you is a trial, not a budget line. Third-party resellers have started listing per-token figures for the preview endpoint, but those are reseller rates, not Alibaba&#8217;s published price.</p><h4>Kimi K3 is the one with a scorecard</h4><p>Moonshot launched Kimi K3 on July 16, 2026, a 2.8 trillion parameter model with native multimodal support, a 1 million token context window, and always-on reasoning. Some outlets date the launch to July 17.</p><p>The difference from Qwen isn&#8217;t size. It&#8217;s that somebody outside the company checked the homework. Artificial Analysis put K3 at 57 on its Intelligence Index, comparable to Opus 4.8 and GPT-5.5, behind Fable 5 and GPT-5.6 Sol. It also took first place on Frontend Code Arena at 1679 Elo, a seventeen-place jump from K2.6.</p><p>The catch is the bill. K3 runs $3.00 per million input tokens and $15.00 per million output, with a $0.30 cache hit rate. During the Artificial Analysis evaluation, K3 generated 130 million output tokens against 70 million for GPT-5.6 Sol and 87 million for Fable 5. Developer Theo Browne&#8217;s widely shared read was blunt about the gap between per-token price and per-task price. Artificial Analysis put K3&#8217;s cost per task at $0.94, close to GPT-5.6 Sol at $1.04 and about half of Opus 4.8.</p><p>One thing to note on sourcing. Artificial Analysis reported K3&#8217;s GDPval-AA v2 Elo at 1668, while VentureBeat reported 1,687 for the same benchmark. Small gap, but it&#8217;s there, and I&#8217;d rather you know than not.</p><p>And the weights? As of July 20, no official K3 Hugging Face repository or license had been published, with Moonshot planning full weights by July 27. That date is three days out from this writing.</p><h4>GLM 5.2 is the one you can hold</h4><p>Z.ai released GLM 5.2 with weights on Hugging Face under an unrestricted MIT license, letting teams download it, fine-tune it, and self-host it for the cost of their own compute. It posts 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1, and Z.ai&#8217;s list price is around $1.40 per million input tokens and $4.40 per million output.</p><p>Two caveats worth your attention. The parameter count is disputed. Both 744B and 753B appear across vendor and reseller pages, with Artificial Analysis using 744B total and roughly 40B active. And route pricing swings hard. OpenRouter has routed GLM 5.2 well below Z.ai direct, and its own listing currently shows $0.77 input and $2.42 output. Check your route before you model your costs.</p><p>It&#8217;s also verbose, burning around 43K output tokens per Index task against 16K for GPT-5.5, which pushes effective cost per task closer to the frontier than the sticker suggests.</p><h4>So, What</h4><p><strong>Do this:</strong> If you need open weights on disk this week, GLM 5.2 is the only one of the three that qualifies. It&#8217;s cheaper, open today, and independently verified.</p><p><strong>Wait on this:</strong> Kimi K3 for self-hosting. The API is live and the scores are real, but the weights and the license land July 27. Prep your vLLM or SGLang staging now and route production through the API in the meantime.</p><p><strong>Skip this for now:</strong> Migrating anything to Qwen 3.8 Max Preview. Test it, absolutely. But treat it as an evaluation target, keep a known-good fallback, and expect outputs to shift without the model ID changing.</p><p><strong>Stealable takeaway:</strong> A parameter count is a marketing number. The active parameter count, the license file, and the third-party score are the buying numbers. If a launch gives you the first and not the other three, you&#8217;re reading an announcement, not a release.</p><div><hr></div><h4>Sources</h4><ol><li><p><a href="https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/">MarkTechPost, Alibaba Previews Qwen3.8-Max</a></p></li><li><p><a href="https://newsletter.towardsai.net/p/tai-214-kimi-k3-brings-open-weight">Towards AI, TAI #214: Kimi K3 Brings Open Weight Closer to the Frontier</a></p></li><li><p><a href="https://docs.qoder.com/events/qwen-max-preview">Qoder, Qwen3.8-Max-Preview Discount Event Page</a></p></li><li><p><a href="https://docs.qwencloud.com/token-plan/overview">QwenCloud, Token Plan Overview</a></p></li><li><p><a href="https://www.eesel.ai/blog/qwen38-max-pricing">eesel AI, Qwen3.8-Max Pricing</a></p></li><li><p><a href="https://coursiv.io/blog/qwen-3-8">Coursiv, Qwen 3.8 Preview Access, Specs, Pricing</a></p></li><li><p><a href="https://www.qwen3coder.com/qwen3-8">Qwen3Coder Guide, Qwen3.8-Max-Preview Status</a></p></li><li><p><a href="https://www.digitalapplied.com/blog/alibaba-token-plan-qwen-pricing-credits-agent-seats-2026">DigitalApplied, Alibaba Token Plan Decoded</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5">Artificial Analysis, Kimi K3 Achieves #3 on the Intelligence Index</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/four-frontier-launches-in-eight-days-six-labs-now-field-a-model-above-50-on-the-artificial-analysis-intelligence-index">Artificial Analysis, Four Frontier Launches in Eight Days</a></p></li><li><p><a href="https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems">VentureBeat, Moonshot AI Releases Kimi K3</a></p></li><li><p><a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">Simon Willison, Kimi K3 and the Pelican Benchmark</a></p></li><li><p><a href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation">Nathan Lambert, Interconnects: Kimi K3, The Open-Weights Escalation</a></p></li><li><p><a href="https://vorplabs.com/models/releases/kimi-k3">Vorp Labs, Kimi K3 Release Review</a></p></li><li><p><a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter, Kimi K3</a></p></li><li><p><a href="https://kylon.io/blog/kimi-k3-benchmark-2026">Kylon, Kimi K3 Benchmarks 2026</a></p></li><li><p><a href="https://explainx.ai/blog/kimi-k3-run-locally-open-weights-desktop-july-2026">ExplainX, Run Kimi K3 Locally, Weights Prep</a></p></li><li><p><a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">VentureBeat, Z.ai&#8217;s GLM-5.2 Beats GPT-5.5 on Long-Horizon Coding</a></p></li><li><p><a href="https://www.morphllm.com/glm-5-2">Morph, GLM-5.2 Overview</a></p></li><li><p><a href="https://www.morphllm.com/glm-5-2-vs-kimi-k3">Morph, GLM-5.2 vs Kimi K3</a></p></li><li><p><a href="https://codersera.com/blog/glm-5-2-complete-guide-2026/">Codersera, GLM-5.2 Complete Guide</a></p></li><li><p><a href="https://openrouter.ai/z-ai/glm-5.2">OpenRouter, GLM 5.2</a></p></li></ol>]]></content:encoded></item><item><title><![CDATA[70+ Languages, a Few Seconds Behind You. Gemini 3.5 LIVE Translate Preview. ]]></title><description><![CDATA[Google Put a Price on Real-Time Interpretation: About 4 cents a Minute.]]></description><link>https://aisignal.veletica.com/p/70-languages-a-few-seconds-behind</link><guid isPermaLink="false">https://aisignal.veletica.com/p/70-languages-a-few-seconds-behind</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Wed, 22 Jul 2026 12:58:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!HNBq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HNBq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HNBq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HNBq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3045931,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/207974160?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HNBq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!HNBq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98e4dd57-d613-4b52-b99c-7b720bcc0f46_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On June 9th, Google attached a billing rate to something that used to require a trained professional wearing a headset in a glass booth: live interpretation. The model is called Gemini 3.5 Live Translate, and in the next five minutes you&#8217;ll know what actually shipped, what it costs, which claims are proven, and whether you should use it, build on it, or hold off.</p><p>The verdict up front: this is a real release, it landed on three surfaces at once, and the pricing is aggressive. You can test it for free today. But nearly every performance claim in circulation traces back to Google itself, early hands-on reports describe real accuracy and turn-taking gaps, and it&#8217;s a preview model with no service guarantees. So, use it now, and trust it slowly.</p><p>Everything below comes from Google&#8217;s pricing tables, the official model documentation, the DeepMind model card, and the launch coverage, cross-checked against each other. The receipts are linked at the bottom.</p><h2>What actually shipped</h2><p>Gemini 3.5 Live Translate is a streaming speech-to-speech model. You talk, and a few seconds later it speaks your words in another language. Google says it auto-detects more than 70 languages and preserves the speaker&#8217;s intonation, pacing, and pitch. It rolled out in three places at once.</p><p>For everyone, it&#8217;s now the engine inside the Google Translate app on Android and iOS, globally and free. One naming detail worth getting straight, because it confuses people: the &#8220;Live translate&#8221; feature itself is not new. It first appeared in the app in August 2025. A headphone-based beta followed in December 2025. What changed in June is the model underneath, and Android also picked up a listening mode that plays the translation quietly through your earpiece.</p><p>For developers, it&#8217;s available in public preview through the Gemini Live API and Google AI Studio under the model ID gemini-3.5-live-translate-preview.</p><p>For enterprises, it&#8217;s headed into Google Meet, currently in private preview for select Workspace customers, with a broader rollout promised later this year.</p><h2>Why this one is different</h2><p>Most translation tools you&#8217;ve used run a relay race. Your speech becomes text, the text gets translated, and a synthetic voice reads the result back. That handoff chain is why old translation apps gave us the international tourist ritual where you type a sentence, rotate your phone 180 degrees, and present it to a stranger like a tiny glowing peace treaty.</p><p>This model skips the relay. It&#8217;s audio in, audio out, and it generates translated speech continuously while you&#8217;re still talking, staying a few seconds behind you instead of waiting for you to finish. Google&#8217;s pitch is that the output keeps your voice&#8217;s character rather than flattening you into robot narrator. That pitch, to be clear, is Google&#8217;s own description of Google&#8217;s product.</p><h2>The spec sheet</h2><p>The official model page is refreshingly blunt about what this thing is and is not. Input is audio only. Output is translated audio plus a text transcript. The context window runs 131,072 tokens in and 65,536 out. Per the DeepMind model card, it&#8217;s built on Gemini 3 Pro.</p><p>Now the off switches, and there are a lot of them: no function calling, no structured outputs, no thinking, no caching, no batch API, no code execution, no search grounding. This is a translation pipe, full stop, and any app logic you want (routing calls, saving transcripts, updating your CRM) has to live in your own backend, triggered by the transcripts it hands you.</p><h2>The pricing math</h2><p>This is where it gets interesting. During the public preview, the free tier costs nothing for both input and output. Google is letting you prototype the entire flow for zero dollars.</p><p>On the paid tier, audio bills at 25 tokens per second. Input runs $3.50 per million tokens, which Google&#8217;s own footnote converts to $0.0053 per minute. Output runs $21.00 per million tokens, or $0.0315 per minute. Combined, Google&#8217;s published effective rate is about $0.0368 per minute of audio. Run a conversation for a full hour and you&#8217;ve spent roughly $2.21.</p><p>Both directions bill, so a two-way conversation counts the speech going in and the translation coming out. Even so, the number is small enough that &#8220;can we afford live translation&#8221; stops being the question. The question becomes whether the output is good enough.</p><h2>The Meet jump</h2><p>Speech translation in Google Meet used to cover five languages. The new model pushes that past 70, and Google says it unlocks more than 2,000 language combinations in a single meeting. It also drops the old limitation of routing everything through English.</p><p>Before your ops team gets excited, note what&#8217;s missing: this is a private preview for selected Workspace customers, and Google has published no pricing, no tier requirements, and no general availability date for it. Grab is testing the model for driver and traveler calls, a channel Google says handles over 10 million voice calls per month, but that figure comes from Google&#8217;s launch post.</p><h2>The receipts nobody reads</h2><p>Here&#8217;s the part that separates this from a press release. Google&#8217;s model card evaluates the system on translation quality (using an automated metric called AutoMQM), latency, and speech naturalness. Those are Google&#8217;s own internal evaluations. The glowing launch quotes from partners like Agora, LiveKit, and Vision Agents appear in Google&#8217;s own announcement, which means they&#8217;re testimonials Google selected.</p><p>Independent testing is still thin six weeks in, and what exists is mixed. Early hands-on coverage flags accuracy problems and awkward turn-taking in real conversations. There&#8217;s also a structural issue reviewers keep raising: some languages save the verb for the very end of the sentence, which means a streaming translator is essentially guessing the punchline before the joke is over. German speakers, you know exactly what I mean.</p><p>Add the standard preview caveats. There&#8217;s no SLA, the model can change before general availability, and preview pricing can change too. And with 2,000+ language pairs, quality will not be uniform. Your pair might be great. It might not.</p><h2>So what</h2><p>Do this: update the Google Translate app and try Live Translate with headphones, because it costs you nothing and the demo value alone is worth it. If you build things, prototype on the free preview tier this month while input and output are free.</p><p>Skip this: putting it in front of customers without a fallback path and without testing your exact language pairs under real conditions, with accents, background noise, and people talking over each other.</p><p>Wait on this: any enterprise planning built around Meet. Private preview, select customers, zero published pricing. There&#8217;s nothing to plan against yet.</p><p>The takeaway you can steal: Google just made live translation cheap and effortless to try but proving it works for your languages is still your job.</p><h3>Sources</h3><ul><li><p><a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/">Google&#8217;s announcement: Fluid, natural voice translation with Gemini 3.5 Live Translate</a></p></li><li><p><a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.5-live-translate-preview">Official model page: gemini-3.5-live-translate-preview</a></p></li><li><p><a href="https://ai.google.dev/gemini-api/docs/pricing">Gemini Developer API pricing (Live Translate section)</a></p></li><li><p><a href="https://deepmind.google/models/model-cards/gemini-3-5-audio/">Google DeepMind model card: Gemini 3.5 Audio (Live Translate)</a></p></li><li><p><a href="https://ai.google.dev/gemini-api/docs/live-api/live-translate">Live Translation developer guide</a></p></li><li><p><a href="https://9to5google.com/2026/06/09/gemini-3-5-live-translate-meet/">9to5Google: Gemini 3.5 Live Translate rolling out to Google Meet and Translate</a></p></li><li><p><a href="https://slator.com/google-ai-live-speech-translation-gemini-3-5-live-translate/">Slator: Google expands AI live speech translation</a></p></li><li><p><a href="https://www.eesel.ai/blog/what-is-gemini-3-5-live-translate">eesel AI: What is Gemini 3.5 Live Translate (early tester reports)</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Kimi K3 Beat Almost Everyone. Then the Bill Showed Up.]]></title><description><![CDATA[The Biggest "Open" Model Ever Released Isn't Open Yet.]]></description><link>https://aisignal.veletica.com/p/kimi-k3-beat-almost-everyone-then</link><guid isPermaLink="false">https://aisignal.veletica.com/p/kimi-k3-beat-almost-everyone-then</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Tue, 21 Jul 2026 20:05:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DmD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DmD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DmD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DmD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:728816,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/207962525?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DmD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 424w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 848w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!DmD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89a06c6b-d83d-4112-8aeb-1bc095ea3716_1376x768.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Moonshot AI shipped Kimi K3 on July 16, and every headline grabbed the same word: biggest. The biggest open-source model in history. 2.8 trillion parameters. Benchmarks nipping at the heels of the best closed models on the planet.</p><p>If you route real work through LLMs, coding agents, research pipelines, anything with a monthly token bill, this breakdown does one specific job for you. By the end you&#8217;ll know exactly which work K3 deserves, which work it would quietly triple your costs on, and what the &#8220;open source&#8221; label is hiding until July 27.</p><p>The verdict up front: K3 is the strongest open-class model anyone has released, and independent numbers back that up. It&#8217;s also slow, spectacularly wordy, and roughly triple the price of the Kimi most users were running last month. And as of today, you can&#8217;t download it. The weights are a promise with a date attached. So treat K3 as an escalation tier. Route it the jobs your cheaper models keep failing, and nothing else.</p><p>Now the receipts.</p><h3>What K3 actually is</h3><p>K3 is a 2.8 trillion parameter mixture-of-experts model. Only a small slice of that fires per token, sixteen experts at a time, so the headline number describes total capacity, not what runs on every request. The context window is 1,048,576 tokens. Reasoning stays on at all times, with adjustable effort. It reads images natively, which matters more than it sounds once your agents start working from screenshots and UI states. The API speaks the OpenAI SDK format, so it drops into existing toolchains without surgery.</p><p>Input costs $3 per million tokens. Output costs $15 per million. A cache hit drops input to $0.30, and that discount ends up doing serious load-bearing work in the cost story. Hold that thought.</p><h3>The receipts, from people who don&#8217;t work at Moonshot</h3><p>Artificial Analysis scores K3 a 57 on its Intelligence Index. Only Claude Fable and GPT-5.6 Sol Max sit above it there. Vals AI ranks it second on its own index. On Arena&#8217;s frontend code leaderboard, K3 opened at number one. Its GPQA Diamond score, the graduate-level science benchmark built to resist Googling, lands at 93.5 percent.</p><p>Moonshot&#8217;s own launch table goes further. The company claims K3 performed competitively with Fable 5 and substantially outperformed Opus 4.8, GPT-5.6 Sol, and GPT-5.5. That&#8217;s the vendor talking, and their table mixes different agent harnesses (KimiCode, Claude Code, Codex, Terminus), which means some rows compare models under different test conditions. Use it to build a shortlist. Don&#8217;t use it to crown anybody.</p><h3>The part the launch posts skipped</h3><p>Artificial Analysis clocks K3 at roughly 39 output tokens per second. The first token takes about 2.6 seconds to arrive. That&#8217;s slow for 2026. And this model talks a lot. Running the Intelligence Index evaluations, K3 generated around 130 million tokens where the average model needs about 63 million. Picture the coworker who answers &#8220;quick question&#8221; with a whiteboard session and then a recap email about the whiteboard session. Now bill that coworker at $15 per million output tokens.</p><p>Verbosity at premium output rates is a compounding problem, because agents don&#8217;t send one prompt. They loop for hours. Moonshot knows this, which is exactly why the 90 percent cache discount exists. Turn it on before you run anything, or the invoice will teach you the lesson instead.</p><h3>About that &#8220;open source&#8221; headline</h3><p>As of this writing, you cannot download K3. Moonshot says the full weights arrive by July 27 under a Modified MIT license. Until the files actually appear, K3 is an API product with a very good press release. And once the weights do land, self-hosting means real hardware. Moonshot recommends 64 or more accelerators for deployment. Your gaming rig is not invited, and neither is your startup&#8217;s single rented GPU.</p><p>For almost everyone, K3 will remain a hosted model. The open label is a story about future customization rights, not about something you&#8217;ll run in the garage this summer.</p><h3>Is it a good agent brain?</h3><p>Yes, with supervision. K3 supports OpenAI-style function calling, JSON schema structured outputs, and dynamic tool loading, which is the exact plumbing frameworks like Hermes Agent and OpenClaw depend on. Hermes already lists moonshotai/kimi-k3 in its OpenRouter catalog, and that&#8217;s the clean path right now, because the direct Kimi provider catalogs still stop at the K2 family.</p><p>Moonshot&#8217;s own guidance carries warnings worth respecting. K3 expects its thinking history preserved, so don&#8217;t swap it into the middle of a thread another model started. The docs also flag it as overly proactive, their words, which in agent land means it may take steps you never asked for. Every credible setup guide says the same thing about that combination: keep shell access, browser control, and external writes behind explicit approvals until you&#8217;ve watched it work on your own tasks.</p><h3>K3 vs GLM-5.2: brains or bargain</h3><p>Z.ai&#8217;s GLM-5.2 arrived in mid-June, and it&#8217;s the other open flagship in this conversation. The split is clean. K3 wins on capability, GLM-5.2 wins on economics.</p><p>On the Artificial Analysis index, GLM-5.2 scores 51 to K3&#8217;s 57. Its weights shipped on Hugging Face on day one under a straight MIT license, no waiting period. Sources list the model at 744 billion total parameters or 753 billion, depending on who&#8217;s counting, a discrepancy I&#8217;m flagging rather than resolving for you. About 40 billion of those fire per token. Input runs $1.40 per million on Z.ai&#8217;s API. Output runs $4.40. Reported throughput is several times faster than K3&#8217;s. And GLM-5.2 is text only, no image input.</p><p>So the routing logic writes itself. Visual agent work and the genuinely hard long-horizon jobs go to K3. High-volume text work goes to GLM-5.2, and your accountant sends Z.ai a card.</p><h3>Coming from K2.6 or K2.7? The upgrade math</h3><p>Don&#8217;t think of K3 as an update to the K2 line. It&#8217;s a new base model with new behavior. The K2 generation runs a 256K context. K3 goes to a full million. K2 was text-focused, and K3 sees images. The frontend coding jump is real too, since the K2 line sat mid-pack on Arena&#8217;s frontend rankings and K3 opened in first.</p><p>The price moved just as hard in the other direction. K2.7 Code costs $0.95 per million input tokens. Its output runs $4 per million. It also isn&#8217;t going anywhere, and for routine coding it remains the rational default. Meanwhile, Moonshot has marked K2.5 and moonshot-v1 for retirement on August 31, and new users already can&#8217;t access them. If production traffic runs on either one, your migration clock started last week.</p><h3>Context worth having</h3><p>The Kimi line had serious customers before this launch. Fortune reports that Cursor used Kimi models to help build its Composer 2 coding agent. DoorDash&#8217;s CTO said in a July post that the company delegates lower-level work to K2.6. Thinking Machines tapped K2.5 to generate early post-training data for its Inkling model. Moonshot also raised $2 billion in May. That round valued the company above $20 billion. A statement from the company&#8217;s financial advisor put annual recurring revenue past $200 million, which is self-reported, so salt it accordingly.</p><h3>So what do you do with this?</h3><p><strong>Do this:</strong> If you run long agent workflows, big-repo coding, or anything visual, test K3 this week through OpenRouter with caching on and a hard token budget. Send it only the jobs your current model fails.</p><p><strong>Skip this:</strong> Self-hosting plans until the weights actually exist, and routine high-volume coding where K2.7 Code or GLM-5.2 already passes your tests. Paying K3 rates for work a $4-output model handles is just tipping Moonshot.</p><p><strong>Wait on this:</strong> Independent head-to-head agent benchmarks, since the vendor table mixes harnesses. Also July 27 itself, when we find out whether the weights ship as promised.</p><p><strong>Steal this line:</strong> K3 is an escalation tier, not a new default. Route it your failures and leave everything else alone.</p><p>More when the weights drop on the 27th. If they drop.</p><h3>Sources</h3><ul><li><p><a href="https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals">Bloomberg: Moonshot Unveils Kimi K3 AI Model, Narrowing Gap With US Rivals</a></p></li><li><p><a href="https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems">VentureBeat: China&#8217;s Moonshot AI releases Kimi K3, the largest open-source model ever</a></p></li><li><p><a href="https://fortune.com/2026/07/16/moonshots-kimi-k3-pushes-chinese-ai-into-fable-level-territory/">Fortune: Moonshot&#8217;s Kimi K3 pushes Chinese AI into Fable-level territory</a></p></li><li><p><a href="https://www.forbes.com/sites/tylerroush/2026/07/17/chinese-ai-startup-moonshot-unveils-kimi-k3-model-will-it-challenge-openai-and-anthropic/">Forbes: Chinese AI Startup Moonshot Unveils Kimi K3 Model</a></p></li><li><p><a href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation">Interconnects (Nathan Lambert): Kimi K3, The open-weights escalation</a></p></li><li><p><a href="https://artificialanalysis.ai/models/kimi-k3">Artificial Analysis: Kimi K3 model page</a></p></li><li><p><a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter: Kimi K3 listing</a></p></li><li><p><a href="https://www.kimi.com/resources/kimi-k3-pricing">Moonshot AI: Kimi K3 pricing</a></p></li><li><p><a href="https://www.verdent.ai/guides/agents/kimi-k3-api-guide">Verdent: Kimi K3 API Guide</a></p></li><li><p><a href="https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index">Artificial Analysis: GLM-5.2 is the new leading open weights model</a></p></li><li><p><a href="https://venturebeat.com/technology/z-ais-open-weights-glm-5-2-beats-gpt-5-5-on-multiple-long-horizon-coding-benchmarks-for-1-6th-the-cost">VentureBeat: Z.ai&#8217;s open-weights GLM-5.2 beats GPT-5.5 on long-horizon coding benchmarks</a></p></li><li><p><a href="https://openclawlaunch.com/guides/hermes-agent-kimi-k3">OpenClaw Launch: Hermes Agent + Kimi K3, OpenRouter Setup and Model Checks</a></p></li><li><p><a href="https://myclaw.ai/blog/kimi-k3-vs-opus-4-8">MyClaw: Kimi K3 vs Claude Opus 4.8</a></p></li><li><p><a href="https://benchlm.ai/moonshot/api-pricing">BenchLM: Kimi API Pricing, July 2026</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Hermes Agent: The Honest Setup Guide (Model, Subscription, and Where to Run It.]]></title><description><![CDATA[Your Agent Forgets You Every Session. This One Writes Down What It Learned.]]></description><link>https://aisignal.veletica.com/p/hermes-agent-the-honest-setup-guide</link><guid isPermaLink="false">https://aisignal.veletica.com/p/hermes-agent-the-honest-setup-guide</guid><dc:creator><![CDATA[Oscar Villa]]></dc:creator><pubDate>Fri, 17 Jul 2026 15:11:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!l2Kz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!l2Kz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!l2Kz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!l2Kz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg" width="1456" height="813" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:813,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2608562,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://aitrendzio.substack.com/i/207434328?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!l2Kz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 424w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 848w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!l2Kz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85ff39f3-cc7e-48a0-a94e-3c3fb2820758_2752x1536.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hermes Agent sits at more than 210,000 GitHub stars as of mid-July, less than five months after its public release. That makes it the highest-ranked dedicated AI agent framework on GitHub, currently #25 among all repositories on the platform. NVIDIA also reports it&#8217;s the most used agent in the world by OpenRouter traffic. So the hype is real, the install base is real, and yet buried in the official documentation is a line most people never read.</p><p>Nous Research tells you not to use their own Hermes 4 models inside their own agent.</p><p>I&#8217;ll get to why, because it&#8217;s the single most useful thing to understand before you spend a dollar on this. But first, the verdict, since you came here for a decision and not a book report.</p><p>If you want the setup that works with the least friction: install Hermes on a cheap virtual private server (VPS), subscribe to Nous Portal Plus at $20 a month, and set Claude Sonnet 4.6 as the brain. That combination covers the model, web search, browser automation, image generation, and text-to-speech under one login. Everything below is the receipts, pulled from the official docs, the GitHub issue tracker, and the pricing pages rather than the launch tweets.</p><h2>What this thing actually is</h2><p>Hermes Agent is an open-source AI agent from Nous Research, released February 25, 2026 under the MIT license. The software itself is free. No premium tier, no locked features.</p><p>The pitch is different from every chatbot you&#8217;ve used. Hermes runs continuously on a machine you control, and when it solves a hard problem it writes the solution down as a reusable skill file. Next time, it doesn&#8217;t re-derive the answer. It also keeps a memory of your projects and preferences across sessions, so you stop re-explaining your life every morning.</p><p>And it lives where you already are. One gateway process connects it to Telegram, Discord, Slack, WhatsApp, Signal, email, and a list of platforms that now tops twenty. You message your agent from your phone while it works on a server somewhere else.</p><p>The July 1 v0.18.0 release added a /learn command that distills a reusable skill from a URL, a directory, or a workflow you just walked it through, plus a /goal mode where a separate judge model grades the work until it actually meets your success criteria. That judge idea matters more than it sounds. &#8220;Done&#8221; stops meaning &#8220;the model got tired.&#8221;</p><h2>How to use it</h2><p>Installation is one command on Linux, macOS, or WSL2:</p><pre><code><code>curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash</code></code></pre><p>Then <code>hermes setup</code> walks you through picking a model provider, and <code>hermes</code> starts the chat. There&#8217;s also a desktop app, shipped June 2, if terminals aren&#8217;t your happy place. Native Windows support exists but Nous labels it experimental, so Windows users should run it through WSL2.</p><p>Two cautions before you paste that command. You are piping a script from the internet into your shell and then handing an AI agent access to a terminal. That&#8217;s the whole value proposition and also the whole risk. Run it on a VPS or a sandboxed machine first, not the laptop with your tax returns. Second, the GitHub README notes that some antivirus engines flag the bundled uv installer as suspicious because it&#8217;s an unsigned binary that downloads packages. Known false positive, but verify you&#8217;re pulling from the official NousResearch repo before you trust anything.</p><p>Day-to-day use looks less like prompting and more like delegating. You connect Telegram, tell it &#8220;every morning at 7, summarize my unread emails and the top three stories on my industry,&#8221; and the built-in scheduler handles it from plain English. No cron syntax. Bigger jobs can fan out to isolated subagents that run in their own sandboxes and report back.</p><h2>The model question, and the answer nobody expected</h2><p>Hermes Agent is model-agnostic. Bring any brain you want. Which raises the obvious question: whose?</p><p>Nous&#8217;s own documentation answers it with unusual honesty. The Hermes 4 model family, the lab&#8217;s namesake, is tuned for chat and reasoning rather than the rapid-fire tool-calling loop the agent depends on. The docs say to skip it for agent work and pick a frontier agentic model instead. Their listed picks: Claude Sonnet 4.6 as the best general-purpose agentic model, GPT-5.5 Pro for heavy reasoning, Gemini 3 Pro for giant context, and DeepSeek V4 Pro when cost matters most.</p><p>A lab telling you its own product isn&#8217;t the right choice inside its own agent is the kind of signal worth trusting. Restraint like that is rare in this market.</p><p>Now, how to pay for that brain. You have five real routes.</p><p><strong>Nous Portal, the default answer.</strong> Launched April 27, 2026, it&#8217;s one subscription covering 300+ models plus the tool gateway (web search, browser automation, image generation, TTS). Tiers per Nous: a free evaluation tier with $0.10 in monthly credits, then Plus at $20, Super at $100, and Ultra at $200 per month. The free tier exists to kick the tires, not to run workloads. For most builders, Plus is the sweet spot because it replaces separate accounts with a search provider, an image provider, a TTS provider, and a browser automation provider. One OAuth, one bill.</p><p><strong>OpenRouter.</strong> Broad model routing with your own key. Flexible, though a Hermes GitHub discussion pegs the platform fee at roughly 5.5% on top of provider rates.</p><p><strong>Direct API keys.</strong> An Anthropic or OpenAI key, pay per token, no middleman. Cleanest billing, most plumbing.</p><p><strong>Your existing subscriptions.</strong> Hermes can route through a GitHub Copilot subscription, and a May 15 integration lets X Premium+ subscribers use their Grok subscription inside the agent. The Claude route is the messy one. Official Hermes docs state the Anthropic OAuth path only works on a Claude Max plan with extra usage credits purchased on top, and that Claude Pro subscribers can&#8217;t use it at all. Worse, users report in an open GitHub issue that the OAuth path still bills against pay-per-token extra-usage credits instead of the subscription&#8217;s included quota. Until that settles, treat &#8220;run Hermes on my Claude plan&#8221; as a science experiment, not a plan.</p><p><strong>Fully local.</strong> Ollama or vLLM with open-weight models, zero API spend, full privacy. NVIDIA&#8217;s coverage points at Qwen 3.6 in the 27B and 35B range as the current local sweet spot, though that assumes serious GPU hardware.</p><p>On cost, the honest math from Hostinger&#8217;s breakdown: expect $5 to $80 per month all-in depending on the model. DeepSeek V4 Flash runs about $0.14 per million input tokens and $0.28 per million output. Claude Opus 4.8 runs $5 and $25 for the same. Roughly a 30x spread for the identical workload. One quiet cost lever: the same analysis found messaging gateways like Telegram send 15,000 to 20,000 tokens of overhead per request, versus 6,000 to 8,000 through the CLI. Where you talk to your agent changes your bill.</p><h2>The Grok 4.5 wildcard</h2><p>The loudest corner of the Hermes community right now is running Grok 4.5 as the brain, and for once the noise has independent numbers behind it.</p><p>xAI shipped Grok 4.5 in early July, and it landed fourth on Artificial Analysis&#8217;s Intelligence Index, behind Claude Fable 5, GPT-5.5, and Claude Opus 4.8. That placement came from a 16-point jump over Grok 4.3, which the firm logged as the largest single-generation leap any lab has posted on that index. On the same firm&#8217;s Coding Agent Index it scored 76, one point below Fable 5 running in Claude Code.</p><p>The price is what turns those rankings into a movement. MindStudio&#8217;s analysis puts Grok 4.5 at $2 per million input tokens, a fraction of frontier pricing for near-frontier agentic scores. xAI also disclosed a cached-input rate of $0.50 per million, a 75% discount that lands especially well in agent loops, where the same context gets re-sent over and over. Benchmark coverage from TechTimes worked that out to roughly $2.49 per completed coding task, against $11.80 for the same work through Claude Code.</p><p>The community reports match the math. Nick Vasilescu, who says he has run Grok 4.5 across more than 100 Hermes agents, calls it the grittiest model he&#8217;s used, meaning it keeps grinding on a task instead of quitting early. Greg Isenberg&#8217;s session with him showed a Hermes agent building a full landing page in about 40 seconds. Worth noting: the writeup of that session flags it as an informal demo with no repeated trials or blind scoring, so treat it as a field report, not a benchmark.</p><p>Now the fine print, because there&#8217;s real fine print. TechTimes&#8217; benchmark coverage reports the hallucination rate roughly doubled to 54% compared with its predecessor. The context window also shrank to 500,000 tokens, down from Grok 4.3&#8217;s 1 million, with no official explanation from xAI. And MindStudio&#8217;s testing found it strong on structured, high-volume work but weaker on hard multi-step reasoning, which is why they position it as the workhorse in a pipeline rather than the orchestrator on top.</p><p>If you want to run it inside Hermes, five tips:</p><p><strong>Pick your route.</strong> Grok 4.5 is reachable four ways: the Nous Portal catalog, OpenRouter, a direct xAI API key, or the X Premium+ OAuth integration that reuses a Grok subscription you may already pay for. Switch with <code>hermes model</code> and pick it from the list.</p><p><strong>Drop the reasoning to low.</strong> Vasilescu&#8217;s recommendation from his testing, and it&#8217;s where the speed reputation comes from. High reasoning mode erases much of the latency advantage on routine tasks.</p><p><strong>Put a judge on it.</strong> The /goal command in v0.18 lets a separate model grade the work against your success criteria. With a reported hallucination rate that high, a cheap verification pass is not optional. It&#8217;s the whole strategy.</p><p><strong>Route by difficulty.</strong> The pattern MindStudio recommends: Grok handles the volume (summaries, drafts, structured coding, research sweeps), and anything needing deep multi-step reasoning escalates to Sonnet 4.6 or Fable 5. You get the speed without betting the hard calls on it.</p><p><strong>Watch long sessions.</strong> The 500K context window is generous until an always-on agent with weeks of memory starts stuffing it. Lean on Hermes skills and memory recall instead of dumping everything into context.</p><p>One more angle worth knowing: because it&#8217;s an xAI model, Grok 4.5 has native reach into X data, which several creators point to as its edge for research built on real-time X content. If your workflows live there, that&#8217;s a genuine differentiator no other brain offers.</p><h2>Every way to run it</h2><p>From cheapest to laziest:</p><p><strong>Your own machine.</strong> Free, but the agent sleeps when your laptop does, which defeats the always-on point.</p><p><strong>A cheap VPS.</strong> The community favorite. Hermes needs about 1 vCPU, 2GB RAM, and 20GB of disk, which fits the bottom tier at Hetzner (the CX22 plan runs about $4.35 a month), DigitalOcean, or Linode. A Raspberry Pi 5 also works if the model lives in the cloud.</p><p><strong>Docker or SSH backends.</strong> Same idea with more isolation. Hermes supports six terminal backends total.</p><p><strong>Serverless.</strong> The Daytona and Modal backends hibernate when idle, so you pay close to nothing between tasks. Clever for an agent that mostly waits for your messages.</p><p><strong>Hermes Cloud.</strong> Nous&#8217;s own hosted option, currently in preview. Ten dollars minimum in credits or an active subscription, and your agent is online in seconds with no server to babysit. There&#8217;s also FlyHermes, a third-party managed option, if you want someone else handling uptime entirely.</p><p><strong>Local GPU.</strong> An RTX box or DGX Spark running everything, model included, on your desk. Maximum privacy, maximum hardware bill.</p><h2>So what</h2><p><strong>Do this:</strong> if you run multi-step workflows more than a few times a week, deploy Hermes on a $5 VPS with Nous Portal Plus and Claude Sonnet 4.6. Total damage lands near $25 a month for an always-on agent that compounds. If speed and cost outrank accuracy for your workload, run Grok 4.5 as the volume worker with a judge model checking its output, and escalate the hard reasoning to Sonnet 4.6 or Fable 5.</p><p><strong>Skip this:</strong> if your AI use is one-off questions. A flat chatbot subscription beats this on price and setup time, and Hostinger&#8217;s analysis says the same.</p><p><strong>Wait on this:</strong> the Claude Pro/Max subscription route. The billing behavior is disputed in open GitHub issues, and Anthropic&#8217;s policy on programmatic subscription use is still settling.</p><p><strong>Steal this:</strong> an agent&#8217;s value isn&#8217;t the model it runs, it&#8217;s what it remembers. Pick the cheapest brain that doesn&#8217;t drop the ball, and let the memory do the compounding.</p><div><hr></div><h3>Sources</h3><ul><li><p><a href="https://hermes-agent.nousresearch.com/docs/">Hermes Agent official documentation</a> &#8212; Nous Research</p></li><li><p><a href="https://github.com/nousresearch/hermes-agent">Hermes Agent GitHub repository</a> &#8212; NousResearch</p></li><li><p><a href="https://hermes-agent.nousresearch.com/docs/integrations/nous-portal">Nous Portal integration guide</a> &#8212; official model recommendations and gateway details</p></li><li><p><a href="https://hermes-agent.nousresearch.com/docs/integrations/providers">AI Providers documentation</a> &#8212; Anthropic OAuth, Copilot, Vertex, Bedrock routes</p></li><li><p><a href="https://portal.nousresearch.com/manage-subscription">Nous Portal subscription plans</a> &#8212; Nous Research</p></li><li><p><a href="https://x.com/Teknium/status/2047442402303226174">Nous Portal tier announcement</a> &#8212; Teknium (Nous Research) on X</p></li><li><p><a href="https://portal.nousresearch.com/cloud">Hermes Cloud</a> &#8212; Nous Research, preview</p></li><li><p><a href="https://blogs.nvidia.com/blog/rtx-ai-garage-hermes-agent-dgx-spark/">Hermes Unlocks Self-Improving AI Agents</a> &#8212; NVIDIA Blog</p></li><li><p><a href="https://www.hostinger.com/tutorials/hermes-agent-cost">Hermes Agent cost: real monthly pricing 2026</a> &#8212; Hostinger</p></li><li><p><a href="https://www.autolearningagents.com/hermes-agent/hermes-pricing">Hermes Agent pricing: free vs Nous Portal</a> &#8212; AutoLearningAgents</p></li><li><p><a href="https://github.com/NousResearch/hermes-agent/issues/40014">GitHub issue #40014: Claude OAuth billing behavior</a> &#8212; user-reported, open</p></li><li><p><a href="https://openclawlaunch.com/guides/nous-portal">Nous Portal guide</a> &#8212; OpenClaw Launch</p></li><li><p><a href="https://www.star-history.com/nousresearch/hermes-agent/">Hermes Agent star history and global rank</a> &#8212; Star History tracker</p></li><li><p><a href="https://www.techtimes.com/articles/320038/20260709/grok-45-cuts-coding-agent-cost-80-near-frontier-speed-higher-hallucinations.htm">Grok 4.5 cuts coding-agent cost 80%</a> &#8212; TechTimes, citing Artificial Analysis benchmarks</p></li><li><p><a href="https://www.mindstudio.ai/blog/grok-4-5-cheaper-sub-agent-multi-model-workflows">How to use Grok 4.5 as a cheaper sub-agent</a> &#8212; MindStudio</p></li><li><p><a href="https://www.mindstudio.ai/blog/grok-4-5-vs-gpt-5-6-sol-agentic-coding-comparison">Grok 4.5 vs GPT-5.6 Sol: agentic coding comparison</a> &#8212; MindStudio</p></li><li><p><a href="https://x.com/nickvasiles/status/2075710380186493338">Grok 4.5 across 100+ Hermes agents</a> &#8212; Nick Vasilescu on X, self-reported testing</p></li><li><p><a href="https://www.ai.joaoqueiros.com/blog/grok-4-5-ai-cofounder-hermes-orgo-agent-stack">Grok 4.5 as an AI co-founder: the Hermes and Orgo stack</a> &#8212; analysis of the Isenberg/Vasilescu session</p></li></ul>]]></content:encoded></item></channel></rss>