<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Kaitchup – AI on a Budget]]></title><description><![CDATA[Weekly tutorials and news on adapting large language models (LLMs) to your tasks and hardware using the most recent techniques and models. The Kaitchup proposes a collection of 180+ AI notebooks regularly updated.]]></description><link>https://kaitchup.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!xY7g!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png</url><title>The Kaitchup – AI on a Budget</title><link>https://kaitchup.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 16 Aug 2026 11:12:37 GMT</lastBuildDate><atom:link href="https://kaitchup.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[The Kaitchup]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[kaitchup@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[kaitchup@substack.com]]></itunes:email><itunes:name><![CDATA[Benjamin Marie]]></itunes:name></itunes:owner><itunes:author><![CDATA[Benjamin Marie]]></itunes:author><googleplay:owner><![CDATA[kaitchup@substack.com]]></googleplay:owner><googleplay:email><![CDATA[kaitchup@substack.com]]></googleplay:email><googleplay:author><![CDATA[Benjamin Marie]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Qwen3.8 27B, Nemotron 3.5, Muse, DeepSeek V4 Pro: A Huge Week for Open-Weight AI]]></title><description><![CDATA[The Weekly Kaitchup #155]]></description><link>https://kaitchup.substack.com/p/qwen38-27b-nemotron-35-muse-deepseek</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-27b-nemotron-35-muse-deepseek</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 14 Aug 2026 20:13:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>This week was unusually rich in open-weight releases.</p><p>Meta released Muse Glimmer. Qwen finally released the weights of Qwen3.8 2.4T. NVIDIA released Nemotron 3.5 Lightning. DeepSeek just pushed a major V4 Pro update.</p><blockquote><p>And, just as I was finishing this article, <strong><a href="https://huggingface.co/Qwen/Qwen3.8-27B">Qwen3.8 27B</a> was released</strong>. It was worth the wait. According to Qwen&#8217;s evaluations, it improves very significantly over Qwen3.6 27B essentially everywhere, with particularly spectacular gains on long-horizon agentic coding: DeepSWE jumps from 13.3 to 42.2, while QwenSWEBench goes from 49.3 to 79.0. The 27B model also beats Qwen3.7-Plus on many coding and agentic benchmarks, and even beats Opus4.6 Max on QwenSWEBench, CoWorkBench, and LiveCodeBench. At this point, Qwen seems so far ahead that catching up is becoming difficult for everyone else.</p><p>There are also a few interesting changes compared with Qwen3.6. <code>preserve_thinking</code>, which Qwen3.6 introduced as an option for keeping reasoning traces across turns, is now <strong>enabled by default</strong>. Qwen3.8 also introduces <code>reasoning_effort</code>, with <code>low</code>, <code>medium</code>, and <code>xhigh</code> levels to trade reasoning depth against cost. Qwen now recommends <code>temperature=1.0</code> and <code>top_p=0.95</code> for thinking mode generally, and, for very long agentic runs, recommends allowing up to 262K tokens for reasoning and 131K for the final answer. I&#8217;ll publish a full analysis of Qwen3.8 27B within the next few days, with a separate look at its quantized versions later.</p><p>Note: <em>I think this &#8220;preserve_thinking&#8221; (which is a feature we can find in other models) should be carefully evaluated, especially its impact on inference cost. Qwen3.8 is a crazy thinker (they recommend allowing &#8220;up to 262K tokens for reasoning&#8221;). So if you preserve reasoning for each turn, the context may grow to millions of tokens. In practice though, for tool calls, reasoning is often much shorter and more/better reasoning may yield fewer turns, so I guess it can work. But I&#8217;m really interested to know how important it is, in terms of accuracy/efficiency, to preserve the reasoning traces. <strong>KV cache quantization will likely be very important</strong>.</em></p></blockquote><p>Finding enough time to cover all of this properly is hard. I don&#8217;t want to publish articles that simply repeat benchmark tables and model cards. A proper analysis means testing accuracy, token efficiency, memory, speed, and then doing it again for several quantized versions.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>For Glimmer, I already published an article covering the model and its architecture:</p><ul><li><p><a href="https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built">Muse Glimmer: Meta&#8217;s 30B Model Built for Efficient Inference</a></p></li></ul><p>I also already have speed measurements for all the non-GGUF quantized versions I&#8217;m testing.</p><p>At the moment, essentially all my compute capacity, several RTX Pro 6000 GPUs provided by <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> is being used to evaluate the quantized versions that run through vLLM: NVFP4, INT4, and my own mixed-precision versions.</p><p>You can already find those here:</p><ul><li><p><a href="https://huggingface.co/collections/kaitchup/quantized-muse-glimmer">My quantized Muse Glimmer collection</a> (they all run, but consider all of them very bad until I publish some official evaluations)</p></li></ul><p>I&#8217;ll probably publish a short article focused only on speed first, since those results are already ready.</p><p>Then, expect perhaps two deeper articles comparing Glimmer with Gemma 4 and Qwen3.6/3.8, and another article focused specifically on the quantized versions. I may include GGUF versions in that one as well, but I&#8217;m not sure yet. There is only so much GPU time, and writing time, in a week.</p><p>So, for this Weekly Kaitchup, I&#8217;ll focus on the three releases that I won&#8217;t cover in more detail:</p><ul><li><p>Qwen3.8 2.4T: we finally know what is inside it, so I can check how well my predictions held up.</p></li><li><p>Nemotron 3.5 Lightning: essentially the new Nano-sized Nemotron.</p></li><li><p>DeepSeek V4 Pro 0813: a very large update that makes the word &#8220;Preview&#8221; much more meaningful.</p></li></ul><h2>Qwen3.8 2.4T: How Good Were My Predictions?</h2><p>A few weeks ago, before Alibaba disclosed the architecture, I tried to estimate what hardware would be required to run Qwen3.8 2.4T:</p><ul><li><p><a href="https://kaitchup.substack.com/p/qwen38-what-hardware-will-you-need">Qwen3.8: What Hardware Will You Need to Run Alibaba&#8217;s 2.4T Model?</a></p></li></ul><p>At the time, we knew only one particularly important number: <strong>2.4 trillion parameters</strong>.</p><p>Everything else had to be inferred.</p><p>Now we have the weights and architecture, so it&#8217;s a good opportunity to see what I got right and what should be corrected.</p><p>Qwen3.8-2.4T-A95B has 2.4T total parameters and 95B activated parameters. It has 92 layers and uses the hybrid architecture introduced with Qwen3.5, alternating Gated DeltaNet blocks with periodic full attention. Its MoE has 512 experts, with 10 routed experts plus one shared expert activated per token. Native context is 262K, extendable to just over one million tokens.</p><p>The downloadable 2.4T checkpoint is also more limited than the hosted Qwen3.8-Max: it is text-only and requires thinking mode, while the hosted Max version adds vision, non-thinking mode, built-in tools, and a 1M default context.</p><p>So, how did my predictions go?</p><h3><strong>&#9989; It is an extremely sparse MoE</strong></h3><p>This one was the most important and easiest assumption.</p><p>I wrote that a dense 2.4T model would make very little sense for inference and that Qwen3.8 was almost certainly going to be a highly sparse MoE.</p><p>That was correct.</p><p>Only 95B of 2.4T parameters are active for each token, or roughly 4% of the entire model.</p><h3><strong>&#9989; Nearly 99% of the parameters are in the routed experts</strong></h3><p>I assumed that approximately <strong>99% of Qwen3.8&#8217;s parameters would live inside routed experts</strong>, based largely on Kimi K2 and other very large MoEs.</p><p>This turned out to be close.</p><p>Using the released dimensions, 512 experts, 92 MoE layers, hidden size 8192, and expert intermediate size 2048, the routed expert matrices account for roughly <strong>2.371 trillion parameters</strong>, or about <strong>98.8% of the entire 2.4T model</strong>.</p><p>My estimate in the original article was 2.374T routed-expert parameters.</p><h3><strong>&#9989; The routing sparsity</strong></h3><p>I used Kimi K2 as a proxy.</p><p>Kimi routes each token through 8 of 384 experts, or about 2.08% of its routed experts.</p><p>Qwen3.8 uses 10 of 512, or about <strong>1.95%</strong>.</p><p>So, while the architectures are different, the routing sparsity I used for the estimate was close enough.</p><h3><strong>&#10060; I estimated ~75B active parameters</strong></h3><p>This is the main miss.</p><p>Using Kimi K2 as a proxy, I estimated that Qwen3.8 might require compute roughly comparable to a 75B dense model per token, before routing and communication overhead.</p><p>The actual number is <strong>95B activated parameters</strong>.</p><p>That is around 27% higher than my estimate.</p><p>So, the general prediction, a 2.4T model with only a small fraction active, was correct, but Qwen3.8 is somewhat more computationally expensive per token than I expected.</p><h3><strong>&#10060; Around 1.4&#8211;1.5 TB with experts-only NVFP4</strong></h3><p>Wrong but close.</p><p>I estimated <strong>1.387 TB</strong> for a version where the routed experts are stored in NVFP4 while the rest of the network remains at higher precision.</p><p>A community <a href="https://huggingface.co/RadixArk/Qwen3.8-2.4T-A95B-NVFP4">NVFP4 checkpoint from RadixArk</a> now does exactly that: only the routed-expert linear layers are quantized to NVFP4 while the other components retain their original precision.</p><p>Its repository is <strong>1.48 TB</strong>.</p><p>So my estimate was around 6&#8211;7% too low, but the overall hardware conclusion was correct.</p><p>A single 8xB300 node is enough to deploy this model.</p><h2>Nemotron 3.5 Lightning: Nano, Updated</h2><p>NVIDIA also released <strong><a href="https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4">Nemotron 3.5 Lightning 30B-A3B</a></strong> this week.</p><p>If the naming feels slightly confusing, Nemotron 3.5 Lightning is effectively the updated model occupying Nemotron 3 Nano&#8217;s position in the family.</p><p>I previously tested the original Nano here:</p><ul><li><p><a href="https://kaitchup.substack.com/p/nemotron-3-nano-fast-cheap-and-surprisingly">Nemotron 3 Nano: A Very Fast Model That Doesn&#8217;t Think Too Much</a></p></li></ul><p>And I covered the much larger Super model here:</p><ul><li><p><a href="https://kaitchup.substack.com/p/nemotron-3-super-1m-tokens-small">Nemotron 3 Super: 1M Tokens, Small KV Cache</a></p></li></ul><p>Lightning still uses the same general idea of interleaving Mamba-2 processing with selected attention layers and sparse experts, and it supports context lengths up to 1M tokens.</p><p>So this is not a new size class. It is much closer to a refreshed Nano.</p><p>But NVIDIA has brought several ideas developed across the rest of the Nemotron 3 family back into this smaller model.</p><p>The biggest additions are around inference and agents.</p><p><strong>Multi-Token Prediction is now part of the model training</strong>, followed by an additional MTP-boosting phase. NVIDIA also releases DFlash and DSpark draft models for speculative decoding, so there are multiple ways to accelerate generation depending on concurrency and hardware.</p><p>This is interesting because MTP was already one of the important additions in Nemotron 3 Super.</p><p>NVIDIA also says that 3.5 Lightning received <strong>harness-optimized training</strong> for agent workloads. Large Nemotron models can do the expensive planning and reasoning, while Lightning is supposed to execute the many smaller calls produced by long-running agents.</p><p>There are BF16 and NVFP4 checkpoints, and NVIDIA says the model can reach up to 4x the output speed of similarly sized models. On its PinchBench test, NVIDIA reports roughly 86% accuracy while completing 10,000 tasks around 30% faster than Qwen3.6 35B at comparable accuracy. Those are NVIDIA&#8217;s numbers, so I would still like to reproduce the speed/accuracy trade-off independently.</p><h2>DeepSeek V4 Pro 0813: &#8220;Preview&#8221; Really Meant Preview</h2><p>Finally, DeepSeek (quietly) released <strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813">DeepSeek-V4-Pro-0813</a></strong>.</p><p>When DeepSeek introduced V4 in April, both V4 Pro and V4 Flash were explicitly called <strong>Preview</strong> models. The architecture was extremely innovative.</p><p>V4 Pro has around 1.6T parameters with 49B active and combines Compressed Sparse Attention with Heavily Compressed Attention. DeepSeek also introduced Manifold-Constrained Hyper-Connections and trained with the Muon optimizer. At a one-million-token context, DeepSeek reported that V4 Pro required only <strong>27% of the single-token inference FLOPs and 10% of the KV-cache memory</strong> of DeepSeek V3.2.</p><p>Accuracy was the problem.</p><p>The original V4 Pro was rather underwhelming for such a huge model, particularly on agentic tasks. Then DeepSeek released Flash-0731, and the Flash model was suddenly beating Pro Preview.</p><p>I wrote about that two weeks ago:</p><ul><li><p><a href="https://kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling">DeepSeek-V4-Flash-0731 and Inkling Small: Smaller, but Better?</a></p></li></ul><p>At the time, I wrote that Flash-0731 was &#8220;a very promising sign for the next V4 Pro update.&#8221; That update is here now. The model now also exposes three reasoning-effort settings: low, high, and max.</p><p>The accuracy jump is much more important.</p><p>A few examples from DeepSeek&#8217;s own evaluation table:</p><ul><li><p>DeepSWE: 12.8 (it was <strong>lower than Qwen3.8 27B!</strong>) &#8594; 62.7<br>Current Flash-0731: 54.4</p></li><li><p>Terminal Bench 2.1: 72.1 &#8594; 87.9<br>Current Flash-0731: 82.7</p></li><li><p>NL2Repo: 38.5 &#8594; 61.5<br>Current Flash-0731: 54.2</p></li><li><p>Cybergym: 52.7 &#8594; 83.3<br>Current Flash-0731: 76.7</p></li><li><p>Toolathlon Verified: 55.9 &#8594; 74.1<br>Current Flash-0731: 70.3</p></li><li><p>AutomationBench: 12.8 &#8594; 31.8<br>Current Flash-0731: 25.1</p></li></ul><p><strong>Pro-0813 is now ahead of the current Flash-0731 model</strong>.</p><p>DeepSWE is the most spectacular improvement: <strong>12.8 to 62.7</strong>.</p><p>As I discussed this week, agentic evaluations can move significantly depending on the harness, configuration, allowed steps, tools, and other details. Independent testing remains necessary.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;497985b7-de19-49c9-b233-74e6f419582c&quot;,&quot;caption&quot;:&quot;Coding benchmarks increasingly test more than whether a model can write a correct function. They ask an agent to explore an unfamiliar repository, run commands, edit files, interpret failures, test its work, recover from mistakes, and decide when the job is actually finished.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Laguna S 2.1: How Agent Harnesses and Inference Budgets Shape Coding Performance&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-08-13T16:11:23.125Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:210853645,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:3,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisioning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>I&#8217;ll probably add a B200 to cover Qwen3.8 27B better. Muse Glimmer has a short native max context length (131K tokens) and a small KV cache. An RTX Pro 6000 is enough. But Qwen3.8 27B&#8217;s KV cache is much larger, with a native context twice as large. 96 GB at 1.7 TB/sec is not enough with high concurrency.</p></div></div>]]></content:encoded></item><item><title><![CDATA[Laguna S 2.1: How Agent Harnesses and Inference Budgets Shape Coding Performance]]></title><description><![CDATA[Testing long-horizon coding performance on DeepSWE and Terminal-Bench 2.1]]></description><link>https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and</link><guid isPermaLink="false">https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 13 Aug 2026 16:11:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hTOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hTOO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hTOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png" width="640" height="360" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:640,&quot;bytes&quot;:2208006,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210853645?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hTOO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!hTOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3b26d3-f664-4437-bed2-85332b804e0d_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Coding benchmarks increasingly test more than whether a model can write a correct function. They ask an agent to explore an unfamiliar repository, run commands, edit files, interpret failures, test its work, recover from mistakes, and decide when the job is actually finished.</p><p>That makes benchmark performance a product of at least three things: the underlying model, the agent harness driving it, and the amount of time and context the system is allowed to consume.</p><p>I evaluated Laguna S 2.1 on two long-horizon benchmarks, DeepSWE and Terminal-Bench 2.1, to understand how much these surrounding conditions affect both accuracy and cost. In particular, I wanted to see how the model behaves outside Poolside&#8217;s native agent harness, how much additional inference-time compute improves results, and what happens when an agent is given enough time to keep working after it gets stuck.</p><p>The results show a capable model whose accuracy can improve substantially when it is given more room to work. They also expose the other side of long-horizon agents: unsuccessful trajectories can consume enormous amounts of computation without getting any closer to a correct solution.</p><p>Perhaps most importantly, the experiments reinforce that model configuration is only part of the story. The agent wrapped around the model matters enormously.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I compare Laguna S 2.1 across DeepSWE and Terminal-Bench 2.1, examine how turn limits, timeouts, and token budgets affect accuracy and cost, analyze the main failure modes, and explain why the agent harness itself has become a critical part of coding-model evaluation.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=lagunas21">Verda</a><span> provided the B300 to run the experiments described in this article.</span></p><p>Verda is the full-stack frontier AI cloud, built for high-performance inference, training, and agentic workloads with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=lagunas21">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisioning it with $5, to try their GPUs.</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p><p><em><span>Note: I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon.</span></em></p></blockquote><h2>Agentic coding evaluation setup</h2><p>I deliberately evaluated <a href="https://huggingface.co/poolside/Laguna-S-2.1">Laguna S 2.1</a> outside Poolside&#8217;s native pool harness.</p><p>For DeepSWE, I used mini-swe-agent.</p><p>For Terminal-Bench 2.1, I used the terminus-2 agent.</p><p>Poolside&#8217;s own Laguna S 2.1 evaluations use its pool agent harness, while the public DeepSWE leaderboard normally uses mini-swe-agent. My numbers should be read as measurements of Laguna operating inside these particular third-party agent setups, not as direct reproductions of Poolside&#8217;s published benchmark results.</p>
      <p>
          <a href="https://kaitchup.substack.com/p/laguna-s-21-how-agent-harnesses-and">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Muse Glimmer: Meta’s 30B Model Built for Efficient Inference]]></title><description><![CDATA[Inside Meta&#8217;s 30B local reasoning model and its tiny KV cache]]></description><link>https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built</link><guid isPermaLink="false">https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 11 Aug 2026 02:01:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tgN2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tgN2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tgN2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tgN2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png" width="672" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:672,&quot;bytes&quot;:2313372,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210643086?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tgN2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!tgN2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab3a1b08-6f7b-4cfb-961c-7249368eaeb3_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">This isn&#8217;t quite as fun as generating images of llamas, but I&#8217;ll take it.</figcaption></figure></div><p>Meta is back, and it&#8217;s great.</p><p>For a while, the center of gravity in open-weight AI moved toward Qwen, DeepSeek, Google&#8217;s Gemma family, and a growing number of Chinese labs. Muse Glimmer gives Meta a genuinely interesting entry again.</p><p>Muse Glimmer is a ~30B-parameter multimodal reasoning model from Meta Superintelligence Labs, released under Apache 2.0.  </p><ul><li><p><a href="https://huggingface.co/collections/meta-models/muse-glimmer">Muse Glimmer</a></p></li></ul><p>Glimmer is distilled from Meta&#8217;s much larger Muse Spark model, supports tool use and multimodal inputs, has a 131K context window, ships with official 4-bit quantization, and targets consumer hardware. It&#8217;s clearly positioned as a competitor to Gemma 4 and Qwen3.6 open dense models.</p><p><em>How good is Glimmer compared with Gemma 4 and Qwen3.6?</em></p><p>The benchmark results put Glimmer surprisingly close to Qwen3.6. But benchmark scores alone tell us nothing about efficiency.</p><p>Glimmer uses an exceptionally lightweight attention design, and its relatively short 131K context window suggests it was not trained around the extremely long reasoning traces Qwen3.6 can produce, sometimes stretching beyond 100K tokens.</p><p>That opens up an interesting possibility: Glimmer may deliver performance close to Qwen3.6 while approaching Gemma 4&#8217;s token efficiency and using significantly less memory overall.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I&#8217;ll first look at what makes Muse Glimmer interesting as a local model, then break down its architecture in detail, including its hybrid attention pattern, extreme GQA, gated attention, positional encoding, multimodal stack, and speculative decoding. I&#8217;ll then dig into its KV-cache efficiency, with a direct 131K-context comparison against Gemma 4 31B and Qwen3.6-27B. </p><p>Finally, I&#8217;ll look at the benchmark results, where Glimmer gets surprisingly close to Qwen3.6, and discuss why token efficiency may turn out to be an even more important metric for agentic workloads.</p><p><em>Note: I&#8217;m also preparing a full analysis detailing the model accuracy, token efficiency, memory consumption, and inference speed for the original and quantized versions. It should take about a week to gather enough data.</em></p>
      <p>
          <a href="https://kaitchup.substack.com/p/muse-glimmer-metas-30b-model-built">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Qwen3.8 Is Almost Here — and Agent Benchmarks Are More Fragile Than They Look]]></title><description><![CDATA[The Weekly Kaitchup #154]]></description><link>https://kaitchup.substack.com/p/qwen38-is-almost-here-and-agent-benchmarks</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-is-almost-here-and-agent-benchmarks</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 08 Aug 2026 00:27:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>We have <a href="https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B">a countdown for the Qwen3.8 releases</a>:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2A10!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2A10!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 424w, https://substackcdn.com/image/fetch/$s_!2A10!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 848w, https://substackcdn.com/image/fetch/$s_!2A10!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1272w, https://substackcdn.com/image/fetch/$s_!2A10!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2A10!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png" width="1299" height="642" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:642,&quot;width&quot;:1299,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:328112,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/210098494?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2A10!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 424w, https://substackcdn.com/image/fetch/$s_!2A10!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 848w, https://substackcdn.com/image/fetch/$s_!2A10!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1272w, https://substackcdn.com/image/fetch/$s_!2A10!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa73a783d-2fb6-41fe-9c88-891fa06b5e9b_1299x642.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Next week is going to be a busy one.</p><p>Here&#8217;s my plan:</p><ul><li><p>Start my own evaluation of Qwen3.8 27B as soon as the model is released. This time, I&#8217;ll also cover agentic coding. I plan to publish the full analysis the following week.</p></li><li><p>Run and report on evaluations of the GGUF versions by the end of the week.</p></li><li><p>Publish a deeper analysis of the quantized models within 14 days of release.</p></li></ul><p>This will be similar to what I have done for Qwen3.6.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;f45f8a39-0604-449f-b6b2-b8a7d5525311&quot;,&quot;caption&quot;:&quot;In a previous article, I found Gemma 4 31B to be superior or comparable to Qwen3.5 27B in most areas, with similar or better accuracy and lower latency.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.6 27B vs Qwen3.5 27B vs Gemma 4 31B: Accuracy, Latency, Memory, and Token Efficiency Tested&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-05T11:39:02.874Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!vAQL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2862fcd9-12bf-4ca0-ba6a-23d6808c8806_1210x783.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen36-27b-vs-qwen35-27b-vs-gemma&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195830510,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:24,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>I&#8217;ll also release my own quantized versions using my new quantization pipeline that brings together AutoRound (for high-quality quantization), LLM Compressor (for packing with a vLLM-friendly format), and a custom repacker (to correct bugs with some layers). The goal is to finally make mixed-precision models, with different layers quantized anywhere from 2-bit to 8-bit, work well with vLLM.</p><p>Conceptually, it brings some of the flexibility that GGUF provides for llama.cpp to the vLLM ecosystem, while targeting much higher throughput under heavy concurrency, for example, when running multiple sub-agents in parallel, by leveraging the <a href="https://github.com/inclusionAI/humming">Humming kernel</a>.</p><p>I&#8217;ve already published Qwen3.6 27B variants built with this pipeline. </p><ul><li><p><a href="https://huggingface.co/collections/kaitchup/qwen36-wna16">Qwen3.6 WNA16</a></p></li></ul><p>The 3.8 and 4.3 bit versions work well. I&#8217;m still evaluating the 3.5 and 3.0 bit versions. I&#8217;ll update the model cards once I&#8217;m sure they all work well. I&#8217;ll also add versions with a quantized language modeling head later this week. I thought about quantizing the token embeddings, but this isn't well supported by vLLM for the Qwen3.5 architecture (it doesn&#8217;t crash, but only generates &#8220;!!!!!&#8221;; that&#8217;s a bug, not a quantization quality issue). </p><p>Now I&#8217;m ready to do the same for Qwen3.8.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The Harness Can Matter Almost as Much as the Model</h2><p>A very interesting test from Composio showed how much the agent harness can change the results, even when the model stays the same.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/composio/status/2085330847951970801?s=20&quot;,&quot;full_text&quot;:&quot;We ran DeepSeek V4 Flash through 4 agent harnesses (Claude Code, Codex, OpenCode, Oh My Pi) on 30 agentic tasks.\n\nA different harness won on each metric: success rate, cost, and speed. &#129525;&#129525;&#129525; &quot;,&quot;username&quot;:&quot;composio&quot;,&quot;name&quot;:&quot;Composio&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2036542629220089856/zc7Eix-q_normal.jpg&quot;,&quot;date&quot;:&quot;2026-08-06T11:43:12.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://pbs.substack.com/media/HPCSzoPXEAAf-bw.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/WFy6QNgv8r&quot;}],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:81,&quot;retweet_count&quot;:51,&quot;like_count&quot;:740,&quot;impression_count&quot;:75062,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Yes, it sounds obvious when stated like this, but I increasingly come across model evaluations claiming superiority over other models while relying on numbers produced with suboptimal agent harnesses for the evaluated models, or with harnesses configured in completely different ways (more turns, higher timeouts, &#8230;).</p><p>At best, these comparisons are useless. More often, they are actively misleading.</p><p>Composio ran DeepSeek V4 Flash through four different harnesses, Claude Code, Codex, OpenCode, and Oh My Pi, across 30 multi-step agentic tasks. The tasks involved live tools such as Gmail, Google Sheets, GitHub, Slack, Notion, Calendar, Airtable, and PagerDuty. A task was considered successful only if it passed every predefined check.</p><p><strong>Observations</strong></p><ul><li><p>Oh My Pi achieved the highest success rate, completing 17 out of 30 tasks. Claude Code and Codex each completed 16, while OpenCode completed 14.</p></li><li><p>OpenCode had the lowest estimated cost per successful task at $0.073, followed by Codex at $0.081. Oh My Pi came in at $0.103, while Claude Code was the most expensive at $0.195.</p></li><li><p>Speed produced yet another ranking. Claude Code had the lowest median completion time at 122.7 seconds, with OpenCode close behind at 129.7 seconds. Codex took 245 seconds, while Oh My Pi was the slowest at 272.4 seconds.</p></li></ul><p>So, <strong>it seems</strong> we have different trade-offs:</p><ul><li><p>Claude Code was the fastest, but also the most expensive.</p></li><li><p>Oh My Pi completed the most tasks, but was the slowest.</p></li><li><p>OpenCode was the cheapest, but also had the lowest success rate.</p></li></ul><p>But we also have another angle not tackled here: <strong>inference-time sampling</strong>. Rerun the same experiments, and you may get a very different picture. I agree with Composio&#8217;s general conclusion, but we need many more runs to conclude which is the cheapest, fastest, and most accurate harness.</p><p>This is also why evaluating models for agentic coding is extremely difficult, and extremely expensive to do properly.</p><p>I&#8217;m currently preparing an article focused on Laguna S2.1 and how I benchmarked it for agentic coding. I ran into far more issues than I would have liked to admit, but the process made one thing very clear to me: a large number of the agentic coding benchmark tables I see online should be treated with extreme caution.</p><p><strong>Without controlling for the harness, its configuration, tools, prompts, retry behavior, and execution environment, comparing model scores can quickly become meaningless.</strong></p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[ThinkingCap-Qwen3.6-27B Review: 2x Fewer Tokens, Same Accuracy?]]></title><description><![CDATA[A faster, more stable Qwen3.6 for local AI inference]]></description><link>https://kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x</link><guid isPermaLink="false">https://kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 05 Aug 2026 01:11:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MbCr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa76d07ac-f9de-490f-8a81-076f8b0bd919_912x522.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Qwen3.6 27B is one of the strongest models for local AI, but its extremely long reasoning traces can make inference painfully slow.</p><p>The community has tried to address this issue, but with limited success. One approach is to introduce a &#8220;thinking cap&#8221; that stops the reasoning process once the model reaches a predefined token budget. Another is to fine-tune the model on shorter, more efficient reasoning traces.</p><p>I examined both approaches in a previous article about Qwen3.5 and Qwen3.6. These approaches either performed poorly or required substantially more expensive fine-tuning.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;dd9a98dc-45de-40b9-b924-9befbadabef1&quot;,&quot;caption&quot;:&quot;LLMs now rely heavily on reasoning traces to improve accuracy. A reasoning trace is the intermediate text generated before the final answer, often delimited by tags such as <think>...</think>.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Reasoning Budgets vs. Structured CoT: Controlling LLM Thinking Tokens&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-25T16:10:24.390Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5HDY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4419269c-1249-4fc1-8766-2f460c5bccb8_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/reasoning-budgets-vs-structured-cot&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195230612,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:10,&quot;comment_count&quot;:6,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;f1e4fd13-172d-4608-b39f-d84286335242&quot;,&quot;caption&quot;:&quot;As we saw in previous articles, Qwen3.6 are very good LLMs for local AI.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwopus and REAP: Custom Qwen3.6 Models for Local Reasoning&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-17T20:22:48.702Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!DVuQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76668f-e137-416c-9826-d6d134dd6a60_1210x753.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201494160,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Hopefully, Qwen3.8 27B, which is expected to be released this week, will be a more efficient reasoner.</p><p>In the meantime, BottleCap AI has released a promising custom alternative that significantly reduces the reasoning cost of Qwen3.6: <strong><a href="https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B">ThinkingCap-Qwen3.6-27B</a></strong>.</p><p>ThinkingCap-Qwen3.6-27B is a post-trained version of Qwen3.6-27B. Its main improvement is behavioral: it has been trained to produce <strong>shorter reasoning traces</strong> before delivering an answer.</p><p>The model also tends to generate <strong>shorter final responses</strong>. According to BottleCap AI, this behavior emerged during training alongside the reduction in reasoning length.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I evaluate the model&#8217;s accuracy and token efficiency across several tasks. I compare it with the original Qwen3.6 and another model known for its efficient reasoning, Gemma 4 31B IT.</p><p>All three models are evaluated using exactly the same evaluation settings, with their recommended hyperparameters, making the results directly comparable. Every number reported here comes from my own testing.</p><p>I also examine the model&#8217;s reasoning stability. In particular, is ThinkingCap-Qwen3.6-27B less prone to excessively long or seemingly endless reasoning than the original Qwen3.6?</p>
      <p>
          <a href="https://kaitchup.substack.com/p/thinkingcap-qwen36-27b-review-2x">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[DeepSeek-V4-Flash-0731 and Inkling Small: Smaller, but Better?]]></title><description><![CDATA[The Weekly Kaitchup #153]]></description><link>https://kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling</link><guid isPermaLink="false">https://kaitchup.substack.com/p/deepseek-v4-flash-0731-and-inkling</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 31 Jul 2026 19:33:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>:</p><ul><li><p>DeepSeek V4 Flash: Now Better than DeepSeek V4 Pro (Preview)</p></li><li><p>Inkling-Small: Better than Inkling (?)</p></li><li><p>Escha-W2: Qwen3.6 35B Compressed to 12.3 GB and Faster</p></li></ul><div><hr></div><p>DeepSeek released an update for its V4 Flash model. As usual, DeepSeek released the weights for this update immediately, rather than imposing the waiting period we have seen from companies such as MiniMax and Qwen.</p><ul><li><p><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a></p></li></ul><p>It now outperforms the much larger V4 Pro (Preview), with particularly strong gains in agentic coding:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eyxR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eyxR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 424w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 848w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1272w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eyxR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png" width="561" height="383.1219512195122" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1066,&quot;resizeWidth&quot;:561,&quot;bytes&quot;:82164,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/209166534?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eyxR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 424w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 848w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1272w, https://substackcdn.com/image/fetch/$s_!eyxR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88a0bb5b-24f3-4e83-83ce-5b9dcfd04b23_1066x728.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>They did not publish results for other task categories. I checked <a href="https://artificialanalysis.ai/models?models=mimo-v2-5-pro%2Cgemini-3-5-flash-lite%2Cinkling%2Cclaude-sonnet-5%2Cminimax-m3%2Ccommand-a-plus%2Cgpt-5-6-luna%2Cnvidia-nemotron-3-ultra-550b-a55b%2Cqwen3-7-max%2Cgemini-3-6-flash%2Cgrok-4-5%2Cclaude-4-5-haiku-reasoning%2Cclaude-opus-5%2Cgpt-5-6-terra%2Cdeepseek-v4-pro%2Cgemma-4-31b%2Cclaude-fable-5%2Cmuse-spark-1-1%2Cgpt-5-6-sol%2Cmistral-medium-3-5%2Cgpt-5-5-pro%2Cgpt-oss-120b%2Cglm-5-2%2Ckimi-k3%2Cdeepseek-v4-flash%2Cdeepseek-v4-flash-0420">Artificial Analysis&#8217; additional results</a> and found that it also improved on other benchmarks, including GPQA Diamond, which is a very different task from agentic coding. In fact, none of the reported benchmarks appear to show a regression. This is not merely a more specialized model, it is a substantial improvement over the preview and a very promising sign for the next V4 Pro update.</p><p>On a related note, Qwen also released Qwen3.7 Flash this week. It is currently available only through APIs and is very inexpensive. The community has speculated that it may be a Qwen3.7 35B-A3B model and that Qwen will eventually release the weights. I am not certain this is correct, but it would make sense.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Inkling-Small: Smaller, but Better on Most Benchmarks</h2><p>Despite the name, <a href="https://huggingface.co/thinkingmachines/Inkling-Small">Inkling-Small</a> is not a lightweight model in the usual sense. It is still an open-weight, decoder-only multimodal model that accepts:</p><ul><li><p>Text</p></li><li><p>Images</p></li><li><p>Audio</p></li></ul><p>It follows the same general architecture as the larger Inkling. Both models use a sparse mixture-of-experts design, routing each token through six of 256 experts, along with two shared experts. Both also support context windows of up to one million tokens.</p><p>So Inkling-Small is mostly the same design, scaled down.</p><h3>How much smaller is it?</h3><p>The larger Inkling has:</p><ul><li><p>66 transformer layers</p></li><li><p>975 billion total parameters</p></li><li><p>About 41 billion active parameters per token</p></li></ul><p>Inkling-Small reduces that to:</p><ul><li><p>42 transformer layers</p></li><li><p>276 billion total parameters</p></li><li><p>About 12 billion active parameters per token</p></li></ul><p>That makes Inkling-Small a little over one-quarter the size of the larger model by total parameter count. Its active parameter count is reduced by a similar amount.</p><h3>The NVFP4 version is much easier to fit</h3><p>Inkling-Small requires around <strong>600 GB of aggregate VRAM in BF16</strong>.</p><p>The <a href="https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4">NVFP4 checkpoint</a> brings that down to roughly <strong>180 GB</strong>.</p><p>For comparison, the larger Inkling&#8217;s NVFP4 checkpoint is around <strong>592&#8211;600 GB</strong>. In other words, the quantized version of the large model takes roughly as much memory as Inkling-Small does in BF16.</p><p>The word &#8220;small&#8221; is doing some work here, but 180 GB is at least in the range of a compact multi-GPU server rather than a large GPU cluster.</p><h3>The smaller model is weirdly better at a lot of things</h3><p>The surprising part is that Inkling-Small beats the larger Inkling across most published head-to-head benchmarks.</p><p>A few examples:</p><ul><li><p><strong>SWE-bench Verified:</strong> 80.2% vs. 77.6%</p></li><li><p><strong>Humanity&#8217;s Last Exam:</strong> 31.6% vs. 29.7%</p></li><li><p><strong>IFBench:</strong> 82.2% vs. 79.8%</p></li></ul><p>The smaller model appears particularly strong in coding, tool use, instruction following, and several reasoning evaluations.</p><p>Thinking Machines attributes the improvement to a revised data mix, an updated training recipe, distillation from the larger Inkling model, and additional reinforcement learning focused on agentic coding. </p><p>The larger model still performs better in some areas, especially factual recall, broad knowledge coverage, and most audio evaluations.</p><p>Still, it is an unusual result: the smaller model uses far fewer parameters, needs much less memory, and yet comes out ahead on a majority of the reported benchmarks.</p><p>Another unexpected result is that it significantly outperforms Qwen3.5 397B despite being much smaller, especially on benchmarks where Qwen3.5 has always been exceptionally good, like GPQA Diamond.</p><p>Is it benchmaxxed?? I&#8217;m waiting for the community feedback!</p><div><hr></div><h2>Escha-W2: Qwen3.6 35B Compressed to 12.3 GB</h2><p>If you primarily use GGUF models locally, a 12.3 GB version of Qwen3.6 35B may not sound especially impressive. After all, that is roughly the size of a heavily quantized GGUF model, like a Q2_K_XL, and we already know those can perform well, as shown by the results I published a few months ago.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6pUp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6pUp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 424w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 848w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1272w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6pUp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png" width="1456" height="730" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b161b14e-e18d-425e-8437-318efff937b5_1680x842.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:730,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6pUp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 424w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 848w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1272w, https://substackcdn.com/image/fetch/$s_!6pUp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb161b14e-e18d-425e-8437-318efff937b5_1680x842.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The key difference with Escha-W2 is that it uses (presumably) a more advanced quantization method and comes with an optimized inference runtime. As a result, it should run faster than the typical GGUF model, while already supporting SGLang for significantly better throughput under concurrent workloads. The developers have also said that vLLM support is coming later.</p><h3>How They Did It?</h3><p>We don&#8217;t really know yet.</p><p>The original model is a mixture-of-experts model with:</p><ul><li><p>35 billion total parameters</p></li><li><p>Around 3 billion active parameters per token</p></li><li><p>256 experts</p></li></ul><p>Escha Labs keeps that underlying architecture but compresses most of the expert weights to roughly two bits. More precisely, it uses a mixture of two- and three-bit quantization for the expert projections, while keeping the dense layers in INT8.</p><h3>The whole model is 12.3 GB</h3><p>The resulting checkpoint takes up <strong>12.3 GB on disk</strong>.</p><p>It can run on:</p><ul><li><p>A single 24 GB GPU under the recommended configuration</p></li><li><p>A 16 GB GPU with reduced context length or fewer concurrent requests</p></li></ul><p>On an RTX 4090, Escha reports around <strong>225 tokens per second</strong> for a single stream. The same setup reaches about 1,321 tokens per second when serving 32 requests at once.</p><h3>Two bits, but roughly FP8-level benchmark scores</h3><p>The more interesting part is how little the compression appears to affect most of the reported evaluations.</p><p>Across six benchmark categories, Escha-W2 averages <strong>100.2% of the FP8 model&#8217;s score</strong>. That does not mean the quantized model is genuinely better, the small gains are mostly within normal evaluation variance, but it does suggest that the overall quality loss is limited.</p><p>A few examples:</p><ul><li><p><strong>MMLU-Pro:</strong> 80.9 versus 82.3 for FP8</p></li><li><p><strong><s>MATH-500:</s></strong><s> 93.8 versus 91.2</s> (ignore this one; this is too old and too easy for Qwen3.6)</p></li><li><p><strong>GPQA-Diamond:</strong> 77.8 versus 74.7</p></li><li><p><strong>BFCL tool use:</strong> 88.9 versus 88.2</p></li><li><p><strong>RULER long-context retrieval:</strong> 89.9 versus 89.4</p></li></ul><h3>Coding is the main place where it loses ground</h3><p>The clearest regression is on longer coding tasks, where quantization always does more damage.</p><p>On LiveCodeBench v6, <strong>Escha-W2 scores 62.6, compared with 67.0 for the FP8 baseline</strong>. <em>Note: I don&#8217;t know which framework they used to get a 67.0, but these scores seem very low. It should be around 85.0. This suggests that they ran it with a limited context length, like 32K max tokens. Not great. This means that the accuracy gap could be greater at longer context length, as quantization tends to be worse as sequence length increases.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oa-Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oa-Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 424w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 848w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1272w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oa-Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png" width="776" height="525" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:525,&quot;width&quot;:776,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:46780,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/209166534?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oa-Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 424w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 848w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1272w, https://substackcdn.com/image/fetch/$s_!oa-Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F389f6842-a7ed-402e-96ef-a32e8c33e021_776x525.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>It is not a normal drop-in quantization</h3><p>There is one practical catch: Escha-W2 needs a custom runtime.</p><p>Its weights use a custom packed format, so the model does not simply load into standard inference software as a regular GPTQ, AWQ, or GGUF checkpoint. Escha currently provides two options:</p><ul><li><p>An SGLang-based runtime for concurrency, tool calling and structured output</p></li><li><p>A standalone ZML runtime aimed at single-user inference</p></li></ul><p>The problem with specialized runtimes is that it&#8217;s very hard work to maintain them. </p><p>The checkpoint is also text-only, even though the original Qwen architecture includes vision components.</p><p>When they release the vLLM runtime, I&#8217;ll double-check the accuracy on longer context and compare it with quantized versions of similar size.</p><p></p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Bonsai 27B Review: Can a 3.9 GB 1-Bit Model Match Qwen3.6 27B?]]></title><description><![CDATA[An in-depth look at Bonsai 27B&#8217;s accuracy, token efficiency, reasoning stability, and production trade-offs.]]></description><link>https://kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit</link><guid isPermaLink="false">https://kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 28 Jul 2026 16:06:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!dlgR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dlgR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dlgR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dlgR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png" width="607" height="607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1254,&quot;width&quot;:1254,&quot;resizeWidth&quot;:607,&quot;bytes&quot;:1283747,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207459899?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dlgR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 424w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 848w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1272w, https://substackcdn.com/image/fetch/$s_!dlgR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffa3862e8-bc3b-4069-8094-b8fa3c787109_1254x1254.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Prism ML has released impressive ternary and 1-bit variants of Qwen3.6 27B. The 1-bit version is only 3.9 GB, making it dramatically smaller than the original, which is roughly 55 GB. </p><ul><li><p><a href="https://huggingface.co/collections/prism-ml/bonsai-27b">Bonsai 27B Models</a> (Hugging Face)</p></li></ul><p>These Bonsai 27B models behave quite differently from the original Qwen3.6 model, particularly in terms of accuracy and token efficiency.</p><p><em>Can you get Qwen3.6 27B-level accuracy from a 3.9 GB model? </em></p><p>With reasoning enabled, Bonsai 27B can nearly match the accuracy of Qwen3.6 27B running without reasoning. Moreover, that dramatic reduction in memory comes at a substantial computational cost: Bonsai needs to generate far more tokens to achieve the same result, up to 14 times more on some coding tasks.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>When reasoning is enabled for both models, Qwen3.6 27B remains clearly ahead. It uses fewer tokens while achieving significantly higher accuracy. Still, what Prism ML has achieved is remarkable.</p><p>In this article, I take a deep dive into Bonsai 27B&#8217;s recipe, benchmark performance (my own numbers), and token efficiency. We will examine where the model struggles most, why its raw accuracy is lower, and why the results are nevertheless highly promising.</p>
      <p>
          <a href="https://kaitchup.substack.com/p/bonsai-27b-review-can-a-39-gb-1-bit">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 ]]></title><description><![CDATA[The Weekly Kaitchup #152]]></description><link>https://kaitchup.substack.com/p/agentic-ai-at-two-different-scales</link><guid isPermaLink="false">https://kaitchup.substack.com/p/agentic-ai-at-two-different-scales</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 25 Jul 2026 04:12:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, I&#8217;m looking at two of the week&#8217;s most interesting releases for agentic workloads: Nanbeige4.2-3B and Laguna S 2.1.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Agentic workloads include multi-step reasoning, tool use, interaction with external environments, and tasks that may require many actions before completion.</p><ul><li><p><a href="https://huggingface.co/Nanbeige/Nanbeige4.2-3B">Nanbeige/Nanbeige4.2-3B</a></p></li><li><p><a href="https://huggingface.co/poolside/Laguna-S-2.1">poolside/Laguna-S-2.1</a></p></li></ul><p>Despite this shared focus, they operate at very different scales.</p><p>Nanbeige4.2-3B is a compact dense model with approximately 4 billion total parameters and 3 billion non-embedding parameters. It is intended to make capable agentic behavior practical on consumer and workstation hardware.</p><p>Laguna S 2.1 is a 118-billion-parameter Mixture-of-Experts (MoE) model. It activates approximately 8 billion parameters for each token.</p><p>The two models represent different approaches to agentic AI. Nanbeige uses repeated computation over a relatively small set of weights, while Laguna uses sparse access to a much larger pool of learned parameters.</p><h2>Nanbeige4.2-3B</h2><p>While Nanbeige4.1 used a basic Llama architecture, Nanbeige4.2-3B uses a Looped Transformer.</p><p>Instead of containing a large number of unique transformer layers, the model has 22 physical decoder layers that are <strong>executed twice</strong>. The same weights are reused during the second pass, giving the model the computational depth of approximately 44 layer executions without storing 44 independent layers.</p><p>This design reduces weight memory, but it does not reduce inference computation to that of a normal 22-layer model. Each token must still pass through the transformer stack twice.</p><p>The model has 48 attention heads with a dimension of 128, and eight key/value heads. This is a lot. For instance, Qwen3.6 35B A3B has only two KV heads. </p><p>Looping introduces a further memory consideration. Unless the implementation explicitly shares KV-cache state between the two passes, each loop may require separate key/value entries. Nanbeige&#8217;s default configuration does not appear to enable loop-level KV sharing.</p><p>Moreover, I think models using the Looped Transformer should be better named. When you use a 3B model, you don&#8217;t expect it to be nearly as slow as a 6B model. small MoE models, like Qwen3.6 and Gemma 4, show the number of active parameters in their name, like 26B-A4B. For Looped Transformers, we should adopt a similar convention, like 3B-2P (for &#8220;2 passes&#8221;),  for instance.</p><h4>Nanbeige4.2-3B&#8217;s memory consumption</h4><p>A 16GB GPU should be able to run the model, without quantization, at moderate context lengths with careful settings. A 24GB GPU provides more room for longer prompts, greater concurrency, and runtime overhead.</p><p>Based on the model&#8217;s 22 physical layers, two loop passes, eight key/value heads, 128-dimensional heads, and BF16 KV values, a conservative estimate places KV-cache consumption at approximately 176 KiB per cached token for each active sequence.</p><p><em>I show in the following article how to estimate the KV cache memory consumption:</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6e951bec-dcc3-4e7e-a5ff-d1b040652017&quot;,&quot;caption&quot;:&quot;Inference efficiency has become one of the main ways LLMs distinguish themselves. Two models can look similar on paper: roughly the same parameter count, similar context length, both marketed as fast, yet behave very differently once you actually try to serve them at scale. In this article, we focus on four recent open models built for efficient deployment using different architectures:&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The KV-Cache of Small MoEs: Qwen3, Qwen3.5/3.6, GLM 4.7 Flash, and Nemotron 3 Nano Compared&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-03-18T20:37:03.129Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!BR8P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00c04d8e-bfbd-4069-b1a3-2be9d9e2e8ea_1751x790.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/the-kv-cache-of-small-moes-qwen3&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:191136984,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:27,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>At 16K tokens, this would require roughly 2.75 GiB of KV-cache memory. Using the complete 256K context could require roughly 44 GiB of KV-cache memory for a single active sequence.</p><p>This is very large. Almost twice as much as what Qwen3.6 27B would consume for the same context length.</p><h3>Using Nanbeige4.2-3B with vLLM</h3><p>At release, Nanbeige provided a dedicated vLLM branch for the model:</p><pre><code><code>git clone -b nanbeige42 https://github.com/Nanbeige/vllm.git
cd vllm
pip install -e .</code></code></pre><p>The model can then be served through vLLM&#8217;s OpenAI-compatible API:</p><pre><code><code>vllm serve Nanbeige/Nanbeige4.2-3B \
  --host 0.0.0.0 \
  --port 8000 \
  --enable-auto-tool-choice \
  --tool-call-parser nanbeige \
  --reasoning-parser nanbeige

 #use --host 127.0.0.1 if you are running it locally</code></code></pre><p>The reasoning and tool-call parsers are important for agent frameworks. They allow vLLM to separate reasoning content and structured tool calls from ordinary assistant text.</p><p>Nanbeige&#8217;s chat template exposes two controls.</p><p>The <code>enable_thinking</code> option determines whether the current response contains an explicit reasoning phase. The <code>preserve_thinking</code> option determines whether reasoning content from previous assistant turns remains in the conversation history.</p><p>For normal chat and question answering, preserved reasoning can usually be disabled. For multi-turn tool use, office workflows, and coding agents, the developers recommend retaining earlier reasoning content but I don&#8217;t understand how this can work well. Reasoning traces can be very long, like 50K+ tokens. Keeping them in the context, even if it&#8217;s just one, means that the model may reach its max context length at next turn.</p><h3>Performance</h3><p>According to the published evaluation, Nanbeige exceeded Qwen3.5-9B on most of the included agent, coding, and reasoning tasks. It also surpassed Gemma 4 12B on benchmarks including GDPval, SWE-Bench Verified, SWE-Bench Pro, and Terminal-Bench 2.0.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!z6QI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!z6QI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 424w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 848w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1272w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!z6QI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png" width="1456" height="993" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:993,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!z6QI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 424w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 848w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1272w, https://substackcdn.com/image/fetch/$s_!z6QI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe8e3b3ee-342e-484e-b3b7-ff2a4e570171_4920x3356.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Agent results are influenced not only by the base model but also by the allowed reasoning budget, tool definitions, prompting strategy, conversation-state handling, retry policy, and maximum number of actions.</p><p><em>Reference: <a href="https://huggingface.co/Nanbeige/Nanbeige4.2-3B/blob/main/Nanbeige42_report.pdf">Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model</a></em></p><h2>Laguna S 2.1</h2><p>Laguna S 2.1 is a much larger coding-focused MoE model.</p><p>It contains 118 billion total parameters but activates approximately 8 billion parameters for each token. A routing mechanism selects a subset of experts during inference, allowing the model to access a large learned parameter pool without executing every parameter for every token.</p><p>Laguna contains 48 transformer layers. Twelve use global attention, while the remaining 36 use sliding-window attention with a window of 512 tokens. The global and local layers are interleaved in approximately a one-to-three ratio.</p><p>The global layers allow information to move across the complete input context. The sliding-window layers restrict attention to nearby tokens, reducing the cost of long-context inference and limiting KV-cache growth.</p><p>The model contains 256 routed experts and one shared expert. For each token, the router selects the top ten routed experts in addition to the shared computation.</p><p>Laguna also uses per-head softplus output gating.</p><blockquote><p><strong>Per-head softplus output gating</strong> gives each attention head its own learned volume control. For every token, the model can reduce, preserve, or amplify the contribution of individual heads before combining their outputs. The softplus function keeps these gates positive while allowing values above one, so useful attention heads can be strengthened rather than merely switched on or off.</p></blockquote><p>It uses interleaved reasoning. The model can reason, issue a tool call, receive the tool result, and resume its reasoning before selecting another action.</p><p>Poolside also provides a <a href="https://huggingface.co/poolside/Laguna-S-2.1-DFlash">DFlash draft model</a> for speculative decoding. These models generate candidate tokens that the main model can verify in parallel, potentially increasing output speed.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;d6c5a31a-2b8c-4985-a343-58e4d7652974&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.333Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h3>Using Laguna S 2.1 with vLLM</h3><p>Laguna support is documented for vLLM 0.25.0 or newer:</p><pre><code><code>uv pip install -U "vllm&gt;=0.25.0"</code></code></pre><p>The model can run on a single B300 GPU:</p><pre><code><code>vllm serve poolside/Laguna-S-2.1 \
  --enable-auto-tool-choice \
  --tool-call-parser poolside_v1 \
  --reasoning-parser poolside_v1 \
  --default-chat-template-kwargs '{"enable_thinking": true}'</code></code></pre><p>The <a href="https://huggingface.co/poolside/Laguna-S-2.1-NVFP4">NVFP4</a> and <a href="https://huggingface.co/poolside/Laguna-S-2.1-INT4">INT4</a> versions consume fewer than 80 GB but you would need a 96 GB GPU, like an RTX Pro 6000, to exploit the context length of the model.</p><pre><code><code>vllm serve poolside/Laguna-S-2.1-INT4 \
  --enable-auto-tool-choice \
  --tool-call-parser poolside_v1 \
  --reasoning-parser poolside_v1 \
  --default-chat-template-kwargs '{"enable_thinking": true}'</code></code></pre><p>For coding-agent workloads, Poolside also recommends retaining prior <code>reasoning_content</code> in the conversation history.</p><p>As for the memory consumption of the KV cache, 256K tokens should consume around 24 GB.</p><h3>Target tasks</h3><p>Laguna S 2.1 is more specialized than Nanbeige.</p><p>Its primary target is long-horizon software engineering. This includes repository-level bug fixing, terminal interaction, shell-based workflows, multilingual code maintenance, codebase exploration, and answering questions that require understanding large software repositories.</p><h3>Performance</h3><p>Laguna&#8217;s sparse architecture and coding-focused training are particularly effective for terminal interaction, repository-level reasoning, and multilingual software engineering.</p><p>It looks like a very good model but so far, I have found community feedback to be mixed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8N_3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8N_3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 424w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 848w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1272w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8N_3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg" width="1456" height="2295" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2295,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;benchmarks&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="benchmarks" title="benchmarks" srcset="https://substackcdn.com/image/fetch/$s_!8N_3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 424w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 848w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1272w, https://substackcdn.com/image/fetch/$s_!8N_3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5d6b9f39-833b-40ba-9ce8-a34b84e635c5_699x1102.svg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>reference: <a href="https://poolside.ai/blog/introducing-laguna-s-2-1">Introducing Laguna S 2.1</a></em></p><p>I&#8217;ll spend some time with Nanbeige and Laguna. If everything goes well, I&#8217;ll probably write an article using them for agentic coding.</p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning your account with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p><p><em>Note: I share this coupon because I really think it&#8217;s a good deal. I don&#8217;t receive any form of compensation from Verda, or any usage information, related to this coupon. Verda also provides compute sponsorship for some of my articles.</em></p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p>]]></content:encoded></item><item><title><![CDATA[Qwen3.8: What Hardware Will You Need to Run Alibaba’s 2.4T Model?]]></title><description><![CDATA[Estimating the memory, storage, and GPU requirements for BF16, NVFP4, Q4, and TQ1 versions of Qwen3.8.]]></description><link>https://kaitchup.substack.com/p/qwen38-what-hardware-will-you-need</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen38-what-hardware-will-you-need</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 22 Jul 2026 16:29:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8fje!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8fje!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8fje!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!8fje!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!8fje!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!8fje!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8fje!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png" width="507" height="253.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:507,&quot;bytes&quot;:804225,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207990470?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8fje!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!8fje!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!8fje!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!8fje!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F072e0383-28d2-41b9-ab71-16c66f1d86fc_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Alibaba&#8217;s Qwen3.8 is a 2.4-trillion-parameter frontier model. The company has indicated that it plans to release the model&#8217;s weights, potentially making Qwen3.8 one of the largest openly downloadable AI models ever produced.</p><p>For comparison, Alibaba&#8217;s largest open-weight model to date, <a href="https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct">Qwen3-Coder-480B-A35B-Instruct</a>, is roughly five times smaller. Qwen3.8 would still have about 400 billion fewer parameters than Kimi K3.</p><p>With both K3 and Qwen3.8 performing nearly at the frontier, <em>was closing the accuracy gap with OpenAI and Anthropic largely a matter of scaling from hundreds of billions of parameters to more than two trillion?</em></p><p>An open-weight release for Qwen3.8 is excellent news. It could give researchers, companies, and the broader AI community unprecedented access to a model at this scale.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>But what could most people realistically do with a 2.4-trillion-parameter model?</p><p>Very little locally. Even with aggressive quantization, compression, or distributed inference, a model of this size would be far too large to run on a typical personal computer. Its release could still be highly significant for research institutions, infrastructure providers, and well-resourced community projects. For casual users, however, direct local deployment would remain largely impractical.</p><p>Let&#8217;s check, with a few assumptions, what you would need to run NVFP4, Q4, and Q1 GGUF versions.</p><h2>Qwen3.8: A Sparse MoE?</h2><p><em>As I write this, Alibaba has not disclosed when it plans to release the model&#8217;s weights. That could change quickly. They may even be released by the time this article is published.</em></p><p>Alibaba has not yet disclosed how many parameters are activated for each token, how many experts the model contains, or whether every layer uses mixture-of-experts routing.</p><p>Nevertheless, a dense 2.4-trillion-parameter transformer would be extraordinarily expensive to serve. It is thus reasonable to assume that Qwen3.8 uses a highly sparse mixture-of-experts, or MoE, architecture.</p><p>Until Alibaba publishes the architecture, Moonshot AI&#8217;s Kimi K2 provides a useful proxy for estimating how a trillion-parameter MoE model might distribute its weights.</p><p>Kimi K2 contains:</p><ul><li><p><strong>Total parameters:</strong> 1.04 trillion</p></li><li><p><strong>Activated parameters:</strong> 32.6 billion</p></li><li><p><strong>Routed experts:</strong> 384</p></li><li><p><strong>Experts selected per token:</strong> 8</p></li><li><p><strong>Routing sparsity:</strong> 8/384, or 1/48</p></li><li><p><strong>Shared experts:</strong> 1</p></li></ul><p><em>Note: These figures come from the <a href="https://arxiv.org/abs/2507.20534">Kimi K2 technical report</a>. </em></p><p><strong>98.9% of Kimi K2&#8217;s parameters are located inside routed experts</strong>, while only about 1.1% are always-active or otherwise non-routed weights. The trend is the same for all the MoE with 500B+ parameters: Nearly 99% of the parameters are in the routed experts.</p><h3>Applying the same ratio to Qwen3.8</h3><p>In other words, Qwen3.8 could store 2.4 trillion parameters while using only a small fraction of them for each token. If its routing sparsity resembled Kimi K2&#8217;s, the computation required per token might be comparable to that of a roughly 75-billion-parameter dense model, before accounting for routing and distributed-communication overhead.</p><p>BF16 stores each parameter using 16 bits, or two bytes.</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.4\\text{T parameters}\\times2\\text{ bytes}\n=4.8\\text{ TB}&quot;,&quot;id&quot;:&quot;NRJXFDHYWW&quot;}" data-component-name="LatexBlockToDOM"></div><p>This does not include the KV cache, inference-engine workspace, CUDA graphs, temporary activations, multimodal components, tokenizer files or checkpoint metadata.</p><p>A practical deployment would consequently need more than 4.8 TB of combined accelerator memory. I won&#8217;t speculate on the KV cache size. There are too many variables that influence it: number of linear layers, number of KV heads, etc.</p><h2>Experts in NVFP4, remainder in BF16</h2><p>I hope Alibaba releases an official NVFP4 version. Producing a high-quality NVFP4 quantization of a model this large would be prohibitively expensive for most community projects. NVIDIA has also been relatively slow to publish its own NVFP4 conversions of open-weight models; for example, <a href="https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4">its version of Qwen3.6 appeared only recently</a>.</p><p>NVFP4 stores each parameter as a 4-bit value, with a single FP8 scaling factor shared across all 16-value blocks. This results in an effective storage cost of approximately 4.5 bits per parameter, plus a negligible FP32 scaling value for each tensor.</p><p>Fortunately, routed experts tend to be relatively robust to quantization. In MoE models, the expert weights can often be stored at lower precision while the comparatively small set of shared, attention, embedding, and other always-active parameters remains at higher precision. This mixed-precision approach can substantially reduce storage requirements while limiting the effect on model quality.</p><p>So, with 99% of parameters quantized to NVFP4:</p><ul><li><p>2.374 trillion routed-expert parameters at 4.5 bits each</p></li><li><p>25.8 billion remaining parameters at 16 bits each</p></li></ul><p>The expert weights consume:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.374\\text{T}\\times\\frac{4.5}{8}\n\\approx1.335\\text{ TB}&quot;,&quot;id&quot;:&quot;JUTJBGJUHL&quot;}" data-component-name="LatexBlockToDOM"></div><p>The BF16 remainder consumes:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;25.8\\text{B}\\times2\n\\approx51.5\\text{ GB}&quot;,&quot;id&quot;:&quot;QAIKYJOHOD&quot;}" data-component-name="LatexBlockToDOM"></div><p>Total:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;1.335\\text{ TB}+0.0515\\text{ TB}\n\\approx1.387\\text{ TB}&quot;,&quot;id&quot;:&quot;ZKFYSZIKTS&quot;}" data-component-name="LatexBlockToDOM"></div><p>That&#8217;s an average of <strong>4.62 bits per parameter.</strong></p><p>Compared with a 4.8 TB BF16 checkpoint, that saves approximately:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;4.8-1.387=3.413\\text{ TB}&quot;,&quot;id&quot;:&quot;SMXBFEKRUQ&quot;}" data-component-name="LatexBlockToDOM"></div><p>or <strong>71.1% of the original weight memory</strong>.</p><h2>What about a typical Q4 GGUF?</h2><p>A common GGUF format such as Q4_K_M is approximately 4.8 bits per weight in representative models, while Q4_K_S is closer to 4.6 bits per weight. Actual ratios vary with architecture and which tensors remain at higher precision.</p><p>For a 2.4-trillion-parameter checkpoint:</p><ul><li><p><strong>Q4_K_S-like:</strong> 4.58 bits; approximately <strong>1.374 TB</strong></p></li><li><p><strong>Q4_K_M-like:</strong> 4.84 bits; approximately <strong>1.452 TB</strong></p></li><li><p><strong>Higher-overhead Q4:</strong> 5.0 bits; approximately <strong>1.50 TB</strong></p></li></ul><p>With runtime allocations and a useful KV cache, a deployment should budget at least 1.6&#8211;1.8 TB of usable memory, and more for long contexts or concurrent users.</p><h2>What would TQ1 weigh?</h2><p>I&#8217;m mentioning TQ1 since Unsloth has previously released very good TQ1 versions of Qwen3.5. We have no guarantee they can do/will do the same for Qwen3.8.</p><p>TQ1_0 uses a compact ternary representation at approximately 1.69 bits per weight.</p><p>Purely as storage arithmetic:</p><div class="latex-rendered" data-attrs="{&quot;persistentExpression&quot;:&quot;2.4\\text{T}\\times\\frac{1.69}{8}\n=507\\text{ GB}&quot;,&quot;id&quot;:&quot;LKVCVHKHQG&quot;}" data-component-name="LatexBlockToDOM"></div><h2>Could Qwen3.8 run &#8220;locally&#8221;?</h2><p>It will not be a normal desktop model. Even the Q4 version would be roughly twenty times larger than a 70-billion-parameter Q4 model.</p><h3>Full BF16</h3><p>A BF16 deployment would need at least 4.8 TB for weights and roughly 5.5 TB after adding a modest 15% operational allowance.</p><p>That points toward configurations such as:</p><ul><li><p>Approximately <strong>24 B300 GPUs</strong>, assuming 288 GB each.</p></li><li><p>Approximately <strong>32 B200 GPUs</strong>, assuming 180 GB each.</p></li></ul><p>NVIDIA&#8217;s eight-GPU DGX B200 provides 1.44 TB of HBM and can be configured with 2&#8211;4 TB of system RAM. NVIDIA&#8217;s Blackwell Ultra B300 provides 288 GB of HBM per GPU, or approximately 2.3 TB in an eight-GPU system.</p><p>BF16 Qwen3.8 is firmly a multi-server deployment. </p><h3>NVFP4, Q4, and TQ1</h3><p>A 1.4&#8211;1.5 TB quantized model is more approachable, but &#8220;approachable&#8221; still means rack-scale hardware.</p><p>An <strong>eight-GPU B300 server with approximately 2.3 TB of HBM</strong> should have enough room for the quantized weights, runtime allocations, and a reasonable KV cache.</p><p>An eight-GPU B200 system has only 1.44 TB of HBM. It might barely load the 1.387 TB mixed-NVFP4 estimate, but it would leave almost no memory for the inference engine or KV cache. A two-node, 16-B200 configuration would be much more practical.</p><p>A 507 GB TQ1 checkpoint would still need roughly 600 GB or more after runtime overhead.</p><p>It could fit in:</p><ul><li><p>A 768 GB or 1 TB RAM server.</p></li><li><p>Four B200 GPUs with sufficient aggregate memory.</p></li><li><p>Three B300 GPUs in a custom configuration.</p></li><li><p>An eight-GPU node with considerable unused capacity.</p></li></ul><p>A current maximum-memory Mac Studio offers up to 512 GB of unified memory, which would be too tight once runtime overhead and the KV cache are included.</p><p>So, running even the most compressed version of Qwen3.8 won&#8217;t be cheap.</p><p>Hopefully, Alibaba will also release smaller versions of Qwen3.8.</p><p></p>]]></content:encoded></item><item><title><![CDATA[Inkling, Gemma 4 Updates, and 1-Bit Qwen3.6]]></title><description><![CDATA[The Weekly Kaitchup #151]]></description><link>https://kaitchup.substack.com/p/inkling-gemma-4-updates-and-1-bit</link><guid isPermaLink="false">https://kaitchup.substack.com/p/inkling-gemma-4-updates-and-1-bit</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 18 Jul 2026 00:29:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, we will discuss:</p><ul><li><p>Inkling and the Return of Global Competition in Open-Weight AI</p></li><li><p>Gemma 4: Small Updates with Significant Improvements</p></li><li><p>Bonsai Binary and Ternary Qwen3.6: Only 3.9 GB and It Still Works!</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Inkling and the Return of Global Competition in Open-Weight AI</h2><p>GLM, Kimi, MiniMax, and many of the largest recent open-weight language models have come from China. Outside China, NVIDIA has been one of the few companies to release a comparably large model recently, with Nemotron 3 Ultra.</p><p>Against this backdrop, the release of Inkling by Thinking Machines is a welcome development. Strong open-weight models from a wider range of organizations and countries are important for maintaining healthy global competition in frontier AI.</p><ul><li><p><a href="https://huggingface.co/thinkingmachines/Inkling">thinkingmachines/Inkling</a></p></li></ul><p>Inkling is a 975-billion-parameter multimodal model with approximately 41 billion parameters active per token. Its most interesting technical contributions are not its mixture-of-experts architecture or hybrid attention, both of which are now relatively common, but its approach to long-context modeling, optimization, and reinforcement learning.</p><h3>Long-context architecture</h3><p>Inkling uses learned relative positional representations instead of rotary positional embeddings. Thinking Machines reports that this approach performed better in its long-context extrapolation experiments.</p><p>The model combines these representations with an attention pattern consisting of five sliding-window attention layers for every global-attention layer. Most computation remains local, while periodic global-attention layers allow information to move across the full context. This idea of global-local layers is somewhat close to what Gemma 4 does.</p><p>Short convolutions are also applied after the attention key and value projections, as well as to the outputs of the attention and MLP branches. This introduces explicit local sequence mixing at several points within each transformer block, rather than relying exclusively on attention.</p><h3>Mixture-of-experts design</h3><p>Each mixture-of-experts layer contains:</p><ul><li><p>256 routed experts</p></li><li><p>Two shared experts (so, one more than most open-weight models)</p></li><li><p>Six routed experts selected for each token</p></li></ul><p>The outputs of the shared and routed experts are normalized together. This allows the routing mechanism to control the relative contribution of general-purpose shared computation and more specialized expert computation.</p><h3>Parameter-specific optimization</h3><p>Inkling uses Muon for large matrix parameters and Adam for other parameter types. Its weight decay is also adjusted alongside the learning-rate schedule instead of remaining fixed throughout training.</p><p>This represents a notably large-scale example of assigning different optimizers according to parameter structure.</p><h3>Learning to reason within a compute budget</h3><p>Thinking Machines reports running more than 30 million asynchronous rollouts while varying both the requested level of reasoning effort and the cost associated with generating additional tokens.</p><p>The objective is to teach the model to adapt its reasoning strategy to a given compute budget. This differs from simply imposing a maximum output length after training. Instead, the model learns to estimate when additional reasoning is likely to improve its answer and when the expected benefit is not worth the additional token cost.</p><p>Thinking Machines also randomized the available tools and their schemas during training. The goal was to prevent the model from becoming overly dependent on fixed function names, argument structures, or a single tool-calling format.</p><p>This could help the model generalize more reliably across unfamiliar tools and changing software environments.</p><h3>Competitive, but not yet the leader</h3><p>Inkling performs competitively with other leading open-weight models. However, on the reported benchmarks, GLM-5.2 remains slightly stronger across most evaluations despite being smaller.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ui-7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ui-7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 424w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 848w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1272w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ui-7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png" width="956" height="1174" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1174,&quot;width&quot;:956,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:147740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207328806?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ui-7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 424w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 848w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1272w, https://substackcdn.com/image/fetch/$s_!ui-7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7dab735-60c3-42d6-b50d-26eb8deaeb40_956x1174.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>They also released an NVFP4-quantized version that requires just 592 GB of memory, compared with roughly 1.9 TB for the original model.</p><ul><li><p><a href="https://huggingface.co/thinkingmachines/Inkling-NVFP4">thinkingmachines/Inkling-NVFP4</a></p></li></ul><p>The smallest GGUF made by Unsloth requires 270 GB + KV cache. I expect it to work very well despite the heavy compression since at that scale most of the parameters are routed experts, which are usually very robust to quantization.</p><ul><li><p><a href="https://huggingface.co/unsloth/inkling-GGUF">unsloth/inkling-GGUF</a></p></li></ul><div><hr></div><h2>Gemma 4: Small Updates with Significant Improvements</h2><p>The recent Gemma 4 changes are mainly runtime and interface updates rather than a new checkpoint.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;edc2f68d-ac52-4b6d-bcbe-71a79a88e727&quot;,&quot;caption&quot;:&quot;Last week, we compared Gemma 4 31B with Qwen3.5 27B and found that Gemma 4 31B outperformed it on most tasks while also being faster and more token-efficient.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Gemma 4 31B Quantization Comparison: Best FP8, NVFP4, and INT4 Models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-04-20T18:00:18.907Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!rbnP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3c894d4-7925-49aa-991c-e71e4114b4e4_1210x843.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/gemma-4-31b-quantization-comparison&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:194307038,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:14,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>One update is broader FlashAttention 4 support on NVIDIA Hopper GPUs. Google reports higher prompt-processing throughput and lower time to first token. The main benefit is during prefill, when the model processes the complete prompt before generating its first output.</p><p>Gemma 4 also now exposes several visual-token budgets per image. Developers can use smaller allocations for classification, captioning, or video frames and larger allocations for OCR, documents, handwriting, or images containing small text.</p><p>This makes visual resolution a direct inference control. More visual tokens preserve more detail but also increase prompt length, prefill time, cache use, and the amount of context consumed by each image. The setting is useful for workloads in which some images require detailed analysis while others do not.</p><p>The other relevant change concerns chat and tool-use templates. These templates convert structured messages into the special-token format expected by the model.</p><p>Recent fixes address turn boundaries, tool-response handling, structured arguments, multimodal messages, and reasoning continuity across multiple tool calls. These are implementation changes, but they can directly affect agent reliability. An incorrect marker may cause the model to treat a tool result as a user message or to restart a reasoning sequence instead of continuing it.</p><p>Google reports accuracy improvements on benchmarks using tools thanks to these updates:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zHo5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zHo5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zHo5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Image&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Image" title="Image" srcset="https://substackcdn.com/image/fetch/$s_!zHo5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 424w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 848w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!zHo5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe638ade2-bb48-47e6-833b-4f4e2c336a07_3200x1800.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Important:</strong> The Gemma 4 repositories were updated in place. As a result, reproducing earlier evaluation scores requires loading the models with the corresponding revision explicitly specified.</p><div><hr></div><h2>Bonsai Binary and Ternary Qwen3.6: Only 3.9 GB and It Still Works!</h2><p>PrismML&#8217;s Bonsai applies binary and ternary weights to most of the language network of Qwen3.6-27B.</p><ul><li><p><a href="https://huggingface.co/collections/prism-ml/bonsai-27b">prism-ml/bonsai-27b</a></p></li></ul><p>The ternary version uses negative, zero, and positive weight states and occupies about 5.9GB. The binary version uses only negative and positive states and occupies about 3.9GB. Both use one half-precision scale for groups of 128 weights.</p><p>The low-bit format covers embeddings, attention projections, MLP layers, and the final language-model head. Many quantized models keep sensitive layers at higher precision, especially embeddings or output heads. Bonsai does not use those higher-precision exceptions in the language network.</p><p>The vision tower remains in four-bit precision. This suggests that the visual encoder is more sensitive to binary or ternary quantization than the language network. Vision performance still declines, likely because errors accumulate across the visual encoder, multimodal projection, and low-bit decoder.</p><p>Bonsai also requires custom CUDA and MLX kernels. A compressed weight file does not automatically produce fast inference if the runtime first expands the weights into a higher-precision matrix.</p><p>The custom kernels operate directly on the packed representation. Binary weights can be handled mainly as sign changes, while ternary weights add a zero state that can represent an absent connection. Accumulation and activations still require higher precision.</p><p>The reported phone deployment should be interpreted mainly as a weight-memory result. A 3.9GB model can fit within the memory range of a high-end mobile application, but runtime memory must also include the KV cache, activations, workspaces, the vision tower, and operating-system overhead. The full advertised context length is unlikely to be practical on a phone.</p><p>Despite the heavy compression, the model remains very capable:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q4jg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q4jg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 424w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 848w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1272w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q4jg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png" width="1201" height="555" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:555,&quot;width&quot;:1201,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:95787,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/207328806?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!q4jg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 424w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 848w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1272w, https://substackcdn.com/image/fetch/$s_!q4jg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02dec9d0-62a5-4fe5-a46c-d2dc04454fc9_1201x555.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The degradation compared with the original model is significant, but the accuracy is far better than all the other sub-4 GB models I have evaluated.</p><p>I&#8217;ll publish full analysis, accuracy, and token efficiency next week!</p><div><hr></div><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p>In collaboration with <a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a>, I&#8217;m sharing a <strong>$50 coupon</strong> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</p><p><strong>Coupon code:</strong> <code>KAITCHUP-50</code></p><p>Follow <a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a>.</p></blockquote><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[Qwen3.6-27B KV Cache Quantization in vLLM: Accuracy, Memory, and Speed]]></title><description><![CDATA[A smaller KV cache enables longer sequences and higher concurrency with virtually no loss in accuracy.]]></description><link>https://kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Thu, 16 Jul 2026 16:24:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!alUj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc5dc43aa-10ea-4534-9c9e-0ba78a52affe_1210x843.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Long-context inference is often limited less by the model&#8217;s advertised context window than by the memory and bandwidth required for its KV cache.</p><p>During generation, attention layers store keys and values for every previous token. As context length and concurrency increase, this cache can consume tens of gigabytes and become a major memory-bandwidth bottleneck.</p><p>KV-cache quantization reduces that cost by storing keys and values at lower precision. Moving from 16-bit to approximately 4-bit storage can shrink the cache by nearly four times, enabling longer contexts and more concurrent requests.</p><p>In practice, however, a method that works well in a research implementation may perform poorly inside an optimized engine such as vLLM or llama.cpp. Quantization can add computational overhead, require specialized kernels, reduce hardware compatibility, and interfere with features such as speculative decoding.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>In this article, I compare three 4-bit KV-cache formats in vLLM, <code>turboquant_4bit_nc</code>, <code>int4_per_token_head</code>, and <code>nvfp4</code>, using Qwen3.6-27B. I evaluate their memory usage, accuracy, token generation, and inference speed.</p><blockquote><p><strong>Acknowledgments</strong></p><p><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=kvquantization">Verda</a> provided the B200s and RTX Pro 6000s to run the experiments described in this article.</p><p>Verda is the full-stack frontier AI cloud, built for high-performance inference, training, and agentic workloads with data privacy and sustainability at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=kvquantization">here</a><span>. There is a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs. </span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p></blockquote><h3>Qwen3.6-27B KV Cache Size</h3><p>I&#8217;m going to compare KV-cache memory usage across different quantization methods. First, let&#8217;s establish the memory required for Qwen3.6-27B&#8217;s native 262,144-token context window.</p><p>Qwen3.6-27B has 64 language layers, but only 16 are Gated Attention layers that maintain a full KV cache. The remaining layers are Gated DeltaNet blocks, so they are not included when calculating the conventional attention KV cache.</p><p>Each Gated Attention layer has four KV heads, with a head dimension of 256. </p><p>In BF16, the KV cache for these full-attention layers consumes approximately 17.2 GB. This is substantially less than for models that use full attention in every layer, but it is still significant relative to the model weights themselves, which occupy roughly 54 GB.</p><p>On a device with 96 GB of memory, that leaves:</p><p>96 &#8722; 54 &#8722; 17.2 = 24.8 GB</p><p>This is not enough to run several full-context requests in parallel, for example, multiple sub-agents, especially once additional overhead is included. CUDA graphs, temporary buffers, runtime allocations, and other inference-engine components all consume extra memory.</p>
      <p>
          <a href="https://kaitchup.substack.com/p/qwen36-27b-kv-cache-quantization">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Efficient and Reasoning AI at the ACL 2026]]></title><description><![CDATA[The Weekly Kaitchup #150]]></description><link>https://kaitchup.substack.com/p/efficient-and-reasoning-ai-at-the</link><guid isPermaLink="false">https://kaitchup.substack.com/p/efficient-and-reasoning-ai-at-the</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 11 Jul 2026 00:46:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of <em>The Weekly Kaitchup</em>, I&#8217;ll highlight some of the most interesting work I saw at ACL 2026, which I attended this week in San Diego.</p><p>I&#8217;ll also discuss OpenAI&#8217;s recent take on benchmarking with SWE-Bench Pro, one of the most widely used coding benchmark. Once again, a SWE-Bench is being declared &#8220;bad.&#8221; I can&#8217;t say I&#8217;m too upset about having never spent compute on it.</p><blockquote><h3><strong>A $50 Coupon to Try Verda&#8217;s GPUs</strong></h3><p><span>In collaboration with </span><a href="https://verda.com/?utm_source=kaitchup.substack.com">Verda</a><span>, I&#8217;m sharing a </span><strong>$50 coupon</strong><span> that you can redeem in your Verda account, after provisionning it with $5, to try their GPUs (B200, B300, RTX Pro 6000, &#8230;).</span></p><p><strong>Coupon code:</strong><span> </span><code>KAITCHUP-50</code></p><p><span>Follow </span><a href="https://docs.verda.com/resources/obtaining-free-credits/how-to-redeem-credits/">these instructions to redeem it</a><span>.</span></p></blockquote><h2>ACL 2026: Why I Went, and Why I Skipped ICML 2026</h2><p>This year, two of the major annual research conferences shaping the development of AI, ACL and ICML, took place at the same time, on opposite sides of the Pacific: ACL in San Diego and ICML in Seoul.</p><p>While there is significant, and increasing, overlap between the two communities, they tend to emphasize different areas. ICML is often the place for deeper discussions about machine learning methods, algorithms, and theoretical foundations. ACL, by contrast, is where I usually find the most relevant work on evaluation, multilinguality, benchmarks, datasets, and language-centered applications.</p><p>That distinction is far from absolute. Many of the key ideas behind today&#8217;s Transformers and LLMs, as well as many of the benchmarks and evaluation tasks we still rely on, originated in papers published within the ACL community.</p><p>I chose to attend ACL this year primarily because it aligns more closely with my own research background. It is also the community where I know more people, having published several papers there during my Ph.D.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><p>ACL 2026 was capped at 3,500 onsite attendees. For reference, ACL 2025 in Vienna had more than 5,000 in-person attendees.</p><p>The smaller scale made ACL easier to navigate, but it was still impossible to see everything I had planned. As I did for NeurIPS 2025, I will mainly focus on the papers I actually saw and discussed at the conference rather than trying to summarize the full program.</p><p>I spent most of my time in the poster sessions, which were by far the best place to have useful conversations with authors. The talks were less crowded, and some large rooms were surprisingly almost empty, especially on the last day.</p><p>To keep this edition useful and focused, I will mostly discuss papers on LLM reasoning and efficiency, two of the major themes of ACL 2026 along with reinforcement learning, as it has been at most recent AI and machine learning conferences. </p><p>The program chairs even noted that papers on some of these topics seemed to have higher acceptance rates, which suggests that they were not only popular among authors, but also particularly appealing to reviewers and area chairs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7KzG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7KzG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7KzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg" width="588" height="783.8653846153846" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:588,&quot;bytes&quot;:4228692,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7KzG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!7KzG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd22678e-263b-4c18-809c-39c032093f67_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Note: For each paper discussed below, I also included a photo of the corresponding poster. The images are high-resolution, but I do not fully trust Substack to display them cleanly (depending on your device), so apologies if some details are difficult to read. ACL may also publish the posters directly on each paper&#8217;s page, or may have already done so by the time you read this. You can check by clicking the paper links.</em></p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.1717">Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models</a><br><em>Hoang Phan, Xianjun Yang, Yuanshun Yao, Jingyu Zhang, Shengjie Bi, Xiaocheng Tang, Madian Khabsa, Lijuan Liu, Deren Lei<br>Meta; New York University; Johns Hopkins University</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!prGW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!prGW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 424w, https://substackcdn.com/image/fetch/$s_!prGW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 848w, https://substackcdn.com/image/fetch/$s_!prGW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!prGW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!prGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg" width="662" height="496.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:662,&quot;bytes&quot;:1662114,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!prGW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 424w, https://substackcdn.com/image/fetch/$s_!prGW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 848w, https://substackcdn.com/image/fetch/$s_!prGW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!prGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F079d4b74-394e-4de2-9bb0-b5101932b248_2560x1920.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper looks at a well-known problem in reinforcement learning for reasoning models: training that improves reasoning can also make the model forget broader capabilities. I often see that when running RL for recent models like Qwen3.5/Gemma 4. </p><p>The authors show that RLVR-style post-training may improve mathematical or multimodal reasoning while degrading skills such as perception, OCR, robustness, and general instruction following.</p><p>They propose RECAP, a replay-based method that keeps general-capability data in the training loop and dynamically reweights objectives. The main contribution is a clearer diagnosis of &#8220;reasoning gain&#8221; as a trade-off problem: a model can become better at narrow reasoning benchmarks while becoming less generally useful. RECAP is presented as a way to preserve the original model&#8217;s breadth while still gaining reasoning ability.</p><p>Very simple and intuitive.</p><div><hr></div><p><a href="https://aclanthology.org/2026.acl-long.1530/">SSSD: Simply-Scalable Speculative Decoding</a><br><em>Michele Marzollo, Jiawei Zhuang, Niklas Roemer, Niklas Zwingenberger, Lorenz K Muller, Lukas Cavigelli<br>Huawei; ETH Zurich</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CCGF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CCGF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CCGF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1894802,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CCGF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!CCGF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F79929376-5f2e-453a-ba97-42ec9423279f_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper proposes SSSD, a training-free speculative decoding method for making LLM inference faster.</p><p>Instead of relying on an additional trained draft model, like EAGLE3 or DFlash, SSSD retrieves likely n-gram continuations from the prompt, previously generated tokens, and a datastore, then verifies them with the main model.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;4d72fac1-b9b9-4e02-bc16-03acf1a974e9&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.333Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>It avoids extra model training and is easier to deploy across changing workloads, domains, and languages. </p><p>The paper reports latency reductions of up to 2.9x over standard autoregressive decoding, making it relevant for production serving where simplicity and robustness matter as much as peak speed.</p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.677/">Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction</a><br><em>Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, Shafiq Joty<br>Salesforce AI Research; University of Wisconsin-Madison; The University of Hong Kong</em></p><p>Early layers of an LLM can already identify many of the tokens that will matter for answering a query. The authors turn this into GemFilter, a training-free method that uses early layers as filters to select a much smaller subset of the input before later computation.</p><p>The method is aimed at long-context inference, where latency and GPU memory grow quickly with input length. By compressing the input token set, GemFilter speeds up inference while keeping the important evidence visible to the model.</p><p>A nice aspect of the work is interpretability: because the method explicitly selects input tokens, humans can inspect what the model kept. The paper reports a 2.4x speedup and 30% GPU memory reduction compared with strong baselines, while performing well on Needle-in-a-Haystack and competitively on LongBench. </p><div><hr></div><p><a href="https://doi.org/10.1162/TACL.a.692">A Survey on Memory-Efficient Fine-Tuning for Large Language Models</a><br><em>Yeachan Kim, Mingyu Lee, SangKeun Lee<br>Hankuk University of Foreign Studies; Korea University</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pVhK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pVhK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pVhK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1750655,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pVhK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!pVhK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F351c069f-b342-4883-add8-9018ac7d41e6_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This survey organizes the landscape of memory-efficient fine-tuning for LLMs. I recommend reading it if you are lost in the recent progress made in this area. </p><p>One point I strongly agree with, although it is not very well developed in the paper, is that these methods are often poorly evaluated.</p><p>The paper mentions GLUE-style tasks. I did not double-check how widely GLUE is still used in this line of work, but even replacing it with old and relatively easy benchmarks such as QNLI, PIQA, or CommonsenseQA, like they did in this paper, would not solve the problem. People use these benchmarks because they are cheap to run.</p><p>As I often argue here, if your benchmarks do not require the model to generate millions of tokens, they are unlikely to tell you much about how a generative language model actually performs, unless of course you thouroughly study what the benchmarks has generated instead of just looking at the scores.</p><div><hr></div><h2>KV Compression</h2><p><a href="https://aclanthology.org/2026.acl-long.1683/">LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning</a><br><em>Haoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang, Song Guo<br>The Hong Kong University of Science and Technology; The Hong Kong Polytechnic University</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OPck!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OPck!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OPck!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OPck!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OPck!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OPck!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg" width="656" height="874.5164835164835" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:656,&quot;bytes&quot;:1905771,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!OPck!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!OPck!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!OPck!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!OPck!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F645c343c-1f0e-4c96-b8b0-ef6537c48b5a_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper studies KV-cache compression for long chain-of-thought reasoning. The key finding is &#8220;Token Importance Recurrence&#8221;: tokens that seem unimportant at one decoding step can become important again later. This makes simple greedy KV eviction risky, because it may remove tokens that the model will need in future reasoning steps.</p><p>LazyEviction addresses this by delaying eviction and observing attention patterns over a window before deciding what to discard. The method is designed for long reasoning sequences, where memory cost grows with every generated token, and it aims to reduce memory while avoiding sudden reasoning failures caused by prematurely evicting useful context.</p><p>This is another interesting piece of work on KV cache compression, one of many presented at the conference.</p><p>As usual, though, the real test will be implementation. Unless a method lands in frameworks such as vLLM or llama.cpp, it is hard to know whether it actually improves inference efficiency in practice, or whether it interferes with other optimizations already implemented in these systems.</p><div><hr></div><p><a href="https://aclanthology.org/2026.acl-long.1542/">Anchoring the Cache: Mitigating Contextual Hallucination in KV-Compressed Long-Context Summarization</a><br><em>Yu Fu, Chen Luo, Josef Valvoda, Xin Zhang, Xuejing Lei, Xiao Pan, Hui Liu, Yue Dong<br>University of California, Riverside; Amazon</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Yn1p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Yn1p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Yn1p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1848955,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Yn1p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Yn1p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e5eb67-cc21-4b17-8726-919e97d28ae1_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This paper studies an important side effect of KV-cache compression in long-context summarization: it can make models hallucinate more. </p><p>The authors show that aggressive compression may preserve surface-level summarization metrics while weakening factual grounding, because retrieval heads drift away from the source context during generation.</p><p>Their method, HalluKV, anchors selected retrieval heads to the source by selectively removing generated KV pairs from those heads during decoding. The idea let the model keep the efficiency benefits of cache compression, but prevent the heads responsible for source retrieval from over-attending to the model&#8217;s own generated text.</p><p>The result is a decoding-time intervention that reduces hallucination while maintaining long-context efficiency. The paper reports that KV compression can increase hallucination scores by up to 3.36&#215;, and that HalluKV reduces hallucination across multiple models and datasets while preserving competitive summary quality. </p><div><hr></div><p><a href="https://aclanthology.org/2026.findings-acl.1314/">Quantize What Counts: More for Keys, Less for Values</a><br><em>Mohsen Hariri, Alan Luo, Weicong Chen, Tianyi Zhang, Qifan Wang, Xiaotian Han, Vipin Chaudhary<br>Case Western Reserve University; Rice University; Meta</em><br></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AyUv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AyUv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AyUv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg" width="1456" height="1941" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1941,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2121396,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AyUv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 424w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 848w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!AyUv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de3cc53-f55d-4a9d-b819-6c2e9beb0bfe_1920x2560.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Keys and values of the KV cache should not necessarily receive the same number of bits: key projections tend to have larger spectral and Frobenius norms than value projections, meaning keys carry higher-magnitude, more quantization-sensitive information. </p><p>The authors use this observation to justify a simple rule: allocate more precision to keys and compress values more aggressively.</p><p>The paper turns this into a geometry-driven design principle for mixed-precision KV-cache quantization. Instead of tuning key/value bit splits heuristically for each model or task, it argues that model weight geometry can predict which cache component needs more protection. Empirically, key-favored allocations such as 4-bit keys and 2-bit values preserve up to 98.3% of the accuracy of uniform 4-bit KV quantization while saving memory, making the result useful for long-context inference where KV cache memory dominates.</p><p>A very informative paper to better understand why keys are more sensitive to quantization.</p><div><hr></div><h2>More Comments on the ACL 2026</h2><p>The social events at the ACL are always great. This year, the organizers reserved the <a href="https://uss-midway-museum.sandiegotourismtickets.com/">USS Midway</a> for an evening dinner. The USS Midway is an aircraft carrier turned museum, and the venue was genuinely impressive.</p><p>These social events are often among the most useful parts of a conference. They make it much easier to meet people outside your usual circle, and this one was no exception.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ko6e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ko6e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ko6e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg" width="1456" height="1456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1456,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:696600,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/206345919?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ko6e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 424w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 848w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!Ko6e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc689162-cd46-40e9-a067-64ff91cdf0eb_2047x2047.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The next ACL will be in Kyoto, so I will probably end up playing local guide: I lived and worked there for six years.</p><p>Hopefully, the conference will not coincide with the worst of Kyoto summer: 35&#176;C or more, 70%+ humidity, and the occasional typhoon. Otherwise, the most important survival tips may have nothing to do with AI.</p><div><hr></div><h2>SWE-Bench Is Not a Good Benchmark, Again</h2><p><em>Unrelated to the ACL, but very interesting findings published by OpenAI this week.</em></p><p>OpenAI has used SWE-Bench Pro to benchmark all its models since GPT 5.2.</p><p>And yet, according to <a href="https://openai.com/index/separating-signal-from-noise-coding-evaluations/">OpenAI&#8217;s recent audit</a>, many SWE-bench Pro tasks fail for reasons unrelated to model capability.</p><p>Examples include:</p><ul><li><p>hidden requirements that never appear in the issue description</p></li><li><p>contradictory specifications</p></li><li><p>unit tests that reject perfectly reasonable implementations</p></li><li><p>grading logic that expects one specific implementation instead of checking whether the issue is actually fixed</p></li><li><p>mismatches between the GitHub issue, the merged pull request, and the evaluation tests</p></li></ul><p>Because SWE-bench Pro is built automatically from real repositories, OpenAI argues that these inconsistencies are common enough to substantially affect reported scores.</p><p>OpenAI&#8217;s main conclusion is:</p><blockquote><p><strong>Approximately 30% of SWE-bench Pro tasks are broken.</strong></p></blockquote><p>Their recommendation is therefore to &#8220;<strong>carefully inspect benchmark results instead of treating leaderboard rankings as ground truth.&#8221; </strong>Yes, of course, but who actually does this? Most AI labs want to report the strongest possible numbers. If you publish a score of 80 and your competitor publishes 90, most people will not ask whether that 10-point gap reflects a real performance difference or a benchmarking artifact. They will simply see the lower score, and it will look bad.</p><p>Ironically, OpenAI itself helped create <strong>SWE-bench Verified</strong> in 2024 because it believed the original SWE-bench contained many impossible or ambiguous tasks. Now, after auditing SWE-bench Pro, it argues that even this newer benchmark still contains enough flawed tasks that raw leaderboard scores should not be over-interpreted.</p><p>My take is that these benchmarks are becoming increasingly complex: automatically generated, frequently refreshed, available in multiple versions, and often difficult to interpret consistently. As a result, I think people will develop more and more trust issues around published scores and may eventually start disregarding them altogether.</p><p>However, this is less of a problem when evaluating different versions of the same model. For example, when comparing a quantized model to the original, the goal is not necessarily to establish the model&#8217;s absolute performance. What matters is whether the quantized version performs as closely as possible to the non-quantized one.</p><p>In that setting, the original score matters less. The main objective is to minimize the accuracy delta between the two versions. Even a flawed benchmark like SWE-Bench can still be useful for this purpose, as long as it is used consistently across both model variants.</p><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[LFM2.5 230M and 350M: How Accurate Are the GGUF Versions?]]></title><description><![CDATA[230M or 350M GGUFs?]]></description><link>https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate</link><guid isPermaLink="false">https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 08 Jul 2026 03:20:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2bDO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tiny language models that can run quickly with very limited memory are one of Liquid AI&#8217;s specialties. And while models with only a few hundred million parameters obviously cannot rival today&#8217;s multi-billion-parameter models, tiny models are now capable enough for many practical tasks. They can make tool calls, retain some basic world knowledge, and follow instructions surprisingly well.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2bDO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2bDO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 424w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 848w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2bDO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png" width="1456" height="978" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:978,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2bDO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 424w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 848w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!2bDO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F640981c7-6dcf-4c9d-a38a-93b0e16c2e93_2800x1880.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://huggingface.co/LiquidAI/LFM2.5-230M</figcaption></figure></div><p>LFM2.5 is a family of small language models. In this article, I will focus specifically on the tiny 230M and 350M versions. I ran both models on a large set of benchmarks with two goals:</p><ol><li><p>Reproduce some of the results published by Liquid AI.</p></li><li><p>Better highlight what these models can and cannot do.</p></li></ol><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>At first glance, the size difference between the 350M and 230M models may seem negligible: only 120M parameters. But that gap feels much larger when you realize that 120M parameters represent roughly 34% of the 350M model&#8217;s total size. On memory-constrained devices, especially when using longer context lengths, this difference can matter a lot. In 16-bit precision, it reduces weight memory consumption from about 0.71 GB for the 350M model to about 0.46 GB for the 230M model.</p><p>And given how accurate is quantization now, it could be possible to make a quantized version of the 350M that is both smaller and better than the 230M.</p><p>This raises another question that I will address in this article:</p><p><em>Should you use a low-bit quantized version of the 350M model, or a higher-precision version of the 230M model?</em></p><p>To answer that, we will compare accuracy versus model size for both models across different GGUF quantization formats released by Liquid AI and Unsloth.</p><h2>LFM2.5 230M and 350M: What They Can and Cannot Do</h2>
      <p>
          <a href="https://kaitchup.substack.com/p/lfm25-230m-and-350m-how-accurate">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[DSpark and NVIDIA's Qwen3.6 NVFP4 Models]]></title><description><![CDATA[The Weekly Kaitchup #149]]></description><link>https://kaitchup.substack.com/p/dspark-and-nvidias-qwen36-nvfp4-models</link><guid isPermaLink="false">https://kaitchup.substack.com/p/dspark-and-nvidias-qwen36-nvfp4-models</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 04 Jul 2026 03:21:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of The Weekly Kaitchup:</p><ul><li><p>Which NVFP4 Version of Qwen3.6 27B Should You Use?</p></li><li><p>DSpark: DeepSeek&#8217;s New Speculative Decoding Method</p></li></ul><p>I&#8217;m at the ACL 2026 and stopped by the Qwen corner to ask the important question:</p><blockquote><p><strong>Should we expect Qwen3.7 27B soon?</strong></p><p>Answer:<br><em>We haven&#8217;t made a final decision yet, but we&#8217;re committed to open-sourcing more models and have some exciting releases coming soon.</em></p></blockquote><p>Not very informative&#8230; It sounds like a well-prepared answer. And they told me they already got this question a lot during just a single morning at the conference.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D0d3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D0d3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D0d3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg" width="550" height="412.5" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:550,&quot;bytes&quot;:4715993,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/204695456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D0d3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 424w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 848w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!D0d3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0f70ed8-526d-4d81-b715-8382317126cb_2560x1920.jpeg 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>Which NVFP4 Version of Qwen3.6 27B Should You Use?</h2><p>NVIDIA released an NVFP4 version of Qwen3.6 27B. It is a good occasion to revisit the other NVFP4 variants that have been released and to put NVIDIA&#8217;s checkpoint into context.</p><p>When I published my evaluations and analysis of quantized Qwen3.6 27B, only a few NVFP4 versions were available. There were so few that I had to make one myself with <a href="https://github.com/intel/auto-round">AutoRound</a>.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;bcf583e6-8211-4164-bd6d-c6ee65102265&quot;,&quot;caption&quot;:&quot;Like I did for Qwen3.5 and Gemma 4, let&#8217;s see how well Qwen3.6 27B holds up when quantized into different formats: FP8, INT4, and NVFP4.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.6 27B Quantization: FP8 vs INT4 vs NVFP4&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-12T06:45:55.916Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5Qvj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b412df9-eef1-4bfe-aa6b-b90105838569_1210x903.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen36-27b-quantization-fp8-vs-int4&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196100250,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p><strong><a href="https://huggingface.co/kaitchup/Qwen3.6-27B-autoround-nvfp4-linearattn-BF16">kaitchup/Qwen3.6-27B-autoround-nvfp4-linearattn-mtp-BF16</a></strong><br>This is my version made with AutoRound. It uses NVFP4 while keeping the linear-attention and MTP layers in 16-bit. The model size is <strong>28.6 GB</strong>.</p><p><strong><a href="https://huggingface.co/Peutlefaire/Qwen3.6-27B-NVFP4">Peutlefaire/Qwen3.6-27B-NVFP4</a></strong><br>This version goes further and quantizes more <code>linear_attn</code> layers too. The model size is <strong>20.6 GB</strong>, with a 19.7 GB main model file and an 849 MB MTP file. Compared with my AutoRound MTP BF16 version, that is roughly 8.0 GB smaller.</p><p><a href="https://kaitchup.substack.com/p/qwen36-27b-quantization-fp8-vs-int4">In my analysis</a>, I showed that quantizing the attention path to NVFP4 significantly degraded the model. It did not break the model, but it made it more prone to endless thinking, and accuracy went down on nearly all benchmarks. Check the full analysis for details.</p><p><strong><a href="https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4">unsloth/Qwen3.6-27B-NVFP4</a></strong><br>Unsloth also released an NVFP4 checkpoint. Its config applies NVFP4 quantization to most Linear layers, while leaving some of the linear-attention projections (linear_attn.out_proj) in BF16. You can see it as a version between mine and a full NVFP4 quantization. The model size is <strong>26.4 GB</strong>.</p><p><strong><a href="https://huggingface.co/nvidia/Qwen3.6-27B-NVFP4">nvidia/Qwen3.6-27B-NVFP4</a></strong><br>The NVIDIA version is interesting because it is mixed precision: attention and linear-attention projections are FP8 W8A8, while the MLP layers and LM head are marked as <strong>W4A16_NVFP4</strong>. The checkpoint is <strong>21.9 GB</strong>, which makes it <strong>6.7 GB smaller</strong> than my 28.6 GB AutoRound MTP BF16 version and <strong>4.5 GB smaller</strong> than Unsloth&#8217;s 26.4 GB version.</p><p>This also means that the NVFP4 parts of NVIDIA&#8217;s model do not use NVFP4 activations. They are W4A16: 4-bit NVFP4 weights, but 16-bit activations. The attention path, meanwhile, is FP8 rather than NVFP4. With Blackwell GPUs, this likely leaves speed on the table compared with W4A4 NVFP4. In my experience, NVFP4 activations can be almost harmless with good calibration.</p><p>Another important detail: NVIDIA&#8217;s <code>config.json</code> includes an FP8 KV-cache scheme. That matters a lot for speed and memory use, especially on devices where memory bandwidth is the bottleneck. I already saw some third-party speed comparisons on X comparing this checkpoint against other NVFP4 models without matching the KV-cache format, which is not an apples-to-apples comparison. At minimum, KV-cache dtype should be controlled explicitly across all models; vLLM, for example, can be run with an explicit FP8 KV cache.</p><p>NVIDIA also reports its own evaluation numbers against an FP8 baseline. In their table, the NVFP4 checkpoint is very close to the FP8 checkpoint across MMLU-Pro, GPQA Diamond, HLE, &#964;&#178;-Bench, MMMU Pro, SciCode, AIME 2025, AA-LCR, and IFBench. This is useful, but it is still not a community-wide comparison against the other NVFP4 checkpoints under identical settings.</p><p>It is also worth mentioning the PrismaQuant family.</p><p><strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaQuant-5.5bit-vllm">rdtand/Qwen3.6-27B-PrismaQuant-5.5bit-vllm</a></strong><br>This version uses a sensitivity-driven per-Linear allocation. PrismaQuant chooses between NVFP4, MXFP8, and BF16 under a 5.5-bit-per-parameter target. According to the model card, bulk dense MLP layers are usually NVFP4 W4/A4, higher-sensitivity dense linears can be MXFP8 W8/A8, and the most sensitive tensors, norms, biases, embeddings, and LM head stay BF16.</p><p>The model size is 22.7 GB, which slighly larger to the one released by NVIDIA, probably due to the LM head staying at higher precision.</p><p>There are newer PrismaQuant-family variants too. <strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm">PrismaAURA 5.5-bit</a></strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaAURA-5.5bit-vllm"> </a>uses an AURA/KL-Fisher allocation over NVFP4, FP8, and BF16, while <strong><a href="https://huggingface.co/rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm">PrismaSCOUT 5.31-bit</a></strong> is described as a Blackwell-oriented successor to the earlier 5.5-bit PrismaQuant artifact. PrismaSCOUT uses NVFP4 plus selected BF16, includes MTP tensors, and its repository is <strong>20.2 GB</strong>.</p><h2>How to Choose?</h2><p>We do not yet have a good comparison. Quantization is cheap; evaluation is expensive.</p><p>What we need is a benchmark suite run on all these checkpoints with exactly the same settings: same inference engine, same chat template, same thinking mode, same KV-cache dtype, same context length, same sampling settings, same MTP/speculative-decoding settings, and same hardware.</p><p>Until then, I would assume the following:</p><ul><li><p>There is no strong reason to keep all activations in 16-bit on Blackwell if NVFP4 activation quantization is well calibrated.</p></li><li><p>Keeping the attention path out of NVFP4 is safer.</p></li><li><p>Unsloth&#8217;s version and my AutoRound version are safer choices if your priority is a more conservative quantization strategy.</p></li><li><p>PrismaSCOUT and the other PrismaQuant-family models are worth testing, but I would not rank them without independent, apples-to-apples accuracy numbers. This could be the best version. We simply don&#8217;t know.</p></li></ul><p>So, for now, I would use either Unsloth&#8217;s version or my AutoRound MTP BF16 version if accuracy safety is the main concern. They are also better documented and easier to reason about than many one-off community checkpoints. If memory footprint matters more, NVIDIA&#8217;s version and PrismaSCOUT are especially interesting.</p><div><hr></div><h2>DSpark: DeepSeek&#8217;s New Speculative Decoding Method</h2><p>DSpark is a speculative decoding method for making LLMs generate faster without changing the final model that verifies the text. </p><p>Instead of asking the full model to produce one token at a time, DSpark, like any other speculative decoding method, lets a smaller draft module propose several future tokens, then asks the full model to verify those tokens in a batch. When the draft is right, the server accepts multiple tokens from one target-model pass. When the draft is wrong, the server keeps the longest valid prefix and continues from there.</p><p>The target model we want to accelerate remains the source of truth, while the drafter tries to guess what the target model is likely to accept next.</p><h2>What DSpark Changes</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VX38!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VX38!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VX38!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VX38!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VX38!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VX38!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg" width="1280" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify." title="DSpark architecture: a parallel draft backbone, a lightweight serial head linking adjacent tokens, a confidence head scoring each token, and a hardware-aware scheduler choosing how much of the block to verify." srcset="https://substackcdn.com/image/fetch/$s_!VX38!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 424w, https://substackcdn.com/image/fetch/$s_!VX38!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 848w, https://substackcdn.com/image/fetch/$s_!VX38!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!VX38!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd524e11b-c63b-4395-92fd-d1d441815e3b_1280x720.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://deepseek.ai/blog/deepseek-dspark-speculative-decoding</figcaption></figure></div><p>The weakness of many speculative decoders is suffix decay. Drafting the first token is easy but drafting the fifth, sixth, or seventh token is much harder because every position depends on earlier guesses. Fully parallel drafters are fast, but later draft positions often become noisy. Autoregressive drafters are more faithful, but they give back too much of the latency advantage because they draft sequentially.</p><p>DSpark tries to sit between those two extremes. It keeps a DFlash-style parallel draft backbone, then adds a lightweight Markov logit-bias head. <em>Note: I explained how DFlash works here:</em></p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;4fc4f756-15cc-4d6a-a2b0-a64c213eaa20&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.334Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>That head lets each draft position condition on previously sampled tokens inside the draft block, which improves the later positions without turning the whole drafter into a slow autoregressive model.</p><p>The other important piece is confidence. DSpark adds a confidence head that predicts the acceptance probability of each draft position. That gives the serving system a practical scheduling signal. When the drafter is confident, the runtime can verify more proposed tokens. When confidence drops, or when the server is under load, it can avoid wasting verification work on tail tokens that are likely to be rejected. In other words, DSpark is not just a better drafter; it is a more serving-aware drafter.</p><h3>How it works in the decoding loop</h3><p>At each generation step, DSpark first prepares a block of draft tokens. The target model then scores that block in one verification pass. The runtime accepts the longest prefix that agrees with the target model&#8217;s sampling path and discards the rest. If several drafted tokens survive, the user sees several tokens produced for roughly one target-model step. If only the first token survives, the system behaves closer to normal decoding for that step.</p><p>The Markov head just helps the drafter keep later draft positions coherent. The confidence head helps the runtime decide how much speculation is worth attempting. Together, they attack the two practical bottlenecks of speculative decoding: low acceptance at later positions and wasted verification work under real serving load.</p><h3>Speed results</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qSuE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qSuE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 424w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 848w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1272w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qSuE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png" width="1297" height="696" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:696,&quot;width&quot;:1297,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:302141,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/204695456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qSuE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 424w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 848w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1272w, https://substackcdn.com/image/fetch/$s_!qSuE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06e5b107-1b63-4702-bc59-ca5f9ffc9a0c_1297x696.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>DSpark improves per-user generation speed by about 60&#8211;85% over the previous MTP-1 baseline on DeepSeek-V4-Flash, and by about 57&#8211;78% on DeepSeek-V4-Pro at matched throughput. Offline accepted-length results also reportedly improve over Eagle3 and DFlash. As usual with speculative decoding, the exact gain depends on the prompt mix, sampling settings, hardware, batch pressure, and how often the drafter&#8217;s later tokens are accepted. Once widely available for other open-weight models, like Qwen3.6 and the larger Gemma 4, it will be interesting to see how fast is it compared with high MTP, like MTP-4 or 5 which are much faster than MTP-1 for these two models.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;35cea535-aba5-400d-96ae-db8ee9bc7d80&quot;,&quot;caption&quot;:&quot;Qwen3.6 inference is faster when MTP layers are enabled to draft tokens. Gemma 4 now supports this option as well. Both model families also have public DFlash speculator checkpoints, which can draft blocks of tokens in a single forward pass.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;DFlash vs MTP: Qwen3.6 Speculative Decoding Benchmarks with vLLM and llama.cpp&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-02T18:08:30.973Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!79OP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a729fb2-7c13-44ac-bad1-08602638a0f9_1530x742.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/dflash-vs-mtp-qwen36-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:198352019,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:8,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>You can find more results in the technical report:</p><p><a href="https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf">DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation</a></p><h3>What DeepSeek released</h3><p>The <a href="https://github.com/deepseek-ai/DeepSpec">DeepSpec</a> repository currently includes three draft-model algorithms: DSpark, DFlash, and Eagle3. </p><p>DeepSeek also released trained draft checkpoints for four target models: <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_4b_block7">Qwen3-4B</a>, <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_8b_block7">Qwen3-8B</a>, <a href="https://huggingface.co/deepseek-ai/dspark_qwen3_14b_block7">Qwen3-14B</a>, and <a href="https://huggingface.co/deepseek-ai/dspark_gemma4_12b_block7">Gemma-4-12B-it</a>. For each target, the repo lists Eagle3, DFlash, and DSpark checkpoints, so the open release covers twelve reference draft checkpoints.</p><p>DSpark checkpoints are also available for <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark">DeepSeek V4 Flash</a> and <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark">Pro</a>.</p><h3>Running DSpark in vLLM</h3><p>DSpark support is landing in vLLM through the speculative decoding path. The relevant vLLM PR for DSpark speculators-format checkpoint support was merged on July 2, 2026, so use a nightly build or a release that includes that change.</p><pre><code><code>uv pip install -U vllm --extra-index-url https://wheels.vllm.ai/nightly

vllm serve deepseek-ai/DeepSeek-V4-Flash-DSpark \
  --tensor-parallel-size 8 \
  --speculative-config '{"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}'</code></code></pre><h3>Training your own DSpark speculator</h3><p>You can train a DSpark speculator on your own data for Qwen3 and GLM 5.2 (only these architectures are support for now) using the <a href="https://github.com/vllm-project/speculators">speculators</a> project. This is very similar to training a DFlash a speculator.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;87ba04d9-9692-4a5f-86bb-a756a43c3148&quot;,&quot;caption&quot;:&quot;Speculative decoding is becoming a popular way to accelerate LLM inference, with approaches such as MTP and DFlash. The idea is simple: a smaller or specialized draft model proposes future tokens, and the full target model verifies them. Matching tokens are accepted and when a mismatch occurs, only the valid prefix is kept and decoding falls back to the target model.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Train and Run DFlash Speculative Decoding&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-18T19:43:12.334Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sbvO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b274a34-f749-4a9f-b68b-0e0f501a9016_1448x1086.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/train-and-run-dflash-speculative&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:196847181,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[MiniMax M3 GGUF Quantization: From 852 GB to ~150 GB Without Breaking Accuracy]]></title><description><![CDATA[Benchmarks, token efficiency, and tensor-level analysis of low-bit M3 GGUFs.]]></description><link>https://kaitchup.substack.com/p/minimax-m3-gguf-quantization-from</link><guid isPermaLink="false">https://kaitchup.substack.com/p/minimax-m3-gguf-quantization-from</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Tue, 30 Jun 2026 20:20:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RopP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RopP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RopP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!RopP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!RopP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!RopP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RopP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png" width="512" height="288" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:512,&quot;bytes&quot;:1449718,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/203645642?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RopP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!RopP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!RopP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!RopP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F726b79c2-1c3a-47a0-96b2-4df2391ff351_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>MiniMax M3 is an excellent model, but with roughly 428B parameters, running it locally is very challenging. In BF16, the weights alone require around 852 GB of memory, so an unquantized setup realistically needs a large multi-GPU server, for example an 8&#215;H200 machine.</p><p>Quantization can dramatically reduce these requirements. However, <a href="https://kaitchup.substack.com/p/lessons-from-gguf-evaluations-ternary">as we saw in previous evaluations, MiniMax M2.5 degraded heavily once quantized, even at 4-bit</a>. The natural question is whether M3 is more robust.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>As we will see in this article, the answer is yes: M3 is much easier to quantize. My hypothesis is that this improved quantization robustness comes from M3&#8217;s shared-expert MoE design, a feature that was absent from M2.5. The shared expert provides a path that can be preserved during quantization at a high precision, while the routed experts, which account for about 97% of M3&#8217;s parameters, can be compressed much more aggressively.</p><p>In this article, I evaluate several low-bit MiniMax M3 GGUFs, including Unsloth&#8217;s UD GGUFs and my own MoQ quantization. </p><ul><li><p><a href="https://huggingface.co/kaitchup/MiniMax-M3-GGUF-MoQ">kaitchup/MiniMax-M3-GGUF-MoQ</a></p></li><li><p><a href="https://huggingface.co/unsloth/MiniMax-M3-GGUF">unsloth/MiniMax-M3-GGUF</a></p></li></ul><p>The main result is that M3 can be compressed from 852 GB to around 150 GB while preserving most of its accuracy. I then analyze the tensor-level quantization choices behind each GGUF to explain why some variants remain strong, and why the smallest ones start to break.</p><blockquote><p> <strong>Acknowledgments</strong></p><p><span>This article would not have been possible without the compute sponsorship generously provided by </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=m3quantization">Verda</a><span>, whose B300 GPUs I used throughout this work.</span></p><p>Verda is a European, AI-focused cloud and GPU infrastructure provider with sovereignty, sustainability, data privacy, and performance at its core.</p><p><span>You can check them out </span><a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=m3quantization">here</a><span>.</span></p></blockquote><h2>Results: Which GGUF Can You Safely Use?</h2><p>First, a note on the benchmarks I ran. For my GGUF evaluations, I usually use large subsets of MMLU Pro for world knowledge, Math 500 for math questions ranging from easy to difficult, LiveCodeBench for challenging coding tasks, and the full GPQA Diamond benchmark for difficult science questions.</p><p>However, running M3 with llama.cpp is very costly, especially when evaluating low-bit versions that can generate many more tokens than the original model to answer the same questions. Running only 100 LiveCodeBench samples was already too expensive for this evaluation, at least $200 for a single GGUF evaluaiton, so I did not include this benchmark.</p><blockquote><p><strong>How should you read the following results?</strong></p></blockquote>
      <p>
          <a href="https://kaitchup.substack.com/p/minimax-m3-gguf-quantization-from">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[This Week in Open Models: Tiny LFM2.5, Ornith-1.0, and GLM-5.2 REAP]]></title><description><![CDATA[The Weekly Kaitchup #148]]></description><link>https://kaitchup.substack.com/p/this-week-in-open-models-tiny-lfm25</link><guid isPermaLink="false">https://kaitchup.substack.com/p/this-week-in-open-models-tiny-lfm25</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Sat, 27 Jun 2026 05:34:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of The Weekly Kaitchup, let&#8217;s discuss some of the open models released this week:</p><ul><li><p>LFM2.5 230M: Tiny Models Can Follow Instructions</p></li><li><p>Ornith-1.0: Self-Scaffolding Coding Models</p></li><li><p>GLM-5.2 REAP: Cutting a 753B MoE Down to a More Deployable 504B</p></li></ul><p>I&#8217;m also currently evaluating the M3 GGUFs I created with MoQ. This has taken more time than expected, as some of them generate more tokens than anticipated. The evaluation is almost finished, though, and you can expect the full analysis on Monday or Tuesday.</p><p>The models are available here:</p><ul><li><p><a href="https://huggingface.co/kaitchup/MiniMax-M3-GGUF-MoQ">kaitchup/MiniMax-M3-GGUF-MoQ</a></p></li><li><p>I already confirmed they are all good, except for the MoQ-2.5 (not broken but much less accurate than the others).</p></li></ul><blockquote><p>Next week, I&#8217;ll be at the <a href="https://2026.aclweb.org/">ACL 2026</a>, one of the main AI conferences, in San Diego. Let me know if you&#8217;ll be there and would like to meet!</p><p><a href="https://kaitchup.substack.com/p/efficient-llms-at-scale-my-neurips">As I did for NeurIPS 2025</a>, I&#8217;ll publish a report on the most interesting papers, presentations, and demos I come across.</p></blockquote><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h2>LFM2.5 230M: Tiny Models Can Follow Instructions</h2><p>Was LFM2.5-350M too large? Liquid AI just released <a href="https://huggingface.co/LiquidAI/LFM2.5-230M">LFM2.5-230M</a>, and for a 230M-parameter model, it performs surprisingly well.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sNAv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sNAv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 424w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 848w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sNAv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png" width="1456" height="978" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:978,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;lfm2_5_230m_benchmarks&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="lfm2_5_230m_benchmarks" title="lfm2_5_230m_benchmarks" srcset="https://substackcdn.com/image/fetch/$s_!sNAv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 424w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 848w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!sNAv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b474c72-4a66-4089-8b02-cca263743117_2800x1880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two numbers that stand out are IFBench and IFEval. LFM2.5-230M is in the same ballpark as models far larger, especially regarding the IFeval score, which looks very close to what we would get with Llama 3.1 8B. That is roughly a 35&#215; parameter gap.</p><blockquote><p><em>Note on GPQA Diamond</em></p><p><em>GPQA Diamond is probably too difficult for models this small. It is a 4-choice multiple-choice benchmark, so a model that simply picks A, B, C, or D at random should get 25% accuracy. When results hover around that baseline, they mostly tell us that the model is not really solving the task. Scores below 25% can also happen in generative evaluations where the model is asked to output exactly one of A, B, C, or D. If it generates anything else, the answer is marked wrong. So at this scale, the GPQA number can look meaningful in a table while mostly reflecting answer-format compliance plus noise rather than scientific reasoning.</em></p></blockquote><p>As for its architecture, LFM2.5-230M is not just a uniformly scaled-down <a href="https://huggingface.co/LiquidAI/LFM2.5-350M">350M</a>. Both models keep the same 1024 hidden size, 16 attention heads, 8 KV heads, and 65,536-token vocabulary.</p><p>The savings come from the blocks. LFM2.5-230M uses 14 layers instead of 16: 8 double-gated LIV convolution blocks + 6 GQA blocks, versus 10 + 6 for LFM2.5-350M. The bigger difference is the FFN/intermediate dimension: 2560 for the 230M model, versus 6656 for the 350M model. That is where most of the parameter reduction comes from.</p><p>In memory terms, the 230M model only consumes 459 MB and is 250 MB lighter than the 350M in bf16.</p><p><em>So now we have a genuinely interesting question: should we use the bf16 230M, or a quantized version of the 350M?</em></p><p>I think that a carefully quantized 350M could win on both quality and memory. But at this scale, quantization is less forgiving: 4-bit quantization can easily erase the quality advantage.  I&#8217;ll try to find time to run the bf16 and quantized variants of these tiny LFM2.5 models and plot accuracy vs memory curves.</p><div><hr></div><h3>Ornith-1.0: Self-Scaffolding Coding Models</h3><p>Ornith-1.0 is a family of open-source language models built specifically for agentic coding: settings where a model must not only write code, but operate inside a tool-using loop, inspect repositories, run commands, interpret failures, and iteratively repair its own solution. </p><ul><li><p>Models: <a href="https://huggingface.co/collections/deepreinforce-ai/ornith-10">The Ornith-1.0</a></p></li><li><p>Blog Post: <a href="https://deep-reinforce.com/ornith_1_0.html">Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding</a></p></li></ul><p>The family spans several deployment and capability tiers, including compact dense models and larger mixture-of-experts models. The announced lineup includes 9B Dense, 31B Dense (not released yet), 35B MoE, and 397B MoE variants, with released Hugging Face artifacts including FP8, and GGUF builds.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6To2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6To2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 424w, https://substackcdn.com/image/fetch/$s_!6To2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 848w, https://substackcdn.com/image/fetch/$s_!6To2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 1272w, https://substackcdn.com/image/fetch/$s_!6To2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6To2!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png" width="1034" height="506.3475274725275" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:713,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1034,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Ornith-1.0-9B evaluation results&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="Ornith-1.0-9B evaluation results" title="Ornith-1.0-9B evaluation results" srcset="https://substackcdn.com/image/fetch/$s_!6To2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 424w, https://substackcdn.com/image/fetch/$s_!6To2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 848w, https://substackcdn.com/image/fetch/$s_!6To2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 1272w, https://substackcdn.com/image/fetch/$s_!6To2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a9e0f43-cb75-47b6-ba17-96231dcf4e8c_3000x1470.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B</figcaption></figure></div><p>The models are post-trained on top of Gemma 4 and Qwen 3.5, then specialized for coding-agent behavior. </p><p>The central training idea is &#8220;self-scaffolding.&#8221; </p><p>In conventional reinforcement-learning setups for coding agents, the scaffold, i.e., the harness, orchestration logic, memory strategy, error-handling routine, or prompting structure that guides the rollout, is usually designed by humans and held fixed. Ornith-1.0 instead treats that scaffold as something the model can improve. During RL, the model first proposes or refines a task-specific scaffold, then uses that scaffold to generate a solution rollout. The reward from the rollout is assigned back to both the scaffold-generation step and the solution-generation step, so the model is trained not only to produce better code, but also to discover better ways of organizing its own problem-solving process.</p><p>This creates a feedback loop: scaffolds that lead to higher-reward coding trajectories are reinforced, while weaker scaffolds are discarded. Over many training iterations, the model learns task-category-specific strategies for agentic coding without depending entirely on hand-engineered harnesses. DeepReinforce frames this as the key distinction of Ornith-1.0: reinforcement learning is used not just to improve answers, but to improve the model&#8217;s internal orchestration of tool use and search.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KNzm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KNzm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 424w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 848w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 1272w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KNzm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png" width="1456" height="698" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:698,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Ornith self-improving training framework&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Ornith self-improving training framework" title="Ornith self-improving training framework" srcset="https://substackcdn.com/image/fetch/$s_!KNzm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 424w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 848w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 1272w, https://substackcdn.com/image/fetch/$s_!KNzm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e87ec32-2265-40f8-9086-080ba353d2d3_2009x963.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://deep-reinforce.com/ornith_1_0.html</figcaption></figure></div><p>The training setup also addresses an obvious risk of self-improving agents: reward hacking. If a model can shape the scaffold that drives its own rollout, it may learn shortcuts that satisfy a verifier without solving the task, such as reading hidden test artifacts, modifying verification scripts, or hard-coding expected outputs. They have implemented three guardrails: keeping the environment and tool boundary outside the model&#8217;s control, using deterministic monitors to block forbidden actions, and adding a frozen LLM judge as a veto layer on top of the verifier for cases where intent-level gaming is harder to specify mechanically.</p><p>The models look very strong on the benchmarks, and early community feedback appears to be positive.</p><blockquote><p>One particularly interesting detail is that the only checkpoint based on Gemma 4 is missing from both the release and the evaluation results. Qwen3.5 is known to be strong at agentic coding, whereas this is an area where Gemma 4 underperforms. That makes the idea of an Ornith-1.0 model based on Gemma 4 especially exciting as it could correct one of its main weaknesses.</p><p>At the same time, its absence leaves room for interpretation. Was Ornith-1.0 31B performing significantly below the 35B model based on Qwen3.5? Was it much harder to train or align? Or is the release simply delayed?</p></blockquote><div><hr></div><h2>GLM-5.2 REAP: Cutting a 753B MoE Down to a More Deployable 504B</h2><p>0xSero&#8217;s GLM-5.2 REAP project is a practical attempt to make Z.ai&#8217;s massive GLM-5.2 mixture-of-experts model easier to serve.</p><ul><li><p><a href="https://huggingface.co/0xSero/GLM-5.2-504B">0xSero/GLM-5.2-504B</a></p></li></ul><p>Instead of quantizing the whole model more aggressively or merging experts together, REAP removes low-saliency routed experts from each MoE layer. The released 504B version keeps 168 of the original 256 routed experts per layer. In other words, it prunes 88 experts per layer, or about 34.4% of the routed expert pool.</p><p>The released checkpoint above is an NVFP4 version.</p><p>After pruning, Router-KD, or router-only knowledge distillation, was applied to recover behavior. The experts, attention layers, embeddings, and most of the network are frozen. Only the router gate matrices are trained, representing around 0.016% of the model&#8217;s parameters. The goal is to re-teach the router how to route tokens through the surviving 168 experts so that the pruned model better matches the unpruned teacher&#8217;s next-token distribution. This makes the recovery step far cheaper than full fine-tuning.</p><p>While it&#8217;s very effective at reducing the model size, as we saw last week, REAP has significant drawbacks: </p><ul><li><p>it significantly degrades the model&#8217;s world knowledge, or more generally, its accuracy on tasks not used to evaluate the expert activations</p></li><li><p>and the resulting model is often generating more tokens which makes it less cost-effective at inference.</p></li></ul><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;1fa90c3a-b980-47c5-80d4-cfbc34e4db7d&quot;,&quot;caption&quot;:&quot;As we saw in previous articles, Qwen3.6 are very good LLMs for local AI.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwopus and REAP: Custom Qwen3.6 Models for Local Reasoning&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-06-17T20:22:48.702Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!DVuQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76668f-e137-416c-9826-d6d134dd6a60_1210x753.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201494160,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:10,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>The published evaluation results are also limited and, in some cases, do not include the original model&#8217;s results as a baseline. As a result, it is hard to tell how close this model comes to the original. What we do know is that it scores around 70 on Terminal-Bench 2.1, which is very encouraging.</p><p>Running more careful evaluations for large models remains prohibitively expensive for the freelance AI community; for example for GLM 5.2 and some pruned/quantized variants, it would cost $20k or more to run <a href="https://kaitchup.substack.com/p/qwen36-27b-vs-qwen35-27b-vs-gemma">the same type of evaluation and analysis I ran for Qwen3.6</a>.</p><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[GLM-5.2: Only a Few Months Behind Commercial Models]]></title><description><![CDATA[The Weekly Kaitchup #147]]></description><link>https://kaitchup.substack.com/p/glm-52-only-a-few-months-behind-commercial</link><guid isPermaLink="false">https://kaitchup.substack.com/p/glm-52-only-a-few-months-behind-commercial</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 19 Jun 2026 23:16:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of The Weekly Kaitchup, let&#8217;s discuss:</p><ul><li><p>GLM-5.2: Only a Few Months Behind Commercial Models</p></li><li><p>VibeThinker-3B: How Far Can Verifiable Training Push an Old Small Model?</p></li><li><p>MoQ GGUF Updates</p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>GLM-5.2: Only a Few Months Behind Commercial Models</h2><p>Many people are calling GLM-5.2 the best model for local AI. I would not go that far.</p><p>My definition of <em>local AI</em> is stricter: it should be something a serious hobbyist, researcher, or small company can run on a local machine with consumer GPUs, CPU RAM, and a realistic budget. By that standard, GLM-5.2 is not really local.</p><p><a href="https://huggingface.co/unsloth/GLM-5.2-GGUF">Unsloth&#8217;s Q1 GGUFs</a> are about 217&#8211;228 GB before the KV cache. The Q2 versions are around 238&#8211;254 GB. That means you are already outside the comfort zone of normal workstations before you even start thinking about context length.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KP0b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KP0b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 424w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 848w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 1272w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KP0b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png" width="529" height="393.43888070692196" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:505,&quot;width&quot;:679,&quot;resizeWidth&quot;:529,&quot;bytes&quot;:143896,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/202631080?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KP0b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 424w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 848w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 1272w, https://substackcdn.com/image/fetch/$s_!KP0b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2623d7b8-1737-4193-a1b1-7ebaed391ed9_679x505.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">https://huggingface.co/unsloth/GLM-5.2-GGUF</figcaption></figure></div><p>Then, we have to add the KV cache&#8217;s memory consumption. With GLM-5.2&#8217;s MLA-style cache:</p><p><code>78 layers &#215; tokens &#215; (512 + 64) &#215; bytes_per_value</code></p><p>So, with BF16 KV cache, that is roughly:</p><ul><li><p>9 GB per 100k tokens</p></li><li><p>18 GB per 200k tokens</p></li><li><p>45 GB per 500k tokens</p></li><li><p>90 GB for the full 1M-token context</p></li></ul><p>With FP8/INT8 KV cache, divide those numbers by about two. Real deployments may need more because of runtime overhead, batching, padding, sparse-attention grouping, and framework-specific memory management.</p><p>So yes, you may be able to load GLM-5.2 on a machine with huge system RAM, or on a 256 GB unified-memory Mac. But it won&#8217;t be the original model, and it&#8217;ll be very slow unless most of the hot path is on serious GPU memory and bandwidth. </p><p>Nonetheless, GLM 5.2 is exciting because it gives individuals, labs, and companies access to a frontier-class model without depending entirely on a closed API.</p><p>And the quality is surprisingly close.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bIN-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bIN-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 424w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 848w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 1272w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bIN-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png" width="1456" height="961" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:961,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;bench_52&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="bench_52" title="bench_52" srcset="https://substackcdn.com/image/fetch/$s_!bIN-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 424w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 848w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 1272w, https://substackcdn.com/image/fetch/$s_!bIN-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441b573d-7213-49ef-8177-37d6cba4640f_4239x2799.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For coding and agentic work, GLM-5.2 is near the commercial frontier on several benchmarks. On SWE-Bench Pro, it scores above GPT-5.5 and Gemini 3.1 Pro, though still behind Claude Opus 4.8. On Terminal-Bench 2.1, it is in the same range as GPT-5.5 and Claude Opus 4.8. On MCP-Atlas, it is only about one point behind Opus 4.8 and slightly ahead of GPT-5.5.</p><p>That is what I mean by &#8220;only a few months behind commercial models.&#8221; Not that GLM-5.2 beats the best closed models everywhere. It does not. Claude Opus 4.8 still has a large lead on some hard software-engineering benchmarks, and the closed models are usually easier to serve, faster to use, and more polished as products.</p><p>But the gap is now small enough to be operationally interesting. A model with open weights, MIT licensing, 1M-token context, strong coding scores, and private deployment is no longer a toy alternative to commercial AI.</p><div><hr></div><h2>VibeThinker-3B: How Far Can Verifiable Training Push a Small Reasoning Model?</h2><p>VibeThinker-3B is a 3-billion-parameter dense reasoning model released by WeiboAI. It is designed for tasks where the answer can be checked with a relatively clear signal: mathematics, competitive programming, STEM reasoning, and instruction-following tasks with explicit constraints.</p><ul><li><p><a href="https://huggingface.co/WeiboAI/VibeThinker-3B">WeiboAI/VibeThinker-3B</a></p></li></ul><p>The model builds on the earlier VibeThinker-1.5B release, a smaller model designed to test whether highly compact language models could acquire useful reasoning abilities through targeted post-training. At the time of its release, VibeThinker-1.5B drew sustained attention in the AI community for several weeks, largely because its benchmark scores were unexpectedly strong given its small size.</p><p>VibeThinker-3B continues that line of work at a larger scale, but it is still very small compared with the large reasoning models usually associated with strong performance on competition math and code.</p><p>The main point of VibeThinker-3B is not that small models can replace larger general-purpose models. It is more specific than that. The model is an example of how far a small model can be pushed when the training tasks have reliable feedback.</p><p>The model is strongest where correctness is measurable and the model card explicitly warns that it was not trained for tool-calling or agent-style programming workflows. For code, it is better understood as a competitive-programming model than as a general software-engineering assistant.</p><p>They published a technical report here:</p><p><a href="https://arxiv.org/abs/2606.16140">VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models</a></p><p>VibeThinker-3B focuses on multi-step reasoning, answer verification, constraint satisfaction, and self-correction.</p><p>The authors describe this through a distinction between reasoning and knowledge. Their argument is that some reasoning procedures may be highly compressible into a relatively small model, provided the task space is structured and the feedback is reliable. In contrast, broad world knowledge and long-tail factual coverage still benefit from much larger parameter counts.</p><h2>Architecture</h2><p>VibeThinker-3B is based on Qwen2.5-Coder-3B, which is almost 2 years old!</p><p>This is a Qwen2-style decoder-only Transformer.</p><p>The Qwen2.5-Coder series is trained for code generation, code reasoning, and code repair, while retaining some general and mathematical ability. Starting from a code-oriented base gives VibeThinker-3B a foundation that is already useful for executable reasoning and competitive programming.</p><h2>How it was trained</h2><p>The training pipeline is built around what WeiboAI calls the Spectrum-to-Signal Principle. The idea is to separate two phases that are often mixed together.</p><p>The &#8220;spectrum&#8221; part means exposing the model to a wide range of possible solution paths. Instead of training only on one canonical solution, the model is trained on multiple valid reasoning traces. This is meant to help it explore different ways to decompose a problem.</p><p>The &#8220;signal&#8221; part means using reliable feedback to reinforce the useful paths. In math, this can mean checking the final answer. In code, it can mean running the generated program against tests. In constrained instruction-following, it can mean checking whether the output satisfies explicit requirements.</p><p>The first major stage is supervised fine-tuning. VibeThinker-3B uses a two-stage curriculum. The first stage covers a broad mixture of math, code, STEM reasoning, general dialogue, and instruction following. This gives the model a broad behavioral starting point. The second stage shifts toward harder and longer reasoning samples, so the model spends more training effort on long-horizon problem solving.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DhVF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DhVF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 424w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 848w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 1272w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DhVF!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png" width="1200" height="347.8021978021978" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:422,&quot;width&quot;:1456,&quot;resizeWidth&quot;:1200,&quot;bytes&quot;:146814,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/202631080?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DhVF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 424w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 848w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 1272w, https://substackcdn.com/image/fetch/$s_!DhVF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6cbdaab-780b-4400-8237-b2c09f8c1534_1723x499.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The report describes query expansion, teacher-generated reasoning traces, answer verification, code execution, majority voting, and filtering of invalid traces. For reasoning-heavy samples, the model is trained with multi-path reasoning distillation. That means strong teacher models generate multiple candidate reasoning paths, and the training process keeps useful verified traces rather than collapsing everything into a single fixed style.</p><p>After supervised fine-tuning, the model goes through multi-domain reinforcement learning. This stage uses MaxEnt-Guided Policy Optimization, the reinforcement-learning method introduced in the earlier VibeThinker-1.5B work. In simple terms, the method focuses training on prompts that are near the model&#8217;s current capability boundary. These are problems where some sampled solutions are correct and others are wrong. Such prompts provide a more useful learning signal than problems the model always solves or always fails.</p><p>The reinforcement-learning stage is applied sequentially across math, code, and STEM reasoning. Math RL is used to strengthen symbolic derivation and long multi-step reasoning. Code RL focuses on executable logic, boundary cases, and program constraints. STEM RL is then used to broaden the reasoning behavior into scientific tasks.</p><p>A notable change from the earlier model is the long-context reinforcement-learning setup. The authors report that aggressive early truncation weakened long-thinking behavior, so VibeThinker-3B instead uses a single 64K long-context window during RL. The goal is to avoid cutting off complete reasoning trajectories.</p><p>The training also includes a Long2Short math RL stage, meant to reduce redundant reasoning among already-correct outputs. In practice, the model is still encouraged to solve the problem, but shorter correct trajectories are preferred over longer correct trajectories when both are valid.</p><p>After the main reinforcement-learning stages, the model goes through offline self-distillation. High-quality traces from math, code, and STEM checkpoints are verified, filtered, and distilled back into a unified student model. The goal is to consolidate the specialized abilities learned during separate RL stages.</p><p>The final stage is instruction-oriented reinforcement learning. This is meant to make the reasoning-enhanced model more reliable for user-facing prompts. For explicit constraints, rewards can be computed by rule-based validators that check format, ordering, item count, keyword constraints, and task completion. For open-ended instructions, rubric-based reward models evaluate qualities such as helpfulness, coherence, instruction adherence, and redundancy.</p><h2>Results</h2><p>The reported results are strong for a 3B model, especially on verifiable math and code benchmarks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HP3v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HP3v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 424w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 848w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 1272w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HP3v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png" width="1456" height="732" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:732,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;alt text&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="alt text" title="alt text" srcset="https://substackcdn.com/image/fetch/$s_!HP3v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 424w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 848w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 1272w, https://substackcdn.com/image/fetch/$s_!HP3v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F45d9ece0-737d-4e6b-8ced-194fbfc0f282_1620x814.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>These numbers should be read carefully. They are concentrated in domains where the model&#8217;s training objective is well matched to the evaluation. VibeThinker-3B looks especially strong on competition-style math and executable coding, where correctness can be checked.</p><p>The report notes a clearer gap on knowledge-heavy evaluation such as GPQA-Diamond. This is consistent with the model&#8217;s own framing: compact models may be able to learn dense reasoning procedures, but broad factual knowledge and general-purpose competence still benefit from larger parameter budgets.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TYXC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TYXC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 424w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 848w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 1272w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TYXC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png" width="1198" height="727" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:727,&quot;width&quot;:1198,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:215549,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/202631080?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TYXC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 424w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 848w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 1272w, https://substackcdn.com/image/fetch/$s_!TYXC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74ec50ca-c1ee-4312-b2ca-55eb168c4d55_1198x727.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><h2>MoQ Updates</h2><p>This week, I had to dedicate all my computing resources to evaluating GGUF versions of M3 and producing MoQ quantizations.</p><p>I am using my own money to evaluate the <a href="https://huggingface.co/unsloth/MiniMax-M3-GGUF">GGUFs produced by Unsloth</a>.</p><p>For the MoQ quantization itself, and for evaluating the resulting GGUFs, I am supported by <a href="https://verda.com/?utm_source=kaitchup.substack.com&amp;utm_medium=referral&amp;utm_content=m3quant">Verda</a>&#8217;s compute sponsorship. I am using two of their B300s for this work, totaling 560 GB of GPU VRAM, backed by a generous amount of CPU RAM. This setup lets me load the full M3 model, and GGUF quantization remains very fast on a hybrid-memory system like this.</p><p>The quantization should be finished this weekend, and the evaluation should be fast afterward. I&#8217;ll publish the models and a full article detailing the results once this is done.</p><p>At the beginning of the week, I also released MoQ GGUFs for Qwen3.6 27B.</p><ul><li><p><a href="https://huggingface.co/kaitchup/Qwen3.6-27B-GGUF-MoQ">kaitchup/Qwen3.6-27B-GGUF-MoQ</a></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0GkM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0GkM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 424w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 848w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 1272w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0GkM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png" width="927" height="608" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:608,&quot;width&quot;:927,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Qwen3.6 27B GGUFs_ Accuracy&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Qwen3.6 27B GGUFs_ Accuracy" title="Qwen3.6 27B GGUFs_ Accuracy" srcset="https://substackcdn.com/image/fetch/$s_!0GkM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 424w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 848w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 1272w, https://substackcdn.com/image/fetch/$s_!0GkM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6a231d4b-84b8-44b2-8d25-9e0693491418_927x608.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I also have GGUFs for Qwen3.6 35B-A3B, but I delayed their evaluation so I could focus on M3. I will evaluate and publish them once my M3 work is complete.</p><p><em><a href="https://huggingface.co/w-ahmad">Waleed Ahmad</a> is the author of the MoQ method and he is still working on optimizing it.</em></p><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[Qwopus and REAP: Custom Qwen3.6 Models for Local Reasoning]]></title><description><![CDATA[Large-scale evaluations beyond accuracy]]></description><link>https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models</link><guid isPermaLink="false">https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 17 Jun 2026 20:22:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DVuQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76668f-e137-416c-9826-d6d134dd6a60_1210x753.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As we saw in previous articles, Qwen3.6 are very good LLMs for local AI. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;0a3d7319-483b-481d-a5a7-0b8e618dbd8e&quot;,&quot;caption&quot;:&quot;In a previous article, I found Gemma 4 31B to be superior or comparable to Qwen3.5 27B in most areas, with similar or better accuracy and lower latency.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Qwen3.6 27B vs Qwen3.5 27B vs Gemma 4 31B: Accuracy, Latency, Memory, and Token Efficiency Tested&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:155699076,&quot;name&quot;:&quot;Benjamin Marie&quot;,&quot;bio&quot;:&quot;Research scientist in NLP/AI.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cad63296-e403-4e10-b54f-a1dc5602f881_1280x1280.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100}],&quot;post_date&quot;:&quot;2026-05-05T11:39:02.874Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!vAQL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2862fcd9-12bf-4ca0-ba6a-23d6808c8806_1210x783.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://kaitchup.substack.com/p/qwen36-27b-vs-qwen35-27b-vs-gemma&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195830510,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:23,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1783977,&quot;publication_name&quot;:&quot;The Kaitchup &#8211; AI on a Budget&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!xY7g!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffbb331d7-37df-408d-9f36-30b3b6369433_1256x1256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>On one side, <strong>Qwen3.6-35B-A3B</strong> is a sparse Mixture-of-Experts model with 35B total parameters, but only around 3B active parameters per token. On the other, <strong>Qwen3.6-27B</strong> is a dense model: smaller in total parameter count, but much heavier for each generated token.</p><p>In practice, the two models target slightly different users. The 35B-A3B is attractive if you want faster inference. The 27B is simpler: no router, no experts, and no MoE-specific surprises, but slower and more accurate.</p><p>Shortly after Qwen3.6 appeared, the community started releasing custom variants, including <strong>Qwopus3.6-35B-A3B</strong>, <strong>Qwopus3.6-27B</strong>, and <strong>Qwen3.6-28B-REAP</strong>.</p><p>They are all based on Qwen3.6, but they are not trying to solve the same problem. Qwopus focuses on reasoning style, answer structure, and distillation from stronger models. REAP focuses on reducing the size of the MoE by pruning experts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">The Kaitchup &#8211; AI on a Budget is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>This raises a practical question: should you use the original Qwen3.6 model, one of the Qwopus fine-tunes, or the smaller REAP version?</p><p>In this article, I compare these Qwen3.6 variants across several dimensions: how they were made, how much memory Qwen3.6-28B-REAP saves compared with the original Qwen3.6-35B-A3B, how token-efficient they are, and how accurate they are across different tasks and domains.</p><p>As we will see, customizing advanced models like Qwen3.6 is difficult. Accuracy alone is not enough. To properly adopt or validate a model, we also need to look at efficiency, failure modes, and behavior at scale.</p><h2>The Custom Qwen3.6 Models Compared</h2><ul><li><p><a href="https://huggingface.co/Jackrong/Qwopus3.6-27B-v2">Jackrong/Qwopus3.6-27B-v2</a></p></li><li><p><a href="https://huggingface.co/Jackrong/Qwopus3.6-35B-A3B-v1">Jackrong/Qwopus3.6-35B-A3B-v1</a></p></li><li><p><a href="https://huggingface.co/0xSero/Qwen3.6-28B">0xSero/Qwen3.6-28B</a></p></li></ul><blockquote><p><strong>Acknowledgements</strong></p><p>Before turning to the analysis, which is largely negative for all these models, I would like to thank <a href="https://huggingface.co/0xSero">0xSero </a>and <a href="https://huggingface.co/Jackrong">Jackrong</a> for releasing them and for investing their time, and money, into exploring new recipes to make local AI better and more affordable. These models are extremely valuable for research purposes.</p></blockquote><div><hr></div>
      <p>
          <a href="https://kaitchup.substack.com/p/qwopus-and-reap-custom-qwen36-models">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[New DiffusionGemma and MoQ GGUFs for Gemma 4 12B and LFM2.5 8B A1B]]></title><description><![CDATA[The Weekly Kaitchup #146]]></description><link>https://kaitchup.substack.com/p/new-diffusiongemma-and-moq-ggufs</link><guid isPermaLink="false">https://kaitchup.substack.com/p/new-diffusiongemma-and-moq-ggufs</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Fri, 12 Jun 2026 22:12:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c7_f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png" width="1456" height="815" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:815,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:245022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://kaitchup.substack.com/i/167992895?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c7_f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 424w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 848w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1272w, https://substackcdn.com/image/fetch/$s_!c7_f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd3f4615-2c66-4a1a-8e7c-7002fe563e0b_1517x849.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Hi everyone,</p><p>In this edition of The Weekly Kaitchup, let&#8217;s discuss:</p><ul><li><p>DiffusionGemma: 4x Faster than Gemma 4</p></li><li><p>MoQ GGUFs: Gemma 4 12B IT and LFM2.5 8B A1B</p></li></ul><h3><strong>30% Off the Subscription to The Kaitchup:</strong></h3><blockquote><p><strong><a href="https://kaitchup.substack.com/30off2026">Get the Coupon and Subscribe</a></strong></p></blockquote><div><hr></div><h2>DiffusionGemma: 4x Faster than Gemma 4</h2><p>DiffusionGemma is Google DeepMind&#8217;s experimental text-diffusion variant of Gemma 4: </p><ul><li><p><a href="https://huggingface.co/google/diffusiongemma-26B-A4B-it">google/diffusiongemma-26B-A4B-it</a></p></li></ul><p>It is based on the Gemma 4 26B A4B Mixture-of-Experts family: roughly 25.2B total parameters, 3.8B active parameters, 30 layers, 128 experts with 8 active experts, 256K context, and<strong> a 256-token diffusion canvas</strong>. </p><p>Unlike a normal autoregressive LLM, it does not generate one token at a time. It generates text in blocks by repeatedly denoising a whole canvas of tokens.</p><p>Architecturally, it is a block-autoregressive discrete diffusion model. The prompt is processed causally into a KV cache, then the model works on a 256-token canvas with more bidirectional interaction inside that canvas. After several denoising iterations, the finalized block is appended to the cache and the model moves to the next block. </p><p>Generation starts from noisy/random token guesses on the canvas. At each denoising step, the model: </p><ul><li><p>predicts tokens for all positions</p></li><li><p>estimates uncertainty</p></li><li><p>keeps the most confident positions</p></li><li><p>re-noises uncertain ones</p></li><li><p>repeats</p></li></ul><p>Google&#8217;s public sampler uses Entropy-Bounded Denoising with Adaptive Stopping: up to 48 denoising steps, temperature decaying from 0.8 to 0.4, entropy-based token selection, and early stopping once the canvas is stable enough. This lets the model revise earlier positions in light of later positions, which is useful for editing, infilling, code, structured text, and constraint-heavy tasks.</p><p>Why diffusion? Because it&#8217;s much faster than a standard autoregressive LLM, even when accelerated with speculative decoding.</p><p>Its speed comes from replacing many tiny sequential token steps with larger parallel GPU workloads. Autoregressive decoding is often memory-bandwidth-bound at low batch size because each token needs another forward pass over the model. DiffusionGemma instead refines 256 positions at once, making better use of GPU parallelism. </p><p>Google reports up to 4x faster generation, with figures such as 700+ tokens/s on RTX 5090, 1000+ tokens/s on H100, and over 1100 tokens/s on H100 FP8 in low-batch settings.<em> Note: The baseline is Gemma 4 26B A4B with MTP, not with the <a href="https://huggingface.co/z-lab/gemma-4-26B-A4B-it-DFlash">DFlash speculator</a> made by z-lab which could be faster.</em></p><p>The advantage is strongest for local or single-user inference, and weaker in high-throughput server settings where autoregressive models can batch many users efficiently.</p><p>However, in terms of accuracy, DiffusionGemma significantly underperforms the original 26B model:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yF-W!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yF-W!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 424w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 848w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 1272w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yF-W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin" width="1000" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;DiffusionGemma Benchmark&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="DiffusionGemma Benchmark" title="DiffusionGemma Benchmark" srcset="https://substackcdn.com/image/fetch/$s_!yF-W!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 424w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 848w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 1272w, https://substackcdn.com/image/fetch/$s_!yF-W!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd2472c72-f9ff-436b-a997-919ec7fcdf61_1000x562.bin 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>reference: https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/</em></p><div><hr></div><h2>MoQ GGUFs: Gemma 4 12B IT and LFM2.5 8B A1B</h2><p>I&#8217;ve continued creating and evaluating new MoQ GGUFs in collaboration with <a href="https://huggingface.co/w-ahmad">Waleed Ahmad</a>, the author of the method.</p><p>This week, I released Gemma 4 12B MoQ GGUFs:</p><ul><li><p><a href="https://huggingface.co/kaitchup/gemma-4-12b-it-GGUF-MoQ">kaitchup/gemma-4-12b-it-GGUF-MoQ</a> </p></li></ul><p>They are, to the best of my knowledge, the strongest sub-7 GB GGUFs of Gemma 4 12B currently available.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ikmn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ikmn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 424w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 848w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 1272w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ikmn!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png" width="808" height="455.05494505494505" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:808,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!Ikmn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 424w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 848w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 1272w, https://substackcdn.com/image/fetch/$s_!Ikmn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbf3455a-5269-42db-848a-ce22e1994a73_1728x973.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3o90!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3o90!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 424w, https://substackcdn.com/image/fetch/$s_!3o90!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 848w, https://substackcdn.com/image/fetch/$s_!3o90!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 1272w, https://substackcdn.com/image/fetch/$s_!3o90!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3o90!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png" width="812" height="348.5576923076923" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:625,&quot;width&quot;:1456,&quot;resizeWidth&quot;:812,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!3o90!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 424w, https://substackcdn.com/image/fetch/$s_!3o90!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 848w, https://substackcdn.com/image/fetch/$s_!3o90!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 1272w, https://substackcdn.com/image/fetch/$s_!3o90!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F541eabcd-ad6a-49f0-9c30-99369d74420f_1728x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Waleed has also been working on <a href="https://huggingface.co/w-ahmad/LFM2.5-8B-A1B-GGUF-MoQ">LFM2.5 8B A1B</a>, and my evaluation results show that these MoQ GGUFs outperform other GGUFs as well, including those made with APEX.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YMCO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YMCO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 424w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 848w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YMCO!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png" width="888" height="500.1098901098901" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:820,&quot;width&quot;:1456,&quot;resizeWidth&quot;:888,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!YMCO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 424w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 848w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 1272w, https://substackcdn.com/image/fetch/$s_!YMCO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F742c095c-1e8f-4d9c-a6be-684a8667bbce_2024x1140.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H6Up!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H6Up!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 424w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 848w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 1272w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H6Up!,w_2400,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png" width="782" height="446.85714285714283" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:false,&quot;imageSize&quot;:&quot;large&quot;,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:782,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:&quot;center&quot;,&quot;offset&quot;:false}" class="sizing-large" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!H6Up!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 424w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 848w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 1272w, https://substackcdn.com/image/fetch/$s_!H6Up!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c00702f-9d17-43de-b63e-d6373d054979_1550x886.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>For me, this is enough empirical evidence to start investing more seriously in the method. I&#8217;m now working on MoQ versions of Qwen3.6, the larger Gemma 4 models, <strong>and M3</strong>. This will take some time, as the method is costly to run. MoQ GGUFs for Qwen3.6 27B should be released this weekend.</p><p>Once we gather enough results, the code will be open-source and we will publish a technical report.</p><div><hr></div><p>That&#8217;s all for this week.</p><p>If you like reading The Kaitchup, consider sharing it with friends and coworkers (there is a 20% discount for group subscriptions):</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share The Kaitchup &#8211; AI on a Budget&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://kaitchup.substack.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share The Kaitchup &#8211; AI on a Budget</span></a></p><p>Have a nice weekend!</p><p></p>]]></content:encoded></item><item><title><![CDATA[Make Your Own Optimized GGUFs with AutoRound]]></title><description><![CDATA[Build optimized GGUF models for llama.cpp and LM Studio using AutoScheme, custom bit-widths, and layer protection.]]></description><link>https://kaitchup.substack.com/p/make-your-own-optimized-ggufs-with</link><guid isPermaLink="false">https://kaitchup.substack.com/p/make-your-own-optimized-ggufs-with</guid><dc:creator><![CDATA[Benjamin Marie]]></dc:creator><pubDate>Wed, 10 Jun 2026 00:11:46 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!n-zd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37676490-4dd2-4e6d-9e98-33a82f341a02_1941x1220.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>GGUF is the standard format for running LLMs locally with tools such as llama.cpp and LM Studio. It stores both model weights and inference metadata in a compact binary format, making GGUF models easy to download, load, and serve.</p><p>GGUF is especially popular because it supports many quantization levels. Instead of using a full BF16 or FP16 checkpoint, users can pick a smaller 2-bit to 8-bit model that better fits their RAM, VRAM, and quality requirements.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Support The Kaitchup and get <strong>30% off your subscription today.</strong></p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The most interesting GGUFs today increasingly rely on mixed-precision quantization. Unsloth Dynamic, or UD, helped popularize this approach by using lower precision for less sensitive tensors and higher precision for important ones.</p><p><a href="https://github.com/intel/auto-round/blob/main/docs/step_by_step.md#autoscheme">AutoRound&#8217;s AutoScheme</a> one of the best practical alternatives for creating your own mixed-precision GGUFs today. Given a target average bit-width and a list of candidate schemes, AutoScheme automatically chooses between GGUF quantization types, for example GGUF:Q2_K_S and GGUF:Q4_K_S, layer by layer.</p><p>This article is therefore a step toward learning how to build, control, and evaluate mixed-precision GGUF recipes. As most open-weight models already exists in some high-quality GGUF version, making your own GGUF is especially relevant if you have a fine-tuned version of a model and want to GGUF it.</p><p>I also expect <a href="https://huggingface.co/w-ahmad/LFM2.5-8B-A1B-GGUF-MoQ">MoQ-like strategies</a> to eventually be adapted and supported by tools such as AutoRound, which could further improve the kind of recipe built in this article.</p><p>In this article, you will learn how to:</p><ul><li><p>create your own GGUF model with AutoRound AutoScheme;</p></li><li><p>control the target average bit-width and candidate GGUF quantization types;</p></li><li><p>protect important layers;</p></li><li><p>compare the quality-size trade-off.</p></li></ul><p>Here is a notebook you can use to make your own optimized GGUFs for Qwen3.5/3.6 models:</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://kaitchup.substack.com/p/notebooks&quot;,&quot;text&quot;:&quot;Get the Notebook (#200)&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://kaitchup.substack.com/p/notebooks"><span>Get the Notebook (#200)</span></a></p>
      <p>
          <a href="https://kaitchup.substack.com/p/make-your-own-optimized-ggufs-with">
              Read more
          </a>
      </p>
   ]]></content:encoded></item></channel></rss>