<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[UncoverAlpha]]></title><description><![CDATA[Deep dives/analyses of the technology companies and tech sub-industries. Mostly about AI, semiconductor, cloud, software, and ad tech sectors.]]></description><link>https://www.uncoveralpha.com</link><image><url>https://substackcdn.com/image/fetch/$s_!YsyF!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f3c972-e191-4eec-857d-3ccfda10a107_227x227.png</url><title>UncoverAlpha</title><link>https://www.uncoveralpha.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 12 Sep 2026 00:53:33 GMT</lastBuildDate><atom:link href="https://www.uncoveralpha.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Rihard Jarc]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[uncoveralpha@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[uncoveralpha@substack.com]]></itunes:email><itunes:name><![CDATA[UncoverAlpha]]></itunes:name></itunes:owner><itunes:author><![CDATA[UncoverAlpha]]></itunes:author><googleplay:owner><![CDATA[uncoveralpha@substack.com]]></googleplay:owner><googleplay:email><![CDATA[uncoveralpha@substack.com]]></googleplay:email><googleplay:author><![CDATA[UncoverAlpha]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Meta just launched its most important product since WhatsApp and Instagram: the Muse AI agent ]]></title><description><![CDATA[I cover the business case for Muse Spark and some economics, as well as technical details on how the agent is set up, what this means for Meta&#8217;s compute, and my research on the latest sentiment/usage around adoption of Muse Spark 1.3, OpenAI GPT 6.]]></description><link>https://www.uncoveralpha.com/p/meta-just-launched-its-most-important</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/meta-just-launched-its-most-important</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Wed, 09 Sep 2026 11:47:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tkKi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0ebd520-2586-4895-81fb-94a68aa26f2d_1242x590.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Hey everyone,</span></p><p><span>Meta released Muse yesterday, its personal AI agent. I have been following Meta in depth for years, and this is a significant milestone for Meta, as I think this is the most important product Meta has shipped since it bought Instagram and WhatsApp. Muse is the first Meta product that can capture significant value outside of the ad impression ecosystem and, over time, possibly become even bigger.</span></p><p><span>In the article, I cover the business case for Muse Spark and some economics, as well as technical details on how the agent is set up, what this means for Meta&#8217;s compute, and my research on the latest sentiment/usage around adoption of Muse Spark 1.3, OpenAI GPT 6, and Google Gemini 3.8.</span></p><div><hr></div><p><em><span>Before we start with the article, I would like to invite everyone to an upcoming talk I will be hosting on Sep 16th. The topic is Token Economics (usage trends, model tradeoffs, and best-practice token optimizations).</span></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!51-j!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!51-j!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 424w, https://substackcdn.com/image/fetch/$s_!51-j!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 848w, https://substackcdn.com/image/fetch/$s_!51-j!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 1272w, https://substackcdn.com/image/fetch/$s_!51-j!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!51-j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png" width="1200" height="708" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a215f92f-8b92-479f-86e7-89660091787b_1200x708.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:708,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:265205,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/214868635?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!51-j!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 424w, https://substackcdn.com/image/fetch/$s_!51-j!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 848w, https://substackcdn.com/image/fetch/$s_!51-j!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 1272w, https://substackcdn.com/image/fetch/$s_!51-j!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa215f92f-8b92-479f-86e7-89660091787b_1200x708.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>I will be joined by Kyle Cheng, a former Anthropic technical staff member; Andy Hock, Chief Strategy Officer at Cerebras; and Chris Ackerson, SVP of Product at AlphaSense.</span></em></p><p><em><span>Each of the guests represents a different angle on token economics (frontier lab view, infrastructure/semi view, and big token consumer customer view) and, because of this, will give us a very interesting perspective on the whole topic and the trends we are heading toward.</span></em></p><p><em><span>Even if you can&#8217;t join us live, make sure to register to receive the link to the recordings of the discussion. You can sign up for free using this link.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.alpha-sense.com/resources/webinars/token-economics-how-heavy-ai-users-are-optimizing-spend-at-scale/?utm_source=pt_rihard&amp;utm_medium=sponsored&amp;utm_campaign=SWB_DG_09-16-26_IMP-GENAI_CORPFS_UncoverAlpha-ExpertPro-TokenEconomics&quot;,&quot;text&quot;:&quot;Sign up for Free using this link&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.alpha-sense.com/resources/webinars/token-economics-how-heavy-ai-users-are-optimizing-spend-at-scale/?utm_source=pt_rihard&amp;utm_medium=sponsored&amp;utm_campaign=SWB_DG_09-16-26_IMP-GENAI_CORPFS_UncoverAlpha-ExpertPro-TokenEconomics"><span>Sign up for Free using this link</span></a></p><div><hr></div><p><span>Now let&#8217;s start with the article.</span></p><p><strong><span>What Muse actually is</span></strong></p><p><span>The simplest way to describe it: every Muse user gets their own dedicated computer in Meta&#8217;s cloud, and an agent that lives on it and works for them around the clock. It&#8217;s the mainstream version of OpenClaw. OpenClaw showed the tech community what a personal agent with full access to your life can do, but it required a weekend of configuration, and I think everyone knows that this version of the product was not what was going to go mainstream.</span></p><p><span>Some data from Meta&#8217;s Muse launch</span></p><ul><li><p><span>Available in the US only for now, available via web, iOS, Android, and inside WhatsApp chats. AI glasses integration is still coming.</span></p></li><li><p><span>Free tier with up to 100 million tokens per week, then two paid tiers: Power at $20/month and Maximum at $100/month.</span></p></li><li><p><span>Connectors to email, calendar, payments (Link by Stripe at launch, Shop Pay), health, smart home, dining, shopping, music, events, plus Instagram and Facebook, including the tools for running a business and buying ads.</span></p></li><li><p><span>If a service isn&#8217;t a built-in connector but has a public API, Muse writes its own connector. If there is no API, it uses its browser.</span></p></li><li><p><span>Powered by Muse Spark 1.3, which Meta recently released.</span></p></li></ul><p><span>The tasks list Meta gives as examples are sending emails, booking travel, lowering your bills, filling out forms, turning a recipe reel into a grocery list, sending party invites, buying things. The assistant is made for the digital chores of a normal household, which is exactly the segment where distribution beats model capability and Meta has an edge.</span></p><p><strong><span>Why this is a TAM expansion and not a feature</span></strong></p><p><span>For the last decade, the value Meta&#8217;s platforms created has been much bigger than the value Meta captured. Influencer marketing, social commerce, Marketplace, WhatsApp business chats. While Meta wanted to expand that value capture, two things blocked those efforts: Meta&#8217;s own product choices and the platform owners. Apple&#8217;s ATT alone cost Meta roughly $10B in 2022 revenue, by the company&#8217;s own estimate. Every attempt at payments, shops, and commerce got squeezed by app store rules, OS-level tracking limits, and the fact that Meta doesn&#8217;t own the device. Meta has many failed projects like the payment project Libra, which got regulatory scrutiny even before launch, and many others.</span></p><p><span>Muse is different because the agent doesn&#8217;t run on your iPhone. It runs on a Linux VM in Meta&#8217;s data center with its own Chromium browser. When Muse shops for you, it is Meta&#8217;s browser hitting the merchant&#8217;s site, Meta&#8217;s wallet issuing a single-use card number via Stripe Link, and Meta&#8217;s Sentinel approving the checkout. Apple is reduced to being the screen on which you tap &#8220;approve.&#8221;</span></p><p><span>And Zuck was unusually explicit on the business model on Sources. The plan is for Muse to pay for itself by helping people earn and save money, with Meta taking, in his words, a very small cut of the transaction, potentially paid by the business on the other side. Add the $20/$100 subscription for power users on top. That is a transactional revenue line and a subscription revenue line, which are exactly two of the four new revenue lines (transactional, compute, APIs, subscriptions) I wrote about after Meta&#8217;s Q2 earnings.</span></p><p><span>Meta makes about $67 per year per daily active person, blended globally. A $20/month Power subscription is $240 per year, 3.6x the current ARPP. A $100/month Maximum tier is 18x. Even if only 1% of DAP ends up on a paid tier, that is 36 million subscribers at a blended, say, $30/month, or about $13B of annual revenue that didn&#8217;t exist before, at software-like margins on top of an ad business that did $59.4B in a single quarter. And I think the subscription line is actually the smaller of the two. The transaction cut has the bigger ceiling, as it not only captures the e-commerce market but also extends into traditional commerce and experiences.</span></p><p><strong><span>Facebook Marketplace might be the first killer use case</span></strong></p><p><span>If I had to bet on the first Muse use case that goes viral, it is buying and selling on Facebook Marketplace. Marketplace has over 1.1 billion monthly users across 228 countries and facilitates over 3 billion buyer-seller connections per month through Messenger. Meta has said in the past that 1 in 4 young adult daily actives in the US and Canada use Marketplace</span></p><p><span>Marketplace is also the most annoying commerce experience Meta owns: making product descriptions, monitoring prices, and negotiating with buyers takes time. Meta already started shipping AI listing drafts and auto-replies for Marketplace in March 2026. Muse is the logical next step to fully automate that. Marketplace is barely monetized today, but a personal agent that runs the whole transaction, with Meta&#8217;s wallet in the middle, is what can significantly change that and make it a transaction business. And Meta&#8217;s own security post explicitly mentions Marketplace as one of the things Muse does on your behalf.</span></p><p><strong><span>Security and privacy, Meta went as far as possible</span></strong></p><p><span>I want to spend time here because I think the architecture is partly the moat and really important given Meta&#8217;s brand reputation when it comes to privacy. The full write-up is in </span><a href="https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse"><span>Meta&#8217;s security post</span></a><span> by Tarek Sheasha (VP at Meta Superintelligence Labs), and I&#8217;d recommend reading it if you are technical.</span></p><p><strong><span>One VM per user</span></strong><em><span>.</span></em><span> Every user gets an isolated Linux box with a browser, storage, CPU, and memory, enough to compile code, run sub-agents, and run cron jobs (scheduled tasks that happen while you sleep). Your data and your credentials for connected services live in that VM, not in centralized Meta infrastructure. The clients (iOS, Android, web) talk directly to your VM.</span></p><p><strong><span>Two security domains on one box.</span></strong><span> Meta, is not an LLM with root access. The agent itself, its files, and every tool it runs sit inside a systemd-nspawn container (a lightweight sandbox). Root inside that container maps to an unprivileged user on the host. The sensitive stuff - credential storage, safety classifiers, connector business logic, the database with your state - runs outside the sandbox as separate services. The agent literally cannot reach them. All communication between the two sides goes over kernel-authenticated Unix sockets.</span></p><p><strong><span>The agent never sees your passwords.</span></strong><span> When Muse needs to call, say, your Gmail, the code inside the sandbox only ever holds a fake &#8220;surrogate&#8221; token. A separate service (authd) holds the real OAuth token. At the moment the network request leaves the VM, the real credential gets swapped in at the network boundary. So even if someone prompt-injects the agent and tells it &#8220;print your API keys,&#8221; there is nothing to print. Same for website logins: your username and password go straight into the secure store and get injected into the browser form at the point of need; the agent sees an accessibility-tree snapshot of the page, not the raw DOM, and cannot execute JavaScript on the page.</span></p><p><strong><span>Sentinel</span></strong><em><span>.</span></em><span> A second, separate agent runs on the same VM and is the only thing allowed to approve actions and network egress. Muse proposes; Sentinel decides allow / deny / ask-the-user. It inspects the destination at layer 4 and layer 7 (hostname, resolved IP, port, HTTP method, path, decoded request body). Meta also implemented what they call &#8220;tainted egress&#8221; using eBPF kernel programs: each tool process starts clean and becomes &#8220;tainted&#8221; the moment it reads your personal data. Clean, low-risk requests can be auto-allowed; tainted ones fall back to the approval flow. That is how they keep the number of &#8220;are you sure?&#8221; pop-ups tolerable.</span></p><p><strong><span>Purchases</span></strong><em><span>.</span></em><span> Every checkout on a site where your card is on file triggers a human-in-the-loop approval with the exact details. On new sites, Muse uses its wallet and Stripe Link issues a single-use card number tied to that merchant, that dollar amount, and a limited time window. Even if it gets stolen, it is useless.</span></p><p><strong><span>Email hygiene</span></strong><em><span>.</span></em><span> The email connector filters out one-time passcodes, password reset links, and magic login links with deterministic filters plus a classifier, because your inbox is the master key to every other account you own.</span></p><p><strong><span>Prompt injection.</span></strong><span> Meta&#8217;s defense is layered: the model is trained against it (they say Spark 1.3 is close to state of the art on their internal evals), external data is labeled untrusted in the harness, an ensemble of independent classifiers runs on everything entering the model&#8217;s context, and then the deterministic boundaries above apply even if the model is fooled. Meta opened the bug bounty to the public yesterday, paying up to $300,000 per valid report and up to $130,000 for a successful prompt injection affecting a single user.</span></p><p><strong><span>Confidential VM (coming later this year)</span></strong><em><span>.</span></em><span> Today, Meta restricts employee access to your VM by policy. The next step is cryptographic: a mode where Meta itself provably cannot read what is inside your VM, with the design and source code going to external auditors and a continuous, publicly inspectable audit after launch. This is the part Zuck and Nat Friedman personally recruited Moxie Marlinspike, the founder of Signal, to build. Zuck&#8217;s claim on Sources was that even Meta cannot see the content, and that he isn&#8217;t aware of anyone offering anything close.</span></p><p><span>All of this matters a lot because the personal agent race is not going to be won on model capabilities. It is going to be won on who people trust with their inbox, their calendar, and their credit card, and on who can serve the workload cheaply enough at scale. Instinct, the hottest startup in this category, raised $250M at a $2.5B valuation two weeks ago, and in the same week got dragged for a &#8220;perpetual and irrevocable&#8221; data license in its terms and for sending an email on a user&#8217;s behalf without asking. Meta, of all companies, is showing up to this fight with the most paranoid architecture in the market, and I think it is the right one given Meta&#8217;s history and brand reputation.</span></p><p><strong><span>The unit economics of 100 million free tokens per week</span></strong></p><p><span>Muse Spark 1.3 is priced at $1.25 per million input tokens and $4.25 per million output tokens on the standard tier, with a $0.15 cached-input rate; the &#8220;contributor&#8221; tier, where Meta can train on your data, is $0.10/$0.20. 100 million tokens per week at Meta&#8217;s own standard list price, assuming a typical agent mix of ~90% input / ~10% output, is about $155 of API value per week, or roughly $8,000 per year, given away for free. Even at the contributor tier price it is ~$11 per week / ~$570 per year. Meta is essentially pricing the category so that Instinct, Town, and anyone else without a big profitable business cannot compete on price.</span></p><p><span>My rough scenario for what this costs Meta. If Muse gets to 100 million monthly actives and an average user actually consumes 5-10 million tokens per week (a handful of real agentic tasks with sub-agents and a browser session or two), that is 0.5-1 quadrillion tokens per week, or 26-52 quadrillion per year. For context, Google was processing 1.3 quadrillion tokens per month across all of its surfaces in October 2025. At an internal serving cost of, say, $0.10-0.20 per million tokens blended (the contributor tier at $0.10/$0.20 has to be at or above cost), that is $2.6B-$10B per year of inference cost, before the CPU, memory, and storage of keeping tens of millions of VMs alive.</span></p><p><span>And it seems usage is off to a strong start already, as Alex, the head of AI at Meta, already tweeted that usage of intially users is 10x than what the testing cohort was:</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kkv0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kkv0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 424w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 848w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 1272w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kkv0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png" width="578" height="111" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:111,&quot;width&quot;:578,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:21191,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/214868635?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kkv0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 424w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 848w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 1272w, https://substackcdn.com/image/fetch/$s_!kkv0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30a17566-63c5-439c-a921-1e37eb4e706e_578x111.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>There are a few implications to this which I outlined here:</span></p><p><span>On the Q2 call, Zuck said Meta was getting a lot of offers for its compute at a significant premium and floated running an auction over compute. Muse is the alternative use of that compute. If Muse usage takes off, I think the odds of a third-party compute deal go down, because the internal return on that compute, measured in transaction take rate and subscriptions, is probably higher than renting it out, plus it adds long-term durable product value. If Muse stalls, the compute deal becomes more likely. Either way, the compute will get used.</span></p><p><span>A personal agent&#8217;s job is 90% integrations, distribution, and harness, and 10% raw intelligence. Meta&#8217;s model needs to be good enough and cheap enough to serve hundreds of millions of people without them hitting a wall after three tasks. That is the exact opposite of the problem OpenAI has right now with Astra. Meta doesn&#8217;t need a frontier model for this use-case.</span></p><p><strong><span>The network effects of AI personal agents</span></strong></p><p><span>An interesting thing that Zuck said on Sources is that the whole industry currently treats agents as a single-player game, and Meta is building for multiplayer. Concretely, Muse has a &#8220;fleet&#8221; concept: the underlying agents can learn anonymized insights from the whole user base, so the product gets better as more people use it. That is a network effect on a product category that has not had one yet.</span></p><p><span>The practical version of this is the idea/suggestion tab on Muse. Most people still don&#8217;t know what to ask an AI agent to do for them. Muse learns from your conversations what matters to you and makes unprompted suggestions, and it does what the team calls &#8220;learning through the night,&#8221; reorganizing and proposing next steps while you sleep. Now scale that with anonymized fleet data: &#8220;people like you who connected their calendar and email mostly use me for X.&#8221; If any company is good at recommendation systems and network effects it&#8217;s Meta so this could give them a durable advantage over competitors, especially if they reach scale.</span></p><p><span>Then we come to Meta&#8217;s distribution. Threads went from zero to 500 million monthly actives almost entirely on Instagram&#8217;s distribution, and Zuck said on the Q2 call that Meta plans to run that playbook repeatedly for AI apps. Muse is already inside WhatsApp chats at launch. And add partners: on day one, Muse ships with Stripe Link, Shop Pay, and 1Password, plus Instagram and Facebook business tools, because everyone wants to be inside the Meta ecosystem. A small startup has a much harder time getting those integrations; with Meta, it's different because everyone wants to work with them and tap into their ecosystem.</span></p><p><strong><span>My Sentiment analysis check: Muse Spark 1.3 vs GPT-6 Astra vs Gemini 3.8 Flash (analyzed X, Hacker News, GitHub, and Reddit)</span></strong></p><p><span>This part is for paid subscribers only, but the results are quite surprising compared with the current consensus.</span></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/meta-just-launched-its-most-important">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Meta's AI Infrastructure: Mispriced, Misread, and Massive]]></title><description><![CDATA[I analyzed how big is Meta&#8217;s AI data center footprint and how much revenue & profit it could generate. I also shared what moves could change sentiment on the company, and on top of it, I also shared why I believe Meta is right now in a unique position.]]></description><link>https://www.uncoveralpha.com/p/metas-ai-infrastructure-mispriced</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/metas-ai-infrastructure-mispriced</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Sun, 30 Aug 2026 12:51:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NqeV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Hey everyone,</span></p><p><span>In this article, I analyzed how big is Meta&#8217;s AI data center footprint and how much revenue &amp; profit it could generate. I also shared what moves could change sentiment on the company, and on top of it, I also shared why I believe Meta is right now in a unique position (with leverage) in this AI race that has morphed into a capital race, but why Meta needs to act is weeks, not years. For Meta Compute, it&#8217;s now or never.</span></p><p>This article is exclusive for paid subscribers.</p><p><span>Let&#8217;s get into it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NqeV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NqeV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NqeV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg" width="1024" height="559" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:559,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:135020,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/213382101?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NqeV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 424w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 848w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!NqeV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2737cf5-56d2-4caf-a3e3-a6327c389caa_1024x559.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>How many gigawatts does Meta actually have?</span></strong></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/metas-ai-infrastructure-mispriced">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Nvidia and the Hyperscalers High Stakes Poker Game]]></title><description><![CDATA[The relationship between Nvidia and the big hyperscalers Amazon, Google, and Microsoft was once very clear and friendly, where one was a supplier and the other three were buyers. I believe that era is clearly over, as each is trying to commoditize the layer the other serves.]]></description><link>https://www.uncoveralpha.com/p/nvidia-and-the-hyperscalers-high</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/nvidia-and-the-hyperscalers-high</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 13 Aug 2026 12:29:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yBig!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>The relationship between Nvidia and the big hyperscalers Amazon, Google, and Microsoft was once very clear and friendly, where one was a supplier and the other three were buyers. I believe that era is clearly over, as each is trying to commoditize the layer the other serves. The chip design and data center infrastructure markets. This clash is critical for the AI buildout age, as it determines where much of the AI value (margin) ultimately ends up.</span></p><p><span>You have Nvidia&#8217;s Jensen Huang with the strongest balance sheet in semiconductor history and a ~75% gross margin business. On the other side are the three biggest Nvidia customers &#8212; Amazon, Microsoft, and Google &#8212; who together will spend around $600B in CapEx this year and have all come to the same conclusion: they do not want to keep paying Nvidia&#8217;s margin forever.</span></p><p><span>In the last few years, the relationship was largely defined by Nvidia being the only game in town for AI chips, and most workloads were used for training. But now things are changing, because:</span></p><ul><li><p><span> we are shifting to an inference-led market,</span></p></li><li><p><span>Nvidia has a historic high 75% gross margin, and it is bothering the data center ecosystem because it&#8217;s long-term hurting their margin</span></p></li><li><p><span>The stakes have just become so much bigger because hundreds of billions, and maybe even soon trillions, per year will be spent on AI chips, and every layer of the AI stack wants to capture as much value as possible.</span></p></li></ul><p><span>Let&#8217;s dive in.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yBig!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yBig!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!yBig!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!yBig!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!yBig!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yBig!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1985620,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/211014765?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yBig!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!yBig!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!yBig!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!yBig!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2e715c7a-fcf9-465a-9c63-7c54e217d72f_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span>Training the land grab. Inference the economy.</span></h3><p><span>Most of you know this well, but just to shortly summarize. Training is when you build the model. You take a giant cluster, you run it for months, and you need the best interconnected hardware because the whole cluster works as one. This is where Nvidia&#8217;s moat is the strongest: CUDA, NVLink, the networking stack. In training, Nvidia&#8217;s share of the merchant market is still estimated to be north of 90%; the only real competitive chip that comes close is Google&#8217;s TPU (mostly because of Google&#8217;s interconnect and clustering capabilities).</span></p><p><span>Inference is when you run the model and is different:</span></p><ul><li><p><span>It is massively parallel but doesn&#8217;t need the same tight cluster coherence as training.</span></p></li><li><p><span>The workload is known and stable, which is exactly the situation in which a custom ASIC (a chip designed for one task) often beats a general-purpose GPU in cost per token.</span></p></li><li><p><span>It is where most of the future volume is going. Microsoft processed over 100 trillion tokens in a quarter in 2025, up 5x year over year. Google cited 22 billion tokens processed per minute across Alphabet&#8217;s APIs.</span></p></li></ul><h3><span>The customers became the competitors</span></h3><p><span>The problem for Nvidia in this inference world is that the market is so big that companies and their biggest customers don&#8217;t want to allow a monopolization of one of the biggest markets in the world; it&#8217;s just too much money and profit to be made in this layer, or value transfer to go above the stack for them to just leave it alone. The problem is also that the biggest customers, the hyperscalers, have very strong operating cash flows and business to actually have significant ASIC labs in their conglomerates.</span></p><p><span>Google</span><strong><span> </span></strong><span>is the furthest ahead in chip design. Google produces +3 million TPUs annually. In October last year, it signed Anthropic to a deal for access to up to 1M TPU</span><strong><span>s</span></strong><span>, bringing well over a gigawatt of capacity online in 2026. In April this year, that was expanded with Broadcom to roughly 3.5 additional gigawatts of next-gen TPU capacity starting 2027. Google is also now selling TPU systems to external customers as hardware. Essentially doing the same thing Nvidia&#8217;s core business is. Because it seems Gemini has fallen behind and the AI lab team is in a refocus period, the likelihood of Google giving even more capacity to GCP and with it TPUs to outside customers has only increased.</span></p><p><span>Amazon told you on multiple recent earnings calls that it sees its semiconductor design business as one of the next pillars of the Amazon empire. And it&#8217;s already significant in size, with big growth rates. Amazon&#8217;s chips business (Trainium + Graviton) now runs at over $25B in annualized revenue, growing triple digits. Graviton is used by 98% of the top 1,000 EC2 customers. Project Rainier - the Anthropic cluster is scaling from 500,000+ Trainium chips toward 1M. The problem is also that Amazon, as a company, often likes to drive margins to zero during adoption periods, as they have an obsession with focusing on customers before reaching scale and passing a lot of the economic benefits to their customers. Trainium, given how young the project is, is highly competitive with Nvidia on a cost-per-token basis; much of this is due to Amazon pushing for very low margins. There is a reason why Anthropic likes Trainium for inference tasks.</span></p><p><span>Microsoft was and still is the laggard, but they, too, have figured out that owning the chip design layer will be key for a hyperscaler. The Information reported that Microsoft will unveil Maia 300 as soon as September and is negotiating with TSMC for capacity for more than 300,000 chips for delivery in 2027, with an eventual target of over 1 million units and, per an Azure Maia manager, &#8220;gigawatts&#8217; worth&#8221; of Maia. For context, Maia 200 (launched in January 2026, TSMC 3nm, 140B+ transistors) has shipped only in the tens of thousands and is housed in two US data centers. Reports mention a targeted cost advantage of 30-40% over Nvidia chips, and Microsoft is actively courting Anthropic as an anchor customer for it.</span></p><p><span>The fact that Anthropic has so far become the biggest success story from the AI labs' perspective is also not great for Nvidia, as Anthropic has notoriously built most of its chip-usage stack on Google&#8217;s TPUs and Amazon&#8217;s Trainium (not Nvidia). Yes, even Mythos (Fable) is rumored to be mostly TPUs and served on Trainium, though that&#8217;s also changing with some new Nvidia deals.</span></p><p><span>Even with Nvidia now trying hard to offer its chips to Anthropic in its deal announcements, it&#8217;s clear Anthropic doesn&#8217;t want to drop off TPUs and Trainiums. Instead, they want as many options as possible and to drive prices down, which is exactly the scenario Nvidia doesn&#8217;t want. Matching workloads to the most suitable, cheapest chip.</span></p><p><span>And I haven&#8217;t even mentioned AMD, where the OpenAI deal ( 6GW of MI-series deployment with equity warrants attached ) and the Broadcom/OpenAI 10GW custom program give buyers even more options, not to mention the SRAM players like Cerebras.</span></p><h3><span>Hyperscaler comments that increased tensions</span></h3><p><span>On the Q2 calls, all three hyperscalers and Meta described their CapEx in a way that Nvidia surely doesn&#8217;t like. They framed hundreds of billions in future AI spending as an option rather than a must.</span></p><p><span>They all committed to long-lead-time CapEx (land, shells, power, cooling &#8212; assets that last 15-25 years) years in advance, but said they are going to buy the chips, the &#8220;short-lived assets,&#8221; only months before a data center comes online. If demand is there, you buy them; if not, you don&#8217;t buy the chips.</span></p><div class="pullquote"><p><span>&#8220;if the demand environment changes, you just slow down what is, in fact, the largest component&#8221; </span></p><p><span>(Amy Hood, Microsoft)</span></p></div><p><span>Amazon, Meta, and Google essentially said the same thing. With Jassy talking in detail about &#8220;different capital cycles&#8221; for data center land, power vs chips, etc.</span></p><div class="pullquote"><p><span>&#8220;Google will keep investing as long as we see an attractive return on that investment&#8221; </span></p><p><span>(Google&#8217;s Anat Ashkenazil).</span></p></div><p><span>The fact is that as more time passes, the more mature the ASIC programs from these companies become, and the more independent they can be in that layer as they get more TSMC allocation or even strike deals with Intel or Samsung for fabrication of these chips. All of the hyperscalers are also going direct to memory manufacturers to secure memory capacity, so their long-term intentions are to essentially copy the AWS playbook on the CPU book with Graviton, where Amazon&#8217;s own Graviton chip now handles most of the CPU workloads in AWS.</span></p><p><span>Nvidia doesn&#8217;t want to wait for that and wants to front-load as much demand as possible today, when they still have a dominant market position, but they also want to secure their future and commoditize the hyperscalers.</span></p><h3><span>Nvidia&#8217;s counter-move</span></h3><p><span>Nvidia is on track to generate roughly $100B+ of annual free cash flow at current run rates. Nvidia is deploying it to manufacture demand outside the hyperscalers&#8217; walls, so that when the hyperscalers exercise their option to buy less, someone else is will buy more.</span></p><p><strong><span>Save the neoclouds at any cost.</span></strong><span> In September 2025, Nvidia signed a $6.3B commitment to buy back any unsold CoreWeave capacity through April 2032; a parallel deal with Lambda was worth $1.5B. Then, in January 2026, with CoreWeave, Nvidia went further by injecting $2B in equity. Nvidia became the equity holder, the demand backstop, and the customer of last resort.</span></p><p><strong><span>Revenue deal.</span></strong><span> On July 1, 2026, Nvidia formalized the model as the &#8220;AI Compute Partnership&#8221;: Nvidia acts as a financial backstop, agreeing to rent back unused GPUs at a fixed rate so lenders will finance neocloud buildouts &#8212; and in exchange, Nvidia takes a share of the neocloud&#8217;s cloud revenue. First neoclouds to join Firmus (up to 170,000 GPUs in Batam, Indonesia) and Sharon AI (40,000 GB300s in Australia). The problem with neoclouds is that their margins and business models are already razor thin, especially because they don&#8217;t have massive free cash flows and they essentially need debt to keep the lights going and to finance their buildouts. Even among them, it seems like the ones taking the Nvidia revenue share deal are seen as the weakest, according to this Former Nebius employee:</span></p><div class="pullquote"><p><span>&#8220;Firmus was almost going bankrupt. It was a cloud computing operator, but they faced substantial challenges, both financially and also regarding the executive board. They had problems regarding customers and also investors. This was the background before. Nvidia stepped in. The ecosystem works like this. When one of the major partners on Nvidia is actually struggling, NVIDIA will try to help the partner and find solutions together&#8221;</span></p><p><span>source: AlphaSense</span></p></div><p><span>He also shared his thoughts about what he thinks the Neoclouds are thinking about this Nvidia revenue share deal:</span></p><div class="pullquote"><p><span>&#8220;NeoCloud, they hate this model, so I can tell you directly. They don&#8217;t like it because Nvidia has a huge leverage on them, and at the same time they paid a far higher price to have, in this case the GPUs. Like I said, you don&#8217;t have the high discounts of buying in bulk. You actually have a far lower discount, and at the same time NVIDIA has also control on how you are using the cloud&#8221;.</span></p><p><span>source: AlphaSense</span></p></div><p><span>He also explained possible tension with Neoclouds and Nvidia, as, according to him, Nvidia would prefer that Neoclouds serve smaller, more diverse customers rather than just have big offloading deals with the hyperscalers.</span></p><p><strong>And another move they did was on Monday with the $500B financing deal,&nbsp;</strong><span>bringing Wall Street in more deeply. Nvidia announced a partnership with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to raise more than $500B in third-party capital for AI infrastructure, with compute itself used as collateral in SPVs, and Jensen saying that Nvidia retains the option to backstop up to $125B, or 25%, of the deals.</span></p><p><span>This is in addition to the reported talks to backstop up to $250B of OpenAI&#8217;s compute leases at the SB Energy 10GW Ohio campus and to finance ~$350B of OpenAI chip purchases. Given the recent commentary from SpaceX about going &#8220;Nvidia exclusive,&#8221; I expect a backstop deal or straight-up equity deal is also brewing in the background when it comes to SpaceX and Nvidia. Elon is smart enough to take advantage of a situation if Nvidia feels it needs to help the ecosystem with financing, as he has never lacked ambition for scale and doesn&#8217;t fear pulling the risk lever to the max.</span></p><p><span>Nvidia already disclosed $18.6B invested into private companies and infrastructure funds in Q1 FY27 alone, with the filing admitting some investees &#8220;may indirectly purchase or use our products&#8221;.</span></p><p><span>The problem with these kinds of deals is that when we have a compute shortage on paper, things mostly work out, but the moment we don&#8217;t have that shortage anymore, a business that pays Nvidia list price for chips and then pays Nvidia a percentage of revenue on top is a business with no margin left. The neocloud business models, in essence, are all a question in terms of whether this is really a sustainable, durable long-term business model or if this is just a byproduct of the current times where there is not enough compute and where some of the biggest compute buyers (hyperscalers) are doing these neocloud deals to buy some short-term capacity needs while they themselves maintain the end-client relationship. Comparing a hyperscaler business model to a neocloud is, in my view, like comparing apples to oranges, as hyperscalers have the data and ecosystem stickiness needed for a durable business. Also, with their scale and now ASIC development, they will control future margins better than someone with only 1 chip supplier.</span></p><p><span>These moves by Nvidia are important because they do change the short-term nature of the compute ecosystem and could help create a real AI bubble and an overbuild in compute capacity. Nvidia is right when it comes to neoclouds: for them to have durable business models, they need smaller, more diversified clients. And even that raises a question: once the compute shortage eases, and clients can get cheap, available compute from the hyperscalers, why would they even use neoclouds?</span></p><p><span>So far, the clients are mostly big companies, AI labs, and the hyperscalers themselves, where the risk is even higher when we do get into a capacity overbuild. The first thing hyperscalers will switch off is Neocloud deals, and the big AI labs are already building their own data centers to become independent of both Neocloud and the hyperscalers. The strategy that Nvidia is trying to encourage with Neoclouds about focusing on non-hyperscalers is a good one for sustainability, but the Neoclouds also like headline-reaching deals with the hyperscalers as it helps their stocks in the short-term.</span></p><p><span>To continue reading this article, you have to subscribe to the paid section. In the paid section, I will cover why Nvidia needs to finance these deals and how Nvidia, Neoclouds, and hyperscalers are positioned for the future.</span></p><h3><span>Why does a potential monopolist need to finance its own demand?</span></h3>
      <p>
          <a href="https://www.uncoveralpha.com/p/nvidia-and-the-hyperscalers-high">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Amazon, Google, Microsoft, Meta Q2 earnings: The AI CapEx ROIC is bad thesis is DEAD ]]></title><description><![CDATA[For more than a year, the number one pushback I hear from investors on big tech is that the CapEx is out of control, and the returns won&#8217;t be there. This was the first earnings season where all four companies directly and, in my view, successfully answered that worry.]]></description><link>https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q2-earnings</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q2-earnings</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 03 Aug 2026 15:56:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pa-A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Hey everyone,</span></p><p><span>We just got Q2 2026 earnings from Google, Microsoft, Meta, and Amazon. I want to share what I think are the most important takeaways.</span></p><p><span>For more than a year, the number one pushback I hear from investors on big tech is that the CapEx is out of control, and the returns won&#8217;t be there. This was the first earnings season where all four companies directly and, in my view, successfully answered that worry. Not by cutting CapEx; Amazon raised its 2026 number to ~$220B, Google raised to $195-205B, Meta lifted the floor of its guide to $130-145B. For the first time, all four laid out the same playbook on how they de-risk that spend: commit early to the long-lived assets (land, data center shells, power), and decide on the short-lived assets (the chips, which are the majority of the cost) only a few months before they need them, once they can actually see the demand. If the demand isn&#8217;t there, they don&#8217;t order the chips.</span></p><p><span>And on top of the playbook, all four went deep into their ROIC math. After I went through the fourth call, it felt like coordinated messaging in the sense that every management team knew exactly what question the market is asking, and they came prepared, maybe even spoke to each other on a common messaging regarding CapEx ROIC.</span></p><p><span>There are a few patterns from this earnings season:</span></p><ul><li><p><span>Nobody is guiding CapEx down, but everyone is now explaining CapEx as a two-part investment: flexible long-lived assets committed early, chips ordered just-in-time against visible demand</span></p></li></ul><ul><li><p><span>AI workloads are still a relatively small % of total cloud revenue (AWS: ~$25B AI run rate vs. $169B total), but they are pulling traditional workloads up with them</span></p></li></ul><ul><li><p><span>Cloud margins at AWS, Azure, and Google Cloud expanded again - the &#8220;AI CapEx has low ROIC&#8221; thesis is basically dead</span></p></li></ul><ul><li><p><span>Microsoft&#8217;s Copilot finally changed momentum, as net adds more than double QoQ</span></p></li></ul><ul><li><p><span>Meta is building four new revenue lines on top of ads: transactional, compute, APIs, and subscriptions</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pa-A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pa-A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pa-A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2527976,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/209648066?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pa-A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!pa-A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d43f7d-912a-4413-9d80-6f9f59a396a4_1408x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Let&#8217;s get into it.</span></p><p><strong><span>Land And power first, chips a few months before you need them</span></strong></p><p><span>I want to start with this theme because it showed up on all four calls and that is the framing of the AI investments.</span></p><p><span>Amy Hood at Microsoft:</span></p><div class="pullquote"><p><span>&#187;A lot of the expense, especially you see it in CapEx, you&#8217;ve seen our CapEx really pivot toward what I would call and do call short-lived assets, which really, right, that CPUs and GPUs that have relatively shorter lead times. And so if the demand environment changes, you just slow down what is, in fact, the largest component, right, and the driver of COGS. The investment into land and data center builds is actually quite flexible, right?&#171;</span></p></div><p><span>Amazon&#8217;s management gave the most detailed version of the same framework:</span></p><div class="pullquote"><p><span>&#187;There are 2 major parts of the investment, the data centers and the servers and networking equipment that go into them. These have different capital cycles. Data center capital is spent starting 2 years before we can put servers into them to start monetizing. Once a data center opens with servers plugged in, we start generating significant revenue right away and then get to monetize these data centers for 30-plus years without having to spend that start-up capital again.&#171;</span></p><p><span>&#187;Servers and networking equipment operate on a shorter cycle. We typically purchase these a few months before putting them into service, so we have strong visibility into customer demand before we trigger the spend. If the demand isn&#8217;t there, we won&#8217;t spend the capital. For servers and networking equipment, on average, it takes a little less than 3 years to break even on that investment. The servers currently have a useful life of at least 5 to 6 years, and most of our AI capacity these days is being contracted for at least 5-year terms.&#171;</span></p></div><p><span>Data center shells monetize for 30+ years. Servers break even in under 3 years, have a 5-6 year useful life, and most AI capacity is contracted on 5-year terms before the servers are even bought.</span></p><p><span>Susan Li at Meta said the same thing on the call:</span></p><div class="pullquote"><p><span>&#187;Therefore, our longer-term capacity strategy aims to give us the flexibility to continue growing compute in 2028 and beyond by laying down data center and network foundations to accommodate future server decisions. The long-lived nature of these assets inherently provides the flexibility that will make it possible to adjust our investment to the pace of AI adoption.&#171;</span></p></div><p><span>And Google&#8217;s CFO Anat Ashkenazi framed the whole thing through an ROIC and pricing lens:</span></p><div class="pullquote"><p><span>&#187;Look, we are working off a disciplined ROIC framework here. Obviously, we, to the extent that our input costs going up to us, we reflect that in our ability to price our solutions and see returns there.&#171;</span></p></div><p><span>The chips are the majority of the CapEx dollar and the fastest-depreciating part, and all four companies are now telling us the chip orders are placed months ahead of deployment, against demand they can already see (Amazon said the lion&#8217;s share of 2027 capacity is already reserved, and &#8220;quite a bit&#8221; of 2028). So the doomsday scenario where hyperscalers wake up with hundreds of millions of stranded GPUs requires demand to disappear inside a one-to-two-quarter ordering window. That is a very different risk profile than the market narrative.</span></p><p><span>The Google comment about pricing is the first hint of what happens if component costs (memory above all - Amazon explicitly raised its CapEx guide from ~$200B to ~$220B because of &#187;the higher cost of memory&#171;) keep rising: prices for cloud and AI services go up. In a supply-constrained market, the hyperscalers have serious pricing power. Ashkenazi even added that Google is more confident on returns than 12 months ago: &#187;And so I think if anything, the dynamics look healthier than where we were about a year ago.&#171;</span></p><p><span>And for those worried about reckless spending and equity issuances, Google was more explicit:</span></p><div class="pullquote"><p><span>&#187;But we also want to make sure we have a resilient, not just growth outlook but also a resilient balance sheet and a strong balance sheet, a healthy balance sheet, which is the rationale behind expanding into the equity markets. At this point, we&#8217;re not planning to go back to the equity markets with the exception of, as you recall, part of our equity offering was the ATM.&#171;</span></p></div><p><strong><span>Google Cloud +82%, Operating Margin from 20.7% to 35.6%</span></strong></p><p><span>Google Cloud revenue grew 82% YoY to $24.8B in Q2 (from $13.6B a year ago), with management noting that &#187;GCP grew faster than Cloud overall&#171;. </span></p><p><span>But the number I care most about is the margin. Cloud operating margin came in at 35.6%, up from 20.7% in Q2 of last year, with operating income roughly tripling YoY. This is the second consecutive quarter where Google Cloud posted significant margin expansion while AI workloads became a bigger share of the mix. If AI workloads had structurally lower margins than traditional cloud workloads, this line would not be expanding this significantly.</span></p><p><span>The backlog also keeps compounding: $514B, up ~$50B sequentially. And, importantly, management repeated the framing that the demand is near term:</span></p><div class="pullquote"><p><span>&#187;We expect to recognize just over 50% of the total backlog as revenue over the next 24 months.&#171;</span></p></div><p><span>On CapEx, Google raised the full-year guide to $195-205B from $180-190B (Q2 CapEx alone was $44.9B, roughly double YoY), and management was direct that the increase is capacity, not inflation:</span></p><div class="pullquote"><p><span>&#187;The increase in the range is primarily due to an acceleration in the delivery of capacity to meet growing demand.&#171;</span></p></div><p><span>Meta and Amazon both called out memory prices as a driver of higher CapEx numbers. Google is saying ours is going up because we are delivering more capacity.</span></p><p><strong><span>Microsoft Azure crosses $100B,  and Copilot finally has momentum</span></strong></p><p><span>Microsoft closed its fiscal year with Microsoft Cloud surpassing $214B in revenue, up 27%, and Azure surpassing $100B for the first time, up 41% for the year. In the June quarter itself, Azure grew 43%, ahead of guidance, and Microsoft guided approximately 45% constant currency growth for the September quarter - meaning Azure is accelerating further.</span></p><p><span>Two numbers from the call deserve special attention. First, this one:</span></p><div class="pullquote"><p><span>&#187;For the full year, our cloud revenue surpassed $214 billion with nearly 90% from customers outside of frontier model companies.&#171;</span></p></div><p><span>The biggest bear case on Microsoft for the past year has been &#8220;it&#8217;s all OpenAI.&#8221; Nearly 90% of a $214B cloud business is not OpenAI.</span></p><p><span>Second, the RPO disclosure, which was unusually detailed this quarter:</span></p><div class="pullquote"><p><span>&#187;Commercial remaining performance obligation grew 84% to $678 billion. All sequential commercial RPO growth was driven by commitments from customers outside of frontier model companies. And RPO increased 25% when excluding OpenAI. RPO, including OpenAI, has a weighted average duration of 2.3 years and roughly 30% will be recognized in revenue in the next 12 months, up 37% year-over-year.&#171;</span></p></div><p><span>A $678B contracted book with a weighted average duration of only 2.3 years, where the portion converting in the next 12 months is itself growing 37% YoY.</span></p><p><span>Now Copilot. I have been fairly critical of the Copilot story over the past two years because seats were growing but engagement was questionable. This quarter, the momentum changed:</span></p><div class="pullquote"><p><span>&#187;When it comes to knowledge work, we now have over 30 million paid Microsoft 365 Copilot seats, with net seat adds more than doubling quarter-over-quarter. Over the last 3 quarters, user satisfaction scores have doubled and are now at an all-time high.&#171;</span></p></div><p><span>For context on the trajectory: 15M paid seats in the December quarter, 20M in March, 30M+ now. Adding 10M+ paid seats in a single quarter is the fastest ramp in the product&#8217;s history. And the engagement data supports the seat data:</span></p><div class="pullquote"><p><span>&#187;The number of conversations per user nearly doubled year-over-year. Average weekly engagement is on par with Outlook and Teams.&#171;</span></p></div><p><span>Engagement on par with Outlook and Teams means Copilot is becoming a habit with users. And the revenue line is following: &#187;Copilot revenue accelerated over 60% quarter-over-quarter.&#171;</span></p><p><span>Microsoft, after two years of trying, finally looks like it is on the right path with Copilot, and I think the E7 suite plus the shift toward usage-based billing (the June GitHub Copilot business model change already improved segment gross margins through the quarter) is a big part of why.</span></p><p><span>One more strategic point that I think is underappreciated. Satya&#8217;s framing of the model layer:</span></p><div class="pullquote"><p><span>&#187;We are building a new model system where the harness, context, memory, and action space are separate from any one model family, thereby moving the frontier on the cost-to-outcome curve. And it&#8217;s not just about cost. It also has the added benefit of business continuity and resilience because every model is substitutable.&#171;</span></p></div><p><span>Combined with the disclosure that customers building with models from multiple providers is up 5x since the start of the year, Microsoft is positioning the orchestration layer, not any single model, as the durable asset.</span></p><p><span>Microsoft is also extending the estimated useful life of its data centers and office buildings from 15 to 25 years. The bigger effect is that more future data center leases will shift from finance leases (which are counted in CapEx) to operating leases (which are not). So reported CapEx optics will change, and the &#187;over $50 billion&#171; CapEx guide for the September quarter already includes this reclassification impact.</span></p><p><span>Even with all this investment, Hood guided full-year operating margins down less than 1 point and said Microsoft expects to remain free cash flow positive in FY27 - which, given what Alphabet&#8217;s and Meta&#8217;s free cash flow just did, is a differentiator.</span></p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong><span>Amazon AWS reaccelerates to 36.7%, but the margin is key</span></strong></p><p><span>AWS grew 36.7% YoY, accelerating for the fifth straight quarter, its fastest growth in 18 quarters - back when, as Jassy pointed out, AWS was less than half its current size. AWS added over $4.6B in revenue quarter-over-quarter, about 80% more than its largest sequential increase ever, and now runs at a $169B annualized revenue rate. Backlog stands at $496B, growing triple digits YoY.</span></p><div class="pullquote"><p><span>&#187;Our chips business now has an annual revenue run rate of over $25 billion, growing triple-digit percentages year-over-year. Our AI revenue run rate climbed significantly quarter-over-quarter, and is now also over $25 billion, growing triple-digit percentages year-over-year.&#171;</span></p></div><p><span>AI is a ~$25B run rate inside a $169B run-rate AWS. Roughly 15% of the business. AI is still small compared to all workloads, but it is accelerating traditional workloads:</span></p><div class="pullquote"><p><span>&#187;We&#8217;re seeing strong growth across both AI and non-AI, what we call core, and growth in one is driving growth in the other. Growth in AI drives core because post-training reinforcement learning and agent tool use is mostly done on CPUs versus AI accelerators.&#171;</span></p></div><p><span>This is the point I keep coming back to: agentic AI doesn&#8217;t just consume GPU cycles. An agent that executes a multi-step workflow hits databases, storage, networking, orchestration, and a lot of plain CPU compute. So the $25B of AI revenue is not just $25B - it is the accelerant for the other $144B. AWS said it directly: </span></p><div class="pullquote"><p><span>&#187;As customers invest in AI, we see a corresponding increase in core consumption. We expect this relationship to strengthen over time as more AI workloads move into full-scale production.&#171; </span></p></div><p><span>And with Graviton used by 98% of the top 1,000 EC2 customers, and Graviton revenue commitments up nearly 3x quarter-over-quarter (Graviton5 ramping nearly 2x faster than Graviton4 did), AWS captures that CPU pull-through on its own silicon economics.</span></p><p><span>Now the margin, which for me was the single most important data point of the entire earnings season. AWS operating income was $16.6B in the quarter at a ~39% operating margin, up 650 basis points YoY. There was a one-off: a ~$600M benefit from fair value changes on energy contracts subject to derivative accounting. The CFO stripped it out: excluding it, margins were still up 520 basis points YoY.</span></p><div class="pullquote"><p><span>&#187;The profitability you&#8217;re seeing from AWS isn&#8217;t random, it&#8217;s a result of disciplined efficiency gains, capacity optimization, which we benefited quite a bit from in Q2, and always closely managing our fixed costs.&#171;</span></p></div><p><span>And on the AI workloads specifically:</span></p><div class="pullquote"><p><span>&#187;We see the margins and returns in AI tracking what we saw with core at the same point of evolution, actually, a little ahead.&#171;</span></p></div><p><span>Sit with that. AWS is in the heaviest investment cycle in its history (Amazon now guides approximately $220B in cash CapEx for 2026, up from the prior ~$200B estimate on higher memory costs), AI is scaling to a $25B+ run rate growing triple digits, and the segment operating margin expanded 520bps YoY excluding one-offs - to a level above where AWS was before the AI cycle even started. The thesis that AI data center CapEx structurally carries low returns is so far being proven wrong.</span></p><p><span>Which is why Jassy felt comfortable saying this:</span></p><div class="pullquote"><p><span>&#187;In fact, the demand we already have for 2028 is striking. And remember, enterprises are still very early in using inference at scale in their current production applications. We long believed AWS could become a few hundred billion-dollar revenue business and now believe it&#8217;ll be at least double that, and very possibly be $1 trillion annual revenue business for us in time, with very appealing accompanying free cash flow and return on invested capital.&#171;</span></p></div><p><span>Two quick additional notes. First, Trainium demand keeps broadening beyond the anchor deals: &#187;In addition to the 2 leading AI labs in the world, Anthropic and OpenAI, making multiyear, multi-gigawatt commitments to Trainium, an increasing number of AI startups are also adopting Trainium&#171; - the list now includes NEURA Robotics, Odyssey, TwelveLabs, Decart, Poolside, plus larger companies like Uber and Pinterest. Second, Jassy repeated his view that &#187;there is not going to be one model to rule the world&#171; and that AWS can be wildly successful without its own frontier model - the Bedrock + SageMaker positioning (technically competent companies building their own smaller models on proprietary data) is the same &#8220;the platform is the moat&#8221; bet Microsoft is making with Foundry.</span></p><p><strong><span>Meta: The ad business expands further and four new revenue lines emerge</span></strong></p><p><span>Meta grew total revenue 28% YoY to $60.8B, with Family of Apps revenue at $60.4B. Instagram reached 2 billion daily actives. The reported operating income optics were messy because of $2.4B in legal charges and $1.18B of severance; excluding those, operating income would have grown 9% YoY per the company&#8217;s own math. CapEx was $31.1B in the quarter (nearly double YoY), and the full-year guide was narrowed upward to $130-145B.</span></p><p><span>The AI-on-the-core-business flywheel is intact:</span></p><div class="pullquote"><p><span>&#187;On Instagram, global time spent this quarter grew double digits year-over-year, largely driven by improvements to our feed and reels recommendations. On Facebook, video time spent increased 9% globally year-over-year and over 10% within the U.S. and Canada, where it was driven by ranking improvements. We are finding that LLMs are increasingly capable of delivering ranking and recommendations gains.&#171;</span></p></div><p><span>But the reason I think this Meta quarter deserves a separate look is not the ad business. It&#8217;s that Meta laid out a concrete map of new revenue lines beyond advertising, and it&#8217;s worth listing them explicitly because I believe several of them will be reported segments within two years:</span></p><ol><li><p><span>Transactional / business services revenue. This is the most Meta-native one. Zuck:</span></p></li></ol><div class="pullquote"><p><span>&#187;And soon, it will go further, including suggesting ways to grow your business, giving you competitive intelligence and real-time insights into what&#8217;s working and what&#8217;s not. And over time, we&#8217;d like to build this into a business-in-a-box service that can help you start and run a whole business using Meta&#8217;s platforms. In terms of how we will monetize these, we have a mix of subscriptions, volume-based pricing, and I expect that we&#8217;re going to continue to evolve more of these products to be like our ad systems where businesses only pay us when we achieve results for them.&#171;</span></p></div><p><span>&#8220;Businesses only pay us when we achieve results&#8221; is the ad auction logic extended to commerce and operations.</span></p><ol start="2"><li><p><span>Compute revenue - direct.</span></p></li></ol><div class="pullquote"><p><span>&#187;We&#8217;re getting a lot of offers for compute at a significant premium over what we paid for it.&#171;</span></p></div><p><span>And the mechanism Zuck laid out is genuinely novel:</span></p><div class="pullquote"><p><span>&#187;Over time, that will let us run an efficient auction over our compute, similar to how we do that for advertisers today.&#171;</span></p></div><p><span>An auction for compute. Meta runs one of the most sophisticated real-time auctions in the world for ad impressions; running the same machinery over GPU/accelerator capacity is a very natural extension. And the commentary about the trade-off between monetizing today versus building for the future </span>reads, to me, like Meta is preparing a direct compute deal with a third party in the near future.</p><div class="pullquote"><p><span>&#187;</span>Now in terms of running the business, obviously, a common trade-off that we need to make is around how much do you monetize something today versus develop future assets for the future. And I think that it&#8217;s always a portfolio, right? It&#8217;s not like you don&#8217;t want to only do long-term things and do no -- like -- and not kind of prove the markets have that exist in the near term, but I also think it would be foolish to basically just sell all of the compute and take a short-term profit<span>&#171;</span></p></div><ol start="3"><li><p><span>Enterprise / API revenue. </span></p><p></p><div class="pullquote"><p><span>&#187;We see a large enterprise opportunity to sell to businesses, including APIs, business agents, potentially selling compute directly and other services that we&#8217;re building for large customers.&#171; </span></p></div><p><span>Business agents are already being used by over 1 million businesses weekly.</span></p></li></ol><ol start="4"><li><p><span>Consumer subscriptions. First evidence it works: Family of Apps &#8220;other&#8221; revenue crossed $1B in a quarter for the first time, up 73% YoY, driven primarily by WhatsApp paid messaging and subscriptions, but the potential here for expansion is enormous as Meta hasn&#8217;t laid out the killer products yet.</span></p></li></ol><p></p><p><span>And the distribution advantage that ties all four together:</span></p><div class="pullquote"><p><span>&#187;I expect it to become a lot easier to ship new apps. So we are planning to build out more ideas and use our recommendation systems to scale them to the people who will find them interesting, as we&#8217;ve done with Threads.&#171;</span></p></div><p><span>Threads went from zero to a top-tier social app almost entirely on Instagram&#8217;s distribution. Meta is telling us it plans to run that playbook a dozen more times in the AI application layer, because AI is redistributing the playing field and shipping apps is getting cheap.</span></p><p><span>The last quote I want to highlight is Zuck on why Meta builds the full stack rather than renting frontier models:</span></p><div class="pullquote"><p><span>&#187;A lot of people view the surface layer of we build some social media apps and we have an ad business. We are really a full-stack technology company. We built our own data centers, our own infrastructure, our own chips, our own low-level software. ... It just seems to me pretty clear that having kind of sovereignty over building your own models is going to be an important part of that stack going forward.&#171;</span></p></div><p><span>Zuck is saying that relying on another lab&#8217;s model is a policy and business risk Meta will not take, and that full-stack ownership enables product experiences others mostly cannot build. You can agree or disagree on the cost of that choice (a substantial amount of Meta&#8217;s compute goes to training to remain a leading lab, which is exactly why the free cash flow line looks the way it does right now), but it is a fair strategy, and Zuck closed the loop on it himself:</span></p><div class="pullquote"><p><span>&#187;I get that this is sort of a big bet across the industry. My personal bet is that the people who invest in this are going to be rewarded and feel very good over time.&#171;</span></p></div><p><span>He said almost exactly the same thing in the 2022-2023 drawdown. That worked out more than okay.</span></p><p><strong><span>My views on these companies going forward</span></strong></p><p><span>I think</span></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q2-earnings">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Q2 2026 Channel Checks & Alternative Data: Cloud, Ad Tech and Software]]></title><description><![CDATA[I covered the cloud hyperscalers and what I see from aggregating multiple alt providers, and share some of my views on the upcoming earnings. In this report, I also share some data on the ad tech and software sector.]]></description><link>https://www.uncoveralpha.com/p/q2-2026-channel-checks-and-alternative</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/q2-2026-channel-checks-and-alternative</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Fri, 17 Jul 2026 13:14:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Msas!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18ab5b48-45a0-4993-bc2e-603548401d4b_751x415.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>Posting my regular channel check &amp; other alternative data report before we start the big tech Q2 earnings.</p><p>In this report, I covered the cloud hyperscalers and what I see from aggregating multiple alt providers, and share some of my views on the upcoming earnings. In this report, I also share some data on the ad tech and software sector, as well as my general view on how things might play out this earnings season.</p><p>Particularly on the cloud side, some of my data points to a significant acceleration from previous levels for one specific cloud provider.</p><p>Let&#8217;s dive in.</p><p><strong>Cloud hyperscalers</strong></p><p>First of all, I should note that cloud capacity remains sold out at the hyperscalers, and alt data continues the trend seen in previous quarters. If we first look at the sellside data, we can see that sellside is optimistic about cloud revenue growth for Q2, with expectations mostly on the high end of the guidance ranges provided by cloud providers. By conducting my analysis of the most relevant expert interviews and channel checks from platforms like AlphaSense, we get more granular data.</p><p>The sentiment from expert interviews this quarter is very bullish, with public cloud growth accelerating substantially in Q2 2026 vs Q1 2026. In many cases, the growth estimates went from 10-20% YoY in Q1 2026 to 25-35% YoY for Q2 2026. Strong usage is seen across the board, but the standout company for this quarter in terms of rate of change from the previous quarter is </p><p>(<a href="https://www.uncoveralpha.com/subscribe">Subscribe to Paid</a> to see the full article, explaining which cloud provider momentum looks the best in Q2, Microsoft Copilot adoption, and Ad Tech growth, particularly Meta) </p>
      <p>
          <a href="https://www.uncoveralpha.com/p/q2-2026-channel-checks-and-alternative">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Why Token Optimization Is a Gift to the Hyperscalers]]></title><description><![CDATA[The shift away from always-buy-the-best-model is, on its surface, bearish for the AI labs and looks like it should compress the whole stack. But it is quietly one of the most bullish structural setups for the three hyperscalers &#8212; Microsoft, Amazon, and Google.]]></description><link>https://www.uncoveralpha.com/p/why-token-optimization-is-a-gift</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/why-token-optimization-is-a-gift</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:55:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!twBl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Hey everyone,</span></p><p><span>A few weeks ago I wrote a piece called </span><a href="https://www.uncoveralpha.com/p/most-of-the-economy-wont-run-on-the"><span>Most of the Economy Won&#8217;t Run on the Best Model</span></a><span>, where I argued that the AI market will bifurcate: the frontier model goes to the small slice of work where intelligence is unbounded in economic value (drug discovery, novel math, the hardest agentic reasoning), and the middle of the economy &#8212; classification, extraction, summarization, routine code, support &#8212; runs on the cheapest model that clears the quality bar.</span></p><p><span>This article is the sequel to that one. Who captures the value when the world switches from token maxing to token optimization?</span></p><p><span>The shift away from always-buy-the-best-model is, on its surface, bearish for the AI labs and looks like it should compress the whole stack. But it is quietly one of the most bullish structural setups for the three hyperscalers &#8212; Microsoft, Amazon, and Google. </span></p><p><span>Let me explain why.</span></p><p><span>Think about a highway toll road. There are two businesses operating on it. The first is the company that manufactures the cars &#8212; they make the actual machine that does the work of getting you somewhere, and there&#8217;s a real margin in a car. The second is the company that owns the tollbooth. The tollbooth owner doesn&#8217;t care what car you drive. Ferrari, Toyota, or a 12-year-old used Honda. Every one of them pays the same toll to cross the bridge. Now imagine a world where, suddenly, everyone realizes they were commuting to work in a Ferrari for no reason, and they all downgrade to the cheaper Honda to save money. The carmaker&#8217;s revenue per vehicle collapses, but people don&#8217;t drive less because they switched to a cheaper car. They drive more, because now it&#8217;s cheap enough to justify trips they&#8217;d never have taken before. And every one of those extra trips still crosses the bridge. The tollbooth&#8217;s revenue goes up.</span></p><p><span>In the AI stack, the AI labs are the carmakers. The hyperscalers are the tollbooths. And in the next months, we are entering a period where most of the economy is trading down from the Ferrari to the Honda, while simultaneously driving 10x more miles.</span></p><p><strong><span>From token maxing to token optimization</span></strong></p><p><span>For the last 18 months, the dominant behavior in enterprise AI was token maxing. You found the single best model on the leaderboard, you pointed every workload at it, and you didn&#8217;t think too hard about cost, because the whole thing was a pilot and the bill was small relative to the perceived upside of &#8220;does this even work?&#8221;</span></p><p><span>That era is ending fast. Companies are essentially blowing through their AI budgets in a quarter. Altman said recently that the &#8220;my company spent my entire 2026 budget in Q1, can you make this more efficient?&#8221; complaint went from something that &#8220;never came up&#8221; to &#8220;all of a sudden a huge issue.&#8221; And it&#8217;s not just him; you can see it across the industry, from companies like Salesforce and Meta to many other smaller ones blowing their planned yearly budget in a matter of days.</span></p><p><span>The natural behavior change from this is that instead of having one model for everything, companies start routing: a small, cheap, often open-weight model handles the 80% of requests that are less complex, and only the genuinely hard requests escalate to the frontier. This is token optimization, and it doesn&#8217;t reduce total token consumption &#8212; it accelerates it. The moment inference gets cheap enough, you stop rationing it. You run the agent in a loop. You let it read the whole codebase. You re-run it five times and vote on the answer. A single coding-agent session now chews through millions of tokens of context where a chatbot query used a few thousand.</span></p><p><span>You can see this in the hard numbers the hyperscalers themselves disclose. Microsoft said it processed over 100 trillion tokens in a single quarter in 2025, up 5x year-over-year, with a record 50 trillion in one month alone. By its fiscal Q3 2026 call, Microsoft said over 300 customers were on track to process more than a trillion tokens each on Foundry this year &#8212; and that this was accelerating 30% quarter-over-quarter. Google went from 480 trillion tokens per month at I/O in May 2025, to 980 trillion by July, to 1.3 quadrillion by October &#8212; and in its Q1 2026 filing disclosed that its first-party models alone were processing more than 16 billion tokens per minute via direct API, up 60% in a single quarter.</span></p><p><span>At the same time, the price per unit of capability is falling substantially&#8212; the Stanford HAI AI Index found inference cost for GPT-3.5-level performance fell more than 280-fold in two years, and a16z pegs the decline at roughly 10x per year for any fixed capability level. And yet, total tokens processed are growing several-fold per year.</span></p><p><strong><span>Where the margin lives in a routed token economy</span></strong></p><p><span>When a company calls the SOTA model directly through, say, an AI lab&#8217;s first-party API, the lab captures the full economic rent of the token. There are many industry reports surfacing around the margins that providers like Anthropic have right now, and many of them claim that the margins went substantially up this year to as high as 70% gross margins. That price embeds the lab&#8217;s R&amp;D, its brand, its leaderboard position &#8212; call it the &#8220;model-provider margin.&#8221;</span></p><p><span>And to be fair, when token optimization happens, there might not be a direct hit to margins in the short-term at these frontier labs (as there will always be use cases for the best AI model), but it does mean that growth acceleration becomes smaller than it would be if everyone stayed for every use case on the frontier path.</span></p><p><span>The company routes that same workload to an open-weight model &#8212; GLM 5.2, DeepSeek V3.2, Qwen3 Coder, Kimi K2.5, Llama, MiniMax M2.5 &#8212; running on a hyperscaler&#8217;s managed inference. The &#8220;model-provider margin&#8221; essentially goes to zero, because the model is open-weight and nobody is charging a brand premium for it. But the token still has to run on somebody&#8217;s GPUs, inside somebody&#8217;s data center, behind somebody&#8217;s managed API with its security, compliance, logging, and SLAs. And that somebody are the hyperscalers. They still charge their full infrastructure margins on the token. AWS has historically run a roughly 35&#8211;38% operating margin. Google Cloud, which lost money for years, posted operating margins above 33% in Q1 2026 and is still climbing. That margin doesn&#8217;t care whether the token came from a $50/million frontier model or a $1/million open-weight model. To return to the analogy that from before, the tollbooth charges the same toll regardless of the car.</span></p><p><span>So here is the structural beauty of it for the hyperscalers, and the structural danger for the labs:</span></p><ul><li><p><span>Per-token economics compress, but the compression lands almost entirely on the model layer, not the infrastructure layer. The AI lab&#8217;s margins on simple workloads get squeezed. The hyperscaler&#8217;s infrastructure margin is sticky.</span></p></li><li><p><span>Total token volume explodes, and almost every single token crosses the hyperscaler&#8217;s tollbooth</span><strong><span>.</span></strong><span> More usage, on a partly-depreciated, increasingly efficient installed base, means absolute infrastructure revenue and gross profit dollars go up even as the price of any individual token falls.</span></p></li></ul><p><span>This is Jevons&#8217; paradox pointed directly at the cloud P&amp;L. The hyperscalers squeeze more revenue out of the infrastructure they already own, and they don&#8217;t need the model-provider margin to do it. Outside of Google, they were never in that business in the first place.</span></p><p><strong><span>The orchestration layer brings real value</span></strong></p><p><span>If the future is multi-model &#8212; and it clearly is &#8212; then someone has to own the orchestration: the layer that decides which model handles which request, holds the fine-tunes, runs the agent loop, manages memory and tool-calling, handles fallback when a model is down. This layer can bring immense value to those who own it, and all three hyperscalers are naturally positioned to become just that.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!twBl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!twBl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!twBl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!twBl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!twBl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!twBl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2611433,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/204108337?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!twBl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!twBl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!twBl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!twBl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb9b27bc9-4200-4e7f-af9d-6dc701b66826_1408x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Amazon&#8217;s Bedrock is the clearest example. It&#8217;s no longer &#8220;a place to call Claude.&#8221; As of 2026, the catalog spans 18 providers and 110+ individually addressable model variants, and AWS bolted on Intelligent Prompt Routing, which automatically routes each request to the cheapest model in a family that can handle it &#8212; AWS claims up to 30% cost savings with no accuracy loss. On top of that sits Bedrock AgentCore, a production agent harness with built-in Runtime, Memory, Gateway, Browser, Identity, and Observability. Bedrock reached a multi-billion-dollar annualized run rate with customer spend growing 60% quarter-over-quarter across 100,000+ customers.</span></p><p><span>Microsoft&#8217;s Azure AI Foundry and Google Vertex are the same play, although Google&#8217;s preferred scenario is where the Gemini model family dominates the workloads.</span></p><p><span>The natural position for hyperscalers to be the orchestration layer is great because companies already have their data, security perimeter, billing, and compliance live in the cloud environments. They can pick from a menu of models.&#8221; The lab moves from the front end &#8212; the thing the customer chooses and bonds with &#8212; to the back end, an interchangeable component sitting behind the hyperscaler&#8217;s harness. Customers also start building the fine-tunes, the RAG pipeline, the agent scaffolding, the eval suite at the cloud orchestration layer, not at the model layer; switching models becomes something normal and not a migration task. One of my readers, who runs his own harness, put it perfectly in the comments of my last piece: the harness is worth as much as the model itself, and you&#8217;re paying for it either way.</span></p><p><span>The moat for the labs becomes shallow with most use cases but still remains important at the most complex ones (the ones where economic upside is not capped). But for hyperscalers, they benefit in all use cases that run the economy.</span></p><p><span>One might argue why companies couldn&#8217;t move even outside of the cloud, but the problem is here that besides the normal reasons that people migrated a big part of workloads from on-prem to cloud (easier to scale, use, collaborate, manage), now on top of those reasons, you also have the problem compute shortages where infrastructure outside of cloud environments is even harder to get and managing that infrastructure becomes a much more difficult job. But an important aspect, I think, that is forming is also the cybersecurity one. Given the recent developments here, it is clear that only a handful of companies will have first-row access to the most capable cyber models, to first shield their systems before the model is then available for everyone else. If you are an enterprise, you want safety, and it seems the only way for you to get it is if you have your infrastructure at one of those preferred partners that get first-row access to cyber-capable models before everyone else does. The hyperscalers are those partners.</span></p><p><span>The hyperscalers&#8217; job, then, is almost simple to state: enable as many models as possible, make routing and fine-tuning and observability frictionless, and let customers optimize tokens to their heart&#8217;s content. Every model they add makes their tollbooth more valuable.</span></p><p><strong><span>The government approval window just made this worse for the labs</span></strong></p><p><span>Now layer on a development from June that I don&#8217;t think the market has connected to this thesis at all: the June 2, 2026 executive order: Promoting Advanced Artificial Intelligence Innovation and Security.</span></p><p><span>The headline version is that it sets up a voluntary framework where developers of &#8220;covered frontier models&#8221; &#8212; designation determined by the NSA via a classified benchmarking process focused on cyber capabilities &#8212; give the federal government access to the model for up to 30 days before release to &#8220;other trusted partners.&#8221; (The earlier draft had a 90-day window; it was cut to 30 to avoid blunting U.S. competitiveness.) The order pointedly does not create mandatory licensing &#8212; but it establishes a structured pre-release evaluation pathway, an AI cybersecurity clearinghouse run out of Treasury, and a government role in selecting which &#8220;trusted partners&#8221; get early access. The catalyst, per CFR&#8217;s reporting, was rising concern about models like Anthropic&#8217;s Claude Mythos being able to autonomously find and exploit software vulnerabilities. Commerce ordered Anthropic to cut off non-U.S. access to its Mythos 5 and Fable 5 models on export-control grounds, and those models were pulled from Bedrock days after launch. Now there is growing concern that, even when models are released to the public, U.S. citizens may get access before people from other countries. Hyperscalers have clients worldwide, so it becomes even more important for them to offer multiple models to companies and for companies to have an orchestration layer that isn&#8217;t just one AI lab, since they will have to juggle geopolitical compliance as well.</span></p><p><span>It strengthens the case </span></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/why-token-optimization-is-a-gift">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Memory per Token Optimizations. Where are we, and how much more room do we have?]]></title><description><![CDATA[Software optimization techniques for lowering HBM usage and give my estimates on how much of that optimization is already captured, how much reduction these optimizations could still offer, hardware angles from SRAM accelerators like Groq and Cerebras, as well as the decode and prefill split up of hardware, and what that opens up.]]></description><link>https://www.uncoveralpha.com/p/memory-per-token-optimizations-where</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/memory-per-token-optimizations-where</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 11 Jun 2026 13:12:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q2-s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>In the AI compute buildout phase, especially for inference, which is where the industry is shifting most of its capex and where the actual revenue gets generated, the binding constraint is increasingly not compute. It&#8217;s memory. Capacity and bandwidth.</p><p>I covered the supply side of this story in my memory cycle article &#8221;<a href="https://www.uncoveralpha.com/p/every-memory-cycle-ends-the-same">Every Memory Cycle Ends the Same. Until It Doesn&#8217;t</a>,&#8221; where I argued that HBM has turned memory from a gadget component into a raw input for intelligence. In that article, I also wrote that the real risk for the memory cycle is a technical breakthrough that would require orders of magnitude less memory. Today&#8217;s article is about the other side of that exact coin: what the labs and inference providers are doing in software and model architecture to need less memory per token, how much of that optimization potential is already captured, and what it means for the hardware stack, including some unconventional setups like using depreciated H100s and A100s as dedicated decode machines.</p><p>The reason this matters now and not in two years is simple: companies are starting to hit their token spend limits. Agentic workloads (coding agents, research agents, computer-use agents) consume tokens at a rate that makes the chatbot era look like a rounding error. A single coding agent session can chew through millions of tokens of context and companies are increasingly becoming frustrated with it.</p><p>In this article, I cover software optimization techniques for lowering HBM usage and give my estimates on how much of that optimization is already captured, give my view on the best solution, and how much reduction it could offer, and also cover hardware angles from SRAM accelerators like Groq and Cerebras, as well as the decode and prefill split up of hardware and what that opens up.</p><p>Let&#8217;s start.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q2-s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q2-s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q2-s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2009669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/201591402?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q2-s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Q2-s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F328e7b20-36b8-4545-85e3-1a6423897e9e_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Prefill and decode: the two jobs inside every AI request</strong></p><p>To understand why memory is the bottleneck, you first need to understand that every LLM request is actually two completely different workloads stapled together: prefill and decode.</p><p><strong>Prefill</strong> is what happens when you send the model your input: the prompt, the document, the codebase, the conversation history. The model reads all of it and processes every input token in parallel, in one big pass. Think of it as an analyst who gets handed a 300-page data room before a meeting. He reads the whole thing in one sitting, takes structured notes on every page, and files those notes away. This is brute-force work; the limiting factor is how fast his brain works, not how fast he can pull pages out of the binder. In GPU terms, prefill is compute-bound: the chip&#8217;s arithmetic units (the FLOPs) are the bottleneck, and the memory system can keep up.</p><p><strong>Decode</strong> is what happens when the model generates its answer, one token at a time. To generate each single token, the model has to read essentially all of its weights from memory, plus all of the notes it took during prefill (the so-called KV cache, more on this later in the article). Then it does a comparatively tiny amount of math, produces one token, and does the entire memory read again for the next token. Our analyst is now in the meeting, and before he speaks each individual word, he has to re-skim his entire stack of notes. The bottleneck is no longer his brainpower. It&#8217;s how fast he can flip pages. Decode is memory-bandwidth-bound.</p><p>You can put numbers on this. An Nvidia H100 delivers roughly 989 TFLOPS of dense BF16 compute against 3.35 TB/s of HBM3 bandwidth. That ratio means the chip needs to perform roughly ~295 floating point operations for every byte it pulls from memory just to keep its compute units fed. Prefill, processing thousands of tokens in parallel, easily clears that bar. Decode doesn&#8217;t come close: roofline analyses of autoregressive generation put its arithmetic intensity at roughly 1 FLOP per byte, about two orders of magnitude below the compute-bound ridge point. In plain English: during decode, the most expensive compute engines on the planet sit idle 95%+ of the time, waiting for memory.</p><p>The simplest illustration: a 70B parameter model in FP16 is ~140GB of weights. On an H100 with 3.35 TB/s of bandwidth, the theoretical single-user decode ceiling is about 3,350/140 = 24 tokens per second. On an A100 with 2 TB/s, it&#8217;s about 14 tokens per second. Notice what&#8217;s not in that equation: FLOPs. You could double the H100&#8217;s compute and single-stream decode speed wouldn&#8217;t move at all. The only lever is bandwidth, or reading fewer bytes.</p><p>This is the &#8220;memory wall,&#8221; and it&#8217;s also why a lot of inference revolves around batching. If reading 140GB of weights produces one token for one user, that&#8217;s terrible economics. If the same 140GB read produces one token each for 200 users simultaneously, your cost per token just dropped ~200x. The weights are read once, amortized across the batch. So the game every inference provider plays is: cram as many concurrent users as possible onto each GPU. And what limits how many users you can cram on? Memory capacity. Because every user brings their own luggage.</p><p><strong>The luggage: KV cache, and why agents made it explode</strong></p><p>That luggage is the KV cache. During prefill, the model stores intermediate &#8220;key&#8221; and &#8220;value&#8221; representations of every token in the context, so that during decode it doesn&#8217;t have to re-process the whole prompt for every new token. Those are the analyst&#8217;s notes. The catch: the notes grow linearly with context length, and they have to sit in the same precious HBM as the model weights.</p><p>In a standard transformer, Llama 3.1 405B needs 516 KB of KV cache per token of context; Qwen-2.5 72B needs 327 KB per token. Run that forward: a single user with a 128K-token context on a 70B-class model is carrying roughly 40GB of KV cache. At one million tokens of context, a Llama-70B-scale model would need ~330GB in BF16 for the KV cache alone, which doesn&#8217;t fit in any single GPU. For reference, the H100 has 80GB total. The B300 has 288GB.</p><p>In the chatbot era, this was manageable because most conversations were a few thousand tokens. Agentic workloads change that. An agent doing a long coding task holds the repo, the tool outputs, the execution traces, the full plan, all of it, in context, for hours. Context lengths of 100K-1M tokens went from research demo to daily production workload. And every one of those tokens occupies HBM for the entire duration of the session. The KV cache, not the model weights, becomes the dominant consumer of memory. Which means batch sizes collapse, which means the weight-read amortization collapses, which means cost per token explodes.</p><p>So what is the industry doing about it? A lot, actually. Let&#8217;s go through the software optimization stack, and, importantly for investors trying to model how much efficiency is still on the table, my estimate of how much of each technique&#8217;s potential has already been captured.</p><p><strong>The optimization stack: where we are on each curve</strong></p><p>A quick framing note: the percentages below are my own estimates of &#8220;captured potential&#8221; at the frontier (the major labs and serious inference providers), based on what&#8217;s publicly documented. The long tail of enterprise deployments is far behind the frontier on all of these, which is itself an investment-relevant point: there is a lot of &#8220;free&#8221; efficiency still sitting unused in corporate AI deployments.</p><p><strong>1. Continuous batching </strong></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/memory-per-token-optimizations-where">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Most of the Economy Won't Run on the Best Model]]></title><description><![CDATA[It&#8217;s a thesis about where the money in AI actually goes once we transition to scaled-up AI workloads and how that might look different from today&#8217;s expectations.]]></description><link>https://www.uncoveralpha.com/p/most-of-the-economy-wont-run-on-the</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/most-of-the-economy-wont-run-on-the</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 04 Jun 2026 12:49:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!TR1s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>I want to share some of my thoughts on why I think a significant change is going to happen in the AI industry, and the market is blind to it. It&#8217;s a thesis about where the money in AI actually goes once we transition to scaled-up AI workloads and how that might look different from today&#8217;s expectations.</p><p>Let me start with an analogy that I keep coming back to.</p><p>When a company hires an accountant, it does not go out and hire a PhD in pure mathematics to reconcile the ledgers. Not because the PhD couldn&#8217;t do it &#8212; they obviously could, and probably faster &#8212; but because it makes no economic sense. The PhD is overqualified, which is just another way of saying they are too expensive for the value the task produces. The economic output of bookkeeping is capped. There is only so much upside in getting the books done. So you hire the cheapest person who clears the quality bar, and you pocket the difference. And you can take this analogy and apply it to multiple other jobs.</p><p>Now flip it. If you are running a drug-discovery program, you absolutely want the PhD &#8212; in fact you want five of them, plus a Nobel laureate consulting on the side. Why? Because the economic output of a single discovery is enormous, almost unbounded. The expected value of a breakthrough is measured in tens of billions, so the cost of the smartest possible person working on it rounds to zero against the prize. Here, intelligence is the only thing that matters, and cost is an afterthought.</p><p>This is, I think, exactly how the AI model market is going to bifurcate. And we are right at the inflection point where it starts to happen.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TR1s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TR1s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TR1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1801680,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/200606531?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TR1s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TR1s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51126957-2f3d-4f00-be4b-a267700e792b_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The metric is not intelligence. It&#8217;s intelligence per dollar</strong></p><p>Today, essentially everyone uses the state-of-the-art (SOTA) model for everything. You want to summarize an email? SOTA model. Classify a support ticket? SOTA model. Extract three fields from an invoice? SOTA model. We do this for one simple reason: the frontier models have only just crossed the threshold of being broadly truly impactful for knowledge work, and when something has only just started working, you reach for the best version of it you can find. You don&#8217;t optimize cost on a capability you weren&#8217;t sure you had last quarter.</p><p>But I believe this is a transitional behavior, not a stable equilibrium. And there are a few reasons for that.</p><p>The Stanford HAI AI Index found that the inference cost for a system performing at GPT-3.5 level dropped more than 280-fold between November 2022 and October 2024 &#8212; from roughly $20 per million tokens to about $0.07. Andreessen Horowitz, looking at the same phenomenon across the whole performance spectrum, concluded that for a model of equivalent performance, cost falls by roughly 10x every year &#8212; faster than compute fell during the PC revolution, faster than bandwidth fell during the dotcom build-out. Epoch AI, slicing it by benchmark, found the price to hit GPT-4-level performance on PhD-level science questions fell by about 40x per year, with the range across benchmarks running anywhere from 9x to 900x annually. We just had Sara Friar, OpenAI&#8217;s CFO on the All in conference say the following:</p><div class="pullquote"><p>&#8220; The good news on compute is that there is a massive deflationary curve on cost, right? From ChatGPT... uh, [GPT] 4 to 5.4, I think the deprecation of cost was something like 97%. It&#8217;s like kind of an amazing curve, actually..but that happened in like two years. &#8220;</p></div><p>Pick whichever number you find least aggressive. They all say the same thing: the capability you are paying a premium for today becomes nearly free in about a year.</p><p>On top of it, many companies and enterprises are starting to burn through their annual planned token consumption in just a few months. This is a trend that is accelerating, and I have been hearing it all across the industry. Yesterday, there was a comment published from Sam Altman saying:</p><div class="pullquote"><p>&#8220;Probably the second biggest theme is around cost. People are really saying, that&#8217;s kind of become a meme now, but &#8220;my company spent my entire 2026 budget in Q1. Can you make this more efficient?&#8221;...that went from at the beginning of this year, an issue that never came up - I know people were totally happy with the amount they were spending - to all of a sudden a huge issue&#8221;.</p></div><p>What this means is that companies have no choice but to optimize costs, and that will soon mean using models other than the SOTA model for specific tasks.</p><p>The second force is that the frontier itself is getting smaller, not just cheaper. Epoch AI has pointed out that frontier models are now roughly an order of magnitude smaller in parameter count than GPT-4 was, because once inference becomes the dominant cost, you stop training huge models and start over-training small ones on far more data. Distillation compounds this: a teacher model&#8217;s capability gets compressed into a student, a fraction of its size. This is exactly what Meta is doing internally when using their AI models to power their ads and content platform. The student model is the one that is applied at scale, and it distilled its knowledge from the teacher model.</p><p>So putting all of these together. We have a rapidly falling price for any given level of capability and frontier that is already shrinking in size in terms of what is actually being deployed, and we have companies burning through their annual token budgets in a matter of months.</p><p>As such, I believe that for the overwhelming majority of economically valuable knowledge work, the correct model is not the SOTA model. It&#8217;s the cheapest model that clears the task&#8217;s quality bar. And as pilots move into full production (which is the stage we are in today) &#8212; where you&#8217;re suddenly paying for millions or billions of tokens a day instead of running a demo &#8212; intelligence-per-dollar becomes the only metric that survives contact with a CFO.</p><p>At the same time, the SOTA model and its use case don&#8217;t disappear. It goes where the economic ceiling is unbounded: frontier R&amp;D, drug discovery, novel mathematics, and the hardest agentic reasoning chains. But that is a smaller slice of the token volume in terms of our current economy. The accountant&#8217;s quadrant &#8212; classification, extraction, summarization, routine code, customer support, the boring profitable middle of the economy &#8212; is where the majority of tokens actually are, and that quadrant is going to run on cheaper, distilled, often fine-tuned, frequently &#8220;older&#8221; models.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong>The investment angle: the money moves to the owners of installed compute</strong></p><p>If the thesis above is right, where does the capital go?</p><p>The intuitive answer &#8212; the one the market is currently screaming &#8212; is &#8220;buy the picks and shovels.&#8221; Buy Nvidia, buy Broadcom, buy the ASIC co-designers, buy memory,  anything that sells new compute. But my view is that while that sector still might do well, there is a different part of the tech stack that will benefit even more, looking at the rate of change from the current state.</p><p>The sellers of new compute (semis) are only winners in a world of continued high-cadence spending on new compute. And my thesis specifically questions whether that cadence is necessary. So let me lay out the two states the world can be in, because the asymmetry between them is the whole argument.</p><p><strong>Scenario 1: Capex falls or stabilizes.</strong> If you can squeeze an order of magnitude more useful tokens out of the hardware you already own &#8212; because models got smaller, cheaper, more efficient and verticalized &#8212; then you no longer need to spend $100bn+ every single year just to stay relevant. In this world, the owners of the installed base win and the sellers of new compute lose. Hyperscaler free cash flow inflects sharply upward, because capex was the one thing suppressing it. Multiples re-rate higher as the cloud business converts from a capex incinerator into a cash machine running largely paid-for, partly-depreciated hardware. And the semis de-rate, because the market finally realizes the upgrade treadmill has slowed.</p><p><strong>Scenario 2: Capex stays high &#8212; and revenue explodes.</strong> This is the Jevons-paradox-on-steroids case. Demand is so strong that hyperscalers do both: they extract enormous output from cheap, long-lived existing hardware and keep buying new gear. Here everyone wins at once &#8212; but the hyperscalers win more, because their incremental revenue now lands on a cost base that is partly depreciated and dramatically more efficient per token. Operating leverage goes vertical.</p><p>The interesting thing is that the market is currently priced for neither. The market is currently pricing only a future in which CapEx continues to grow for the foreseeable future and the semiconductor industry benefits, but at the same time, the hyperscalers are making a losing bet with spending on this CapEx, as the market is questioning the return on that spend.</p><p>I think this market premise is very wrong, as we are actively transitioning to production-scale AI workloads where the economics are different from those in the pilot world, where we mostly lived for the last few months.</p><p>As always, I hope you found this article valuable. I would appreciate it if you could share it with people you know who might find it interesting. I also invite you to become a paid subscriber, as paid subscribers get additional articles covering both big tech companies in more detail, as well as mid-cap and small-cap companies that I find interesting.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe&quot;,&quot;text&quot;:&quot;Subscribe to Paid&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe"><span>Subscribe to Paid</span></a></p><p>Thank you!</p><p><strong>Disclaimer:</strong></p><p>I own hyperscaler Meta (META), Amazon (AMZN), Microsoft (MSFT), Google (GOOGL) stock.</p><p>Nothing contained in this website and newsletter should be understood as investment or financial advice. All investment strategies and investments involve the risk of loss. Past performance does not guarantee future results. Everything written and expressed in this newsletter is only the writer&#8217;s opinion and should not be considered investment advice. Before investing in anything, know your risk profile and if needed, consult a professional. Nothing on this site should ever be considered advice, research, or an invitation to buy or sell any securities.</p>]]></content:encoded></item><item><title><![CDATA[The Harness: The Moat for AI Model Providers?]]></title><description><![CDATA[The real moat that the frontier labs are building right now is not the model. It is the harness around the model and the cost of serving the model.]]></description><link>https://www.uncoveralpha.com/p/the-harness-the-moat-for-ai-model</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-harness-the-moat-for-ai-model</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Fri, 22 May 2026 14:03:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aul4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>I want to walk through what is, in my opinion, an important structural shift happening in AI right now.</p><p>For years, the way we value AI model providers is via benchmarks based on the raw performance of these models. The narrative has been &#8220;whoever has the best benchmark model wins&#8221;</p><p>I think this framing is increasingly becoming wrong. The real moat that the frontier labs are building right now is not the model. It is the harness around the model and the cost of serving the model. In this article, I will focus on the harness. And once you understand what a harness is, you start to see why Anthropic&#8217;s enterprise revenue keeps compounding even when its raw benchmark scores are not always the best, and why the labs that own both the model and the harness are setting themselves up for the kind of platform lock-in that historically produced high gross margin businesses.</p><p>Let start.</p><p><strong>What is a harness?</strong></p><p>You can picture a harness as this entire system that wraps the model and turns it from a text generator into something that can do real work. Tools, memory, system prompts, permission policies, sandboxes, subagent dispatch, context management, the agent loop itself.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aul4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aul4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 424w, https://substackcdn.com/image/fetch/$s_!aul4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 848w, https://substackcdn.com/image/fetch/$s_!aul4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 1272w, https://substackcdn.com/image/fetch/$s_!aul4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aul4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png" width="953" height="828" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e7c689d9-6577-4899-aa84-70bac90561a3_953x828.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:828,&quot;width&quot;:953,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73021,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/198827305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aul4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 424w, https://substackcdn.com/image/fetch/$s_!aul4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 848w, https://substackcdn.com/image/fetch/$s_!aul4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 1272w, https://substackcdn.com/image/fetch/$s_!aul4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c689d9-6577-4899-aa84-70bac90561a3_953x828.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The cleanest one-line definition comes from LangChain: &#8220;Agent = Model + Harness. If you&#8217;re not the model, you&#8217;re the harness.&#8221;</p><p>A useful analogy is to think of a model as a very smart contractor who can only communicate through notes. A harness is the tool setup &#8212; the computer, the phone, the filing cabinet, the rulebook, the permissioning &#8212; that lets the contractor actually do the job.</p><p>Mechanically, even the most sophisticated harnesses are just a while loop. The model emits a tool call, the harness executes it, the result gets fed back, the loop continues. Anthropic literally calls their runtime a &#8220;dumb loop&#8221; where all the intelligence lives in the model. But the complexity is in everything that loop manages: which tools exist, what their schemas look like, what gets injected into context on each turn, what gets compacted away, what gets remembered across sessions, how subagents are spawned, how errors are recovered, when the model is allowed to take a destructive action without asking. None of that is &#8220;the model.&#8221; All of it determines whether your agent actually finishes the task or burns 100,000 tokens going in circles. This has effects both on the quality of the output and on costs.</p><p><strong>The empirical evidence: the harness is moving benchmarks even more than the model is in some cases</strong></p><p>We now have some hard data showing that the harness can move model performance by more than a full model generation upgrade.</p><p>Three independent data points from the last six months, all converging on the same conclusion:</p><p><strong>Data point 1 - LangChain on Terminal-Bench 2.0.</strong> Terminal-Bench 2.0 is now the standard benchmark for evaluating coding agents on real terminal tasks (89 tasks across machine learning, debugging, biology, infrastructure). LangChain kept the model fixed (GPT-5.2-Codex) and only changed the harness. Score went from <strong>52.8% to 66.5%</strong>. That moved them from outside the Top 30 to <strong>Top 5</strong> on the leaderboard. Same model, same weights, different scaffolding. A 13.7 percentage point jump from harness alone.</p><p><strong>Data point 2 - Cursor on Terminal-Bench 2.0.</strong> Cursor&#8217;s research team published a piece on April 30, 2026 reporting that they took their own coding agent from <strong>Top 30 to Top 5</strong> by only changing the harness. Same conclusion, different team, different harness. A 25-position jump on a public leaderboard, attributable to scaffolding alone.</p><p><strong>Data point 3 - Claude Opus 4.6, same weights, very different harnesses.</strong> This is the cleanest one. Look at the Terminal-Bench 2.0 leaderboard from late April 2026:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HNYZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HNYZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 424w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 848w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 1272w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HNYZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png" width="787" height="140" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/848847aa-07f2-48ee-947b-f03541587073_787x140.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:140,&quot;width&quot;:787,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:14503,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/198827305?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HNYZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 424w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 848w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 1272w, https://substackcdn.com/image/fetch/$s_!HNYZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F848847aa-07f2-48ee-947b-f03541587073_787x140.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Same model. Different harnesses. A 4.5 percentage point spread<strong> </strong>between ForgeCode and Capy, on a benchmark where teams are fighting for tenths of a point.</p><p>The interesting observation is that the<strong> </strong>labs that trained the models do not always have the best harness for their own models, which I believe shows just how much untapped potential there still is in this area. ForgeCode is a third-party harness, and it lands three of the top six entries on Terminal-Bench 2.0 by routing across model families. LangChain summarized this as: &#8220;Opus 4.6 in Claude Code scores far below Opus 4.6 in other harnesses.&#8221; Anthropic&#8217;s flagship model in Anthropic&#8217;s flagship harness gets beaten by the same weights running in third-party scaffolding.</p><p>Now compare this to model generation upgrades. Going from Claude Opus 4.5 (80.9% on SWE-bench Verified) to Opus 4.7 in April 2026 took SWE-bench Verified from 80.8% to 87.6% &#8212; about a 6.8-point upgrade. The harness can move you 13.7 points on Terminal-Bench from scaffolding alone. The harness is moving the score by more than a model generation upgrade.</p><p>Don&#8217;t get me wrong, I am not trying to say that raw AI model performance doesn&#8217;t matter. The model floor is rising, and that floor matters a lot. The point is that on top of any given model, the harness is now the largest single performance variable in agent quality. Anyone shipping a coding agent in 2026 who picks the model first and the harness second is leaving a lot of performance and cost-effectiveness.</p><p>This section from an interview with a former high-ranking Meta employee is very interesting. It explains just how much more room there is for improvement around the harness, and that because raw model improvements are so fast right now, the resources are not yet focused on it enough:</p><div class="pullquote"><p>&#187;The biggest barrier, I would say it&#8217;s just the modeling evolution is so quick. We want to push the model architecture faster but we don&#8217;t really have the time to really do parameter tuning to find fast architecture for our data or for different use cases. Right now, it&#8217;s more like we have one safe home. That&#8217;s our goal at that time. We have the foundation model. The foundation model powers roughly four or five different orgs&#8217; ranking models.<br><br> Although it&#8217;s a foundation model, it&#8217;s a Wikipedia, knows everything, but still, how to find an optimal maybe adapter or optimization, tuning for each of the five use cases that are under its cloud. I think that&#8217;s one of the biggest challenges. We have to chase our targets and there are some ways to hit a target a little bit easier than really going deep into understanding this model, what the model does and what&#8217;s the best parameter.&#171;</p><p>Source: Former Meta employee found on AlphaSense</p></div><p><strong>Why the harness forms a moat</strong></p><p>It is important to understand that models are post-trained against a specific harness. They are not generic. Every modern frontier model was fine-tuned, with reinforcement learning, against a specific tool surface, a specific schema, a specific memory ritual, a specific citation format, a specific system prompt structure. The model&#8217;s instincts &#8212; what it reaches for when it needs to edit a file, what tag it wraps citations in, how it structures a plan, how it handles subagent dispatch &#8212; are baked into the weights during post-training. And they are baked in against one specific harness.</p><p>The clearest concrete example is the file-editing tool. From Cursor&#8217;s harness team:</p><div class="callout-block" data-callout="true"><p>&#8220;OpenAI&#8217;s models are trained to edit files using a patch-based format, while Anthropic&#8217;s models are trained on string replacement. Either model could use either tool, but giving it the unfamiliar one costs extra reasoning tokens and produces more mistakes. So in our harness, we provision each model with the tool format it had during training.&#8221;</p></div><p>Cursor is saying that the wrong tool format produces a measurable cost in reasoning tokens and an observable increase in error rate, recorded at scale across millions of agent turns. The model performs worse because the wire format it sees at runtime does not match the wire format it was trained on.</p><p>This is becoming a challenge for everyone trying to become an orchestration layer on top of AI model providers like Microsoft. I am not saying it is an impossible task, but you need to have different harnesses for each AI model you use. It also suggests that the products that these AI labs are offering are not just the model, but the model together with the custom harness, as the model and the harness were fused over months of post-training, and you cannot pull them apart without giving up performance. In the end, this might mean the moat becomes stronger for AI labs, as switching AI models becomes increasingly complex.</p><p><strong>Why this is a co-evolution, not a one-time thing</strong></p><p>The more powerful version of the moat argument is that the post-training and the harness feed back into each other on every model generation. This is what is starting to create compounding lock-in over time.</p><p>An example of this is Anthropic&#8217;s harness team ships a new primitive &#8212; say, a smarter subagent dispatch verb, or a new way to compact context, or a memory file convention. By month three, that primitive shows up in millions of real agent traces from Claude Code users. By month six, those traces are training data for the next model generation. By month twelve, the next model has the primitive baked into its instincts, and the harness can now lean on it as something the model does natively.</p><p>Anthropic puts it this way in their March engineering blog: &#8220;every component in a harness encodes an assumption about what the model can&#8217;t do on its own. Those assumptions go stale.&#8221; When a model upgrade kills an old assumption, that piece of scaffolding gets retired, freeing up the harness to push on the next ceiling.</p><p>A concrete example from Anthropic&#8217;s own work: when Claude Sonnet 4.5 was the frontier model, the agent had &#8220;context anxiety&#8221; &#8212; it would wrap up tasks prematurely as it approached what it thought was its context limit. The harness compensated by aggressively resetting context between sessions and using structured handoff artifacts to carry state across the boundary. When Opus 4.6 shipped, that behavior was largely gone in the model itself, so Anthropic dropped the entire context-reset machinery and ran continuous sessions over two hours. The harness shrank because the model swallowed a chunk of its work.</p><p>The matched pair is not static. It moves with each model generation. And the labs that own both sides of the pair are the only ones who can move it cleanly. A third-party harness builder is always reacting to a model release; the labs are designing the next model with the next harness in mind.</p><p>This is genuinely structural moat behavior. It looks a lot like the way Microsoft built lock-in through tightly coupling Windows + Office + Exchange in the 1990s and 2000s - each layer made the others stickier, and competing on any single layer in isolation got harder every year.</p><p>There is a counter-example worth knowing because it sharpens the moat argument. In December 2025, Vercel published an engineering post-mortem: their internal text-to-SQL agent had 16 specialized tools - schema lookup, query validation, error recovery, intent clarification, join-path finders, syntax validators. They deleted 80% of them and replaced everything with a single capability: execute arbitrary bash commands against a file system.</p><p>Results: success rate went from 80% to 100%. Speed improved<strong> </strong>3.5x. Token usage dropped 37%. The worst-case run improved from 724 seconds, 100 steps, and 145,463 tokens (failing) to 141 seconds, 19 steps, 67,483 tokens (succeeding).</p><p>The lesson is more than just that fewer tools are better. It is that the right harness depends on the model&#8217;s current capabilities, and that gap shifts every model generation. When the model gets smart enough to use bash + grep + cat directly, your custom tool layer becomes overhead. GitHub&#8217;s Copilot team hit the same wall from the opposite direction: they cut Copilot&#8217;s tool count from 40+ to 13 core tools, and pre-expansion accuracy jumped from 19% to 72%.</p><p>This is exactly why owning both the model and the harness is so valuable. You can retire scaffolding as your model gets smarter, faster than anyone else can. A third party building on top of your model is always reacting to your release notes. You are designing both sides.</p><p><strong>Where the harness wars are headed</strong></p><p>A few things I am watching that I think will define the next 12-18 months of this space:</p><ol><li><p><strong>The harness becomes the product, not the model.</strong> This is already happening. Anthropic does not sell &#8220;Claude&#8221; anymore as the main enterprise product &#8212; they sell Claude Code, Cowork, the Claude Agent SDK, and Claude Managed Agents. OpenAI sells Codex CLI, the Agents SDK, and the Codex cloud agent. Google just announced that they are deprecating Gemini CLI entirely and replacing it with Antigravity CLI, which shares a unified server-side harness with their Antigravity desktop IDE. The model is the engine; the harness is the car. Customers buy the car.</p></li><li><p><strong>&#8220;Harness-as-a-Service&#8221; (HaaS) is the next API layer.</strong> The Claude Agent SDK, the OpenAI Agents SDK, and Google&#8217;s new Antigravity SDK all point the same way: you do not get a model API anymore, you get a harness API &#8212; the loop, the tools, the context management, the hooks, the sandbox primitives, the memory layer, all out of the box. This is a bigger and stickier surface than &#8220;completions.&#8221;</p></li><li><p><strong>These are important implications for the SaaS companies. </strong>The harness will absorb more of what is currently sold as SaaS. A managed harness with built-in code execution, browser control, file storage, memory, and orchestration is the runtime for AI-native software. That competes directly with a stack of SaaS tools previously bought separately: developer environments, browser automation, vector databases, observability layers, sandboxing services. Anthropic just shipped Claude Managed Agents, which puts the entire harness behind an API.</p></li><li><p><strong>Harnesses will increasingly diverge by vertical.</strong> The highest-quality harnesses today are all coding harnesses, because the ROI is most obvious. The same primitives apply to legal, finance, healthcare, and support. Whoever builds the best vertical harness for a domain will own that domain&#8217;s spend, even if the underlying model is shared. The providers that specialize and build harnesses for specific verticals can overcome even AI model raw performance deficiencies because the harness is better. An example of this would be someone like Meta, specializing in shopping, social media, and the healthcare vertical.</p></li></ol><p><strong>The investment angle</strong></p><p>This is where the moat argument actually matters for capital allocation in terms of sectors and for which public company this might be most important.</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/the-harness-the-moat-for-ai-model">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Market Is Pricing Meta Like It's the AI Loser. Big Mistake.]]></title><description><![CDATA[In this article, I want to walk through why I think the market is making one of its bigger mistakes of this cycle with Meta. The stock is down roughly 24% from its high of $796.25, trading around $602, with a forward P/E of 19x &#8212; well below its 10-year average of around 26-27x and below the S&P 500 multiple. The dominant narrative is that Meta is &#8220;spending too much&#8221; on AI data centers without a clear path to monetization.]]></description><link>https://www.uncoveralpha.com/p/the-market-is-pricing-meta-like-its</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-market-is-pricing-meta-like-its</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 11 May 2026 14:14:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E-vS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>In this article, I want to walk through why I think the market is making one of its bigger mistakes of this cycle with Meta. The stock is down roughly 24% from its high of $796.25, trading around $602, with a forward P/E of 19x &#8212; well below its 10-year average of around 26-27x and below the S&amp;P 500 multiple. The dominant narrative is that Meta is &#8220;spending too much&#8221; on AI data centers without a clear path to monetization. Every quarter, the market punishes the stock harder on CapEx prints &#8212; the post-Q1 2026 reaction was a 10% drawdown on a 57% EPS beat, the worst price reaction Meta has had in its last six earnings reports, despite delivering the largest earnings surprise in that window.</p><p>The thesis I lay out has four parts:</p><p>- Meta&#8217;s AI CapEx is already showing up in revenue and engagement in a significant way and more importantly, will continue to do so (expert interview with a Former Meta employee on this field)</p><p>- The data centers are not &#8220;moonshots&#8221; (and far from the metaverse spend analogy) &#8212; they are Meta&#8217;s new workforce, and the more compute Meta has, the more it can improve products and ship new ones;</p><p>- Meta has the single most underappreciated asset in tech right now, which is distribution, and the market is assigning zero value to it;</p><p>- The valuation has compressed to a point where the bar for an aggressive re-rate is much lower than people think.</p><p>Let&#8217;s dive in.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E-vS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E-vS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E-vS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2503432,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/197218225?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E-vS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!E-vS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74d35c16-727c-4e21-b16e-dff47d7ed3b6_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The CapEx that the market hates is already showing up in the P&amp;L</h2><p>Meta did $200.97 billion in revenue in 2025, up 22% YoY, with operating income of $83.28 billion, up 20%. In Q1 2026, they did $56.31 billion of revenue, up 33% YoY &#8212; an acceleration from a $200 billion base. Ad impressions grew 19% YoY and average price per ad grew 12% YoY in Q1 2026, with the Q2 guide of $58-61 billion implying continued acceleration.</p><p>Now think about what that means in context. The bear narrative is that Meta is spending $125-$145 billion in 2026 CapEx with nothing to show for it. But the company is putting up 33% growth on a $200 billion base while also compounding the underlying business with double-digit price-per-ad gains. When advertisers pay more per impression, and impression volume also grows 19%, that shows that the system is getting better at allocating attention to higher-value placements.</p><p>The AI in the P&amp;L is most visible in three specific areas that I want to walk through, because management actually quantified them on the calls.</p><p>First, the ad-ranking models. Meta has been rolling out a model called GEM (Generative Ads Recommendation Model), which is essentially their LLM-style foundation model for ads. On the Q4 2025 call, management said this directly:</p><div class="pullquote"><p><em>In Q4, we doubled the number of GPUs we used to train our GEM model for ads ranking. We also adopted a new sequence learning model architecture, which is capable of using longer sequences of user behavior and processing much richer information about each piece of content. The GEM and sequence learning improvements together grew a 3.5% lift in ad clicks on Facebook and a more than 1% gain in conversions on Instagram in Q4.</em></p></div><p>A 3.5% lift in ad clicks on a base of roughly $200 billion is multiple billions of incremental dollars, and that&#8217;s from one model improvement in one quarter.</p><p>Meta also said on the Q3 2025 call that GEM is now &#8220;4x more efficient at driving ad performance gains&#8221; compared to the original ranking models. This is exactly the scaling law dynamic you want to see &#8212; more compute thrown at the model translates to more revenue, and the elasticity of that conversion is improving, not deteriorating.</p><p>A <a href="https://www.alpha-sense.com/uncoveralpha/">recent interview</a> with a high-ranking former Meta employee who worked in this field was very useful. He explained some details about Meta&#8217;s internal metric called Internal Revenue per Engagement (iREV). According to him, Meta&#8217;s internal goal is a minimum 1.5-2% improvement of that metric every 6 months. So far, they have been delivering on this metric, and he is very confident that going forward, Meta will continue to be able to deliver that as they have a lot of room for improvement in three areas that affect iREV: model architecture, more data, and transfer learning. He even quantified it: 30-40% from model architecture changes, 20-30% from more data, and 30-40% from transfer learning. While model architecture and more data are quite straightforward, transfer learning is something many of you might not be familiar with. To explain the context as simply as possible, because serving a large LLM model across the scale of Meta is too expensive, they had to figure out a structure where they had a teacher LLM (the big one) and a student LLM who basically distilled knowledge from the teacher model to be smaller and more effective to run inference on at the scale Meta needs it. So improving data transfer between the teacher and the student model for Meta is key, as it will be for any other company running large-scale production workloads (the phase of AI adoption we are now in). So every time Meta makes improvements in this realm, it translates to better ad performance and rankings, as performance is essentially determined by the quality of the student model.</p><p>What I found really insightful was when the person was asked about what the biggest barrier to improving even faster in terms of ad ranking and performance beyond the 2% per 6 months was:</p><div class="pullquote"><p>&#187;The biggest barrier, I would say it&#8217;s just the modeling evolution is so quick. We want to push the model architecture faster but we don&#8217;t really have the time to really do parameter tuning to find fast architecture for our data or for different use cases. Right now, it&#8217;s more like we have one safe home. That&#8217;s our goal at that time. We have the foundation model. The foundation model powers roughly four or five different orgs&#8217; ranking models.<br><br>Although it&#8217;s a foundation model, it&#8217;s a Wikipedia, knows everything, but still, how to find an optimal maybe adapter or optimization, tuning for each of the five use cases that are under its cloud. I think that&#8217;s one of the biggest challenges. We have to chase our targets and there are some ways to hit a target a little bit easier than really going deep into understanding this model, what the model does and what&#8217;s the best parameter.<br><br>Even what I said, what&#8217;s the teacher model capacity and the student model capacity ratio? What&#8217;s the optimal ratio between the two models? That even what was studied during my time there. I feel that at that time, the big challenge, we are just chasing the whole statistically or too aggressively and ignore those foundational things, all those long-term things a little bit less.&#171;</p><p>source: <a href="https://www.alpha-sense.com/uncoveralpha/">AlphaSense</a></p></div><p>What this means is that because the model performance is moving so fast, internally, Meta hasn&#8217;t even had enough time and resources to pull or spend time to optimize other levers, which shows us just how early we still are in these improvements and how much more growth from this ad AI tailwind Meta has for its core business, not just from scaling laws and model improvements but also from the data and student teacher mechanizm.</p><p>Second, engagement. On the Q3 2025 call, Zuckerberg said:</p><div class="pullquote"><p><em>Across Facebook, Instagram, and Threads, our AI recommendation systems are delivering higher quality and more relevant content, which led to 5% more time spent on Facebook in Q3 and 10% on Threads. Video is a particular bright spot, with video time spent on Instagram up more than 30% since last year. As video continues to grow across our apps, Reels now has an annual run rate of over $50 billion.</em></p></div><p>A 5% lift in time spent on Facebook &#8212; a 21-year-old app that everyone wrote off as dead &#8212; is huge. The platform was supposed to be in decline. AI ranking systems brought it back. In Q4 2025, the optimizations drove a 7% lift in views of organic feed and video posts on Facebook, which Susan Li called &#8220;the largest quarterly revenue impact from Facebook product launches in the past two years&#8221;.</p><p>Third, the end-to-end AI ad tools. The annual run-rate of revenue going through Meta&#8217;s fully AI-powered ad tools (Advantage+) passed $60 billion on Q3 2025. The video generation tools alone hit a $10 billion run rate by Q4 2025, with quarter-over-quarter growth outpacing the broader ad revenue increase by nearly 3x. Click-to-WhatsApp ads grew revenue 60% YoY in Q3. None of this exists without the AI infrastructure that the market is currently punishing the company for building.</p><p>The way to think about this is the same logic I laid out in my <a href="https://www.uncoveralpha.com/p/the-market-hates-big-cloud-spending">February article</a> about the hyperscalers: the CapEx Meta is spending this year doesn&#8217;t show up in this year&#8217;s revenue. A data center takes around 2 years to build and operationalize, so the revenue acceleration Meta is showing today is the return on 2023 CapEx (~$28 billion), not 2025 CapEx ($72.2 billion). When 2025&#8217;s CapEx starts showing up in 2027 revenue, the operating leverage will be far more aggressive than what we&#8217;re seeing now.</p><h2>Data centers are Meta&#8217;s new workforce</h2><p>Here is the part I think most investors are missing, and it&#8217;s where the analogy needs to shift.</p><p>For the last 15 years, Meta&#8217;s growth engine was: hire engineers, ship product, get more users, monetize via ads. Headcount was the input that scaled output. That framework is now different, because the marginal unit of &#8220;intelligence&#8221; inside Meta is no longer an engineer. It&#8217;s a GPU.</p><p>Zuckerberg essentially said this on the Q3 2025 call when he framed Meta&#8217;s strategy around three &#8220;giant transformers&#8221; running Facebook, Instagram, and ads, with the goal of merging them into one unified system:</p><div class="pullquote"><p><em>At the same time, we&#8217;re also working on combining these three major AI systems into a single unified AI system that will effectively run our family of apps and business &#8212; using increasing intelligence to improve the trillions of recommendations that it will make for people every day.</em></p></div><p>Meta is openly saying that the entire company &#8212; the feeds, the ads, the recommendations across 3.56 billion daily users &#8212; is going to be run by AI systems whose performance scales with compute. The CapEx number is the headcount number for the AI era.</p><p>And here is the kicker &#8212; Meta is compute-starved on the current business. Zuckerberg said it directly on Q3 2025:</p><div class="pullquote"><p><em>We are sort of perennially operating the Family of Apps and ads business in a compute-starved state at this point, which is on the one hand sort of an odd thing to say, given the compute that we built up. But we really are taking a lot of the resources and using them to advance future things that we&#8217;re doing. And we think that there&#8217;s a lot more compute that we could put towards these that would just unlock a huge amount of opportunity in the core business as well.</em></p></div><p>Meta CFO Susan Li doubled down on this on the same call:</p><div class="pullquote"><p><em>We&#8217;re certainly seeing that we wish we had more capacity today than we do. We would be able to put it towards good use, certain not only would the MSL team appreciate having more capacity, but we&#8217;d be able to put it towards good and ROI positive use in the core business as well.</em></p></div><p>This is not a company building speculative infrastructure for products that might monetize in 2030. This is a company that has more profitable use cases for compute than it has compute, and is rationing GPU hours between training the next frontier model and improving the ranking systems that drive a quantifiable lift in conversions every quarter. Investors who treat this CapEx like a moonshot are misreading the situation.</p><h2>Meta&#8217;s distribution is Slept on</h2><p>There is a more general thesis floating around in the market right now that says: &#8220;If AI commoditizes, then everyone with compute can build products, from software to something like a Meta platform.&#8221; I think this gets the second-order logic backward. If models commoditize and anyone with enough GPUs can ship a product, then the question becomes: who can get that product in front of users? Distribution becomes the bottleneck, not the model. And Meta&#8217;s distribution machine is arguably the single biggest in the world.</p><p>Meta&#8217;s Family of Apps had 3.56 billion daily active people in March 2026. Instagram crossed 3 billion monthly active users in September 2025. WhatsApp also has over 3 billion users across 180+ countries. Facebook still serves billions of people daily. That is unmatched at this scale anywhere in tech &#8212; Google has Search, but the engagement-per-user profile is fundamentally different (people come to Search, do a query, leave; people stay on Meta apps for 30+ minutes a day).</p><p>Here&#8217;s how that distribution muscle has shown up historically. Meta bought Instagram for $1 billion in 2012. It is now a 3-billion-MAU asset that is the cultural center of gravity for an entire generation. They bought WhatsApp for $19 billion in 2014. It has more than 6x to over 3 billion users. They built Threads from scratch in mid-2023 &#8212; a product that, frankly, was not particularly differentiated from X &#8212; and rode Instagram&#8217;s social graph to 400 million MAUs and 150 million daily actives in roughly 2.5 years. Similarweb data shows Threads passed X in daily mobile active users in January 2026.</p><p>Threads was a clone. The product was almost identical to X. There was nothing technically novel about it. And in 2.5 years, by being plugged into Instagram&#8217;s distribution graph, it overtook a 19-year-old product with deep cultural roots. The question every investor should ask themselves: what other company on earth could have done that?</p><p>Now apply this to AI products. Meta AI hit 1 billion monthly active users by May 2025, doubling from 500 million in roughly 8 months, and let&#8217;s be honest, the product wasn&#8217;t even good. ChatGPT took roughly 2 years to reach similar scale. Meta did it by embedding the assistant into search bars and chat interfaces inside WhatsApp, Instagram, Facebook, and Messenger. Roughly 63% of Meta AI&#8217;s usage comes from WhatsApp alone. Meta did not need to convince anyone to download an app, learn a new interface, or change a habit. The distribution infrastructure was already there.</p><p>If you believe &#8212; and I do &#8212; that the next phase of AI is going to produce a wave of consumer products (AI-generated content, personalized AI assistants, business AI, voice agents, creator tools, AI shopping experiences), then the company that can ship each of those products to 3.56 billion people on day one has a structural advantage that the market is not pricing in. Zuckerberg said it himself on the Q3 2025 call:</p><div class="pullquote"><p><em>I would guess that Meta has the best track record of any company out there of taking a new product that people love and getting it to billions of people in terms of usage. So I think that the ability to plug in leading models is going to, I would predict, lead to a very large amount of use of these things over the coming years.</em></p></div><p>The market is essentially treating Meta as if distribution is free. It&#8217;s not free. It is the single hardest moat to build in consumer technology, and Meta is the only company that has built three of them in parallel (Facebook, Instagram, WhatsApp), then bolted on a fourth (Threads) using the first three as the launchpad.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><h2>The AI sentiment changes on a dime</h2><p>The market right now is running on AI sentiment more than fundamentals. Companies are being bucketed as &#8220;AI winners&#8221; or &#8220;AI losers&#8221;, and the valuation gaps between those buckets are enormous. The reason is that the marginal flows of capital in public markets are still controlled by investors with financial-domain expertise but a relatively shallow understanding of how AI actually works at the technical level. So the signal that gets weighted most heavily is: did this company ship a frontier model? Did they show up on the benchmark leaderboards?</p><p>This is exactly the gap that creates opportunity. Meta released Muse Spark on April 8 &#8212; the first model from Meta Superintelligence Labs (MSL). Muse Spark scored 52 on the Artificial Analysis Intelligence Index, behind Gemini 3.1 Pro and GPT-5.4 (both at 57) and Claude Opus 4.6. On absolute benchmark terms, it&#8217;s not SOTA.</p><p>But look at what it is good at. Muse Spark used only 58 million output tokens on the Intelligence Index evaluation, versus 157 million for Claude Opus 4.6 and 120 million for GPT-5.4 &#8212; meaning Meta is delivering near-frontier intelligence with less than half the inference compute of competitors. Meta also said the model achieves the same capability level as the older mid-size Llama 4 Maverick using an order of magnitude less compute. For a company that&#8217;s about to deploy this model to 3 billion daily users, inference efficiency at this scale is a multi-billion-dollar economic advantage. And the model is particularly strong in vision, health, and what Meta is calling &#8220;personal intelligence&#8221; use cases &#8212; exactly the domains that map onto consumer apps.</p><p>Now think about what happens when the larger frontier model lands. Meta has been working on a next-generation flagship, a bigger model internally named the &#8220;Watermelon&#8221; model. If Meta lands a model that is genuinely competitive on benchmarks with frontier models from Anthropic, OpenAI, and Google, the market will re-rate the company aggressively. The current setup is that Meta is being priced as if it can&#8217;t compete at the frontier. The forward P/E of 19x reflects that. Compare that to Google trading at a meaningfully higher multiple post the TPU/Gemini story repricing in late 2025. The asymmetry is real. If Meta merely catches up to where the market is already pricing Google and other labs like Anthropic, OpenAI, the implied upside is substantial&#8212; and that&#8217;s without assigning incremental value to the distribution moat or to any of the new AI products.</p><h2>The Market hates the data center spend. Meanwhile, everyone else is desperate for compute.</h2><p>The market is currently assigning negative value to Meta&#8217;s data center buildout. Every time CapEx goes up, the stock goes down. Meta&#8217;s CapEx jumped from $39 billion in 2024 to $72 billion in 2025 to a guided $125-145 billion in 2026.</p><p>At the same time, the rest of the AI ecosystem is screaming that we don&#8217;t have enough compute. Anthropic just signed a deal to take all 300+ MW of compute capacity at xAI&#8217;s Colossus 1 data center in Memphis &#8212; roughly 222,000 Nvidia GPUs including H100, H200, and GB200 systems. That deal is worth billions. xAI has effectively pivoted to a neocloud model, renting GPUs to Anthropic. The CoreWeave and Nebius backlogs continue to grow. Oracle&#8217;s cloud business is being capacity-constrained. AWS, Google Cloud, and Azure are all selling everything they have available and have multi-year backlogs.</p><p>So the market believes simultaneously that (a) there is a multi-year compute shortage that will continue at least through 2027, and (b) Meta is wrong to be building data centers and putting negative value to them. These two beliefs cannot both be true. If there is a compute shortage, then Meta&#8217;s buildout&#8217;s terminal value in the worst case should be based on the value of those data centers if Meta sells that compute on the market. Zuckerberg made this point explicitly on the Q3 2025 call:</p><div class="pullquote"><p><em>To date, we keep on seeing this pattern where we build some amount of infrastructure to what we think is an aggressive assumption and then we keep on having more demand to be able to use more compute&#8230; any compute that we don&#8217;t need for that, we feel pretty good that we&#8217;re going to be able to absorb a very large amount of that to just convert into more intelligence and better recommendations in our Family of Apps and ads in a profitable way. Now, I mean, it&#8217;s of course possible to overshoot that, right&#8230; If we do, this is what I mentioned in my comments then we see that there&#8217;s just a lot of demand for other new things that we build internally, externally. Like almost every week, people come to us from outside the company asking us to stand up an API service or asking if we have different compute that they could get from us. And we haven&#8217;t done that yet, but obviously if you got to a point where you overbuilt, you could have that as an option.</em></p></div><p>The fact that the market is ignoring this and assigning a negative value to the CapEx is, in my view, a significant mistake.</p><p>Here&#8217;s why this matters even more: Meta is arguably the only company outside the three hyperscalers (AWS, Microsoft, Google) that has the operational capability to run hyperscale data centers for both training and inference at the level required. They&#8217;ve been operating planetary-scale infrastructure for over a decade. They know how to manage multiple AI accelerators: Nvidia GPUs, AMD GPUs, and custom ASIC (their MTIA). If Meta wanted to offer compute externally tomorrow, the renters lining up would include some of the largest AI labs and enterprises in the world, and the unit economics would look more like AWS than like a cap-on-cost neocloud.</p><h2>Valuation: What should the real value be?</h2><p>The company is </p>
      <p>
          <a href="https://www.uncoveralpha.com/p/the-market-is-pricing-meta-like-its">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Amazon, Google, Microsoft, Meta Q1 earnings: AI profits are here, custom silicon is winning]]></title><description><![CDATA[I argued that the market was wrong to punish big tech for raising CapEx because the returns on the 2023-2024 spend were already showing up in the P&L. This quarter, that argument got significantly stronger. Margins on cloud businesses expanded again, the core ad businesses at Meta and Google went into a higher gear]]></description><link>https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q1-earnings</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/amazon-google-microsoft-meta-q1-earnings</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 30 Apr 2026 11:59:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oQvP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>We just got Q1 2026 earnings from Meta, Microsoft, Google, and Amazon, and after spending hours on the calls and the prints, I want to share what I think is the most important takeaway.</p><p>I argued that the market was wrong to punish big tech for raising CapEx because the returns on the 2023-2024 spend were already showing up in the P&amp;L. This quarter, that argument got significantly stronger. Margins on cloud businesses expanded again, the core ad businesses at Meta and Google went into a higher gear (Meta +33% YoY, Google Search +19% YoY - both fastest in years), and we got hard data on what custom silicon is doing to the unit economics of inference.</p><p>There are a few patterns from this earnings season:</p><ul><li><p>The core ad businesses at Meta and Google are accelerating because of AI, not in spite of it</p></li><li><p>Operating margins on cloud are expanding even as AI workloads scale</p></li><li><p>Custom ASICs are no longer a side project - they are the next big business segment</p></li><li><p>The era of &#8220;subsidized&#8221; compute is ending, and we are seeing the first hints of pricing power coming through</p></li><li><p>Compute supply is still the binding constraint</p></li></ul><p>Let&#8217;s get into it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oQvP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oQvP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oQvP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/81438849-f526-4f23-a56b-a192a1657845_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2346019,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/195985479?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oQvP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!oQvP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F81438849-f526-4f23-a56b-a192a1657845_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Google: Cloud +63%, Search +19%, And The Margin Story</strong></p><p>Google delivered a really strong quarter. Two key numbers: Google Cloud up 63% YoY and the &#8220;old left for dead&#8221; Search up 19% YoY.</p><p>Google Cloud revenue hit $20.0B, growing 63% YoY, an acceleration from the 48% growth we saw in Q4 2025. To put that in context, this is the fastest growth rate Google Cloud has ever posted, and they are doing it on a $70B+ annual run-rate base. Google&#8217;s Cloud backlog also nearly doubled sequentially to $462B, with management telling us:</p><div class="pullquote"><p>&#187;The majority of the backlog is related to typical GCP contracts and we expect to recognize just over 50% of the backlog as revenue over the next 24 months&#171;</p></div><p>That last part is important. A lot of the bear case on backlog numbers is that they are stuffed with one mega-deal that won&#8217;t translate into revenue for years (the Microsoft/OpenAI dynamic). Google is essentially telling investors: half of $462B is coming through the P&amp;L in the next 8 quarters.</p><p>But the bigger story for me on the Google call was the margin commentary on Cloud:</p><div class="pullquote"><p>&#187;Cloud operating income was $6.6 billion, tripling year-over-year and operating margin increased from 17.8% in the first quarter of last year to 32.9%&#171;</p></div><p>The bear thesis on hyperscaler AI workloads has been: &#8220;Yes, revenue is growing, but margins on AI workloads will be lower than legacy cloud workloads&#8221;. If that thesis were correct, you would expect to see operating margin compression at Google Cloud, given that AI workloads are now the majority of new revenue growth there. Instead, we saw the opposite: operating margin nearly doubled YoY, from 17.8% to 32.9%, in a single year. Now it is important to understand that not all of this is direct GCP infrastructure a lot of it is now Gemini products:</p><div class="pullquote"><p>&#187;Our enterprise AI solutions have become our primary growth driver for cloud for the first time. In Q1, revenue from products built on our GenAI models grew nearly 800% year-over-year.&#171;</p></div><p>The part of Cloud that is growing fastest (GenAI products, +800% YoY) is also the part driving the margin expansion. The reason Google can do this is threefold: scale-driven optimization (Google said earlier this year that it reduced Gemini&#8217;s serving cost by 78% in 2025), custom silicon (TPUs), and owning its own frontier model.</p><p>On TPUs, we got another important update:</p><div class="pullquote"><p>&#187;we&#8217;ll begin to deliver TPUs to a select group of customers in their own data centers in the hardware configuration to expand our addressable market opportunity.&#171;</p></div><p>This is a massive strategic shift that I have been writing about for over a year now. Google is essentially saying: the TPU is so valuable to certain customers (think Anthropic, but also potentially Apple, Meta, etc.) that we will sell or lease the chips outside of GCP. This is the same direction Amazon is going with Trainium. The implication is that the cloud providers are starting to compete with NVIDIA directly at the chip layer, not just at the cloud layer.</p><p>And on Search, the Search is dead narrative is looking more distant and distant:</p><div class="pullquote"><p>&#187;Turning to Search. AI continues to drive search usage, and queries are at an all-time high&#171;</p></div><p>19% YoY growth in Search advertising on a $240B+ annual run rate is, candidly, an absurd number.</p><p><strong>Microsoft: The AI Business Hit $37B ARR, Up 123%, And Margins Are Holding</strong></p><p>Microsoft&#8217;s print was a strong quarter, and one of the most important numbers was this:</p><div class="callout-block" data-callout="true"><p>&#187;Our AI business surpassed $37 billion ARR, up 123%.&#171;</p></div><p>Just on this number alone: Microsoft&#8217;s AI business is now larger than ServiceNow, Workday, or Datadog as standalone businesses. And it is growing at 123% YoY. There is no public software company at scale growing faster.</p><p>But the more important nuance came from the management commentary on margins, which is essentially the same story Google told. From Amy Hood:</p><div class="pullquote"><p>&#187;Thanks, Brent. I think -- we&#8217;ve been talking about sort of where this AI business of ours has been in the cycle compared to even the cycle we saw with the cloud, which now seems very long ago. And how margins were actually better and they remained better in our AI business versus where we saw in the cloud transition, looking back.&#171;</p></div><p>I want to highlight this because it is being completely missed by the market. Microsoft is telling us that AI workload margins are BETTER than the early cloud workload margins were.</p><p>On Copilot, the seat numbers continued to ramp and were much better than the quarter before:</p><div class="pullquote"><p>&#187;In knowledge work, it was another record quarter for Microsoft 365 Copilot seat ads, which increased 250% year-over-year, representing our fastest growth since launch. Quarter-over-quarter, we continue to see acceleration and now have over 20 million Microsoft 365 Copilot paid seats. The number of customers with over 50,000 seats quadrupled year-over-year and Accenture now has over 740,000 seats, our largest Copilot win to date.&#171;</p></div><p>That is 20M paid seats up from 15M last quarter (~33% sequential growth) and 250% YoY. At a $30/user/month list price, that is a roughly $7B+ ARR business just from M365 Copilot.</p><p>Engagement on Copilot is the other piece of the story:</p><div class="pullquote"><p>&#187;We have seen a surge in usage of our first-party agents with monthly active usage up 6x year-to-date. Copilot queries per user were up nearly 20% quarter-over-quarter. To put this momentum in perspective, weekly engagement is now at the same level as Outlook, as more and more users make Copilot a habit.&#171;</p></div><p>Weekly engagement at the level of Outlook is a wild data point.</p><p>GitHub Copilot also continues to scale:</p><div class="pullquote"><p>&#187;We see this even with GitHub Copilot. Nearly 140,000 organizations now use GitHub Copilot and enterprise subscribers have nearly tripled year-over-year.&#171;</p></div><p>And here is where it gets interesting - Microsoft is moving GitHub Copilot to a usage-based pricing model:</p><div class="pullquote"><p>&#187;And earlier this week, we announced our move to usage-based pricing model for GitHub Copilot as we align pricing to actual usage and cost&#171;</p><p>&#187;Microsoft Cloud gross margin percentage should be roughly 64%, down year-over-year, driven by continued investments in AI and increased GitHub Copilot usage. Just this week, we announced a business model transition in GitHub Copilot that will align pricing with usage and value that takes effect on June 1 of this year.&#171;</p></div><p>Microsoft is admitting that GitHub Copilot adoption is now so heavy that it is dragging down Microsoft Cloud gross margins because they were charging a flat per-seat fee while costs scale with usage. So they are switching to usage pricing on June 1. This is a small story right now, but it is a leading indicator for the entire industry: per-seat pricing on AI products is going to be replaced with consumption pricing, because the cost-to-serve scales with intensity of use, not with seat count. The bigger investor takeaway is what management said elsewhere on the call:</p><div class="pullquote"><p>&#187;Bookings growth was impacted by weaker renewals as customers balance spend between the traditional per-seat and the emerging seats-plus-consumption model.&#171;</p></div><p>So enterprises are already adjusting their procurement around this. Hybrid pricing models are becoming the standard.</p><p>On capacity, Microsoft was, again, very direct:</p><div class="pullquote"><p>&#187;Even with these additional investments and continued efforts to bring GPU, CPU and storage capacity online faster, we expect to remain constrained at least through 2026. Despite these constraints, and the continued need to balance incoming supply, we expect Azure growth to show modest acceleration in the second half of the calendar year compared with the first half.&#171;</p></div><p>And then this:</p><div class="pullquote"><p>&#187;I think in so many ways, this just reminds us of the last cycle. And when the TAM is so expansive and when shortages are generally, I think, growing seems to be the sentiment between supply and demand&#171;</p></div><p>The &#8220;supply shortage growing&#8221; line is important because it is the third quarter in a row where Microsoft has said this.</p><p>The other big tell came on Foundry, Microsoft&#8217;s model marketplace:</p><div class="pullquote"><p>&#187;Over 10,000 customers have used more than one model on Foundry. 5,000 have used open source models, and the number who have used Anthropic and OpenAI models increased 2x quarter-over-quarter.&#171;</p><p>&#187;The majority of users leverage multiple models.&#171;</p></div><p>Note the Anthropic mention here. Microsoft is now a meaningful Anthropic distribution channel. This matters because a year ago, the bull case for Microsoft was &#8220;Microsoft + OpenAI.&#8221; Today, the company is hosting Anthropic, OpenAI, and open-source models in the same product. The orchestration layer could become the moat, not the model and Microsoft could benefit from it.</p><p>The CFO also gave us perhaps the most important signal for what to expect in fiscal Q4 (calendar Q2):</p><div class="pullquote"><p>&#187;We&#8217;re guiding for that to be better again in Q4. I think that&#8217;s where you&#8217;re starting to see, right? I think the thing that investors have been asking and Mark, you&#8217;re asking about is when we&#8217;ll start to see that show up in revenue growth. And I think that&#8217;s the first place you point to. We can also point to it, and I think you&#8217;ll start to see it in GitHub, right, where you see revenue growth rates and usage consumption models result in acceleration in the top line.&#171;</p></div><p>In other words, the next two quarters are when Microsoft expects AI products to start showing up as accelerating revenue, not just bookings.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong>Amazon: AWS Reaccelerates To 28% (Fastest In 15 Quarters), And The Chips Business Is Now Top-3 In The World</strong></p><p>AWS, which a lot of people had been writing off as the &#8220;loser&#8221; of the cloud AI race, posted its strongest growth rate since 2022:</p><div class="pullquote"><p>&#187;Starting with AWS, growth continued to accelerate, up 28% year-over-year, the fastest growth rate in 15 quarters, up $2 billion quarter-over-quarter, the largest Q4 to Q1 AWS revenue increase ever. AWS is now a $150 billion annualized revenue run rate business&#171;</p></div><p>So AWS is now growing 28% YoY (vs 24% last quarter) on a $150B run-rate base. The sequential add of $2B is the largest Q4-to-Q1 increase in AWS history.</p><p>Bedrock is the key to it:</p><div class="pullquote"><p>&#187;high-performance inference with the leading selection of frontier models in Bedrock, which saw 170% growth in customer spend quarter-over-quarter and processed more tokens in Q1 than all prior years combined.&#171;</p></div><p>Bedrock processed more tokens in Q1 2026 than in all prior years combined. Customer spend on Bedrock grew 170% QoQ. There has been a real perception in the market that AWS was behind on the AI inference story, and that is just no longer true based on these numbers.</p><p>Second, Trainium. This is the segment of the call that I think investors really need to focus on, because it is reshaping the unit economics of AWS:</p><div class="pullquote"><p>&#187;We saw nearly 40% quarter-over-quarter growth in Q1, and our annual revenue run rate is now over $20 billion and growing triple-digit percentages year-over-year, but this somewhat masks the size. If our chips business was a stand-alone business and sold chips produced this year to AWS and other third parties as other leading chip companies do, our annual revenue run rate would be $50 billion. As best as we can tell, our custom silicon business is now one of the top 3 data center chip businesses in the world, the speed at which we&#8217;ve gotten here is extraordinary. And we have momentum.&#171;</p></div><p>Jassy already mentioned this in his recent shareholder letter, and I have written about it, but still, Amazon is now a top-3 data center chip business globally. That is a statement most investors are not pricing in, because they think of Amazon as a retailer + cloud company, not a chip company.</p><p>And the demand signal on Trainium specifically is just enormous:</p><div class="callout-block" data-callout="true"><p>&#187;And we now have over $225 billion in revenue commitments for Trainium.&#171;</p></div><p>$225B in commitments just for Trainium is insane and was the most shocking number for me among all the earnings calls yesterday.</p><div class="pullquote"><p>&#187;Amazon Bedrock, which is used expansively by over 125,000 customers, runs most of its inference on Trainium and almost 80% of the Fortune 100 companies are using Bedrock.&#171;</p><p>&#187;While the largest number of AI chips we&#8217;re bringing in are Trainium, we continue to have a deep partnership with NVIDIA&#171;</p></div><p>The fact that AWS is now bringing in more Trainium chips than NVIDIA chips on a unit basis is also an important shift. It is going to flow through to AWS margins over time. The CFO basically said this directly:</p><div class="pullquote"><p>&#187;Different companies will offer different benefits for customers and the uniquely strong price performance that Trainium offers is compelling to our external and internal customers. For perspective, at scale, we expect Trainium will save us tens of billions of dollars of CapEx each year and provide several hundred basis points of operating margin advantage versus relying on others&#8217; chips for inference.&#171;</p></div><p>Several hundred basis points of operating margin advantage. AWS operating margin in Q1 was already strong at 37.7% (up from 35% the last quarter), and management is telling us Trainium adds several hundred more bps over time. This is exactly the same dynamic Google has been showing with TPUs, where Google Cloud operating margin went from 17.8% to 32.9% YoY. The custom silicon advantage will be crucial in the coming years.</p><p>The other under-appreciated comment was on the relationship between AI and core cloud spend:</p><div class="pullquote"><p>&#187;We continue to see customers increase cloud migrations and scale their use of AWS core services. Customers seeking the full benefit of AI are accelerating their transition to the cloud. We also see a strong correlation between AI spend and core growth. As customers spend more on AI, we see a corresponding demand increase in core.&#171;</p></div><p>This is the &#8220;AI lock-in&#8221; effect I have written about before. Companies that adopt AI workloads end up moving more of their non-AI workloads to the cloud as well, because the data needs to be co-located. This is why core (non-AI) cloud spend is also accelerating - which is something a lot of people miss when they look at AWS AI growth in isolation.</p><p>Backlog at AWS:</p><div class="pullquote"><p>&#187;On the backlog, the backlog for Q1 is $364 billion. That does not include the recent deal that we announced with Anthropic for over $100 billion. There&#8217;s reasonable breadth in that as well. It&#8217;s not just 1 customer or 2 customers.&#171;</p></div><p>$364B in backlog ex-Anthropic deal. The &#8220;reasonable breadth&#8221; comment is important - this is not a one-customer book like Microsoft&#8217;s relationship with OpenAI. The AWS book is more diversified.</p><p>And on selling chips outside the cloud, similar to the Google TPU pivot:</p><div class="pullquote"><p>&#187;On the question about Trainium and the notion of our selling racks over time, I do think that&#8217;s very much a possibility. Always, we have to balance -- we have such demand right now for Trainium, and we have such demand from various companies who will consume as much as we make that we have to decide how much we&#8217;re going to allocate to the existing demand and customers and how much we&#8217;re going to save to sell as racks. ... But I expect over time, there&#8217;s a good chance we&#8217;re going to sell racks over the next couple of years.&#171;</p></div><p>The most important framing on AI workloads from the entire AWS call:</p><div class="pullquote"><p>&#187;Most of the value companies derive from AI will be through agents&#171;</p></div><p>The chatbot era was really just us trying to use a new technology in a way that we knew so far (information retrieval via Google Search). Agentic AI - where models execute multi-step workflows on behalf of users autonomously - is where the real economic unlock of LLMs lies and requires orders of magnitude more compute per task. We are still very early in this shift.</p><p><strong>Meta: Revenue Up 33%, The Fastest In 4 Years, Because The Ad Engine Is An AI Engine Now</strong></p><p>Revenue growth 33% is the fastest in the last 4 years.</p><p>$56.3B in Q1 revenue, +33% YoY. To put that in context, in Q1 2025, Meta grew 16%. So the company has roughly doubled its growth rate in 12 months, on a $200B+ annualized base. That is essentially impossible without a structural change in the underlying engine, and the structural change is that Meta&#8217;s ad ranking and content recommendation systems are now LLM-scale.</p><p>On the engagement side:</p><div class="pullquote"><p>&#187;On Instagram, the ranking improvements that we made in Q1 drove a 10% lift in reel time spent.&#171;</p><p>&#187;On Facebook, total video time increased more than 8% globally in Q1, the largest quarter-over-quarter gain in 4 years. Within the U.S. and Canada, ranking improvements we made drove a 9% increase in video watch time on Facebook in Q1.&#171;</p></div><p>Note the Facebook number specifically. Facebook is a 20-year-old product, and it is posting the largest QoQ gain in video time in 4 years. That doesn&#8217;t happen organically. That happens because Meta deployed a new ranking model that materially improves the relevance of what users see.</p><p>The under-appreciated piece here is what Zuck said about how ad ranking actually works now:</p><div class="pullquote"><p>&#187;In the second half of last year, we began rolling out our new adaptive ranking model, which is an LLM scale adds recommender model that we use for inference. This model improves our inference ROI by routing requests to more compute-intensive inference models when it determines there is a higher probability of conversion.&#171;</p></div><p>And then this longer segment, which I think is one of the most important pieces of technical commentary:</p><div class="pullquote"><p>&#187;Historically, we haven&#8217;t used larger model architectures like GEM for inference, because their size and complexity would make them too cost prohibitive. And the way we drive performance from those models is by using them to transfer knowledge to smaller, more lightweight models that are used at run time. The inference models are bound by strict latency requirements since they need to find the right ad within milliseconds, and that has, again, historically prevented us from meaningfully sizing up -- scaling up their size and complexity. But in the second half of last year, we introduced a new adaptive ranking model, which enables us to leverage LLM scale model complexity of 1 trillion parameters, and we made advances in the model architecture and codesigned the system with the underlying silicon, so it maintains the sub-second speed that is required to serve ads at scale. We also developed an approach that intelligently routes requests more compute-intensive inference models if it determines that there is a higher probability of conversion and that lets us drive both better performance and increase inference ROI.&#171;</p></div><p>This is important for the whole industry.</p><p>For years, Meta could not use LLM-scale models (think GPT-style models with hundreds of billions to trillions of parameters) for ad ranking because the latency would be too high. When you load a Facebook feed, the ad selection has to happen in well under 100ms. LLM-scale models can take seconds to respond, which is way too slow.</p><p>What Meta has now done is two things: (1) they have re-architected the model and co-designed it with their custom silicon (this is the Broadcom-partnered ASIC Zuck referenced) so that a 1 trillion parameter model can run within sub-second latency, and (2) they have built a &#8220;router&#8221; that decides when to use the big expensive model vs. the small cheap one, based on the predicted probability that the ad will convert. So if you are clearly not going to click on an ad about a car, Meta won&#8217;t waste compute running the big model on you. But if you are showing strong purchase intent, they will use the most expensive model to find the absolute best ad to show you.</p><p>This is a fundamental change in the unit economics of advertising. It is why the average price per ad on Meta increased 12% YoY while ad impressions grew 19% YoY - Meta is showing more ads AND each ad is more valuable.</p><p>Now, on the macro AI strategy, Zuck is doubling down. From the call:</p><div class="pullquote"><p>&#187;we are increasing our infrastructure CapEx forecast for this year. Most of that is due to higher component costs, particularly memory pricing, but every sign that we&#8217;re seeing in our own work and across the industry gives us confidence in this investment.&#171;</p></div><p>The memory pricing call-out is something I have written about extensively. HBM is sold out through 2026, prices are spiking, and Meta is essentially saying: yes, our CapEx is going up, but it is partly because the price of the input is going up, not because we are buying more capacity than we expected.</p><p>On the custom ASIC story:</p><div class="pullquote"><p>&#187;That said, we are very focused on increasing the efficiency of our investments, and as part of that, we are rolling out more than 1 gigawatt of our own custom silicon that we&#8217;re developing with Broadcom, as well as a significant amount of AMD chips to complement the new NVIDIA systems that we&#8217;re rolling out as well.&#171;</p></div><p>A gigawatt of custom silicon is meaningful - that is a very serious deployment. This is the same playbook Google ran with TPUs and Amazon ran with Trainium. Meta is now in the custom silicon game in a real way, with Broadcom as the partner. That has implications both for Meta&#8217;s long-term margin profile and for Broadcom&#8217;s revenue trajectory.</p><p>On the broader AI investment thesis from Zuck:</p><div class="pullquote"><p>&#187;So you&#8217;re getting to a point where today, the models are still able to learn from people and then I think at some point, the models will have to improve themselves. And that&#8217;s how the growth is going to -- an improvement in the models is going to happen. And if you don&#8217;t -- if we don&#8217;t have an ability to do that, then we or anyone else, I think the companies that don&#8217;t do that are not going to be leading labs, then they&#8217;re not going to produce leading products. So I think that, that&#8217;s like -- that is a table stakes thing that we are focused on.&#171;</p><p>&#187;but then the model improvement, I think, is going to be something that&#8217;s going to go on for a very long time.&#171;</p></div><p>Zuck&#8217;s view that model improvement continues for a long time is a meaningful statement against the &#8220;scaling laws are over&#8221; thesis. He is essentially saying the opposite - he sees a long runway of improvement, and Meta is committing capital accordingly.</p><p>One thing I do want to flag for investors is the nuance of how Zuck is positioning Meta&#8217;s AI strategy versus other labs:</p><div class="pullquote"><p>&#187;That&#8217;s why we believe that we need to be a company that builds frontier models in addition to building the agents. And then in order to do that, you, of course, need to build your infrastructure in order to be able to do that well. So we&#8217;re undertaking this large investment to be able to do that top to bottom.&#171;</p><p>&#187;But I don&#8217;t hear any other labs out there talking about how they&#8217;re building an AI that&#8217;s really good at shopping. And I think that the reason for that is like not because shopping is the most important thing by itself, but because like empowering people to do the things that matter in their lives, whether that&#8217;s local or understanding social context, or shopping or personal health things or understanding what&#8217;s going on around them visually&#171;</p></div><p>This is Meta differentiating its AI strategy from OpenAI/Anthropic. Those labs are building horizontal foundation models. Meta is building vertical AI that is really good at the things people do on Meta&#8217;s surfaces - shopping, social, content discovery. It is a different bet, and arguably a much more defensible one given Meta&#8217;s distribution.</p><p><strong>The Pattern: AI Workloads Are Now Margin-Accretive At Scale, And Custom Silicon Is The Reason</strong></p><p>If I step back and try to find the single most important pattern across all four prints, here it is:</p><p>The bear case on hyperscaler AI spending - that AI workloads have structurally lower margins than legacy cloud workloads, and therefore CapEx returns will be poor - has been wrong.</p><p>Look at the hard data:</p><ul><li><p>Google Cloud operating margin: 17.8% Q1 2025 &#8594; 32.9% Q1 2026 (+1,510bps YoY)</p></li><li><p>AWS Q1 operating margin: 31.4% (Q1 2025) &#8594; 37.7% (Q1 2026)</p></li><li><p>Microsoft saying explicitly &#8220;margins were actually better and they remained better in our AI business versus where we saw in the cloud transition&#8221;</p></li><li><p>Amazon saying Trainium gives them &#8220;several hundred basis points of operating margin advantage versus relying on others&#8217; chips for inference&#8221;</p></li></ul><p>One of the structural reasons for the margin&#8217;s being stable is custom silicon: TPUs, Trainium. The cloud providers have figured out that they can&#8217;t allow NVIDIA to keep 75% gross margins on the most valuable workloads of the next decade, so they are vertically integrating into chips.</p><p>The other pattern is that AI is a significant core growth driver in the ad business at Meta and Google.</p><p>Both companies now run their ad ranking on LLM-scale models with adaptive routing. Both companies are showing direct evidence that AI-driven ranking improvements are translating into both more ad inventory consumed (impressions up) and higher revenue per impression (price per ad up). The ad business is no longer a separate thing from the AI business - the ad business is the AI business now.</p><p>For Microsoft and Amazon, the equivalent flywheel is happening in cloud + agents. Microsoft&#8217;s AI ARR hit $37B at +123% YoY. Amazon&#8217;s Bedrock token volume in Q1 alone exceeded all of 2025 combined.</p><p><strong>The Bigger Picture: The Era Of Subsidized Compute Is Ending</strong></p><p>A theme I want to leave you with, because I think it is the most important macro shift for the next 12-18 months:</p><p>It really comes down to what the end cost of intelligence is and how much companies are willing to pay for it. The whole industry would benefit from an architectural change that would bring new efficiencies to serving these models. It&#8217;s not who has the best model but who can solve the economic task with the least cost (using hardware, cloud infra, models, harness everything). The era of subsidizing compute is over - you can even see it from GitHub. Companies will have to shift budgets from other OpEx items towards compute, the moment is here, and this will only increase as Anthropic and OpenAI go IPO and have to produce &#187;passable&#171; gross margins.</p><p>The GitHub Copilot pricing change Microsoft announced is the canary in the coal mine here. You can&#8217;t have flat per-seat pricing for a product whose cost-to-serve scales with usage intensity. Either prices go up for heavy users or the seller bleeds gross margin. Microsoft chose to raise prices on the heavy users via consumption pricing. Anthropic is doing the same thing - their API pricing has been moving in this direction, and Claude Code&#8217;s heavy users have been hitting rate limits and seeing throttling for months now. The only real question that I believe we will get the answer to soon is how much value end users have and how much they are willing to pay for it.</p><p>As always, I hope you found this article valuable. I would appreciate it if you could share it with people you know who might find it interesting. I also invite you to become a paid subscriber, as paid subscribers get additional articles covering both big tech companies in more detail, as well as mid-cap and small-cap companies that I find interesting.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe&quot;,&quot;text&quot;:&quot;Subscribe to Paid&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe"><span>Subscribe to Paid</span></a></p><p>Thank you!</p><p><strong>Disclaimer:</strong></p><p>I own Google (GOOGL), Amazon (AMZN), Microsoft (MSFT), Meta (META) stock.</p><p>Nothing contained in this website and newsletter should be understood as investment or financial advice. All investment strategies and investments involve the risk of loss. Past performance does not guarantee future results. Everything written and expressed in this newsletter is only the writer&#8217;s opinion and should not be considered investment advice. Before investing in anything, know your risk profile and if needed, consult a professional. Nothing on this site should ever be considered advice, research, or an invitation to buy or sell any securities.</p>]]></content:encoded></item><item><title><![CDATA[Q1 2026 Channel Checks & Alternative Data: Cloud is on Fire]]></title><description><![CDATA[For this report, I covered cloud providers Google, Microsoft, and Amazon in terms of their cloud business, and some insightful signals on Microsoft Copilot, which is a pressure point for Microsoft.]]></description><link>https://www.uncoveralpha.com/p/q1-2026-channel-checks-and-alternative</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/q1-2026-channel-checks-and-alternative</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Fri, 24 Apr 2026 12:02:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4VQ9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a4daaec-27a8-4395-a938-9dc63c6d872e_682x441.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>I am posting my regular channel check &amp; other alternative data report before we start the big tech earnings.</p><p>For this report, I covered cloud providers Google, Microsoft, and Amazon in terms of their cloud business, and some insightful signals on Microsoft Copilot, which is a pressure point for Microsoft.</p><p>Let&#8217;s dive in.</p><p><strong>Cloud is on fire</strong></p><p>Let&#8217;s start with my alt data on the most relevant channel-check interviews from clients, cloud consultants, integrators, and former employees.</p><p>When bulk analyzing these interviews, demand is high for Q1 2026 across all three hyperscalers: AWS, Azure, and GCP. Looking at the future demand pipeline for the next 3-6 months, the results are even more impressive: almost 60% of these experts see demand exceeding their expectations, 26% see it meeting expectations, and 15% see it below expectations. Keep in mind that expectations were already high going into this year and quarter, so for 60% of experts seeing higher-than-expected demand, the signal is very strong. The driver of demand, as expected, is AI workloads and the move from test to production environments, especially with Agentic AI starting to roll out.</p><p>Looking deeper, let&#8217;s look at the individual hyperscaler level and what the data shows. As always, I made the % breakdown of experts who think AWS, Azure, or GCP is accelerating the fastest. Here are the results:</p><p>62% think GCP is growing the fastest, 41% think Azure is growing the fastest, and 27% think AWS is growing the fastest (important note: the sum is greater than 100% because some experts mentioned two platforms as growing at a faster pace than the other).</p><p>Now, this data doesn&#8217;t add much value until we compare it to my historical data from past quarters, as we did in the last reports, to truly understand whether anything shifted significantly in Q1. Here is the data:</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/q1-2026-channel-checks-and-alternative">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Left for Dead on AI, Meta and Amazon Are About to Have the Last Laugh]]></title><description><![CDATA[I break down some significant fundamental shifts when it comes to AI efforts from Amazon and Meta, and why I think both are the next two big AI beneficiaries. Based on what we are seeing, both companies are on a path to reaccelerating their efforts, while still perceived by the market as &#8220;AI laggards&#8221;. We believe this premise will be proven wrong in the coming months.]]></description><link>https://www.uncoveralpha.com/p/left-for-dead-on-ai-meta-and-amazon</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/left-for-dead-on-ai-meta-and-amazon</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Fri, 17 Apr 2026 13:01:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-Jhj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88752802-60c3-4379-ac50-7468e34b7f81_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>In this article, I break down some significant fundamental shifts when it comes to AI efforts from Amazon and Meta, and why I think both are the next two big AI beneficiaries. Based on what we are seeing, both companies are on a path to reaccelerating their efforts, while still perceived by the market as &#8220;AI laggards&#8221;. We believe this premise will be proven wrong in the coming months.</p><p>Let&#8217;s start.</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/left-for-dead-on-ai-meta-and-amazon">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Era of Subsidized AI Model Usage is Over, the IPOs are coming]]></title><description><![CDATA[There are four interconnected themes I want to walk through today, and they all converge on a single conclusion: the era of subsidizing AI model usage is coming to an end.]]></description><link>https://www.uncoveralpha.com/p/the-era-of-subsidized-ai-model-usage</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-era-of-subsidized-ai-model-usage</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Fri, 10 Apr 2026 14:38:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Qcs3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>The AI industry is approaching an inflection point that will reshape the priorities of AI model companies and the entire space. There are four interconnected themes I want to walk through today, and they all converge on a single conclusion: the era of subsidizing AI model usage is coming to an end.</p><p>Here&#8217;s what I cover in this article:</p><ul><li><p>Anthropic is taking over the enterprise &#8212; but the curse of the best model is real</p></li><li><p>OpenAI is losing the enterprise race to Anthropic and facing structural problems heading into its IPO</p></li><li><p>The era of subsidized AI model usage is ending as both companies prepare for public markets</p></li><li><p>The IPO race: who lists first matters more than most people realize</p></li></ul><p>Let&#8217;s get into it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qcs3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qcs3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qcs3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2176992,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/193796773?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qcs3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Qcs3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14f027-f39a-445a-8dd1-62668b239c10_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Anthropic Is Taking Over Enterprise: And the Curse of the Best Model Is Coming for Them</strong></p><p>Anthropic announced that its revenue had surpassed $30 billion, up from $9 billion at the end of 2025. That&#8217;s more than tripling in roughly four months. Anthropic has now supposedly surpassed OpenAI&#8217;s run-rate revenue of approximately $25B.</p><p>The enterprise composition is what separates Anthropic from the rest. Approximately 80% of revenue comes from business customers. The number of customers spending over $1 million annually has doubled to more than 1,000, up from 500+ where it was just 2 months ago. Business subscriptions to Claude Code have quadrupled since the start of 2026.</p><p>And then there&#8217;s Mythos. Just a few days ago, Anthropic announced Claude Mythos Preview - a new general-purpose model that sits in an entirely new tier above Opus. The draft described it as &#8220;by far the most powerful AI model we&#8217;ve ever developed&#8221; and said it is &#8220;very expensive for us to serve, and will be very expensive for our customers to use.&#8221;</p><p>The benchmarks are quite telling: 93.9% on SWE-bench Verified (vs. Opus 4.6&#8217;s ~80.9%), 77.8% on SWE-bench Pro, 82% on Terminal-Bench 2.0, 97.6% on USAMO 2026, and 83.1% on CyberGym vs. Opus 4.6&#8217;s 66.6% &#8212; a 16.5 percentage-point jump on cybersecurity tasks. On Anthropic&#8217;s internal zero-day exploit benchmark, Opus 4.6 had a near 0% success rate at autonomous exploit development. Mythos succeeded 181 times out of several hundred attempts on the same Firefox vulnerability task. It also found thousands of zero-day vulnerabilities across every major operating system and browser, including a 27-year-old bug in OpenBSD and a 16-year-old bug in FFmpeg that automated testing had missed across 5 million test runs.</p><p>Mythos is not being made generally available because Anthropic wants to first roll it out to selected companies because of cybersecurity risks. It&#8217;s deployed through Project Glasswing to 12 partner organizations (Amazon, Apple, Broadcom, Cisco, CrowdStrike, Linux Foundation, Microsoft, Palo Alto Networks, and others) plus about 40 additional organizations, with Anthropic committing $100 million in usage credits. While I am not dismissing any cyber risks that a model like this could bring, it is also convenient that now model providers will &#187;release&#171; these models to a small group of companies, because the reality is that with their current compute, they can&#8217;t even serve Opus 4.6 to their user base let alone Mythos, which is even more expensive to run. The current Mythos Preview, for which these companies got access, is around 5x more expensive than Opus 4.6 after the initial $100M credit commitment from Anthropic based on their pricing. It is also rumored that OpenAI will also &#8220;release&#8221; their newest model in a similar fashion, again citing cybersecurity risks.</p><p>There has been a lot of frustration from Claude users lately, as many have started to hit their rate limits much faster in their subscription plans, as Anthropic is having to manage this surge in demand with the amount of compute that they have.</p><p>This is what I call the Inference Trap, and both OpenAI and Anthropic have now been caught in it.</p><p>The pattern is simple: build the best model &#8594; users surge &#8594; inference compute explodes &#8594; you either throttle users, raise prices, or cannibalize training compute. OpenAI experienced it during the Ghibli moment in March 2025, when ChatGPT gained 1 million new users in a single hour and 100 million signups in a week. Sam Altman admitted they were &#8220;forced to do a lot of unnatural things,&#8221; specifically borrowing compute capacity from OpenAI&#8217;s research division and slowing down the release of new features<strong>.</strong></p><p>Anthropic is living through its own version right now. In March 2026, the company experienced five major platform outages in a single month. Claude Code users reported burning through 5-hour sessions in under 90 minutes. The problem for Anthropic is that if you are a model provider in this AI race, you don&#8217;t want to cannibalize training compute, as it means that you can lose the race for the next model.</p><p>This brings me to another point I want to make: pricing increases on frontier AI models are inevitable.<strong> </strong>When Anthropic eventually deploys Mythos-class models at scale, the inference cost per query will be higher than Opus. And they already can&#8217;t serve Opus at current demand levels without throttling. The math only works if prices go up, or if the compute infrastructure grows fast enough to meet demand, which it can&#8217;t in the short term.</p><p>Moving to OpenAI. </p><p><strong>OpenAI seems to be losing the Enterprise Race And Heading Into an IPO with some headwinds</strong></p><p>While Anthropic is sprinting ahead on enterprise revenue, OpenAI is dealing with a set of problems that are becoming hard to ignore.</p><p>The revenue gap has flipped.<strong> </strong>A year ago, OpenAI was at roughly $6 billion ARR, and Anthropic was at $1 billion. The gap looked huge. Today, Anthropic is at $30 billion, and OpenAI is at $25 or similar to Anthropic, but the pace of growth is slower. Anthropic added roughly $21 billion in net new annualized revenue in just three months. OpenAI&#8217;s enterprise business now makes up 40% of revenue (up from ~30% last year) and is &#8220;on track to reach parity with consumer by the end of 2026&#8221; &#8212; but Anthropic has been enterprise-first from the start, with 80% enterprise revenue and structurally higher retention.</p><p>To add to this, SensorTower data now show that ChatGPT&#8217;s monthly active users in the US have started to fall slightly, adding to the headwinds.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lGin!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lGin!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 424w, https://substackcdn.com/image/fetch/$s_!lGin!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 848w, https://substackcdn.com/image/fetch/$s_!lGin!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!lGin!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lGin!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg" width="970" height="549" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:549,&quot;width&quot;:970,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53591,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/193796773?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lGin!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 424w, https://substackcdn.com/image/fetch/$s_!lGin!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 848w, https://substackcdn.com/image/fetch/$s_!lGin!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!lGin!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a1ee6f2-9457-4399-b285-8009e38ff46d_970x549.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Even without this data, we have been waiting for quite some time for OpenAI to make the update of reaching 1 billion weekly active users, as the 800M mark was announced 6 months ago. Based on this users growth has slowed down.</p><p>The funding structure also isn&#8217;t perfect. OpenAI closed a $122 billion round at an $852 billion valuation at the end of March&#8212; the largest private funding round in history. But the structure is telling. Amazon committed $50 billion, but only $15 billion arrived as upfront cash. The remaining $35 billion is conditional, tied to milestones that some indicate may include achieving certain AI capability thresholds or pursuing an initial public offering by the end of 2026.</p><p>SoftBank pledged $30 billion structured in three equal tranches of $10 billion each, arriving in April , July, and October. SoftBank&#8217;s structure essentially assumes a liquidity event within that window. About $3 billion came from retail investors through bank channels, and OpenAI was included in ARK Invest ETFs.</p><p>In other words, a significant portion of OpenAI&#8217;s headline $122 billion raise is conditional on an IPO actually happening. This means OpenAI is under pressure to go public regardless of whether the timing is optimal. There are now reports from The Information that Altman and OpenAI&#8217;s CFO are on different sides over the IPO, as Altman is pushing for an IPO this year, while the CFO believes OpenAI is not ready yet. OpenAI has already denied this, but it's not expected that any company would confirm such rumors, even if they were true.</p><p>The profitability picture is the key thing.<strong> </strong>According to different reports, OpenAI&#8217;s gross margins sit at approximately 40%, constrained by variable compute costs. The company is generating +$2-3 billion per month but losing +$14 billion per year. Reports of internal documents project that compute costs will reach $121 billion by 2028, with a cumulative loss trajectory that doesn&#8217;t reach breakeven until 2029-2030. Compare this to Anthropic, which projects positive free cash flow by 2027-2028 while spending roughly 4x less on training.</p><p>OpenAI also has an alternative to Claude Code called Codex, but adoption there, although growing, doesn&#8217;t seem to be at the same pace as Claude Code. It&#8217;s telling that OpenAI is even offering users more token usage, while Anthropic is limiting it.</p><p>As this data from Ramp shows, the AI model share of first-time enterprise customers has heavily tilted towards Anthropic in the last few months:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!iYpU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!iYpU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 424w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 848w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 1272w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!iYpU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png" width="999" height="692" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:692,&quot;width&quot;:999,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:56198,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/193796773?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!iYpU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 424w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 848w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 1272w, https://substackcdn.com/image/fetch/$s_!iYpU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F671f46db-dbea-4572-ac97-ab2539f0ea67_999x692.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The key moat that OpenAI has is the ChatGPT brand, which, as a first mover, created a verb similar to Google when it comes to consumers. In the last months, however, Claude and Anthropic have become &#8220;the verb&#8221; when it comes to enterprises and AI use cases for work. On the consumer end, OpenAI will probably have to shift hard towards an ad-supported business model to get some revenue from its big free user base. Building an efficient ad platform is much more complex than most think and requires time. At the same time, OpenAI is trying to stop Anthropic in the enterprise market, but so far, it doesn&#8217;t seem to be working as Anthropic is capturing the market at a faster pace. The question is whether OpenAI&#8217;s strategy of trying to capture both markets at once and &#8220;doing everything&#8221; is really the right one. I would argue that it is not. Now you even have Meta entering the AI arena again, with its first AI model since the formation of Meta's superintelligence unit. While their model is not SOTA, there are specific use cases where it is very competitive. Meta focused on use cases like health, social media, games, and shopping. This pattern will become more dominant in the coming years as the AI model market matures and you see model specialization rather than just general models. In this environment, the importance of having a narrow focus becomes even bigger.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p>According to The Information, OpenAI is now communicating to investors that they believe one of their important advantages going forward vs Anthropic is in the availability of compute. OpenAI said it believes Anthropic had 1.4 GW of capacity at the end of last year, while OpenAI had 1.9 GW. But OpenAI said it plans to ramp its capacity more steeply, with total gigawatts in the mid-single digit range at the end of this year and more than 10 GW in 2027.</p><p>In contrast, it believes Anthropic will have 3 to 4 GW in 2026 and 7 to 8 GW by the end of 2027. While I agree that availability of compute is a big factor going forward, I believe the main one and even more important is inference economics and the costs of serving the models to your clients. As users shift workloads to production, availability and reliability become key. Nobody wants to have a product that is unstable and sometimes works great, but other times is not available.</p><p><strong>The End of Subsidized AI Model Usage</strong></p><p>Both Anthropic and OpenAI are preparing for IPOs. And an IPO changes everything about how an AI company thinks about compute costs.</p><p>When you&#8217;re private and burning venture capital, you can subsidize inference. You can run models at a loss. You can offer $20/month unlimited plans that cost you +$100/ month to serve. You can double the rate limits as promotions. You can hand out credit packages. The goal is growth at all costs, because the next funding round values you on revenue, not margins.</p><p>When you&#8217;re public, the scrutiny shifts to unit economics.<strong> </strong>Gross margins, operating margins, cash burn trajectory, path to profitability &#8212; these become the metrics that determine your stock price. An S-1 filing forces you to disclose all of this in audited detail.</p><p>Both OpenAI and Anthropic are probably operating at approximately 40% gross margins, constrained by the variable cost of running inference. And while no company will show net profit as they are projected towards the end of 2030, the gross margin will be something that investors will particularly keep an eye on, especially the gross margin on inference. While the margin profile math works on API usage, the one in the subscription packages is still often subsidized by the model companies, and this is something they will look to tweak in the coming months before they file their S-1s.</p><p>This has a ripple effect across the entire supply chain industry.<strong>  </strong>For the past three years, the question has been: &#8220;Which chip delivers the most FLOPS?&#8221; The answer was always Nvidia, and companies paid whatever Nvidia charged because performance was the bottleneck, and money was abundant.</p><p>Going forward, the question becomes: &#8220;Which chip gives me the cheapest tokens and the best total cost of ownership (TCO)?&#8221;<strong> </strong>It&#8217;s not even about watts anymore, as you can see from recent podcast comments from Nvidia&#8217;s Jensen and Google Sundar &#8212; it&#8217;s about cost-per-token, because that&#8217;s what directly determines your gross margin as a public company, and that is what investors will be laser focused on. This shift also means the end of subsidizing usage in these subscription packages, with either usage limits or higher prices. We will also see even more resources being focused on software optimizations to run the model. Savings on memory and getting more from existing hardware will be the focus in the coming months for both labs, as they are constrained by compute.</p><p>Both of these companies also need to make the hard math of the IPO being the &#8220;last&#8221; funding round, and after the IPO, have enough capital that will be able to support their growth and cash burn for the coming years. Issuing additional stock for raising capital once you are a public company is never looked at positively by the market, so nobody wants to go down that route.</p><p>Anthropic is uniquely well-positioned here<strong> </strong>because it runs Claude on a diversified hardware stack across three suppliers: Nvidia GPUs, Google TPUs, and Amazon Trainium. This gives it real negotiating leverage and the ability to route workloads to whichever chip offers the best price-performance for each model tier. The company just announced a deal with Google and Broadcom for approximately 3.5 gigawatts of next-generation TPU capacity starting in 2027. This is on top of its existing AWS Trainium partnership and Nvidia GPU deployments. It is worth noting that AWS&#8217;s CEO just mentioned in an interview on CNBC yesterday that all of Anthropic&#8217;s AI models were trained on Amazon Trainium (even Mythos).</p><p>OpenAI, by contrast, has been more dependent on Nvidia through its Azure partnership with Microsoft, though it has been diversifying toward custom silicon.</p><p><strong>The IPO Race: Who Lists First Gets the Biggest Check</strong></p><p>There&#8217;s one final dynamic that ties all of this together: both Anthropic and OpenAI know that whoever goes public first has a significant advantage, and the window for both is narrowing. On top of those, SpaceX (which now includes xAI) is also racing towards a +$1T IPO. OpenAI is targeting a +$1T IPO, Anthropic just closed a funding round valued at $380 billion, but because of the surge in usage and revenue is already valued at around $500-$700 billion in secondary listings, so the IPO could be in the $800B-$1T range as well.</p><p>Between these three companies alone, we&#8217;re looking at potentially +$200 billion in capital being raised from public markets within a 6-12 month window. That&#8217;s an enormous liquidity event. For context, the entire US IPO market raised approximately $33 billion in 2024. Even in the hot 2021 market, total US IPO proceeds were around $140 billion.</p><p>This is why the race to go first matters so much. The first to market captures the freshest investor capital and sets the valuation benchmark. The second has to compete for the same institutional allocation. The third might struggle if the market has indigestion from the first two.</p><p>OpenAI&#8217;s board is reportedly concerned that if Anthropic lists first, it could set a valuation benchmark that makes OpenAI&#8217;s $1 trillion target look stretched &#8212; especially now that Anthropic has higher revenue, better enterprise concentration, and a more credible path to profitability. On the other hand, if OpenAI lists first, it establishes itself as the &#8220;AI category-defining IPO&#8221; and benefits from a first-mover premium in public market pricing.</p><p>Both companies know this. Both are preparing in parallel. And both are racing against time, because every month that passes, compute costs pile up, margins need to improve, and the public market window could shift with macro conditions.</p><p><strong>Summary</strong></p><p>We are entering a new key period in AI where unit economics take front stage. At the same time, I expect we will see a rapid pace of software optimizations to more efficiently serve these models in the coming months as AI labs put their best talent towards solving this task because it has now become the most important thing that is limiting growth and profitability. The software optimization will focus on resolving key bottlenecks, such as memory (KV cache, context window) and wafer availability. Model distillation and the trend toward smaller models will also grow faster than before because of this. If I were to speculate, the hardware companies might have a &#8220;less golden&#8221; time than the era they have had so far, while cloud providers might benefit the most as these software optimizations mean that they get more juice out of their existing infrastructure, while demand for compute still keeps on surging, because of wider adoption of AI.</p><p>Until next time,</p><p>Next week, we are publishing an article on some key developments in the AI space when it comes to Meta and Amazon, exclusive for paid subscribers. If you are not yet a paid subscriber, consider signing up.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe&quot;,&quot;text&quot;:&quot;Become paid subscriber&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe"><span>Become paid subscriber</span></a></p><p>Thank you!</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/p/the-era-of-subsidized-ai-model-usage?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/p/the-era-of-subsidized-ai-model-usage?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p><strong>Disclaimer:</strong></p><p>I own Google (GOOGL), Amazon (AMZN), Microsoft (MSFT), Meta (META) stock.</p><p>Nothing contained in this website and newsletter should be understood as investment or financial advice. All investment strategies and investments involve the risk of loss. Past performance does not guarantee future results. Everything written and expressed in this newsletter is only the writer&#8217;s opinion and should not be considered investment advice. Before investing in anything, know your risk profile and if needed, consult a professional. Nothing on this site should ever be considered advice, research, or an invitation to buy or sell any securities.</p>]]></content:encoded></item><item><title><![CDATA[Every Memory Cycle Ends the Same. Until It Doesn't.]]></title><description><![CDATA[I&#8217;ve studied every major memory cycle of the last 30 years. In this article, we look at them and the numbers. But then I am going to make a case for why the AI era may fundamentally break that pattern.]]></description><link>https://www.uncoveralpha.com/p/every-memory-cycle-ends-the-same</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/every-memory-cycle-ends-the-same</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 12 Mar 2026 12:19:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nCTv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>For three decades, the memory semiconductor industry has followed a brutal and predictable pattern: prices boom, manufacturers over-invest, supply floods in, prices crash, everyone bleeds red ink, and then the whole thing starts over. It&#8217;s been one of the most reliably cyclical businesses in all of technology. The cycle has destroyed shareholder value, bankrupted companies, and taught every investor the same lesson: never trust the words &#8220;this time is different&#8221; when it comes to DRAM.</p><p>And yet, here I am, writing an article arguing exactly that.</p><p>Let me be clear, I know the history. I&#8217;ve studied every major memory cycle of the last 30 years. In this article, we look at them and the numbers. But then I am going to make a case for why the AI era may fundamentally break that pattern, not because demand will be infinite (it won&#8217;t), but because the nature of what memory serves has changed in a way that most investors haven&#8217;t fully internalized.</p><p>Memory is no longer just a component inside your gadget. Memory is becoming a raw input for intelligence. And the demand curve for intelligence looks a lot more like the demand curve for energy, electricity, than it does the demand curve for smartphones.</p><p>Let&#8217;s start.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nCTv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nCTv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nCTv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2475077,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/190710023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nCTv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nCTv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc048e138-3836-4705-8e7f-9d9b255003f7_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The history of memory economics</strong></p><p>For those less familiar with the space, the memory semiconductor market is dominated by three players: Samsung Electronics (South Korea), SK Hynix (South Korea), and Micron Technology (United States). Together, these three companies control approximately 95% of global DRAM production. This is an oligopoly, but not one that has historically behaved like one. Unlike OPEC, these companies can&#8217;t (legally) coordinate output. And unlike logic chips, memory is essentially a commodity&#8212;a bit is a bit. The differentiation comes from process technology, cost structure, and increasingly, product mix (more on HBM later).</p><p>The fundamental problem with memory economics is the mismatch between demand elasticity and supply inelasticity. Building a new DRAM fab costs $15-20 billion and takes 2-3 years. Once built, the economics favor running it at maximum utilization because fixed costs are enormous. So when demand rises, prices spike because supply can&#8217;t respond quickly. When manufacturers finally bring new capacity online, they tend to overshoot, because everyone is building at the same time based on the same rosy demand signals. Prices crash, margins collapse. Some companies go bankrupt or get acquired. The survivors cut capex, and the cycle begins anew.</p><p>This is the pattern. And it has repeated with remarkable consistency.</p><p><strong>Cycle 1: The Windows PC supercycle (1993-1996)</strong></p><p>The first modern memory supercycle was driven by the explosion of Windows PCs and graphical operating systems. Average DRAM content per PC jumped from roughly 1-2MB to 4-8MB&#8212;a 4x increase per device&#8212;while PC unit shipments were growing at double-digit rates.</p><p>During 1993 and 1994, DRAM demand outpaced supply despite most fabs running at full utilization. Spot and contract prices for 4Mb and 16Mb DRAM rose sharply, and gross margins for leading suppliers surged well above 50%. Korean memory makers like Samsung and Hyundai (now SK Hynix) posted record profits. Semiconductors accounted for 13.4% of Korea&#8217;s total exports. It was hailed as the greatest boom in Korean industrial history.</p><p>Then reality hit. Roughly 50 fab construction plans were announced during 1995-1996 alone. Capex as a percentage of semiconductor production exceeded 30%. The inevitable happened: DRAM prices peaked in late 1995 and then collapsed&#8212;falling 51% in 1996 and another 65% in 1997. Korea&#8217;s Big Three chipmakers suffered from overexpansion, and the resulting shock contributed to the Asian Financial Crisis that pushed Korea into a deep recession. Stock prices of memory companies fell 60-80% from peak to trough.</p><p>Looking at the data cycle duration (peak to trough): around 2 years. Price declines 51% in year one, 65% in year two, and the stock declines around 60-80%.</p><p><strong>Cycle 2: The cloud and smartphone era (2016-2019)</strong></p><p>Fast forward two decades, and the cast of characters had changed, but the script was the same. By 2016, the DRAM market had consolidated from roughly 20 players to just three. This was supposed to introduce discipline. And for a while, it seemed like it did.</p><p>The 2016-2018 &#8220;supercycle&#8221; was driven by a convergence of factors: smartphone storage capacity upgrades, the early cloud buildout, and a supply-side twist where manufacturers were shifting capacity to 3D NAND production, which temporarily constrained conventional DRAM output.</p><p>The numbers were spectacular, especially for Micron, the only publicly traded pure-play memory company in the U.S.:</p><p>Micron 2016:<strong> </strong>Revenue of $12.4 billion, gross margin of 20.2%, operating income of just $168 million (1.4% operating margin). The company was barely above breakeven.</p><p>Micron 2017:<strong> </strong>Revenue surged 64% to $20.3 billion. Gross margin expanded to 41.5%. Operating income hit $5.87 billion (28.9% margin).</p><p>Micron 2018:<strong> </strong>Revenue jumped another 50% to $30.4 billion. Gross margin peaked at 58.9%. Operating income reached an astonishing $15.0 billion&#8212;a 49.3% operating margin. From barely profitable to printing nearly 50 cents of operating profit on every dollar of revenue in two years.</p><p>SK Hynix followed a similar trajectory. At its Q3 2018 peak, SK Hynix posted an operating profit of 6.47 trillion Korean won, which at the time was a record.</p><p>DDR4 retail RAM prices doubled over the course of 2017 into early 2018. Industry inventories fell to 3-4 weeks, well below the normal 8-week average.</p><p>Micron&#8217;s stock peaked at roughly $64 in May 2018.<strong> </strong>But notice- revenue and margins didn&#8217;t peak until Q4 of calendar 2018. The stock topped out approximately two quarters before the fundamental peak. This is a classic pattern in cyclical stocks: the market discounts the turn before it shows up in the numbers.</p><p>Then came the crash:</p><p>Micron 2019<strong>: </strong>Revenue fell to $23.4 billion (-23%). Gross margin compressed to 45.7%.</p><p>Micron 2020:<strong> </strong>Revenue dropped further to $21.4 billion. Gross margin fell to 30.6%. Operating income was $3.0 billion, down 80% from the 2018 peak.</p><p>By December 2018, Micron&#8217;s stock had fallen to approximately $28&#8212;a 56% decline from the May high. The stock was pricing in the downturn even as the company was still reporting near-peak earnings.</p><p>Cycle duration (peak to trough in fundamentals): ~6-7 quarters. Revenue decline (peak to trough): ~30% Gross margin decline: from 59% to 27% (at the Q1 FY2020 low) stock decline (peak to trough): ~56%.</p><p><strong>Cycle 3: The COVID cycle (2020-2023)</strong></p><p>The pandemic created an unexpected demand surge. PC shipments exploded as the world went remote. Server demand spiked as cloud usage accelerated. 5G phones launched with higher per-device memory content. The upcycle lasted approximately 14 months before the familiar reversal kicked in.</p><p>By 2022-2023, the downturn was severe. Bloated inventories from pandemic over-ordering met weakening consumer demand. SK Hynix posted a full-year 2023 net margin of approximately negative 28%. Micron&#8217;s 2024 revenue dropped to around $25 billion with gross margins compressing toward the low 20s.</p><p>Memory stocks cratered. Micron fell from around $98 in early 2022 to roughly $49 by late 2022&#8212;a 50% haircut. SK Hynix fell similarly.</p><p>Cycle duration (peak to trough): 6-8 quarters of margin compression. Operating margins went from 30%+ to deeply negative for SK Hynix. Stock decline: ~50%</p><p>The pattern across all three cycles is strikingly consistent: a demand-driven boom lasting 4-7 quarters, followed by an oversupply-driven bust lasting 4-8 quarters, with revenue declines of 25-40%, margin compression from peak levels above 50% to the low 20s or even negative, and stock price declines of 50-60% that lead the fundamental downturn by 1-2 quarters.</p><p>The history is clear, but now let me tell you why I think this cycle might be structurally different.</p><p><strong>From gadget component to intelligence input</strong></p><p>In every previous memory cycle, the demand driver was the same: humans buying devices.<strong> </strong>PCs in the 1990s. Smartphones in the 2010s. Laptops during COVID. The demand function was ultimately capped by the number of humans and the number of devices each human needs. One person buys one phone. Maybe one laptop. Perhaps a tablet. The DRAM content per device grows, but the number of endpoints is bounded.</p><p>This meant that once the initial adoption or upgrade wave passed&#8212;once everyone who needed a new PC had bought one, or every smartphone had been upgraded to the latest generation&#8212;demand would flatten. Supply, which was ramped during the boom, would overshoot. Prices would crash.</p><p>In the AI era, the demand function for memory has fundamentally changed. Memory is no longer predominantly serving a fixed number of &#187;human endpoints&#171;. Memory, especially HBM, is now a critical input for generating intelligence.</p><p>Think about what HBM (High Bandwidth Memory) actually does inside an AI accelerator. When you ask ChatGPT a question or run an inference on a large language model, the model&#8217;s parameters&#8212;billions or trillions of numerical weights&#8212;need to be loaded from memory into the GPU&#8217;s compute cores. The KV cache, which stores the context of your conversation, grows linearly with context length, with Grouped Query Attention (GQA) consuming roughly 0.06 - 0.12 MB per token in a 7B parameter model. A model with 70 billion parameters requires more than a single 80GB GPU worth of HBM just for the weights alone.</p><p>Here&#8217;s the simplified version: More memory = the ability to run larger models, with longer context, serving more users simultaneously.<strong> </strong>Memory is not a peripheral component in AI&#8212;it is the binding constraint. The so-called &#8220;memory wall&#8221; is the single biggest bottleneck limiting AI inference performance today. GPUs often sit idle, waiting for data to be fetched from memory. More bandwidth, more capacity means more intelligence output per second.</p><p>This is where the analogy to energy becomes powerful. Think about oil. When oil prices drop, what happens? Demand for oil increases because cheaper energy enables more economic activity. The demand curve for energy is downward-sloping- lower prices stimulate consumption. There&#8217;s always more work that could be done, more goods that could be transported, more heat that could be generated, if only energy were cheaper.</p><p>I believe AI inference demand behaves similarly. If memory costs drop and inference becomes cheaper, that doesn&#8217;t mean demand for inference drops. It means more applications become economically viable. More AI agents get deployed. More models get served. More context windows get extended. The demand for intelligence, like the demand for energy, is essentially elastic in response to price declines. Cheaper intelligence leads to more consumption of intelligence, not less.</p><p>This is the polar opposite of the gadget cycle. When DRAM prices dropped after the 2018 boom, it didn&#8217;t cause people to go buy a second smartphone. The number of endpoints was fixed. But when the cost of running an AI inference call drops by 50%, you can bet that the number of inference calls per day will more than compensate. Every enterprise that was waiting on the sidelines because of cost will deploy its AI project. Every startup that couldn&#8217;t afford the compute will spin up their service.</p><p>Here&#8217;s a human analogy I think captures this well.<strong> </strong>Imagine two people: one is a genius with poor memory, and the other is of average intelligence but has extraordinary memory and recall. In many real-world tasks&#8212;medicine, law, engineering, customer service&#8212;the person with superior memory will outperform the genius. Why? Because most practical work isn&#8217;t about raw reasoning power. It&#8217;s about retrieving the right piece of information at the right time. An AI model with more memory (longer context, more parameters accessible, faster retrieval) will outperform a theoretically smarter model that is memory-constrained. Memory is intelligence in many practical applications.</p><p>This is not a theoretical argument. The industry data supports it. HBM capacity per GPU has been scaling aggressively: NVIDIA&#8217;s A100 had 80GB of HBM2e. The H200 moved to 141GB of HBM3e. The upcoming Blackwell Ultra configurations push toward 288GB. And the Rubin Ultra platform is targeting 288GB - 576GB of HBM4E per GPU. The trajectory is exponential, and every generation of GPU is constrained by memory, not compute.</p><p><strong>Where we are today</strong></p><p>The current memory cycle is already historic in scale.</p><p>DRAM prices have surged dramatically. By Q4 2025, DRAM spot prices were nearly triple their level from a year earlier. DDR5 prices jumped 30-50% per quarter through H2 2025. Samsung raised memory prices by up to 60% since September 2025. DRAM inventories at major suppliers fell to just 3.3 weeks by the end of Q3 2025&#8212;matching the 2018 supercycle lows. SK Hynix and Micron had roughly 2 weeks of inventory each.</p><p>AI is expected to consume nearly 20% of global DRAM wafer capacity in 2026 when adjusted for HBM&#8217;s 4x wafer intensity.</p><p><strong>The valuation: The market doesn&#8217;t believe in the durability of this cycle</strong></p><p>Here&#8217;s where it gets really interesting from an investment perspective.</p><p>Despite the strongest fundamental setup the memory industry has ever seen&#8212;sold-out HBM capacity through 2026, record margins, structural demand from AI, and a three-player oligopoly with pricing discipline&#8212;the market is still pricing these stocks as if a classic downturn is imminent.</p><p>Micron trades at a forward P/E of about 10x, SK Hynix trades at approximately 5.2x forward P/E, and Samsung<strong> </strong>trades at a forward P/E of roughly 5x-7x&#8212;although this includes the total company, which includes much more than just memory.</p><p>The PEG ratio makes the mismatch even clearer. Micron&#8217;s PEG is approximately 0.16x, Samsung is at 0.17, and SK Hynix is at 0.10&#8212;meaning the market is pricing almost zero growth premium into the stocks.</p><p>But at these valuation levels, the question is not whether these companies will continue to grow; it&#8217;s more about how long the current demand signals will last. If these memory demand levels and margins stay here for a few more years, that would be a scenario that markets are not pricing in.</p><p>Why? Because the market has been burned by memory cyclicality before. Investors remember that in the 2017-2018 supercycle, Micron stock peaked at ~$64 with a forward P/E of about 4-5x at the top, and then the stock fell 56% even though earnings were still rising. The conditioned response is &#8220;memory is peaking, get out before the crash.&#8221;</p><p>But this framing assumes the old cycle repeats.<strong> </strong>It assumes that the demand driver (AI infrastructure buildout and inference scaling) behaves like the demand driver in previous cycles (consumer device upgrades). And I believe that assumption could be wrong.</p><p><strong>Why the downturn when it comes might be shallower</strong></p><p>I&#8217;m not arguing that memory prices will never decline. They will. At some point, new fab capacity from current investment plans will come online. At some point, HBM4 yields will improve, and supply will catch up. The 2017-2018 cycle teaches us that supply response is inevitable.</p><p>But I believe the depth and duration of the downturn will be structurally different this time (dangerous words I know):</p><p><strong>1. The end market is not bounded by human endpoints. </strong>In the PC cycle, once every household had a PC, demand plateaued. In the smartphone cycle, once penetration hit saturation, annual unit growth went to zero. But the number of AI inference calls per day is growing exponentially and is nowhere near saturation. Every enterprise, every consumer app, every autonomous vehicle, every AI agent is an incremental consumer of memory bandwidth.</p><p>This view is also shared by many industry experts. Here is a former high-ranking employee from ASML on this topic:</p><div class="pullquote"><p>&#187;The current conditions actually have made us move away from cyclicality simply because the ratio of the chips that go into laptops and cell phones and other personal-use devices is getting lower each day as the capacity gets transferred to AI-related infrastructure. We may not be able to predict the condition or state of these memory manufacturers based on cyclicality anymore.&#171;</p><p>Source: <a href="https://www.alpha-sense.com/uncoveralpha/">AlphaSense</a></p></div><p><strong>2. Memory content per AI unit is growing exponentially, not linearly. </strong>DRAM content per PC grew from maybe 4GB to 16GB over a decade&#8212;a 4x increase. HBM content per GPU is going from 80GB (A100) to 288GB - 576GB (Rubin Ultra) in just a few years&#8212;a 7x increase. And the number of GPUs being deployed is also growing at 30-40% annually. The compounding effect of more units &#215; more memory per unit is producing demand growth rates the industry has never seen.</p><p><strong>3. HBM is structurally supply-constrained. </strong>One gigabyte of HBM consumes approximately 4x the wafer capacity of standard DRAM. HBM also requires advanced packaging (CoWoS or its equivalents), which has its own supply bottleneck. You can&#8217;t just flip a switch and convert commodity DRAM lines to HBM production. The manufacturing complexity acts as a natural supply governor that didn&#8217;t exist in previous cycles.</p><p><strong>4. Long-term contracts are dampening volatility. </strong>In a major shift from past cycles, memory companies are increasingly locking in multi-year supply agreements with hyperscalers. SK Hynix has finalized its 2026 HBM supply plan with major clients and expects supply to remain tight through 2027. Micron has sold out its 2026 HBM capacity and has pricing agreements already in place. These contracts reduce the spot market&#8217;s influence and provide revenue visibility that the memory industry has never had before.</p><p>On top of the long-term contracts, the memory providers are much more careful with investing in new capacity this time, as the past cycle scars are a strong reminder. Here is a comment from a current Microsoft employee on what they expect in terms of memory supply coming online:</p><div class="pullquote"><p>&#187; I don&#8217;t think anyone on the buying side assumes memory suppliers will automatically rush to add unlimited supply just because demand is strong. The history of boom-bust cycles is very real, and suppliers remember that just as well as buyers do.<br><br>From my perspective, the expectation isn&#8217;t that all suppliers aggressively overbuild, but that they add capacity in much more controlled stages way than in the past cycles. What is different this time is the nature of demand. A lot of AI-driven demand is tied to long-lived infrastructure programs rather than short consumer cycles, which gives the suppliers more confidence but not enough to blindly overspend.&#171;</p><p>Source: <a href="https://www.alpha-sense.com/uncoveralpha/">AlphaSense</a></p></div><p>Perhaps the even more telling comment is this one made by a Fromer high ranking Micron employee on the internal cultural scars that the memory cycles have made:</p><div class="pullquote"><p>&#187;Micron has always positioned themselves as not the cheapest. Like I said, in the past, yes, when it was under Steve Appleton, Mark Durcan, Mark Adams, they&#8217;ve been trying to gain market share by reducing prices, but with the new CEO Sanjay, he is more focused on profitability rather than market share. Market share also is important, but if you were to choose between market share and profitability, he chooses profitability.&#171;</p><p>Source: <a href="https://www.alpha-sense.com/uncoveralpha/">AlphaSense</a></p></div><p><strong>5. The price elasticity of AI demand works in memory&#8217;s favor. </strong>If DRAM prices decline 20-30% (as they inevitably will at some point), the cost of running AI inference drops proportionally. This makes AI deployment cheaper, expanding the addressable market, which in turn supports memory demand. The demand floor is higher than in past cycles because cheaper memory creates new demand, rather than simply being absorbed by a fixed number of devices.</p><p>At some point, we will see a correction, but one that looks more like a 15-25% revenue decline and margins compressing to the 35-40% range, rather than the historic 30-40% revenue declines and sub-25% margins of previous busts. And crucially, I think the trough will be shorter, because AI inference demand will continue growing even during the cyclical correction, providing a demand floor that didn&#8217;t exist in the consumer device era.</p><p><strong>The bottom line</strong></p><p>The memory industry has spent 30 years teaching investors the same lesson: the cycle always turns, the crash always comes, and &#8220;this time is different&#8221; are the four most expensive words in investing. I respect that history deeply, and I&#8217;ve laid out the data to show you exactly how brutal those turns have been.</p><p>But I&#8217;m willing to bet against that lesson&#8212;partially&#8212;because the underlying demand driver has genuinely changed. That is why I also own stakes in SK Hynix and Samsung. Memory was a component in your gadget. Now it&#8217;s a substrate for intelligence. And the demand for intelligence&#8212;like the demand for energy, for computing, for connectivity&#8212;doesn&#8217;t follow the same saturation dynamics as consumer electronics.</p><p>The real risk for the memory cycle at the current stage is a technical breakthrough that would require orders-of-magnitude less memory and HBM, or a change that would bypass memory altogether. The chances of that happening today are low, but it is something to keep a close eye on all the time.</p><p>In the next section of this article for paid subscribers, I analyzed in detail how long I think this memory shortage and cycle will last, the timing of memory supply coming online for memory makers, including Chinese memory providers, and their possible effect on the market. Here is my take:</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/every-memory-cycle-ends-the-same">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Amazon's value in the Age of AI Agents]]></title><description><![CDATA[The changes caused by AI on Amazon: E-Commerce, AWS & Ads. How valuable each looks in a world where AI agents sit between humans and the services they use.]]></description><link>https://www.uncoveralpha.com/p/amazons-value-in-the-age-of-ai-agents</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/amazons-value-in-the-age-of-ai-agents</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Thu, 05 Mar 2026 13:51:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JkjN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hi everyone,</p><p>In this article, I&#8217;m breaking down my current thinking on Amazon. My goal here is to explain in detail the changes caused by AI on three pillars of the business: E-Commerce, AWS, and Advertising, and specifically how valuable each looks like in a world where AI agents increasingly sit between humans and the services they use. At the end, I&#8217;ll do a sum-of-parts valuation that I think gives a useful anchor for where the stock sits today.</p><p>Let&#8217;s start.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JkjN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JkjN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JkjN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png" width="1408" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1408,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2826536,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/189993139?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JkjN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 424w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 848w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 1272w, https://substackcdn.com/image/fetch/$s_!JkjN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3645c396-e7ab-4c14-845b-d53d4fb45e7c_1408x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>E-Commerce - the agentic threat and the logistics moat</strong></p><p>Amazon captures roughly 40% of all U.S. e-commerce spending. It has 240+ million Prime subscribers globally (analyst estimates; Amazon last officially disclosed &#8220;over 200 million&#8221; in 2021), of which approximately 180&#8211;185 million are in the United States, representing penetration in about 80% of U.S. households. The Prime flywheel is well-documented: members spend on average $1,400/year, compared with $600 for non-Prime customers, and the retention rate after the first year is 99%, according to CIRP data. Amazon delivered over 8 billion items same or next day to U.S. Prime members in 2025, a 30%+ increase year-over-year.</p><p>This is the business everyone knows. But here&#8217;s the question that matters for the next 3&#8211;5 years: what happens when AI agents start shopping for consumers?</p><p><strong>The agentic shopping risk</strong></p><p>I strongly believe that in the future, most e-commerce shopping will be done through AI agents acting as personal assistants to consumers, instead of direct consumers. I am not alone in those expectations. McKinsey projects agentic commerce could generate $1 trillion in U.S. retail revenue by 2030. Morgan Stanley expects nearly 50% of American shoppers will use AI agents by then, potentially adding $115 billion in e-commerce spending. Bain research shows that 30&#8211;45% of U.S. consumers already use GenAI for product research and comparison. During Cyber Week 2025, roughly 1 in 5 orders on Shopify involved an AI agent. AI-driven traffic to retailer sites has surged 7x since January 2025, according to Shopify data, with AI-driven orders up 11x.</p><p>What&#8217;s happening is this: instead of opening the Amazon app, a consumer tells ChatGPT, Claude, or Gemini what they need. The agent searches across retailers, compares prices, checks reviews, and either completes the purchase or presents a shortlist. OpenAI has already embedded checkout directly into ChatGPT. Perplexity launched its Comet browser agent. Google is rolling out agentic AI shopping tools.</p><p>This is a major shift in consumer behaviour, and Amazon knows it. In November 2025, Amazon sued Perplexity for its AI browser agent making purchases on Amazon&#8217;s marketplace. The company has blocked 47 AI bots from crawling its site. But at the same time, CEO Andy Jassy acknowledged on their recent earnings call that agentic commerce &#8220;has a chance to be really good for e-commerce.&#8221; Amazon recently even posted a job for a principal corporate development officer specifically for &#8220;agentic commerce&#8221; partnerships.</p><p>Forrester retail analyst Sucharita Kodali captured the tension perfectly: </p><div class="pullquote"><p>&#8220;With an agent on ChatGPT, retailers risk relinquishing transactions on their site to pay a toll on someone else&#8217;s highway.&#8221;</p></div><p><strong>Amazon&#8217;s shot at owning the application layer</strong></p><p>That said, Amazon isn&#8217;t conceding the front-end. They have several assets that give them a legitimate shot at being a surface where agentic shopping enters:</p><p>Rufus &#8212; Amazon&#8217;s AI shopping assistant, used by more than 300 million customers in 2025. Customers using Rufus complete purchases at a 60% higher rate. It can now auto-purchase items when prices hit thresholds.</p><p>The most interesting recent project is the &#187;Buy For Me&#171; project. This is Amazon&#8217;s experimental agent that can purchase from other retailers within the Amazon app. This is a smart flip from Amazon: instead of being the store that other agents shop, Amazon becomes the agent that shops everywhere else. Amazon does have some unique assets that make it valuable as the front-end touchpoint, and the key is around Prime Subscriptions.</p><p>Prime Video &#8212; 315 million ad-supported viewers globally, up from 200 million in early 2024. This is a massive surface for product discovery and agentic commerce integration, especially through interactive shoppable ads during live sports (Thursday Night Football averaged 15.3 million viewers, +16% YoY). Twitch &#8212; 105+ million monthly users, heavily Gen Z. An engaged, commerce-friendly audience. Alexa &#8212; still the most widely deployed voice assistant in smart home devices. If agentic commerce moves to a voice-first or ambient-first paradigm, Alexa has a head start.</p><p>The risk here is that those surfaces might not be enough and that Amazon might not be aggressive enough in the early days of where we are today. From today&#8217;s vantage point, the dominant surfaces if I had to choose would still be the smartphone assistant, or a standalone AI app (similar to ChatGPT, Gemini, Claude), and later on the AI glasses and personal assistant given by that provider. Prime Video and Twitch will still serve as important discovery platforms and could turn out to be much more valuable in terms of ads in a world where it will become increasingly hard to reach a human via digital channels, as internet usage will be dominated by AI agents instead of humans. Still, it doesn&#8217;t solve the fact that the application layer, where most of the e-commerce starts, moves to other providers. Even if Amazon were to launch an independent AI shopping assistant app, I don&#8217;t think in the long-term that would be &#187;moaty&#171; enough. My view is that the dominant provider will be the one that can offer a full AI personal assistant, with shopping as one of its features, not the only or main one. For that to be Amazon, they would need to make an aggressive pivot from current levels and a possibly strong shift into consumer hardware, which I don&#8217;t think is their plan.</p><p>With all that said, my base case is that Amazon will not be the application layer of agentic shopping and that its e-commerce business will move to the backend part of the shopping experience (still being important). Even in this scenario, Amazon still makes a decent margin given the logistics, payment, and fulfillment infrastructure that it offers at scale.</p><p><strong>Advertising</strong></p><p>Amazon&#8217;s advertising revenue hit $68.6B in 2025, growing 22% YoY in Q4. This is now 9.6% of Amazon&#8217;s total revenue, up from 5.9% in 2021. To put it in context, Amazon&#8217;s ad business alone is larger than the total revenue of companies like Netflix, Uber, or Salesforce.</p><p>But here&#8217;s the nuance that most analysts don&#8217;t discuss: Amazon&#8217;s ad business is really two very different businesses glued together.</p><p><strong>Search ads</strong></p><p>The vast majority of Amazon&#8217;s advertising revenue comes from Sponsored Products: essentially search ads within Amazon&#8217;s marketplace. When you search for &#8220;wireless headphones&#8221; on Amazon, the first several results are paid placements. Amazon doesn&#8217;t break this out precisely, but based on WARC data, the retail media component (primarily search ads) accounts for roughly $60.6B of the estimated total, with Prime Video and other upper-funnel formats making up the incremental portion.</p><p>Here is my concern: search ads on Amazon are fundamentally tied to humans browsing Amazon&#8217;s website and app. If an AI agent shops for you, it doesn&#8217;t look at sponsored listings. It doesn&#8217;t scroll past display ads. It skips right to the product that best matches your criteria and places the order. As Bain research noted, about 65% of retail media spending still occurs onsite, and that entire bucket is at risk if product discovery shifts to AI-driven search.</p><p>This is why I think the search ad portion of Amazon&#8217;s advertising business is on a disruption clock. Not tomorrow, not next quarter, but over a 3&#8211;5 year horizon, the economics of Sponsored Products face a structural headwind as agentic interfaces capture more of the purchase journey and as we talked in the previous section I give it a low probabiliticy chance that Amazon is able to capture the AI agent assistant application layer so the eyeballs switch from amazon&#8217;s site and apps towards the AI assistant owners.</p><p><strong>Prime Video ads</strong></p><p>The other side of Amazon&#8217;s ad business is Prime Video advertising, and this is the piece I think is defensible. Amazon introduced ads on Prime Video in January 2024. S&amp;P Global Market Intelligence Kagan estimated Prime Video&#8217;s ad revenue at $433M in 2024 and forecast it to reach $806M in 2025. This is still a small fraction of total ad revenue, but it&#8217;s growing fast and serves a different function: brand advertising through streaming video is not susceptible to agentic disintermediation the same way search ads are.</p><p>Prime Video reaches 315 million monthly ad-supported viewers globally. That&#8217;s larger than Netflix&#8217;s ad-supported tier at 190 million. Thursday Night Football alone averaged 15.3 million viewers with 16% growth YoY, and the Packers-Bears wild-card playoff game drew 31.6 million viewers, the most-streamed NFL game in history. Amazon has also integrated Netflix and Spotify inventory into its Amazon DSP, giving advertisers a broader programmatic buying platform.</p><p>My estimate is that by 2027&#8211;2028, Prime Video ads could reasonably be a $3&#8211;5B annual revenue stream, growing at 40%+ rates as ad loads increase and live sports inventory expands (NBA deal kicks in, international sports expansion). This business is much more structurally defensible because people watch content &#8212; AI agents don&#8217;t.</p><p>But even that revenue doesn&#8217;t materially change my thesis that the majority of Amazon&#8217;s ad business is at risk of serious disruption.</p><p>For the sum-of-parts analysis in the last part of this article, I&#8217;m splitting the ad business into two buckets. For the search/retail media portion (~$60&#8211;63B), I&#8217;m assigning it a terminal value as if profits only last 4 more years with zero terminal value after that. That&#8217;s deliberately punitive - I&#8217;m assuming this revenue stream is structurally impaired. For Prime Video ads, I&#8217;ll fold it into the e- commerce/subscription ecosystem, where it has long-term durability.</p><p><strong>AWS - the cloud business</strong></p><p>AWS is the most important reason why I own Amazon stock and why it has now become my biggest portfolio position.</p><p>The biggest fear around AWS has been that AI-related capital expenditures would permanently compress margins. And yes, there was a dip: AWS&#8217;s operating margin fell to 32.9% in Q2 2025 as the company ramped up spending aggressively. But by Q4, it had recovered to 35.0%, and the full-year margin was 35.4%.</p><p>Here is my core argument: we are severely compute-constrained for the foreseeable future. Amazon has invested $131.8B in capex for 2025 and has guided to approximately $200B for 2026, predominantly for AWS infrastructure. The company added more than 1 gigawatt of data center capacity in Q4 alone and 3.9 gigawatts in the trailing 12 months, which is double what AWS had in total in 2022. And Andy Jassy expects to double power capacity again by the end of 2027.</p><p>Despite this massive buildout, demand continues to outstrip supply. Jassy noted on the Q1 call that GPU and motherboard shortages were limiting the pace of AI workload onboarding. Bedrock (Amazon&#8217;s managed AI service) reached a multi-billion-dollar annualized run rate with customer spend growing 60% quarter-over-quarter to a base of over 100,000 customers. Trainium2 is fully subscribed with 1.4 million chips deployed.</p><p>In this environment, there is no incentive for hyperscalers to engage in a pricing war. When every chip you install is immediately monetized, you don&#8217;t cut prices &#8212; you add capacity. Until compute supply catches up with demand (which I don&#8217;t expect before 2029 at the earliest), AWS can maintain mid-30%+ operating margins without sacrificing growth. The margin should hold around pre-AI era levels (AWS operated in the 28&#8211;35% range historically, with 2024 averaging 37%) because the scarcity dynamic supports pricing power.</p><p><strong>Trainium and Custom Silicon are key things for long-term margins</strong></p><p>This is a point I don&#8217;t think gets enough attention. NVIDIA&#8217;s gross margin sits at roughly 73&#8211;75%. Every cloud provider that is 100% dependent on NVIDIA for AI compute is paying that tax on every GPU. That cost flows through to the cloud provider&#8217;s cost of revenue and structurally limits the margin they can earn on AI workloads.</p><p>Amazon, through its Annapurna Labs subsidiary, has developed Trainium and Inferentia custom ASICs, as well as Graviton CPUs for general compute. Combined, these custom chips have surpassed a $10B annualized revenue run rate, growing at triple-digit percentages YoY. According to Amazon, Graviton provides 40% better price-performance than x86 processors and is adopted by 90% of AWS&#8217;s top 1,000 customers.</p><p>Trainium2 powers Project Rainier, the world&#8217;s largest operational AI compute cluster with 500,000+ Trainium2 chips, which Anthropic uses to train its Claude models. Trainium3 is in preview with broader volumes expected in early 2026, and Trainium4 is targeted for 2027.</p><p>I am sharing here the chart that we made some months ago in our detailed Amazon <a href="https://www.uncoveralpha.com/p/amazon-trainium-scaling-ai-without">Trainium piece</a>, where we calculated the manufacturing costs of Amazon Trainium, Google TPUs, and Nvidia&#8217;s Blackwell B200:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H8dN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H8dN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 424w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 848w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 1272w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H8dN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp" width="761" height="162" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:162,&quot;width&quot;:761,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:15010,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/189993139?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!H8dN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 424w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 848w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 1272w, https://substackcdn.com/image/fetch/$s_!H8dN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fff26a195-f827-4296-8139-ccd03db52b7b_761x162.webp 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>You can see the most significant difference: if it costs Amazon $ 3,000-$ 3,500 to produce a Trainium3 chip, it costs them $35k-$40k to buy an Nvidia B200 chip. Even though B200 is much more performant from a cost-of-ownership perspective, Trainium3 gives B200 a run for its money.</p><p>The margin math is straightforward. When you design and manufacture your own silicon (using TSMC for fabrication and a design partner like Broadcom, Marvell, MediaTek), your cost per unit of compute is significantly lower than buying merchant silicon from NVIDIA at a 73&#8211;75% gross margin. This gives AWS a structural margin advantage for AI workloads vs. a competitor that sources 100% from NVIDIA. It doesn&#8217;t mean AWS abandons NVIDIA (it still offers NVIDIA instances), but having an alternative lets AWS capture more of the AI value chain and maintain margin in ways that someone who is entirely dependent on NVIDIA simply cannot.</p><p>This key difference will prove even more important in the coming years, especially once demand/supply for compute is more in balance and the hyperscalers&#8217; focus shifts from capturing revenue growth to profitability and customer optimization.</p><p><strong>Traditional cloud demand is actually accelerating because of AI agents</strong></p><p>There&#8217;s a narrative that AI is all that matters for AWS growth. That misses something important: AI agents themselves create enormous demand for traditional cloud services, as we already discussed in part in our <a href="https://www.uncoveralpha.com/p/the-forgotten-chip-cpus-the-new-bottleneck">The Forgotten Chip: CPU the New Bottleneck of the Agentic AI er</a>a article. Every AI agent needs storage (S3), compute (EC2, powered increasingly by Graviton), databases, networking, and monitoring. The more AI agents there are in production, the more traditional cloud infrastructure gets consumed.</p><p>The number of AI agents and their deployment is rapidly surging right now. Here is an alt provider that tracks the Model Context Protocol (MCP), an open-source standard for connecting AI applications to external systems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2fB-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2fB-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 424w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 848w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 1272w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2fB-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png" width="808" height="411" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:411,&quot;width&quot;:808,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:62841,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/189993139?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2fB-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 424w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 848w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 1272w, https://substackcdn.com/image/fetch/$s_!2fB-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F413d3f3e-93e1-45df-bab0-17352c13ef4b_808x411.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">source: <a href="https://bloomberry.com/">Bloomberry</a></figcaption></figure></div><p>The number of MCP servers being set up every month is growing exponentially, and the MoM pace is accelerating. The market is still in very early stages, as the current number of MCP servers is probably less than 1% of the API market. But the interesting thing was comparing where these MCP servers were being deployed with where API deployments are. Both Azure and GCP % of these MCP server deployments were lower compared to their API deployments, while AWS MCP deployments actually rose compared to API deployments:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Z5Pt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 424w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 848w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 1272w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png" width="821" height="522" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:522,&quot;width&quot;:821,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37571,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/189993139?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 424w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 848w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 1272w, https://substackcdn.com/image/fetch/$s_!Z5Pt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98469a4a-a09e-4294-ac99-e58080c5c9bc_821x522.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">source: <a href="https://bloomberry.com/">Bloomberry</a></figcaption></figure></div><p>While this data is not large enough yet, it could indicate that more smaller companies are on AWS than on the other two hyperscalers and that they are early adopters. The data in some form also shows the importance of AWS&#8217;s &#187;legacy&#171; cloud infrastructure, which is very much needed in the Agentic AI phase.</p><p>The demand for traditional AI infrastructure is skyrocketing, and you can see it from comments from CEOs of AMD and Intel, where they are basically sold out of CPUs, and you can see it when talking to other industry experts.</p><p>This is a former Amazon employee talking about the usage of the AWS S3 service (storage):</p><div class="pullquote"><p>&#187;Right now, S3, nobody is thinking about how S3 is exploding. It&#8217;s quite an explosion because what are AI systems doing? They&#8217;re generating embeddings, they&#8217;re storing the prompts and the responses. Where are they storing this? S3. They&#8217;re logging every interaction for auditing, tuning, safety. All of this is going into S3. Remember when we used to think of S3 as just their cheap storage? The storage is still cheap, but you&#8217;re using more of it&#171;.</p><p>source: <a href="https://www.alpha-sense.com/uncoveralpha/">AlphaSense</a></p></div><p>The key takeaway from this comment is the need to store this data for audit and safety purposes. Companies running these AI agents need clear, auditable trails of what an AI agent has done, so they can track and monitor as problems arise and fix them. Nobody wants to give an AI agent full permission to run freely across the company's stack and make changes and tasks that nobody has visibility into. This human visibility means services like Storage grow even more in usage.</p><p>A big tell was also the November OpenAI- AWS deal. The press release stated that OpenAI would access &#8220;hundreds of thousands of state-of-the-art NVIDIA GPUs, <strong>with the ability to expand to tens of millions of CPUs </strong>to rapidly scale agentic workloads.&#8221;</p><p>The GPU part is known, but the CPU part is the most interesting one. We need &#187;legacy&#171; cloud workloads and CPUs to enable the AI agent economy; that is just the way it is, and this is a big uplift for AWS, which has the largest fleet of optimized cloud services out there.</p><p>Amazon noted that more of the top 500 U.S. startups use AWS as their primary cloud provider than the next two providers combined. That startup and scale-up cohort is building AI-native applications that are heavily cloud-intensive.</p><p><strong>The on-prem fallacy and SMB lock-in</strong></p><p>Some investors argue that AI inference will eventually move to the edge or on-prem, killing the cloud growth story. Let me push back on this.</p><p>First, even if 90% of personal AI assistant use cases eventually run on edge devices (phones, laptops, local hardware), the remaining 10% that stays on cloud or on-prem infrastructure is still an enormous market. These are the &#8220;god-like AI&#8221; use cases: complex enterprise reasoning, multi-step agentic workflows, financial modeling, drug discovery, code generation at scale. These require the kind of compute density and model size that doesn&#8217;t fit on a phone. And these use-cases are the most profitable, as their outputs are the most valuable.</p><p>Second, on-prem AI infrastructure is radically more complex than anything businesses have managed before. Running an AI inference cluster on-prem means managing GPU and CPU servers, networking fabric, cooling systems, model deployment pipelines, and monitoring at a level of sophistication that most IT departments have never dealt with. For any small or medium-sized business, the cost and complexity of running your own AI infrastructure to have your &#8220;AI accountant&#8221; or &#8220;AI customer service agent&#8221; simply doesn&#8217;t make sense when you can rent it from AWS for a fraction of the upfront cost with zero operational hassle.</p><p>The cloud is the natural home for AI workloads for the vast majority of companies, and that reality isn&#8217;t changing anytime soon. If anything, as AI becomes more central to knowledge work, more companies will move to the cloud specifically to access AI capabilities they can&#8217;t build or run themselves.</p><p><strong>The revenue trajectory of AWS</strong></p><p>With all that in mind, I believe AWS, with its power capacity availability, which I already discussed in my previous articles, is well-positioned for multiple quarters of accelerating growth. I believe, despite AWS&#8217;s size, we will soon see the segment grow by +30% YoY. AWS also exited Q4 2025 with a backlog of $244B with a weighted average remaining life of 4.1 years. Capacity is being installed and monetized as fast as it comes online.</p><p>If AI agents truly absorb a meaningful portion of knowledge work over the next 5&#8211;10 years &#8212; and companies like Anthropic ($19B ARR rate up from $9B just two months ago) and OpenAI are building the models to do exactly that &#8212; then the total demand for cloud inference is going to be multiples of what it is today. Every AI-powered accountant, lawyer, engineer, customer service agent, and analyst running in the cloud creates recurring compute demand.</p><p><strong>The Anthropic and OpenAI stakes are hedges</strong></p><p>Besides the already mentioned segments, Amazon also has other important aspects such as Project Kupier, Subscription business, and stakes in Anthropic and OpenAI, which are now becoming increasingly important.</p><p>Amazon has invested approximately $8B in Anthropic (capped below 33% ownership) and recently announced a strategic partnership with OpenAI that includes an investment of up to $50B (starting with an initial $15B commitment, with the remainder tied to milestones and a potential OpenAI IPO).</p><p>Anthropic just closed a $30B funding round at a $380B post-money valuation in February 2026. If Amazon holds roughly 20% of Anthropic (estimates vary given the cap structure), that stake is worth $76B on paper. But in the last few weeks, Anthropic has accelerated its adoption and revenue growth so much that a $500B valuation for a company that will probably exit 2026 at a $50B ARR growing 5x YoY and disrupting the whole knowledge work economy is nothing extraordinary, which would add $100B of value or almost 5% of Amazon&#8217;s current market cap.</p><p>For OpenAI, the proposed $100B funding round would value the company at approximately $830B. Amazon&#8217;s $50B investment at those terms would represent roughly a 6% stake.</p><p>Combined, these stakes could be worth +$145B. And here&#8217;s the real value: in a world where Anthropic, OpenAI, and Gemini become the application layer, having significant stakes in two of those companies isn&#8217;t just financial investments. They are Amazon&#8217;s guarantee that the biggest AI consumers remain AWS customers. OpenAI has committed to spending $100B on AWS over the next eight years. Anthropic is using Project Rainier (500,000+ Trainium2 chips) for training. Both are locked in as massive cloud customers.</p><p><strong>Valuation</strong></p><p>Now let&#8217;s put the numbers together. I&#8217;m deliberately being conservative in places and factoring in serious disruption risk. Here are my numbers:</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/amazons-value-in-the-age-of-ai-agents">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Forgotten Chip: CPUs the New Bottleneck of the Agentic AI Era]]></title><description><![CDATA[For three years, GPUs have been the only chip that mattered in AI. CPUs? The boring, commodity chip that just sat next to the GPU and passed data along. Nobody cared. That&#8217;s changing fast.]]></description><link>https://www.uncoveralpha.com/p/the-forgotten-chip-cpus-the-new-bottleneck</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-forgotten-chip-cpus-the-new-bottleneck</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 23 Feb 2026 13:50:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zo6V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba3c69d-b335-4e71-94ba-046238903fac_1540x1809.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>For three years, GPUs have been the only chip that mattered in AI. Every investor pitch, every earnings call, every CapEx headline was about who could get more Nvidia GPUs. </p><p>CPUs? An afterthought. The boring, commodity chip that just sat next to the GPU and passed data along. Nobody cared. That&#8217;s changing fast. And if you&#8217;re not paying attention to the &#8220;CPU renaissance&#8221; happening right now, you&#8217;re missing what I believe is one of the more important infrastructure shifts in this AI cycle.</p><p>In this article, I will break down exactly why agentic AI is changing the CPU demand, how exactly CPUs are used in agentic AI, how big the CPU market can become because of AI agents, and which public companies stand to benefit. I&#8217;ll also discuss whether we&#8217;re heading into a genuine CPU bottleneck and how long it could last.</p><p><strong>Why Agentic AI Changes Everything for CPUs</strong></p><p>To understand why CPUs suddenly matter, you need to first understand how agentic AI workloads are fundamentally different from the &#187;classic&#171; chatbot-style AI we&#8217;ve been running for the past three years.</p><p><strong>The old workflow &#8212; chatbot:</strong></p><p>When you use ChatGPT or any standard AI chatbot, the process is straightforward. You type a question, the CPU tokenizes it (converts your text into numerical tokens the model can process), ships it over to the GPU, the GPU runs the tokens through the model and generates a response, then ships the output back to the CPU, which de-tokenizes it and delivers the answer. In this workflow, the CPU does very little. Maybe 5-10% of the total compute. The GPU is doing all the heavy lifting with its matrix multiplications, attention calculations, and token generation. This is why, for three years, the entire industry was laser-focused on GPUs.</p><p><strong>The new workflow &#8212; agentic AI:</strong></p><p>Agentic AI is fundamentally different. Instead of a simple question-answer loop, you&#8217;re dealing with autonomous systems that plan, execute, use tools, browse the web, query databases, make API calls, write and run code, and then reflect on whether they did a good job before deciding what to do next. A single user request can spin off dozens or even hundreds of sub-agents, each running their own loops of reasoning and action in parallel.</p><p>All of that orchestration, tool calling, API handling, memory management, and coordination between sub-agents happens on the CPU, not the GPU.<strong> </strong>The GPU still handles the inference (the &#8220;thinking&#8221; part), but between each inference call, the CPU is doing an enormous amount of work. It&#8217;s parsing responses, deciding which tool to call next, managing the execution plan, handling file I/O, running code, making network requests, and coordinating which sub-agents depend on which other sub-agents&#8217; results.</p><p>In an interview, a VP at Intel explained:</p><div class="pullquote"><p>&#8220;Agentic AI is nothing but a combination of independent agents... If there are in workflow, say, 10, 20, 30, 40, 100 agents, and they all need to talk to them, then they need different locations to operate. When I say location, I talk about CPUs.&#8221;</p><p>source: AlphaSense</p></div><p>A Georgia Tech and Intel research paper from November 2025 quantified this, and the findings are striking: tool processing on CPUs accounts for between 50% and 90% of total latency in agentic workloads.<strong> </strong>In many agentic workflows, the CPU is responsible for the majority of the wait time, not the GPU. The GPU sits idle, waiting for the CPU to finish its work before it gets the next batch of tokens to process.</p><p>This completely inverts the infrastructure economics we&#8217;ve been operating under. In the chatbot era, you needed a small number of high-end CPUs paired with massive GPU clusters. In the agentic era, you potentially need more CPUs than GPUs, and the CPU-to-GPU ratio in a rack or cluster needs to go up significantly.</p><div class="pullquote"><p>&#8220;For every GPU workload, there is a supporting CPU demand. The CPU is going to handle the data processing, the orchestration, the API layers, post processing.&#8221;</p><p>Source: AWS employee on AlphaSense</p></div><p><strong>Breaking Down the CPU Workload in Agent Systems</strong></p><p>Let me walk through what the CPU actually does in an agentic workflow, because I think understanding the details here is important for appreciating why this demand is structural and not a temporary blip.</p><p><strong>Step 1: Planning: </strong>The user gives a broad instruction (e.g., &#8220;Research the competitive landscape of the DRAM industry and write me a report&#8221;). The CPU tokenizes this and sends it to the GPU for an initial inference call. The GPU generates a plan of execution, not a final answer. That plan comes back to the CPU.</p><p><strong>Step 2: Orchestration: </strong>The CPU now breaks that plan into sub-tasks and assigns them to multiple agents. This is pure CPU work. It&#8217;s managing a directed acyclic graph of tasks, determining which ones can run in parallel, which depend on others, and in what order they should execute. If you have 10 research sub-topics, you might have 10 sub-agents that can all run simultaneously.</p><p><strong>Step 3: Tool execution: </strong>Each sub-agent starts working. This is where CPUs get extremely busy. Sub-agent 1 might make a web search API call, wait for results, parse the JSON response, extract relevant text, and package it for another inference call. Sub-agent 2 might query a database, run a SQL query, and process the results. Sub-agent 3 might open a file, read its contents, and prepare them for analysis. All of this &#8212; the API calls, network I/O, file handling, data parsing, JSON processing &#8212; is CPU work. The GPU is idle during these operations.</p><p><strong>Step 4: Inference loops: </strong>Each sub-agent may also run its own chain-of-thought reasoning, sending multiple inference requests to the GPU. Between each inference call, the CPU processes the output, decides if the agent is done, and either feeds the next prompt or moves to the next step.</p><p><strong>Step 5: Reflection: </strong>Once all sub-agents complete, the CPU gathers all their outputs and sends them to the GPU for a reflection inference loop &#8212; essentially asking the model, &#8220;did we answer the original question well enough?&#8221; If not, the whole cycle restarts. The key characteristics a CPU needs for this kind of workload are: high single-core clock speed (to minimize orchestration latency), high core count<strong> </strong>(to run many agents in parallel), fast memory access and large caches<strong> </strong>(to manage all the context and intermediate state), and strong I/O connectivity<strong> </strong>(PCIe lanes for network and storage, because agents are constantly hitting APIs and databases).</p><p>The AI server factories sitting above your general-purpose compute infrastructure don&#8217;t replace those traditional CPU servers. They create more demand for them. Because now, instead of one human slowly browsing the web and running a few apps, you have hundreds of AI agents aggressively consuming CPU resources at machine speed.</p><p><strong>The demand for CPUs is already showing up in earnings calls</strong></p><p>This new CPU demand has already been shown in recent earnings calls.</p><p>On AMD&#8217;s Q4 earnings, AMD&#8217;s data center segment posted record revenue of $5.4 billion in Q4 2025, up 39% year-over-year and 24% sequentially.</p><p>But the key wasn&#8217;t the GPUs but the CPUs. Lisa Su explicitly called out CPUs as a major growth driver, stating:</p><div class="pullquote"><p>&#8220;demand for EPYC CPUs is surging as agentic and emerging AI workloads require high-performance CPUs to power head nodes and run parallel tasks alongside GPUs.&#8221;</p></div><p>AMD&#8217;s 5th Gen EPYC Turin CPUs accounted for more than half of total server CPU revenue by the end of Q4, and the number of EPYC cloud instances grew more than 50% year-over-year to nearly 1,600 instances. The number of large enterprises deploying EPYC on-premises more than doubled in 2025.<strong> </strong>Su specifically highlighted that in agentic workflows, when AI agents spin off work in an enterprise, &#8220;they&#8217;re actually going to a lot of traditional CPU tasks.&#8221; She expects the server CPU market to grow by &#8220;strong double digits&#8221; in 2026.</p><p>Su also noted that &#8220;x86 processors have a particular edge in agentic workloads where AI agents spin off work to traditional CPU tasks, with the vast majority of such tasks running on x86 today.&#8221;</p><p>Looking ahead, Su guided for data center segment revenue to grow more than 60% annually over the next three to five years and for AMD&#8217;s AI business to scale to tens of billions in annual revenue by 2027. CPUs are a meaningful piece of that equation, not just GPUs.</p><p>And it&#8217;s not just the earnings call, you can also see it from multiple conversations with industry experts.</p><p>A former CTO of a HP competitor highlights that infrastructure is moving from static policy-based routing to &#8220;inference-based&#8221; routing. An AI-powered controller layer, running on CPUs, dynamically analyzes incoming workloads to determine whether they require expensive GPU cycles or can be offloaded to traditional x86 CPUs, optimizing resource allocation.</p><p>Agentic AI often involves deterministic tasks&#8212;such as following a specific rule set or executing a defined API call&#8212;that do not require the probabilistic power of a GPU. A Director at a Global Consultancy notes that these deterministic aspects of agentic workflows are most efficiently executed by CPUs, reinforcing the need for a balanced infrastructure where GPUs handle the &#8220;thinking&#8221; and CPUs handle the &#8220;doing&#8221;</p><p><strong>The CPU demand was a shock for Intel</strong></p><p>If AMD saw the CPU demand wave coming, Intel was genuinely surprised by it. Intel&#8217;s Q4 revenue came in at $13.7 billion, above guidance, with data center and AI revenue rising 15% sequentially &#8212; the fastest sequential growth this decade. But here&#8217;s the key: Intel admitted it couldn&#8217;t meet all the demand.</p><p>CEO Lip-Bu Tan said the company &#8220;delivered these results despite supply constraints, which meaningfully limited our ability to capture all of the strengths in our underwriting markets.&#8221; CFO David Zinsner was even more direct, admitting that Intel &#8220;misjudged&#8221; the pace of data center CPU demand and that the company is now &#8220;shifting as much as we can over to the data center&#8221; by reallocating wafer capacity from client (PC) CPUs to server CPUs.</p><p>Zinsner acknowledged that Intel is &#8220;absolutely constrained&#8221; and is deprioritizing the low-end client market to push capacity into data center products. Intel expects its supply to hit a low point in Q1 2026 before improving in Q2, but in the meantime, revenue &#8220;would have been higher if we had more supply. Management explicitly positioned CPUs as &#8220;central to AI orchestration and scaling inference.&#8221;</p><p><strong>The AWS-OpenAI Deal was the tell</strong></p><p>The most interesting data point on CPU demand came not from a chip company but from a cloud infrastructure deal back in November 2025. AWS and OpenAI announced a $38 billion, seven-year strategic partnership.<strong> </strong>The press release stated that OpenAI would access &#8220;hundreds of thousands of state-of-the-art NVIDIA GPUs, <strong>with the ability to expand to tens of millions of CPUs </strong>to rapidly scale agentic workloads.&#8221;</p><p>People wrongly focused on the Nvidia GPU part, but the CPU part is far more interesting. Tens of millions of CPUs. For agentic workloads. They didn&#8217;t have to include that detail. The fact that it&#8217;s in the official announcement tells you how seriously the frontier AI labs are thinking about CPU compute as a scaling requirement. All capacity under this agreement was targeted for deployment before the end of 2026, with options to expand into 2027.</p><p><strong>Nvidia &#8212; The Vera CPU</strong></p><p>Nvidia itself is making a big bet on the CPU side. Its upcoming Vera CPU, part of the Rubin platform announced at CES 2026, is specifically designed for agentic reasoning workloads. Vera delivers up to 2x the performance of the previous Grace CPU, with 88 cores per die and significant uplifts in memory and chip-to-chip bandwidth.</p><p>What&#8217;s particularly notable is that Nvidia announced Vera can be deployed as a standalone platform for agentic processing, separate from the GPU. CoreWeave is set to use standalone Vera CPUs, and Jensen hinted in a Bloomberg interview that &#8220;there are going to be many more&#8221; standalone CPU deployments. And it didn&#8217;t take long the Meta &amp; Nvidia deal was announced a few days ago:</p><div class="pullquote"><p>&#187;This partnership will enable the large-scale deployment of NVIDIA CPUs and millions of NVIDIA Blackwell and Rubin GPUs, as well as the integration of NVIDIA Spectrum-X&#8482; Ethernet switches for Meta&#8217;s Facebook Open Switching System platform&#8230;The collaboration represents the first large-scale NVIDIA Grace-only deployment.&#171;</p></div><p>This is Nvidia essentially confirming the thesis: in agentic AI, the CPU-to-GPU ratio needs to go up, and some workloads may be purely CPU-bound.</p><p><strong>Are We Heading Into a CPU Bottleneck?</strong></p><p>We&#8217;re already in one. The server CPU supply chain is under significant stress, and the constraints are coming from multiple directions simultaneously.</p><p>Intel is struggling with yield issues at some of its fabs, slowing the production ramp for newer Xeon parts. The company has admitted it cannot meet demand and is reallocating capacity from PC CPUs to server CPUs, meaning the PC segment will take a hit. Intel expects supply to improve starting Q2 2026, but the situation remains &#8220;acute&#8221; in Q1.</p><p>TSMC is prioritizing AI accelerators, which means less capacity for CPUs. AMD&#8217;s server CPUs are manufactured by TSMC, but TSMC is aggressively prioritizing its advanced node capacity for higher-margin AI accelerator chips (GPUs and custom ASICs). TSMC chairman C.C. Wei publicly stated that advanced-node capacity is &#8220;about three times short&#8221;<strong> </strong>of what major customers plan to consume. When TSMC&#8217;s 3nm process is running at 160,000 wafers per month and that&#8217;s still not enough, and when CoWoS advanced packaging capacity is sold out through 2026, CPU wafer allocation gets squeezed as a collateral effect.</p><p>Intel has also already warned Chinese customers of delivery lead times of up to six months<strong> </strong>for certain server CPUs. AMD&#8217;s lead times have stretched to 8-10 weeks<strong> </strong>for some products. Intel server chip prices in China have risen more than 10%.<strong> </strong>China represents over 20% of Intel&#8217;s total revenue, and major customers like Alibaba and Tencent are affected.</p><p>An additional problem to supply is the memory-driven pull-forward.<strong> </strong>The severe global memory shortage is creating a rush effect on CPU purchases. When memory prices started rising in China late 2025, customers accelerated CPU purchases to lock in system-level pricing before costs spiraled further. This pull-forward exacerbated the existing supply tightness.</p><p>A cloud computing materials manager reports:</p><div class="pullquote"><p>&#8220;Our supply chain was a constraining factor... GPU, CPU, and RAM were the top three drivers for us being constrained&#8221; as customers convert to &#8220;more powerful CPUs that can run higher AI workloads.&#8221;</p><p>Source: AlphaSense</p></div><p>A global IT distributor reports CPU shortages are &#8220;directly driving a 30% increase in average selling prices (ASPs) during the fourth quarter of 2025&#8221; with &#8220;increased backlogs&#8221; as order intake exceeds expectations.</p><p>So the CPU bottleneck is already here; the question now is how long it will last.</p><p>In the next section, I analyzed how many CPUs we will need in this agentic AI and gave a timeline of when supply could meet the demand, on top of which companies stand to benefit most from this trend:</p><p></p>
      <p>
          <a href="https://www.uncoveralpha.com/p/the-forgotten-chip-cpus-the-new-bottleneck">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Market Hates Big Cloud Spending. The Data Says The Market Is Wrong.]]></title><description><![CDATA[There has been an emergence of fear after big tech earnings related to CapEx AI spending. I decided to share my views on this topic and why I believe the fears around it are wrong at this point.]]></description><link>https://www.uncoveralpha.com/p/the-market-hates-big-cloud-spending</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-market-hates-big-cloud-spending</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Wed, 11 Feb 2026 14:17:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tWTP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>Because there has been an emergence of fear after big tech earnings related to CapEx AI spending, I decided to share my views on this topic and why I believe the fears around it are wrong at this point.</p><p>We had earnings from Meta, Microsoft, Google, and Amazon, and all of them increased AI CapEx substantially. Microsoft CapEx went from $63 billion in 2025 to a guide of over $100 billion for 2026; Google (Alphabet) went from $91 billion to a range of $175 billion to $185 billion; Meta went from $72 billion to a range of $115 billion to $135 billion; and Amazon went from $131 billion to a $200 billion guidance for 2026.</p><p>At this point, given the massive increases in these CapEx number as an investor, you are either on the side of the group of investors who don&#8217;t believe the companies will be able to deliver revenues and profits on these investments, or you are on the side who do believe, and based on that, the future revenue and profit outlooks are very high. Given the stocks mostly sold off on this CapEx news, the &#187;bears&#171; took over, but I don&#8217;t agree with them, and this article explains why. In the last part of the article, I also break down which hyperscaler looks best positioned to further accelerate their growth in 2026 and 2027.</p><p>The CEOs of all these businesses are not only telling you they believe the profits will be there, they are already showing you growth that came from 2023 AI investments (that keep in mind often were also criticized for being outlandish), and more importantly, they are showing you PROFITS on these AI investments.</p><p>Here is what the market has missed.</p><p><strong>The Q4 2025 AI Earnings profits were overshadowed by future CapEx numbers</strong></p><p>Before we go into the actual numbers, the fact is that we got really strong commentary from nearly all the big tech CEOs on AI revenue and returns from these investments. One might think that they are saying this because it is in their interest, but that is not really true. For most big tech companies like Google and Microsoft, it is actually in their interest for AI progress to grow at a more gradual rate than the exponential one it has today. The reason is that a lot of their business lines face disruption risk (Google Search, Microsoft software business, etc.). So these strong commentaries from these CEOs should be taken differently than comments coming from startups like OpenAI, Anthropic, and xAI, who, in some way, have to project the fast growth curve of AI as they need to raise new capital rounds, so they naturally have to project confidence both in terms of companies as well as the market in general.</p><p>We got some interesting comments from Amazon, which has a history of being very strict and efficient in its data center business. Andy Jassy confirming multiple times the confidence in the return on invested capital:</p><div class="pullquote"><p>&#187;We have deep experience understanding demand signals in the AWS business and then turning that capacity into strong return on invested capital. We&#8217;re confident this will be the case here as well.&#171;</p><p>&#187;We have, I think, a fair bit of experience over the years in AWS of forecasting demand signals and doing it in such a way that we don&#8217;t have a lot of wasted capacity and that we also have enough capacity to serve the demand that&#8217;s there.&#171;</p><p>&#187;And I think we&#8217;ve also proven with AWS over the years in how we build data centers and how we run them and how we invent in there, if you think about our chips and our hardware and our networking gear and how we&#8217;ve invented in power that this isn&#8217;t some sort of quixotic top line grab, we have confidence that we -- that these investments will yield strong returns on invested capital. We&#8217;ve done that with our core AWS business. I think that will very much be true here as well.&#171;</p></div><p>Jassy even confirmed that as soon as they bring new capacity online, it&#8217;s essentially sold out:</p><div class="pullquote"><p>&#187;And what we&#8217;re continuing to see is as fast as we install this capacity, this AI capacity, we are monetizing it. And so it&#8217;s just a very unusual opportunity. And so we see that following the same sorts of patterns we saw in the early days of our core AWS investment. I&#8217;m very confident we&#8217;re going to have strong return on invested capital here.&#171;</p></div><p>From the historic understanding of Amazon in terms of words, they often underhype, so a comment like this was very telling:</p><div class="pullquote"><p>&#187;I think this is an extraordinarily unusual opportunity to forever change the size of AWS and Amazon as a whole.&#171;</p></div><p>Remember, just before AI, there was a big trend of companies moving workloads from the cloud back to on-prem because they thought many cloud workloads were too expensive. Now companies are realizing that AI workloads will need to be on cloud, because companies don&#8217;t have the resources or even the possibility to manage complex data centers with liquid cooling requirements (most data centers don&#8217;t have the option of liquid cooling), GPU utilization rates and managing multiple AI accelerators (Nvidia GPUs, AMD GPUs, ASICs like TPUs, Tranium). Because of this, they have also started moving non-AI workloads to the cloud, as the data needs to be close to the AI workloads for them to run properly.</p><div class="pullquote"><p>&#187;We&#8217;re continuing to see strong growth in core non-AI workloads as enterprises return to focusing on moving infrastructure from on-premises to the cloud&#171;</p></div><p>Because of AI, the cloud providers have increased their cloud &#187; lock-in &#171; and are growing even non-AI workloads.</p><p>I already talked about this trend just a few weeks ago in my <a href="https://www.uncoveralpha.com/p/q4-2025-channel-checks-and-alternative">Q4 alternative report article</a>, where we showed this chart confirming that companies are going to move to the cloud at an accelerated pace again over the next 2 years:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5JAM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5JAM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 424w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 848w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 1272w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5JAM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png" width="861" height="563" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:563,&quot;width&quot;:861,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:49197,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/187624962?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5JAM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 424w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 848w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 1272w, https://substackcdn.com/image/fetch/$s_!5JAM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb11d3cc0-5b95-4b2b-bdfc-f03f471ead9d_861x563.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>But it wasn&#8217;t just Amazon talking about AI returns; the other Big Tech companies were, too. Meta gave a lot of color on how AI investments are already showing up in their business results:</p><div class="pullquote"><p>&#187;In Q4, we doubled the number of GPUs we used to train our GEM model for ads ranking. We also adopted a new sequence learning model architecture, which is capable of using longer sequences of user behavior and processing much richer information about each piece of content. The GEM and sequence learning improvements together grow a 3.5% lift in ad clicks on Facebook and a more than 1% gain in conversions on Instagram in Q4.&#171;</p><p>&#187;Instagram Reels had another strong quarter with watch time up more than 30% year-over-year in the U.S. Engagement is benefiting from several optimizations we made to improve the quality of recommendations including simplifying our ranking architecture to enable more efficient model scaling.&#8221;</p></div><p>On Facebook, video time continued to grow double digits year-over-year in the U.S., and we&#8217;re seeing strong results from our ranking and product efforts on both feed and video surfaces.&#171;</p><div class="pullquote"><p>&#187;The optimizations we made in Q4 drove a 7% lift in views of organic feed and video posts on Facebook, resulting in the largest quarterly revenue impact from Facebook product launches in the past two years.&#171;</p></div><p>Meta is seeing results from AI in both better ad targeting and engagement trends. The results actually &#187;revived&#171; Meta&#8217;s core and oldest platform, Facebook, which is seeing growth rates it hasn&#8217;t seen in years. But AI is opening up other avenues of growth at Meta:</p><div class="pullquote"><p>&#187;Another area we&#8217;re deploying AI to improve performance is ad creative. The combined revenue run rate of video generation tools hit $10 billion in Q4, with quarter-over-quarter growth outpacing the increase in overall ads revenue by nearly 3x.&#171;</p></div><p>The returns are not only affecting their revenue but also the productivity of their teams:</p><div class="pullquote"><p>&#187;Since the beginning of 2025, we&#8217;ve seen a 30% increase in output per engineer with the majority of that growth coming from the adoption of agenetic coding, which saw a big jump in Q4. We&#8217;re seeing even stronger gains with power users of AI coding tools, whose output has increased 80% year-over-year. We expect this growth to accelerate through the next half. &#171;</p></div><p>But despite these gains, Meta is telling us that it&#8217;s still very early as they are still using a limited amount of LLMs, as they have to either optimize them with SLMs because of compute limitations, or are still in the early stages of deploying these LLMs through their product stack:</p><div class="pullquote"><p>&#187;We&#8217;re also working on merging LLMs with the recommendation systems that power Facebook, Instagram, Threads and our ad system. Our world-class recommendation systems are already driving meaningful growth across our apps and ads business, but we think that the current systems are primitive compared to what will be possible soon.&#171;</p><p>&#187;We don&#8217;t typically use our larger model architectures like GEM for inference because their size and complexity would make it too cost prohibitive. So the way that we drive performance from those models is by using them to transfer knowledge to smaller lightweight models used at run time. But I would say that we think that there is room for our larger models to benefit from having more compute.&#171;</p></div><p>All of this resulted in Meta giving the highest revenue growth guide in almost 5 years. And despite the higher CapEx guide and costs stemming from both OpEx (new AI team costs + compute costs on public cloud providers) and higher amortization costs, Meta confirmed that they expect 2026 to deliver operating income above 2025.</p><p>In terms of Google, Google Cloud grew 48% YoY, one of the highest growth rates among businesses of this scale. Google Search actually grew 17% YoY, which is another growth rate for Search that hasn&#8217;t been seen for quite some time. On the call, management even commented that Search saw more usage in Q4 than ever before, as &#187;AI continues to drive an expansionary moment for Search&#171;.</p><p>But even ignoring all the commentary from these companies&#8217; management, let&#8217;s look at the hard numbers.</p><p>First, starting with revenue. All three hyperscalers are essentially selling all the compute they have available; if they had more, they would grow revenue even faster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tWTP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tWTP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 424w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 848w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 1272w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tWTP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png" width="1456" height="1129" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1129,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:753267,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/187624962?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tWTP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 424w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 848w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 1272w, https://substackcdn.com/image/fetch/$s_!tWTP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0c8fea0-2c07-46ea-9bee-99b880ebee18_5370x4163.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>AWS grew 24% YoY, Azure grew 39% YoY, and Google Cloud grew 48% YoY. Their backlogs are growing even faster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s877!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s877!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 424w, https://substackcdn.com/image/fetch/$s_!s877!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 848w, https://substackcdn.com/image/fetch/$s_!s877!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 1272w, https://substackcdn.com/image/fetch/$s_!s877!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s877!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png" width="1456" height="699" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:699,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:344389,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/187624962?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!s877!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 424w, https://substackcdn.com/image/fetch/$s_!s877!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 848w, https://substackcdn.com/image/fetch/$s_!s877!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 1272w, https://substackcdn.com/image/fetch/$s_!s877!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cee2578-ee2c-466b-87be-270f68fb43c1_4099x1968.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>From the current revenue growth on top of the backlogs, we can clearly see that the hyperscalers are again growing significantly due to AI workloads. The standout in the quarter, as we correctly pointed out in our alternative data report before earnings, was Google Cloud. It is clear that the AI spend in the past is translating to real revenue growth. So the notion that these companies are spending only on CapEx and we can&#8217;t see revenue from it is false. Now, the questions and the narrative in the market are that the profits won&#8217;t come from this revenue stream.</p><p>The main argument for this thesis is that AI workloads will have a lower long-term profile margin, and, secondly, that people are calculating returns based on projected CapEx guides and comparing them to current revenues.</p><p>If we first tackle the CapEx argument. It is important to understand that the CapEx a hyperscaler spends on a data center this year will be utilized over a 2-year period, as it takes around 2 years to build and operationalize a data center. So when people look at 2025 revenue growth for the hyperscalers, they should translate that into CapEx spent in 2023, not in 2024 or 2025. When we are in a period like we are today, when YoY CapEx growth (estimates for 2026) are +53% (AWS), +93% (Google Cloud), +59% (Microsoft Cloud), the math doesn&#8217;t make much sense when we compared to 2025 revenues, because we should be really comparing 2023 CapEx to 2025 revenue growth.</p><p>If we look at 2023 CapEx numbers, we can see that both Microsoft and Google increased CapEx in 2023 by 17.5% to $32.3B and $28.1B vs 2022 levels, while Amazon reduced CapEx by 17% YoY to $52.7B, although based on my calculations, only a -10% reduction of CapEx in AWS to $24.8B. Now, if we compare those CapEx numbers to the revenues generated by hyperscalers in 2025, the math makes a lot of sense, as yearly revenue additions are outpacing CapEx spending.</p><p>Even for the most conservative investors out there, we can take the example of Google Cloud and even take the 2024 CapEx and compare it to the Q4 2025 results:</p><p>Google&#8217;s 2024 CapEx was $52.5 billion, with roughly $42 billion going to technical infrastructure (cloud/AI). Google Cloud grew from $48 billion (2024) to $70.8 billion (2025)&#8212;a $22.8 billion increase.</p><p>At the new 30.1% operating margins:</p><p>$6.9 billion in first-year operating income from 2024 CapEx</p><p>Add depreciation (as operating margin already includes that): +$7.0 billion (6-year schedule at Google)</p><p>Total first-year cash: $13.9 billion</p><p>First-year return: 33%</p><p>But here&#8217;s where Google&#8217;s trajectory gets interesting. They went from 5% margins (2023) to 17.5% (Q4 2024) to 30.1% (Q4 2025). If margins stabilize at 30% (which I actually think will grow even further) and they run that 2024 infrastructure for five years:</p><p>Cumulative OI:  around $45 billion</p><p>Add depreciation: +$42 billion</p><p>Residual value (data center shell): +$8 billion</p><p>Total: $95 billion on $42 billion invested</p><p>ROI: 126% over 5 years, or a 18% IRR</p><p>And that still assumes growth moderates significantly from the current 48% YoY pace, while the margin stays at the 30% level and doesn&#8217;t improve.</p><p>With the increased pace of 2026 CapEx growth, the hyperscalers are essentially telling us what the revenue additions and, with it, growth rates will be for 2028.</p><p>Moving to the argument that the long-term margin on AI workloads will not be good compared to the pre-AI period. The numbers so far do not suggest this at all. Here is a look at AWS and Google Cloud&#8217;s operating margins over the last quarters, where AI workloads accounted for the majority of growth.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6Dpv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6Dpv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 424w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 848w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 1272w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6Dpv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png" width="1456" height="908" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:908,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:368786,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/187624962?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6Dpv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 424w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 848w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 1272w, https://substackcdn.com/image/fetch/$s_!6Dpv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cc20176-2d56-4dac-b89f-9d18b144c617_4768x2973.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The operating margin either held up in the same range as AWS (where the % of AI workloads compared to others is still smaller) or increased significantly at Google Cloud, where AI workloads are a bigger piece of the pie. Here, we have to acknowledge that Google Cloud is not only GCP, but nonetheless, commentary from management in all the latest quarters has been that GCP growth rates are even higher than total Google Cloud growth rates, so we should have seen a trend of lower operating margin, not higher, if AI workloads carried a low margin profile. An additional point to consider is also that a lot of the AI workload spend at GCP in this period were coming from Anthropic, which is one client that has much more negotiating power in terms of pricing then a bunch of smaller clients where the cloud providers are moving now as inference AI workloads start to take up more space as companies move their AI use-cases to production. Important in the context of margin is also the statement made by Google on its last earnings call:</p><div class="pullquote"><p>&#8220;We were able to lower Gemini serving unit cost by 78% over 2025 through model optimizations, efficiency and utilization improvements.&#8221;</p></div><p>What this tells us is that as these hyperscalers get even larger, they can optimize and squeeze more out of existing infrastructure. While some of those cost optimizations will be passed on to the cloud client, it is very clear that the ones with the most scale will also be able to use them to further expand their margin profile. Scale, but also custom ASICs play a key role here.</p><p><strong>Custom ASIC is the key</strong></p><p>Another strong argument that I already laid out in many of my previous articles is the custom silicon that cloud providers are designing. I continue to believe that this will be a critical element for any cloud provider to maintain healthy margins in the long term and avoid becoming overly dependent on a provider like Nvidia, which now has gross margins of almost 75%. In terms of custom ASICs, Google is best positioned with its TPUs, as we already laid out in the <a href="https://www.uncoveralpha.com/p/the-chip-made-for-the-ai-inference">TPU article</a>, followed by Amazon with Tranium. While Microsoft&#8217;s efforts here lag those of the other two, it is important to note that Microsoft also owns full IP rights to the custom ASICs that OpenAI will develop.</p><p>No surprise that on the Amazon earnings call, Tranium was mentioned 27 times, while Nvidia was not mentioned at all. We got even so far that the CEO called out specifically Amazon&#8217;s chip business and segmented revenue for us as a separate category:</p><div class="pullquote"><p>&#187;I think people know about our chips capability and our chips business, but I&#8217;m not sure folks realize how strong a chips company we&#8217;ve become over the last 10 years.</p><p>If you look at what we&#8217;ve done with Trainium, if you look at what we&#8217;ve done with Graviton, which is our CPU chip, which is about 40% better price performance than comparable x86 processors, 90% of the top 1,000 AWS customers are using Graviton very expansively. If you combine Trainium and Graviton, it&#8217;s well over a $10 billion annualized run rate business, and it&#8217;s still very early there.&#171;</p></div><p>Even though they lag from a product perspective, Microsoft also talked about its custom ASIC business very early in the call:</p><div class="pullquote"><p>&#187;Earlier this week, we brought online our Maia 200 accelerator. Maia 200 delivers 10-plus petaFLOPS at FP4 precision with over 30% improved TCO compared to the latest generation hardware in our fleet. We will be scaling this starting with inferencing and synthetic data gen for our Superintelligence Team as well as doing inferencing for Copilot and Foundry.&#171;</p></div><p>Custom silicon is what ensures hyperscalers can control their margin profile and market share, even in a more heated market where neoclouds and companies like Oracle have entered.</p><p><strong>Investors are questioning the AI compute demand, but in reality, we are just getting started</strong></p><p>A lot of investors are looking at the +$600 billion in combined hyperscaler CapEx projected for 2026  and questioning whether this is too much.  What most people are missing is that we are still in the very early innings of AI compute demand, and the data backs this up. Right now, coding and developer tools have emerged as the single breakout vertical for AI. For those who don&#8217;t follow the industry closely or took a break in January, the difference in usage in 1 month is staggering. Daily install counts on VS Code basically more than doubled in just one month, whereas usage is growing even faster.  Here is data from the usage of VS Code for Anthropic&#8217;s Claude Code and OpenAI Codex. The demand is going off the charts as developers are now not using these LLMs as tools anymore, but as junior to mid programers, where they now only review the code after the AI:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KgI7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KgI7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 424w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 848w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 1272w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KgI7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png" width="601" height="485" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:485,&quot;width&quot;:601,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66451,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/187624962?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KgI7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 424w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 848w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 1272w, https://substackcdn.com/image/fetch/$s_!KgI7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3b97732-fafd-4b27-82cb-e12704db06e3_601x485.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here&#8217;s the thing, though: coding is essentially one vertical. And it&#8217;s already consuming an enormous share of the available inference compute.  Now think about what happens when finance, legal, healthcare, customer operations, and other enterprise verticals start scaling their AI workloads to the same degree. According to Menlo Ventures, enterprise AI investment tripled from $11.5 billion to $37 billion in just one year, yet only 16% of enterprise deployments today qualify as true AI agents&#8212;most are still fixed-sequence workflows. We are nowhere close to saturation. McKinsey&#8217;s data shows 78% of organizations are now using AI in at least one business function, but the actual conversion to heavy inference workloads across non-coding departments is still nascent.  These numbers are tiny compared to where coding already is.</p><p>The market is pricing in CapEx as if coding-level adoption is the ceiling, when in reality, it is the floor.</p><p>These aren&#8217;t businesses lighting money on fire. These are businesses generating 30-35% operating margins on the largest infrastructure buildout in corporate history.</p><p>The custom chip businesses (Trainium, Graviton, TPUs) are growing triple-digits and creating structural moats that compound over time.</p><p>The market is treating this like the 2000 fiber glut. That was infrastructure built for demand that didn&#8217;t exist.</p><p>This is infrastructure being absorbed as fast as it&#8217;s deployed. Hyperscaler CapEx isn&#8217;t irrational exuberance. It&#8217;s the most rational investment decision these companies can make. Amazon, Microsoft, and Google aren&#8217;t hoping for AI to work out. They&#8217;re reporting the P&amp;L that shows it already has.</p><p><strong>Not all hyperscalers will be able to capture market share this year, though. The limiting factor is availability.</strong></p><p>Based on the past capacity commitements I calculated which cloud provider should grow the fastest in 2026 and beyond, and here are the numbers:</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/the-market-hates-big-cloud-spending">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[The Great SaaS Unbundling: Why AI Will Destroy Half the Industry and Supercharge the Other Half]]></title><description><![CDATA[Everyone&#8217;s talking about how AI will &#8220;transform&#8221; software, but I think most people are getting it wrong. The real story isn&#8217;t about transformation&#8212;it&#8217;s about bifurcation]]></description><link>https://www.uncoveralpha.com/p/the-great-saas-unbundling-why-ai</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/the-great-saas-unbundling-why-ai</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 02 Feb 2026 14:46:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VTCB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>I&#8217;ve been thinking a lot about the AI disruption narrative in SaaS. Everyone&#8217;s talking about how AI will &#8220;transform&#8221; software, but I think most people are getting it wrong. The real story isn&#8217;t about transformation&#8212;it&#8217;s about bifurcation. Some SaaS companies are about to get absolutely demolished, while others will emerge stronger than ever. It&#8217;s not about looking at valuation levels for some of these SaaS companies and buying what is cheap on a valuation metric.</p><p>The determining factor for survival isn&#8217;t the brand, or even the data that the SaaS companies have&#8212;it&#8217;s whether their core system is deterministic or probabilistic.</p><p>Let me explain what I mean and why this matters for investors.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VTCB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VTCB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VTCB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1731291,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/186616042?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VTCB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!VTCB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F263462e3-7978-4a86-880e-fcb111dcf621_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>The Core Thesis: Deterministic vs. Probabilistic Systems</strong></p><p>Deterministic systems<strong> </strong>are those where precision is critical, state management is complex, and errors cascade into serious consequences. Think accounting software, ERP systems, compliance platforms, healthcare system, payment processors, and sophisticated workflow engines. These systems need to be right 100% of the time&#8212;not 95%, not 99%, but 100%. When you&#8217;re reconciling a billion-dollar balance sheet or processing payroll for 50,000 employees, &#8220;close enough&#8221; isn&#8217;t acceptable.</p><div class="pullquote"><p>&#8220;Traditional enterprise functions, such as HR, are inherently deterministic; decisions like employee termination are binary and require rigid logic where specific inputs trigger precise, unvarying sequences. In contrast, LLMs are inherently probabilistic, determining the confidence level of the next token rather than following a hard-coded decision tree.&#171; </p><p>Source: Employee at Rippling (AlphaSense)</p></div><p>Probabilistic systems<strong> </strong>are those where the core value proposition is pattern recognition, content generation, basic automation, or simple decision-making. Think chatbots, content recommendation engines, basic customer support automation, simple workflow tools, and generic productivity software. These systems can tolerate errors and are often based on &#8220;good enough&#8221; outputs.</p><p>More likely, AI is going to eat the probabilistic category, while some deterministic systems will become more valuable by integrating AI as a complementary layer and start expanding into other layers.</p><p><strong>Why Deterministic Systems Are Actually Strengthened by AI</strong></p><p>This might seem counterintuitive. If AI is so powerful, why wouldn&#8217;t it disrupt the complex systems?</p><div class="pullquote"><p>&#8220;AI succeeds when autonomy is constrained, execution is owned, and determinism is treated as an asset rather than a limitation.&#8221;<strong> </strong></p><p>Jens Eriksvik</p></div><p>When you look at how enterprises are actually deploying AI agents in 2025/2026, they&#8217;re not replacing their systems of record&#8212;they&#8217;re building orchestration layers on top of them. As a Former Microsoft Manager put it: </p><div class="pullquote"><p>&#187;A &#8216;reality check&#8217; is occurring among CIOs as they realize LLMs lack the deterministic consistency required for critical industries like financial services. For use cases such as underwriting, a system that provides a correct answer &#8216;six out of ten times&#8217; is insufficient; these processes demand 100% consistency, which current probabilistic models struggle to guarantee without extensive re-engineering.&#171;</p></div><p>LLMs interpret human intent, deterministic systems execute the actual work. The deterministic systems are not being disrupted; the operator is (you).<strong> </strong>This is the architecture that&#8217;s winning in production environments.</p><p>Why does this matter? Because the companies that own these deterministic platforms become more valuable in an AI world, not less. They become the essential execution layer that AI needs to actually accomplish tasks.</p><p>The use of these deterministic platforms should rise substantially as more people gain valuable information from them with the help of AI. Usage goes up, but only with platforms that integrate these AI tools well into their deterministic platform cores.</p><p>But even with deterministic systems, there are challenges. The seat-based pricing must be converted to usage pricing. SaaS companies right now have to aggressively cut costs, specifically labour costs, SBC, etc., and get in front of the curve. As a deterministic platform, you can charge a premium for your deterministic core offering. On top of that, you will be able to offer probabilistic tools that complement the deterministic core. Here, the pricing logic is simple: you price it at inference cost + 30% margin. Over time, as you build out your sticky offering as a platform, you can gradually try to expand that margin once again, but right now, that time is not there yet.</p><p>As a deterministic platform provider, the goals are clear: provide a clear deterministic core, execute great probabilistic offerings that enhance the core, and cut cost AGGRESSIVELY in terms of labour as you increase OpEx spend on cloud infrastructure to reach mass scale and then negotiate better inference costs because of that scale. </p><p>The companies that do this will come out as big winners, as they will be able to consolidate and offer probabilistic features on top of their offering, at inference +30% margins, and, with it, expand their TAM.</p><p><strong>The Probabilistic SaaS Bloodbath</strong></p><p>Now let&#8217;s talk about the other side of this equation&#8212;the SaaS companies that are in trouble.</p><p>If your core value proposition can be replicated by an LLM with 90% of the quality at 1% of the cost, and you provide a probabilistic product, you don&#8217;t have a sound business model anymore. The problem becomes if your core value proposition is pattern matching, content generation, recommendations, or simple automation. Foundation models have gotten so good at these exact tasks that they can replicate your entire product in a few lines of code. The problem is not only the costs (which is a big one), but the problem also extends to the user interface, data, integration, and brand moats.</p><p>Having a &#187;great UX&#8221; as a SaaS provider is irrelevant when natural language becomes the interface. Users would rather type &#8220;generate 10 marketing emails for our Q1 launch&#8221; into ChatGPT than navigate through HubSpot&#8217;s 47-screen workflow builder. While some call out proprietary data as the strong moat for these kinds of businesses, I would argue otherwise. Modern LLMs can learn from a small example set and perform as well as a model that has thousands of examples. The accelerating nature of LLMs and the emergence of synthetic data also hurt the incumbent data holders. Research from Meta in 2024 showed that models trained on synthetic data generated by GPT-4 perform within 2% of models trained on real data for most classification tasks. And this was in 2024, till today, this only gotten better. Even if the proprietary data gives you your own model with 2-3% better accuracy, because of the probabilistic nature of your business, customers are not willing to pay 100x premiums for 2-3% better outcomes. They might do that if you had a deterministic system, however.</p><p>Moving to the &#187;integration moat&#171;. The key emphasis here is that these SaaS solutions are integrated with thousands of other apps and that this ecosystem is hard to replicate. Most SaaS products have well-documented APIs. AI excels as an integration layer without the need for pre-built connectors. With AI agents, these integrations and connections will become even more seamless as adoption accelerates and the SaaS companies want to stay &#8220;useful&#8221; in the age of agentic AI, making their APIs even more open and clear.</p><p>Now to the moat called the brand. There is some merit that enterprises, to some extent, are loyal to brands as they build trust in those brands. But with probabilistic systems, that trust is less strong and loyal than it is with deterministic systems, where you know you get those results 100% accurate. Enterprises are loyal to a degree until the cost gap becomes too big. If the discount is 20-30%, most won&#8217;t switch, but if that discount grows to 50-+70% switching starts. The trust factor is also something very fluid. AI startups with low-cost probabilistic system solutions gain trust via media coverage, raising billions in new VC funds, and hiring high-profile people from the incumbents.</p><p>The cutting of probabilistic SaaS is already underway, and this is more than just cutting seats.</p><p>Publicis Sapient reports actively reducing traditional SaaS licenses by approximately 50%&#8212;including major platforms like Adobe&#8212;by substituting them with generative AI tools and chatbots. An executive at the firm in an expert interview explains that AI agents are &#8220;10x faster, 100x smarter&#8221; than junior staff, creating a redundancy that directly cannibalizes the seat-based revenue underpinning commercial SaaS models.</p><p>For probabilistic SaaS, the only viable model is to cut costs to a minimum and price your product with a 30%+ margin on inference, but even that might not be sticky enough, especially if you don&#8217;t have any deterministic offering and if your clients are primarily SMBs. These companies will not be disrupted directly by AI, but by deterministic systems competing with them, offering AI-generated probabilistic offerings and bundling them into a single offering with a core deterministic system holding it together. If your ERP provider starts offering you a customer service system that works flawlessly with your ERP and uses your inference credits for both use-cases, you will likely switch over rather than have a separate customer service offering even if it is at the same cost.</p><p><strong>Valuation Compression is Already Here but it&#8217;s across the board</strong></p><p>Right now, the market is hitting SaaS across the board as it sees the risk of AI disruption. As of December 2025, the median EV/Revenue multiple for public SaaS companies stands at 5.1x, down from the pandemic peak of 18-19x and much lower than the historic average.</p><p>The thing the market hasn&#8217;t fully priced in yet is the deterministic and probabilistic platform differences that I laid out here, so the opportunity to own a deterministic SaaS platform at reasonable prices is definitely here.</p><p>Based on the criteria laid out in this article, I made a list of public SaaS companies and ranked them in deterministic/probabilistic order, and some of the ones I would highlight as being the least at risk of AI disruption:</p>
      <p>
          <a href="https://www.uncoveralpha.com/p/the-great-saas-unbundling-why-ai">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Anthropic's Claude Code is having its "ChatGPT" moment]]></title><description><![CDATA[I am posting an article on Anthropic Claude Code, which has been growing very significantly lately and, I believe, has developed an important product fit in its category.]]></description><link>https://www.uncoveralpha.com/p/anthropics-claude-code-is-having</link><guid isPermaLink="false">https://www.uncoveralpha.com/p/anthropics-claude-code-is-having</guid><dc:creator><![CDATA[UncoverAlpha]]></dc:creator><pubDate>Mon, 26 Jan 2026 16:34:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Yr4h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Hey everyone,</p><p>I am posting an article on Anthropic Claude Code, which has been growing very significantly lately and, I believe, has developed an important product fit in its category.</p><p>Claude Code is going from just another AI coding assistant to a fundamental new architecture that developers need to stay competitive.</p><p>In the final months of 2025 and opening weeks of 2026, Claude Code reached a $1 billion annualized run rate just six months after launch&#8212;a velocity that even ChatGPT didn&#8217;t match. Based on my analysis and data, which I will share in this article, I believe that Claude Code is today closer to $2B ARR than $1B, as it has accelerated significantly in January.</p><p>At the same time, Anthropic&#8217;s overall annualized revenue jumped from approximately $1 billion at the start of 2025 to $5 billion by August&#8212;a 5x increase in eight months&#8212;with projections reaching $9 billion by year-end 2025.</p><p>But raw revenue growth, while impressive, misses the deeper structural shift. Claude Code has achieved what competitors couldn&#8217;t: it&#8217;s become the tool developers reach for when facing their hardest problems. At a Seattle meetup in mid-January 2026, over 150 engineers packed the house to trade use cases. One Google principal engineer publicly acknowledged that Claude reproduced a year of architectural work in one hour. Microsoft&#8212;which sells GitHub Copilot&#8212;has widely adopted Claude Code internally across major engineering teams, with even non-developers reportedly encouraged to use it.</p><p>Let&#8217;s dive in.</p><p><strong>Anthropic is building a defensible moat in enterprise AI.</strong></p><p>Anthropic reached 300,000+ business customers by August 2025, up from fewer than 1,000 businesses two years prior. According to Thunderbit, Claude&#8217;s enterprise AI assistant market share rose from 18% in 2024 to 29% in 2025&#8212;a 61% year-over-year increase&#8212;closing the gap with ChatGPT.</p><p>Anthropic just recently signed a term sheet for a $10 billion funding round at a $350 billion valuation&#8212;nearly double the $183 billion valuation from September 2025. That September round itself represented a massive step up from the $61.5 billion valuation in March 2025. The valuation has grown nearly six-fold in ten months&#8212;a trajectory that few technology companies have ever achieved, and the success is mostly tied to their developer clients.</p><p><strong>Why are developers choosing Claude?</strong></p><p>The market is littered with AI coding tools&#8212;GitHub Copilot, Cursor, Amazon CodeWhisperer, Tabnine, Codex, and dozens more. Yet Claude Code captured the developer community in ways its competitors haven&#8217;t.</p><p>The Architecture!</p><p>Claude Code&#8217;s distinguishing characteristic isn&#8217;t its AI model&#8212;though Claude 4&#8217;s coding capabilities are state-of-the-art. It&#8217;s the architectural decision to operate directly in the terminal with full file system and command-line access. This matters because it changes the fundamental relationship between developer and AI.</p><p>Traditional coding assistants like GitHub Copilot work as IDE extensions, offering autocomplete suggestions and chat interfaces. They&#8217;re stateless&#8212;every interaction starts fresh, with limited context beyond the current file. Claude Code operates differently. It reads and writes files directly, executes bash commands, maintains state across sessions, and coordinates multi-step processes spanning days. </p><p>As Noah Brier, an early LLM adopter who discussed the tool on Bloomberg&#8217;s Odd Lots podcast explained:</p><div class="pullquote"><p>&#8220; <em>it&#8217;s more like hiring a junior developer than using autocomplete.&#8221;</em></p></div><p>The terminal-native design solves two problems that plague competing tools. First, it enables persistent state management. Claude Code stores information in files, building up context and knowledge over time. When working on a multi-day refactor, it remembers architectural decisions, maintains to-do lists, and tracks completed work&#8212;capabilities that chat-based assistants simply can&#8217;t match. Second, it leverages composable Unix commands. Instead of reinventing wheels, Claude Code chains together grep, sed, git, and other standard tools that developers already trust.</p><p>This architectural choice has profound implications for adoption. Developers don&#8217;t need to learn new interfaces or workflows. They work in the environment they already use&#8212;the terminal&#8212;with a tool that speaks their language. And because Claude Code operates as a true agent rather than an assistant, it can handle entire projects autonomously while developers focus on architecture and business logic.</p><p><strong>The model advantage: Claude 4 and Sonnet 4.5</strong></p><p>Ofcourse the underlying AI models matter enormously as well. Anthropic released Claude 4 (Opus and Sonnet) in May 2025, introducing what the company called &#8220;the world&#8217;s best coding model.&#8221; The benchmarks backed up the claim:</p><p>Claude Opus 4: 72.5% on SWE-bench (measuring ability to solve real GitHub issues), 43.2% on Terminal-bench (command-line tasks). Claude Sonnet 4: 72.7% on SWE-bench, balancing performance with cost-efficiency. Extended thinking with tool use: Models can now alternate between reasoning and tool use (like web search) during extended thinking sessions. Memory capabilities: When given file access, Claude 4 creates and maintains &#8216;memory files&#8217; to store key information, dramatically improving performance on long-running agent tasks</p><p>Then in September 2025, Anthropic released Claude Sonnet 4.5, which became their most powerful model to date. The improvements were dramatic:</p><p>&#8226; 77.2% on SWE-bench Verified (82.0% with parallel compute)</p><p>&#8226; Code editing error rate: Dropped from 9% to 0% on Anthropic&#8217;s internal benchmarks</p><p>&#8226; Long-horizon task performance: Maintains focus for more than 30 hours on complex, multi-step tasks (vs. ~7 hours for Opus 4)</p><p>&#8226; 61.4% on OSWorld (desktop/browser interaction), up from 42.2% just four months prior</p><p>In November 2025, Anthropic released Claude Opus 4.5, which achieved 80.9% on SWE-bench Verified while using up to 65% fewer tokens than previous models. This efficiency translates directly to cost savings for developers running complex workflows.</p><p>Critically, these weren&#8217;t just benchmark improvements&#8212;they showed up in production. GitHub integrated Claude Sonnet 4 to power GitHub Copilot&#8217;s new coding agent. Cursor called Opus 4 &#8220;state-of-the-art for coding and a leap forward in complex codebase understanding.&#8221; Replit reported &#8220;dramatic advancements for complex changes across multiple files.&#8221; Block noted it was &#8220;the first model to boost code quality during editing and debugging.&#8221;</p><p>Bloomberry conducted research on over 45k companies, and the results are very insightful into which industries Anthropic dominates vs OpenAI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0jcA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0jcA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 424w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 848w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 1272w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0jcA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png" width="1011" height="674" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:674,&quot;width&quot;:1011,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:81315,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/185856188?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0jcA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 424w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 848w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 1272w, https://substackcdn.com/image/fetch/$s_!0jcA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4655461-dd0e-429c-b995-5b8fdddb490b_1011x674.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">source: <a href="https://bloomberry.com/blog/we-analyzed-44k-companies-to-see-who-uses-claude-or-perplexity/">Bloomberry</a></figcaption></figure></div><p>The software development vertical is especially interesting as companies are 2.3 times more likely to be Claude only than OpenAI only.</p><p>On the other hand, the industries where OpenAI dominates Anthropic are Marketing services, real estate, advertising, and business consulting.</p><p>In another developer survey conducted by UC San Diego and Cornell University in January, from 99 professional developers, Claude Code (58 respondents) appeared alongside GitHub Copilot (53) and Cursor (51) as one of the three most widely adopted platforms, with 29 respondents using multiple agents simultaneously.</p><p>In 2026, Claude is accelerating even faster with the launch of Cowork. The Cowork launch proved particularly significant. Users had been using Claude Code for non-coding tasks (vacation research, spreadsheet work via Slack, oven control). By launching Cowork, Anthropic showed that Claude Code&#8217;s total addressable market extends far beyond the 28 million professional developers globally.</p><p>Now, in addition to Cowork, we have a new trend of a personal assistant called Clawd bot. While Clawd bot is not owned by Anthropic but rather an open-source project, it has become the &#187;ChatGPT&#171; moment for personal intelligence, and for most users, Clawd works best when used with Claude, causing a surge in usage of Claude Code.</p><p>This is the most eye-opening chart from this article. This shows the daily install counts of AI Coding Assistants in Visual Studio Core. For those non-technical, VS Code is the industry standard for code editors and the primary host of AI coding agents:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Yr4h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Yr4h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 424w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 848w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 1272w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Yr4h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png" width="1306" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1306,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:169401,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.uncoveralpha.com/i/185856188?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Yr4h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 424w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 848w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 1272w, https://substackcdn.com/image/fetch/$s_!Yr4h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21f157c5-6b80-47c6-b480-3774aa5b90ac_1306x628.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Since the start of 2026, Claude Code has been surging! It went from 17.7M of daily installs (30-day moving average), similar to where OpenAI&#8217;s Codex was, to 29M and continues to rise exponentially. This really shows that Claude Code is having its own &#187;ChatGPT&#171; moment TODAY.</p><div><hr></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe?"><span>Subscribe now</span></a></p><div><hr></div><p><strong>Why does coding matter so much as an AI vertical?</strong></p><p>In short, because the results are measurable, and companies can put serious investment behind these productivity gains. The academic research and enterprise case studies paint a consistent picture: AI coding tools deliver 26-55% productivity improvements, with experienced developers seeing the largest gains.</p><p>GitHub Copilot baseline: A 2022 controlled experiment found that developers using GitHub Copilot completed tasks 55.8% faster (95% confidence interval: 21-89%) than control groups. Subsequent enterprise deployments confirmed these gains:</p><p>&#8226; GitHub&#8217;s own research: Developers code up to 51% faster for certain tasks</p><p>&#8226; Accenture randomized trial: 8.69% increase in pull requests per developer, 11% increase in merge rates, 84% increase in successful builds</p><p>&#8226; Developer satisfaction: Up to 75% higher job satisfaction, 88% code retention rate (developers keep nearly all AI-generated suggestions)</p><p>&#8226; Success rates: 78% of developers complete tasks using Copilot vs. 70% without it, with 53.2% more likely to pass all unit tests</p><p>Claude Code&#8217;s reported gains exceed Copilot&#8217;s: Internal data from Anthropic and partner companies suggests even stronger performance for complex, long-horizon tasks:</p><p>&#8226; Developers report running 5-15 Claude Code instances concurrently&#8212;multiple in terminals, plus additional browser sessions</p><p>&#8226; Rakuten validated capabilities with a demanding open-source refactor running independently for 7 hours with sustained performance</p><p>&#8226; Boris Cherny (Head of Claude Code at Anthropic): <em>&#8220;Claude Code generated roughly 80% of its own code&#8221;</em> (with human direction, review, and architectural decisions)</p><p>A software engineer in the US costs $200,000-+$400,000 annually. If AI coding tools deliver even conservative 20-30% productivity gains, that translates to $40,000-$90,000 in annual value per developer. For a company with 1,000 engineers, we&#8217;re talking $40-90 million in annual productivity gains, justifying substantial spending on AI coding infrastructure.</p><p><strong>Anthropic&#8217;s Business Momentum</strong></p><p>The AI industry&#8217;s narrative has fixated on OpenAI&#8217;s consumer dominance&#8212;ChatGPT&#8217;s 800 million weekly active users, 2.5-3 billion daily prompts, and $500 billion valuation. But an important story for investors is playing out in enterprise adoption, where Anthropic is systematically outmaneuvering its larger rival.</p><p>This growth trajectory is unprecedented. For context, OpenAI&#8217;s 2025 revenue is estimated at $10-12 billion&#8212;larger in absolute terms but growing more slowly from a higher base. More critically, Anthropic is projected to break even by 2028, while OpenAI isn&#8217;t expected to turn a profit until 2030, according to November 2025 WSJ reporting. OpenAI faces approximately $74 billion in projected losses in 2028 due to massive compute costs, while Anthropic&#8217;s enterprise focus and efficiency gains position it for profitability much sooner.</p><p>While ChatGPT dominates consumer attention, Anthropic systematically captured the enterprise market where switching costs are high, and revenue is sticky. </p><p>According to Views4You, Claude has high penetration rates in different industries:</p><p>&#8226; Healthcare: 61% usage growth in early 2025, with Claude assisting in medical documentation and patient communication</p><p>&#8226; Legal: 18% of AI-enhanced litigation tools rely on Claude</p><p>&#8226; Finance: 24% of major banks use Claude, with 34% of enterprise AI research teams integrating it</p><p>&#8226; Retail/E-commerce: 38% of chatbots employ Claude</p><p>&#8226; Real Estate: 25% of listing analysis tools powered by Claude</p><p>This enterprise penetration is what separates Anthropic from consumer-focused competitors. Enterprise customers sign multi-year contracts, integrate deeply into workflows, and face high switching costs. Revenue from these customers is predictable, recurring, and premium-priced.</p><p>Anthropic with Claude Code is having its own ChatGPT moment, and it&#8217;s important, as coding is a big part of the economy and the job market, especially given the salaries. If there are 36M developers worldwide and their average salary is $48k per year, that would translate to $1.75T in developer salaries each year. If we only take the 20-30% production gains, we are talking about $350B to $525B of value created each year from these tools, and I would argue that the productivity gains are much higher than the 20-30%. </p><p>Anthropic&#8217;s TAM is bigger than many imagine, and its narrow focus on the enterprise and coding markets could prove to be a great strategy as things become more specialized, and it has built a strong head start and developer brand.</p><p>If you enjoyed this article, please consider subscribing to the paid subscription, where I share more in-depth analysis of AI companies and industry trends that I am seeing:</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/subscribe&quot;,&quot;text&quot;:&quot;Subscribe to Paid&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/subscribe"><span>Subscribe to Paid</span></a></p><p>Until next time,</p><p>I hope you found this article valuable. I would appreciate it if you could share it with people you know who might find it interesting.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.uncoveralpha.com/p/anthropics-claude-code-is-having?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.uncoveralpha.com/p/anthropics-claude-code-is-having?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p>Thank you!</p><p><strong>Disclaimer:</strong></p><p>I own Google (GOOGL) &amp; Amazon (AMZN), and Microsoft (MSFT) stock, which all have stakes in Anthropic.</p><p>Nothing contained in this website and newsletter should be understood as investment or financial advice. All investment strategies and investments involve the risk of loss. Past performance does not guarantee future results. Everything written and expressed in this newsletter is only the writer&#8217;s opinion and should not be considered investment advice. Before investing in anything, know your risk profile and if needed, consult a professional. Nothing on this site should ever be considered advice, research, or an invitation to buy or sell any securities.</p>]]></content:encoded></item></channel></rss>