UncoverAlpha

UncoverAlpha

The End of Best AI Model Wins era. Where Does the AI Value Go Next?

UncoverAlpha's avatar
UncoverAlpha
Oct 10, 2026
∙ Paid

Hey everyone,

For most of the last three years, the way the market valued AI was simple: whoever has the best model on the leaderboard wins, and the value flows to them as margins this year for frontier labs have gone to 60,70 or even 80%. I think we have now crossed a threshold where that framing is wrong for the majority of the economy.

For most knowledge work and most personal use, the improvement a user actually feels no longer comes from a better raw model. It comes from the harness, the integrations, the context the model gets, the alignment with what the user wants, and the product wrapped around it. The model is turning into one component of the product, not the product itself.

At the same time, competition at the model layer is getting brutal. OpenAI and Anthropic cut prices on the same day in September, open-weight models now carry the majority of tokens on some of the biggest gateways, and enterprises are routing work to whatever model is good enough and cheap enough. That is compressing margins at the model layer, and in my view, that value is spilling into two places: the data center layer (the hyperscalers and neoclouds) and, increasingly, the application layer.

In this article, I cover:

1. Why we have reached a good enough level of model capability for most work, and what that means for where improvements come from

2. What heavy token consumers like AlphaSense and Vercel are actually seeing in production

3. Why Meta’s Muse is the best consumer proof point of this shift

4. The model layer price war, with data from Ramp, Vercel and the latest price cuts

5. Why low-margin model tokens are a gift to the hyperscalers, and why negotiating leverage is the part most investors miss

6. Why I think we are entering an explosion of the AI application layer


Before we start with the article

Many of the insights from this article are from my recent conversation with Andy Hock (Cerebras Head of Strategy), Chris Ackerson (SVP at AlphaSense) and Kyle Cheng (Former Anthropic employee), which you can listen in full here:

Listen to conversation


Let’s dive in.

We have reached good enough for most of the economy

Think about how a company staffs a call center. Nobody hires a math PhD to handle refund requests, update a customer’s billing address, or qualify an inbound sales lead. They hire someone competent, train them on the company’s processes, give them access to the right systems, and measure them on resolution rate. Raw IQ above a certain bar adds almost nothing to that job. What adds value is knowing the policies, having the customer’s history in front of you, and being allowed to actually issue the refund.

AI models have now crossed that bar for a very large part of knowledge work. A few data points:

• The jump on real work has already happened. On OpenAI’s GDPval benchmark, which grades model deliverables against work from industry experts across 44 occupations, GPT-4o scored only 13.7% (wins or ties vs. the human expert). About 15 months later, GPT-5-high was at 40.6% and Claude Opus 4.1 at 49%. One leaderboard tracker now shows the top model at 87.9% as of September.

• The top of the leaderboard is crowded. On the Artificial Analysis Intelligence Index, Claude Fable 5.1 is at 66, Claude Opus 5 sits at 63, and Muse Spark 1.3, GPT-5.6 Sol and Grok 4.6 are tied at 61. Open-weight models, Kimi K3 and GLM 5.3, score 60, just three points behind Opus 5. Six index points now separate one of the best models in the world from some of the best models you can download for free.

• The cost gap is much wider than the capability gap. Muse Spark 1.3 costs about $0.55 per Intelligence Index task vs. $3.69 for Fable 5.1. That is roughly 6.7x cheaper for a model that is 5 points behind ( I even left out the more compelling contributor pricing, which would extend the gap even more). For a customer service workflow, that 5-point gap is hard to justify given the 6.7x pricing difference.

The GDPval paper found that more reasoning effort, more task context and more scaffolding all improve model performance on the benchmark. Two of those three levers have nothing to do with the model itself, but are the harness.

Andy Hock, Head of Strategy at Cerebras, put this well in the Token Economics conversation I hosted in September. Frontier models are like broad general experts: they handle complex reasoning, coding, and chat all in one. Where Cerebras sees enterprises leaning into open models is where domain specificity is required. In other words, the generalist PhD is still valuable for the hardest, most open-ended work. But most of the economy runs on specialists who know one domain very well.

This is also why I think we are sitting on a massive product overhang. The capability of the models that exist today is well ahead of what products are actually extracting from them. The bottleneck has moved from “can the model do it” to “has someone built the system that lets the model do it reliably, cheaply, and with the right data.” The next leg of improvement for most users will come from filling that overhang and less about what the next frontier model’s raw performance will be like.

The harness is becoming the product

For the non-technical readers, a quick definition. The harness is everything around the model: the system prompt and instructions, the tools the model can call, the retrieval system that pulls the right documents in, memory, the router that decides which model handles which request, the evals that check the output, and the guardrails. If the model is the engine, the harness is the rest of the car: the transmission, the steering, the brakes and the navigation system.

I wrote in May about how Cursor took its coding agent from Top 30 to Top 5 on Terminal-Bench 2.0 by only changing the harness. Since then, the evidence from heavy production users has only gotten stronger.

AlphaSense: 35 trillion tokens a year and no single winning model

AlphaSense is one of the biggest token consumers in enterprise software. Chris Ackerson, SVP of Product at AlphaSense, told me in our Token Economics conversation that the company consumes around 35 trillion tokens per year. For context, Google said that over the past 12 months, only about 375 Google Cloud customers each processed more than 1 trillion tokens. AlphaSense runs at roughly 35x that threshold.

Three things he said stood out to me:

1. Routing and post-training beat the frontier. Model routing and post-training their own models are two of AlphaSense’s most important areas of investment. Chris said that combination delivered about 3x better quality, at a lower cost, than applying frontier models to an unoptimized system with their content. AlphaSense has publicly shared that routing each question to the right model alone lifted answer quality by up to 2.8x over the baseline.

2. No single model wins across task types. Not too long ago, the consensus hypothesis was that models would converge, that they would all become roughly the same. AlphaSense’s testing found the opposite. Because each lab focuses its post-training on different things, models are diverging in what they are good at. That is exactly why AlphaSense needs to route different models to different tasks.

3. Context is now the bottleneck. In Chris’s words:

»The current constraint on output quality is no longer model quality, the actual bottleneck is context, it is can you get the right information to these models to actually do the work.«

That third point is, in my view, the most important one. If the bottleneck is context, then value accrues to whoever owns the context (proprietary data, permissions, workflow position) and whoever owns the system that delivers it to the model. Not to whoever has the model with the highest benchmark score.

Vercel: everyone is multi-model now

Vercel’s AI Gateway sits between applications and model providers, so it sees which models are used in real production traffic. Vercel CEO Guillermo Rauch has said in the Astra Vercel summit that when a team adopts AI, it always adopts multiple models, and the average enterprise adopts over five models. That lines up with an F5 report finding organizations rely on an average of seven AI models, with 52% chaining or orchestrating multiple models together.

More striking is what those models are. Open-weight models handled 56% of all tokens routed through Vercel’s gateway in August, the first month they were a majority, up from 7% in December 2025. On September 18, Rauch flagged a daily snapshot where open-weight models carried 78.4% of gateway tokens, up from 29% in June. Rauch has described models from OpenAI, Google, and Anthropic as steadily losing ground to open models, with openness and speed as the key trends going forward when it comes to adoption.

Microsoft is building for a world where every model is substitutable

The most telling comment from the hyperscalers came from Satya Nadella on Microsoft’s fiscal Q4 FY26 call (Apr–Jun 2026):

»We are building a new model system where the harness, context, memory, and action space are separate from any one model family, thereby moving the frontier on the cost-to-outcome curve. And it’s not just about cost. It also has the added benefit of business continuity and resilience because every model is substitutable.«

Microsoft also disclosed that customers building with models from multiple providers are up 5x since the start of the year.

Satya has also talked in the All in Summit about wanting the KV cache to be transferable between models. That sounds very technical, so here is why it matters. The KV cache is the model’s working memory of everything it has already read in a task: the documents, the conversation, the tool outputs. Today it is tied to one specific model, so if you switch models mid-task, the new model has to re-read everything, which costs time and tokens. If that memory becomes portable, switching models in the middle of a task becomes nearly free. That makes the harness, which owns the memory and decides who does what, even more valuable, and the individual model even more interchangeable.

The consumer proof point: Muse is not the best model, and it doesn’t matter

If you want one product that proves the model is no longer the product, look at Meta’s Muse.

Muse launched in the US on September 8. Sensor Tower projects it crossed 5 million US downloads in 22 days. ChatGPT needed 56 days to hit the same mark, Grok needed 103 and Claude needed 492. Only 23 apps in history have reached 5 million US downloads within a 22-day window, and roughly 85% of those were games. Muse also held the #1 spot among free apps on the US App Store for many days now and continues to show strong daily downloads.

Downloads are not usage, so the more important number is engagement. Per internal data reviewed by The Information, more than 3 million people prompt Muse at least once a week, more than 1 million do it every day, and over 4 million use it in some way weekly. According to multiple alt sources like TickerTrends and SimilarWeb, Muse now has over 2.5M daily active users.

Keep in mind that this is three weeks after launch, in only the US and Canada.

For this article, looking at the model underneath it proves our point. The shipping version, Muse Spark 1.3 (xhigh), scores 61 on the Artificial Analysis Intelligence Index. Claude Fable 5.1 is at 66 and Claude Opus 5 at 63. Muse is a very good model, but it is not the best model. Yet it is the fastest-growing AI consumer product we have seen on these metrics.

So what is driving it? As I wrote when it launched, a personal agent’s job is roughly 90% integrations, distribution and harness, and 10% raw intelligence:

• Harness. Every user gets a dedicated Linux VM in Meta’s cloud, with its own browser, storage and cron jobs, plus a separate “Sentinel” agent that approves every action and network request. The agent never sees your passwords.

• Integrations. Connectors to email, calendar, payments (Stripe Link, Shop Pay), health, smart home, dining, shopping, and Instagram and Facebook business tools. If a service has a public API but no connector, Muse writes its own. If there is no API, it uses the browser.

• Distribution. Muse lives as a separate app and inside WhatsApp chats and is promoted across a Family of Apps with 3.56 billion daily active people. To be fair, Meta pushed hard here: Sensor Tower says Muse took up to 50% of Meta’s daily house ad impressions in the two weeks to September 27. But that is exactly the point. Distribution is a product advantage a frontier lab cannot buy by simply having the best model on a benchmark.

• Alignment and trust. Meta delayed Muse for several months to work through privacy and security. Zuckerberg was explicit that »People aren’t going to adopt it if they don’t trust it«.

Zuckerberg framed the whole shift very directly in his interview with Joanna Stern:

»At a certain point, especially for consumer tools, it won’t matter so much how good the AI is at math, or how much better we could make it at math. What will really matter is how much more aligned can we possibly make it.«

In a post on X a few weeks later, he added that »people won’t want to use agents that are misaligned with them and that don’t do what they ask,« and that trust and reliability are »quickly becoming the most important capabilities« separating good AI products from bad ones.

Here, aligned means: does the agent understand what I actually want, does it do that and only that, and does it not embarrass me in front of my boss by sending the wrong email. That is a post-training, product, and harness problem, not a raw-intelligence problem.

Jeremy Stern’s Colossus profile of Zuckerberg summed up the strategic logic nicely. The potential install base for Meta’s AI is 3.6 billion people who don’t care whether a given model is six months behind the frontier. If AI commoditizes, value accrues to the complements Meta already dominates: distribution, attention, personalization and commerce.

The economics only work because the model is cheap enough. Meta gives free users up to 100 million tokens per week. At Muse Spark 1.3’s own standard API pricing of $1.25 / $4.25 per million input / output tokens, and assuming a typical agent mix of ~90% input, that is about $155 of API value per week, given away.

The model layer price war is here

The cleanest data I have seen on this comes from the Ramp AI Index, which tracks transaction data from more than 70,000 US businesses. The chart below shows that from January to September, token volumes across Ramp businesses grew roughly 7x on an indexed basis, while token spend grew only about 5.5x, peaked in July, and has been falling since.

Some of the specific numbers from Ramp’s lead economist Ara Kharazian:

• AI spend fell 5.2% in a single week, even as token volumes hit a new record.

• Token usage is up about 50% since spending peaked in July.

• The average effective token price dropped from $1.15 per million in March to $0.68 per million in early September, a 41% decline.

• Median AI spend among the top 1% of spenders fell 9.7% in a single month, from $7,976 to $7,205 per employee. That cohort drives the vast majority of enterprise revenue for the model companies.

Kharazian attributes the decline almost exclusively to price competition between OpenAI and Anthropic, plus cheaper and more efficient standard and lite models. On September 22, both labs moved on the same day with price cuts. The output price of OpenAI’s headline model has been cut by two-thirds in about ten weeks. OpenAI is even marketing Sol as beating Claude Opus 5 on its AutomationBench business-workflow test for 9% of Opus 5’s spend per task.

Where I differ slightly from Ramp’s read

I agree the frontier price cuts are the direct driver. But I think they are not the full story, and the reason is what is pushing the labs to cut.

Ramp notes that open-source models remain under 5% of business spend. That sounds like open-weight is irrelevant, but spend is the wrong lens. Open-weight tokens are so cheap that they barely register in dollars even when they dominate volume. On Vercel’s gateway in August, open-weight models processed 56% of tokens but only 14% of estimated spend. Closed-weight tokens cost about 7.8x as much as open-weight tokens on average. On OpenRouter, open-weight models accounted for 60% of US token consumption in August, up from 4.5% of enterprise tokens in early 2025 per one report cited by RedMonk. Ramp’s own token data also comes from a subsample that skews toward large API customers, which are exactly the companies with the biggest frontier contracts.

So my read is that three forces are working together. First, OpenAI and Anthropic are competing with each other. Second, open-weight models are setting a price ceiling for any workload that doesn’t need the frontier. GPT-6 Luna at $0.10 / $0.50 is not priced against Claude, it is priced against Qwen and DeepSeek. Third, people are simply finding that non-frontier models are good enough: Kharazian’s own explanation includes the “standard and lite” tiers, which are the labs’ own good-enough models.

The model layer is starting to look like what you’d expect from a market where the product is good enough, and there are many suppliers. Volume up, price down, and the gross margin pressure lands on the supplier.

Why low-margin model tokens are a gift to the hyperscalers

In June, I wrote about why token optimization is a gift to the hyperscalers, using the tollbooth analogy: the hyperscaler charges the same toll whether you drive a Ferrari or a Honda, and when people switch to the Honda, they drive a lot more miles. I want to push that argument one step further, because I think the part the market is still missing is the negotiating leverage.

A cheap token is not a small token.

The first thing to understand is that the price of a token and the compute behind it are two very different things. A large share of the price of a frontier token is the model provider’s margin: R&D, brand, leaderboard position. Strip that out, and the infrastructure cost of serving a large open-weight model is not small at all.

Take Kimi K3, one of the leading open-weight models. It has about 2.78 trillion parameters. Even compressed to 4-bit precision, its weights alone need about 1,390 GB of memory, which is 96.5% of the 1,440 GB of GPU memory in an eight-GPU DGX B200 system, before a single token of working memory is allocated. When a company routes a workload from Claude to Kimi, the price per token drops by multiples, but the GPUs, memory, networking, power and data center needed to serve it are in the same league. The hyperscaler still gets paid for all of it.

Cheaper tokens actually make more use cases economically viable, so they get adopted deeper into workflows.

Who is on the other side of the table

When the hyperscalers sell compute to OpenAI and Anthropic, they are negotiating with two of the largest, most sophisticated compute buyers on the planet. These buyers commit gigawatts at a time, sign multi-year deals, with prices that are much smaller than spot prices.

• Anthropic now has agreements with more than 10 cloud services, is building its own data centers, and is working with Fluid Stack to deploy TPUs in its own facilities. According to Dylan Patel (Semianalysis), it is buying fewer TPUs from Google over time and is expected to eventually build its own chips.

• Patel also noted that Google sells compute at roughly $20-30M per megawatt, when it could train its own models at about $50M per megawatt with that same capacity. That gap is, in effect, the discount a buyer of that size can command.

• On the price side, an H100 on a big public cloud costs roughly $5-7 per hour on demand, while the same chip on a one-year neocloud contract costs about $2.60. Committed, large-scale buyers pay roughly half of what the long tail pays.

• And the concentration shows up in the backlogs. Microsoft’s commercial RPO grew 84% to $678B, but only 25% excluding OpenAI.

Now compare that with the open-weight token. When a bank, a retailer or a SaaS company calls Kimi, Qwen or Nemotron through Bedrock, Vertex or Foundry, there is no model lab in the middle with bargaining power. The hyperscaler sets the price per token, bundles in security, compliance, logging and SLAs, and sells it to tens of thousands of customers, none of whom has meaningful leverage.

In 2026, AI model provider gross margins rose from the 30-40% range to the 70-80% range. With IPOs coming from the big AI model companies the emphasis on those margins will only increase. That is why this trend matters for the hyperscaler/data center layer and the application layer: it pressures AI model-builder margins and leaves more value to be made in other AI layers.

The evidence that serving open-weight tokens has become a very good business is already here. Patel says open-model inference providers like Baseten went from roughly 10% gross margins in early 2026 to around 60% today, with Fireworks, Together and others in the same range. Inference is now so profitable that GPU access is the hardest it has ever been for new startups, because anyone with capacity can make money immediately and outbid new entrants.

The token volume from a company using a cheap, good-enough model is similar to (or much larger than) what it would have been on the frontier, because cheaper tokens unlock more use cases. The hyperscaler charges the same or a higher infrastructure margin on those tokens, because it is no longer selling to two buyers with enormous leverage. The margin that disappears is the model provider’s, while the cloud providers keep it.

The hyperscalers’ job is almost simple: host as many models as possible and make switching between them frictionless. AWS now offers about 24 fully managed open-weight models on Bedrock, and every model it adds weakens the pricing power of any single lab while strengthening its own. The stickiness and lock in is the harness that the customer has in a service like Bedrock and with the hyperscalers.

The coming explosion of the AI application layer

In my 2026 forecasts in January, I wrote that I expected renewed interest in the application layer of AI this year. I think we are now at the point where the economics finally support it.

An application company sets its price based on the value of the outcome it delivers to the customer. Its cost of goods is set by the cheapest model that clears the quality bar for that task. For the last two years, those two things were stuck together, because the only models good enough for production were frontier models priced like frontier models. That is no longer true.

The margin problem that is now being solved

The AI application layer has had a gross margin problem. ICONIQ’s January 2026 data put the average AI-native gross margin at 52%, up from 41% in 2024, but still 20-30 points below the SaaS baseline. Some of the best-known names were much worse:

• Cursor had a reported gross margin of negative 23% in January 2026. After shipping its own Composer models, the second version built on Moonshot’s open-weight Kimi K2.5, its enterprise contracts turned gross-margin positive by April.

• Harvey saw its gross margin fall from about 50% at the start of the year to minus 50% by June, as customer token usage rose 20x. It turned positive again after Harvey released its own model in August, post-trained on Kimi K3.

• Decagon now routes 80% of customer queries through its own models.

• Ramp’s co-CEO Karim Atiyeh said the company used to think building its own model made no economic sense, but that changed as open-weight models improved dramatically.

Take a strong open-weight base, post-train it on your own domain data and traces, route the easy majority of work to it, and only escalate the hard tail to the frontier.

Customer service: the case study for you don’t need a math PhD

The best example of where this is going is Fin, formerly Intercom. Fin launched in 2023 on OpenAI’s GPT-4, later leaned on Anthropic’s Claude, and then built Apex, its own model post-trained specifically for customer support. Apex reportedly runs at about one-fifth the cost of frontier models and responds 0.6 seconds faster than the next-fastest competitor.

The resolution rate went from 23% at launch to 76% across 12,000+ customers by June 2026, with about 2 million resolutions per week. That is roughly the output of more than 6,500 human agents. The price to the customer stayed at $0.99 per resolution the whole time.

Let me illustrate why that combination is so powerful with a simple, hypothetical example. Assume that when running on a frontier model, inference costs equal 40% of the $0.99 price. Moving to a post-trained model at one-fifth of the cost drops that to 8%, which adds roughly 32 points of gross margin without changing the price. At the same time, a higher resolution rate means more billable outcomes from the same conversation volume. The value-based price stays, the cost collapses, and the margin moves to the application. Salesforce clearly agreed: it agreed to buy Fin for about $3.6B in June.

Customer service, accounts payable, sales development, claims processing, onboarding and back-office operations all look like this. They are bounded tasks with a clear definition of success. What they need is the customer’s history, the company’s policies, access to the right systems and a model that follows instructions reliably. None of that needs a model that can solve novel math problems.

Even large enterprises building internally are moving the same way. AT&T routes about 40% of internal employee queries to open models like Nemotron, Llama and Gemma, and targets 60-70% in the coming years.

Why I think this becomes an explosion, not a trend

Three things are lining up at the same time:

1. Inputs are getting cheaper fast. Effective token prices on Ramp fell 41% in six months, open-weight tokens cost about 7.8x less than closed ones, and post-training a strong open base is now within reach of a well-funded startup.

2. Pricing is moving to outcomes. Per-resolution, per-task, and per-outcome pricing ties revenue to the value of the work. When the token cost falls, the application keeps the difference.

3. The product overhang is large. Gartner expects 40% of enterprise applications to embed task-specific AI agents by the end of 2026, up from less than 5% in 2025. Each of those agents is a new application layer revenue line built on increasingly cheap models.

The application companies that win will be the ones that own something the model cannot get by itself: proprietary data and context, a position inside the customer’s workflow, distribution, or trust. Those are exactly the things that AlphaSense’s Chris Ackerson describes as the real bottleneck now.

What this means for different AI layers

If I simplify the whole article: the value in AI is migrating away from “who has the smartest model” and toward “who owns the infrastructure, the context, the workflow, and the distribution.” Here is how I think about each AI layer (hyperscaler, neocloud, AI model provider, semis, application), starting with the hyperscalers:

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Rihard Jarc · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture