AI to ROI is a podcast that shares how enterprises translate AI investments into measurable business value. Hosted by Ray Rike, Founder and CEO of Benchmarkit, the show features senior enterprise leaders and AI software executives who share how AI initiatives move from pilots to production, and how ROI is actually measured and achieved. In addition, each week, we publish a bonus episode with AI to ROI Newsletter co-author, Peter Buchanan to discuss the Big Story of the Week.
The AI to ROI podcast is the evolution of the original "Metrics to Measure Up" podcast.
Site
RSS
Apple
Données mises à jour le 25/09/2026
Classements récents
Dernières positions dans les classements Apple Podcasts et Spotify.
Liens partagés entre épisodes et podcasts
Liens présents dans les descriptions d'épisodes et autres podcasts les utilisant également.
AI to ROI, podcast de Ray Rike - Statistiques, épisodes et classements - My Podcast Data
Podcasts Similaires Basées sur le Contenu
Découvrez des podcasts liées à AI to ROI. Explorez des podcasts avec des thèmes, sujets, et formats similaires. Ces similarités sont calculées grâce à des données tangibles, pas d'extrapolations !
Six Challenges to Scaling Agentic AI in the Enterprise
mercredi 23 septembre 2026 • Durée 38:51
Deloitte research shows 85% of enterprises are working on agentic AI use cases, yet only 5% have put one into production that delivers meaningful ROI.
In this episode of the AI to ROI Big Story, Ray Rike and Peter Buchanan walk through the six challenges that separate the enterprises scaling agentic AI from those stuck in pilot mode: data readiness, production reliability, governance, cost visibility and ROI measurement, orchestration, and workforce readiness. Drawing on research from Gartner, Deloitte, McKinsey, KPMG, VentureBeat, and the Benchmarkit and Mavvrik 2026 State of AI Cost Governance report, they share how Amazon, FedEx, Lowe's, Cisco, Walmart, and Petrobras are approaching each challenge, and why the winners treat agentic AI as an operating model problem before a technology problem.
Covered in this episode:
Why clean data and a shared semantic layer are the foundation, with Google finding agent accuracy above 90% when data is standardized compared to 60% to 70% without it
How Amazon defines agent reliability through consistency, robustness, predictability, and safety, and why it builds an undo path into every agent
The governance gap: 85% of IT teams believe every agent is accounted for, but only 42% can say who owns them, and only 13% of enterprises believe they have adequate governance for the agent volumes Gartner projects
Why 98% of enterprises track AI infrastructure spend but only 11% can forecast it within 10%, and why cost per successful outcome is the metric that matters, illustrated by Uber exhausting its annual AI budget by April and Petrobras finding $120 million in tax savings in three weeks
Orchestration sprawl across multiple vendor platforms, and why workforce resistance is a myth when only 2% of technology leaders report significant employee pushback
Why every agentic AI pilot should have kill criteria agreed before development starts, with only three valid outcomes: scale, redesign, or stop
Read the full September 15 Big Story and subscribe to the AI to ROI newsletter at ai2roi.substack.com
Building an AI-First organization with Francis Brero, VP AI and Strategy HG Insights
mercredi 16 septembre 2026 • Durée 34:22
Francis Brero, VP of AI Strategy at HG Insights, joins Ray Rike to unpack what it actually takes to move a established software company toward an AI first operating model. Francis founded MadKudu, which was acquired by HG Insights, and was given a mandate to infuse AI into both the product and the internal operating processes. The conversation covers how HG Insights defines AI first, how the company organized around it, who owns execution across the functional groups, and who is accountable for the return on those AI investments.
What Ray and Francis Covered
The buyer is becoming an agent. Francis rebuilt the product assumption from an analyst consuming a data file to an agent consuming data on demand. Shipping an MCP interface was the first build in his first three weeks, and it opened the door to every customer already standing up agentic go-to-market systems.
Velocity is the separator between legacy and AI native. An agentic software development lifecycle changed shipping pace, and Francis makes the case that bolting an engine onto a bicycle only gets you so far before you have to build the motorcycle.
Pricing has to be re-architected for agent discoverability. Unique, high value data assets get priced down so an agent will actually reach for them, commoditized assets absorb more of the price, and data delivered through MCP is leased for a single workflow rather than sold into the customer warehouse. Both changes alter the shape of gross margin.
The $5, $50, $500 decomposition test. Every job to be done gets broken into atomic tasks, and each task gets a price the business would pay to outsource it. Five dollar tasks get automated. Francis notes the hardest part is that most operators have never decomposed their own work that far.
Centralized ownership of AI ROI. Francis owns the leading indicator and the productivity gain, the functional leader still owns the lagging indicator, and the two present the business case jointly to the CFO and CEO. Centralization is also the control point that prevents engineering spend from exploding when everyone gets model access.
Where Are the Killer AI-Native Application Companies?
lundi 14 septembre 2026 • Durée 36:15
Our co-hosts, Ray Rike and Peter Buchanan, open this AI to ROI: Big Story edition with a confession. Reviewing their own newsletter coverage over the past several months, more than half of the Friday news and analysis stories were about AI model companies; another twenty to thirty percent covered semiconductors, data centers, and SaaS to AI incumbents; and AI native application companies accounted for less than five percent.
That imbalance is the starting point for the question behind this episode: where are the killer AI native application companies, and what is standing between them and the breakout status their funding levels imply?
What Ray and Peter Covered
Capital concentration is setting the narrative. An August analysis from The Information found Anthropic and OpenAI now capture 89 cents of every dollar spent across the 35 largest AI startups, up four and a half points year over year. Ray pushes back on the forecast of a one trillion dollar AI software market by 2030, noting that SaaS took roughly twenty years to build a three hundred billion dollar base and the full cloud stack took twenty years to reach roughly eight hundred fifty billion.
A three-front squeeze on the application layer. Model companies are moving up the stack because the model itself will not be the durable moat, the same way Oracle and the client-server database vendors moved into applications in the 1990s. Systems of record and data platforms are re-architecting as AI-first and buying what they cannot build fast enough. And coding agents have made build versus buy credible again, with McKinsey reporting 32 percent of organizations have already decided against purchasing at least one software product or feature because they could build it internally.
COGS is the new CAC. AI native applications running on third-party models are delivering 50 to 65 percent gross margins rather than 80 percent, which pulls capital away from customer acquisition and raises the dependence on outside funding. Usage-based pricing amplifies the problem when the pricing architecture and guardrails were not designed to protect margin.
Why every AI conversation should be an ROI conversation - with Ketan Karkhanis, CEO ThoughtSpot
jeudi 10 septembre 2026 • Durée 31:56
Ketan Karkhanis, CEO of ThoughtSpot joins Ray Rike to make the case that the AI conversation has been stuck on models, tokens, and pilots when it should have started with a KPI.
Drawing on his time building Einstein Analytics and running Sales Cloud at Salesforce, and now leading ThoughtSpot, Ketan lays out how enterprises move from AI aspiration to measured outcomes by rewiring the operating model, not the org chart.
Topics covered in this episode:
Every AI investment conversation is an ROI conversation. Why AI projects should begin with the KPI to be improved rather than the model to be deployed, and why "AI saves you two hours a day" is the lazy version of the value case
Why pilots are where value goes to die. The case for starting with the customer and working backwards into process redesign, and ThoughtSpot's Spot30 program that targets one measurable outcome in 30 days instead of an open-ended proof of concept
Functional KPIs as the buildup to income statement impact. Financial ROI is a derivative of functional gains, so the practical path runs through metrics like NPS, average deal size, cycle time, and DSO before it reaches revenue and margin
Token anxiety and the gross margin problem. Ray shares benchmarking data showing 49% of software companies embedding generative AI had to reprice to protect gross margin and 25% halted an AI initiative over cost overruns. Ketan explains ThoughtSpot's credit-based pricing and the architecture behind it, using LLMs only for intent resolution rather than as a wrapper that resells tokens
Pricing predictability for the budget holder. Why CFOs do not need certainty, they need a unit of consumption they can forecast, and why tying pricing to the customer's own business model is what makes the spend defensible
The OpenAI vs Anthropic Battle for the Enterprise
mercredi 9 septembre 2026 • Durée 37:53
For the first time, the two leading model companies can be compared on an apples to apples basis. In Q2 2026, Anthropic booked $11.6B against OpenAI's $6.7B, and posted roughly $300M in operating profit while OpenAI's operating loss widened to $12.3B.
Ray Rike and Peter Buchanan unpack what that actually signals to an enterprise CFO signing a multi-year platform commitment, and why one profitable quarter does not settle a market where both companies carry hundreds of billions in infrastructure obligations.
The bigger argument: OpenAI and Anthropic are running near-identical playbooks. Comparable frontier models, comparable pricing, the same enterprise logos, competing coding agents, parallel vertical pushes, and mirrored forward-deployed engineering ventures. When strategy converges, execution becomes the differentiator.
What Ray and Peter cover in this episode:
The market being chased: $64B in AI model platforms in 2026 per Gartner, agentic coding tools growing from $4B to $30B by 2030, and implementation services from $18B to $76B by 2031
Claude Code economics, including median enterprise spend tripling from $69 to $219 per month, and why Ray wants to see gross and net revenue retention before calling it durable
Codex bundled inside ChatGPT, 100,000 migration signups in 10 days, and why signups without retention remain a vanity metric
The go to market gap: Anthropic's stable commercial leadership bench versus four sales leaders in two years at OpenAI, and what that churn costs in enterprise continuity
Channel conflict as both companies ship vertical products that compete with the partners embedding their models, and Harvey's move away from Claude as the early warning
Distribution bets: the Salesforce and Claude Force deal with $300M in committed token spend, Ode at $1.5B, and Deploy Co at $4B
The insurgents that make this a multi front war: Google's install base and invisible AI distribution, Microsoft pushing lower cost MAI models with 13% orchestration share, NVIDIA's IBM-style ecosystem play from the 1980s, and open-weight models now at 72% of OpenRouter tokens
A rapid-fire close on the five moves each company needs to make to win
Enterprises Can See Their AI Bill, But They Can't Predict It
jeudi 3 septembre 2026 • Durée 37:00
Ray Rike sits down with co-host Peter Buchanan to unpack Benchmarkit's August 11th research edition, based on a 396-company survey conducted in partnership with Mavvrik. The conversation moves past the "is AI adoption happening" question and into the harder one: do enterprises actually understand what AI is costing them, and can they see it coming.
Tracking isn't the same as forecasting. 98% of companies say they track AI costs and 44% call themselves advanced, but only 11% can forecast spend within 10% of budget, and 54% miss their forecast by more than 26%.
AI costs break down into three distinct patterns: product AI (impacts cost of goods sold and pricing), workflow automation (impacts operating margin), and true agentic AI (non-deterministic, harder to predict, and prone to costly retry loops).
Data platforms, not LLM tokens, are the top driver of budget misses. 47% of companies cite data storage and platform costs as the leading cause, ahead of token costs, with GPU, network, and human-in-the-loop costs frequently left out entirely.
Granular attribution remains rare. Only 36% of organizations running agentic workloads can attribute cost by individual agent, and just 29% can attribute AI coding tool spend down to the individual developer.
The business consequences are real and already happening. 49% of AI product companies have had to reprice, 25% have halted an AI initiative outright, and 40% have had to report an AI cost overrun to their board.
Multi-cloud and multi-model complexity is compounding the problem. The average company now uses 2.4 models and 68% run hybrid hosted and on-prem AI workloads, adding new layers of capital, depreciation, and orchestration cost that finance often isn't capturing.
Stripe and OpenRouter Combination Wants to be #1 in Routing your AI Workloads
mardi 1 septembre 2026 • Durée 35:18
Stripe agreed to pay $7.5 billion for OpenRouter, roughly six times the $1.3 billion valuation the company raised at just 85 days earlier. Ray Rike and Peter Buchanan break down why a payments company bought the plumbing that routes AI model traffic, what OpenRouter's 5.5 percent pass through economics actually look like, and why model routing has become the control point for enterprise token spend. The bigger story is not the headline multiple. It is that CFOs have run out of patience on AI cost, and the routing layer is where token maximization turns into token cost optimization.
The deal math. A three year old company with fewer than 100 employees and 8 million developers on the platform went from a $1.3 billion round in May to a $7.5 billion acquisition in mid August. Annualized revenue was roughly $140 million in July, up 3x since April, then up another 15 percent to $160 million within weeks. Reported split: about $1.5 billion to the founders, $6 billion to investors.
The business model. No subscriptions, no site licenses, no annual contracts. OpenRouter takes 5.5 percent of prepaid credit value and passes 94.5 percent through to the model provider with no markup. Bring your own key customers pay a 5 percent overage fee above thresholds of roughly $25,000 per month on standard plans and $200,000 per month on enterprise.
The contrarian bet that paid off. Founded in early 2023 when consensus said a handful of frontier labs would win outright, OpenRouter bet no single model would win and that developers would need a neutral layer in the middle. With roughly 70 percent of token traffic now flowing to open weight models and 55 trillion tokens per week crossing the platform, that bet looks prescient. The routing data is the moat, not the software.
Why OpenAI subsidized a gateway it does not own. OpenAI cut prices on three models by half on OpenRouter to win developer share on neutral ground. Distribution is the scarce commodity right now, and at a 50 percent discount the effective customer acquisition cost approaches zero. The discount is almost certainly temporary. Anthropic, with an IPO closer in, did not match it.
AI Cost Governance to AI Cost Optimization with Sundeep Goel, CEO, Mavvrik
mercredi 26 août 2026 • Durée 31:43
Ray Rike sits down with Sundeep Goel, co-founder and CEO of Mavvrik, for a deep dive into the second annual AI cost governance research report from Maverick and Benchmarkit. With 30+ years in enterprise software, Sundeep brings a practitioner's view of why AI spend is proving far harder to forecast and govern than cloud spend ever was, and what finance and technology leaders need to do about it.
Forecast accuracy is getting worse, not better. Only 11% of companies can forecast AI spend within plus or minus 10% in 2026, down from 15% the prior year, even as aggregate AI spend has climbed sharply.
Unexpected AI costs are already changing business decisions. 49% of companies with AI-powered digital products have had to modify pricing, and 25% have stalled or shelved an AI project due to cost overruns.
Traditional per-seat and percentage-of-spend pricing models are breaking down. Variable, high AI cost of goods sold is pushing companies toward consumption-based and outcome-based pricing.
Cost tracking and cost allocation are two different problems. 98% of companies say they track AI costs at an aggregate level, but only about 5% can allocate that spend down to the customer, application, or agent level where it actually drives decisions.
Cost governance is the precursor, cost optimization is the destination. Drawing a parallel to the 30%+ waste long documented in cloud spend, Sundeep argues a similar magnitude of waste exists in AI spend today, and that dynamic model selection, prompt caching, and rate optimization are the next frontier.
ROI has a solvable half and a harder half. Cost is measurable today; value remains subjective and requires a company-specific framework to quantify.
Microsoft is Running an AI Marathon
mardi 25 août 2026 • Durée 26:24
For 24 months, the market narrative held that Microsoft was losing the AI race to the lab it had funded. Then came the July 30th earnings call, and the stock added roughly $450 billion in market capitalization the next day, the largest single-day gain by any public company on any exchange in history.
In this week's Big Story, Ray Rike and Peter Buchanan work through why the reassessment happened and what it says about where enterprise AI value is actually accruing.
The thesis is not that Microsoft builds the best model. It does not. The thesis is that Microsoft has figured out how to monetize the gap between model quality and market value at the exact moment the novelty premium on frontier AI is wearing off and enterprise buyers are starting to ask price-performance questions instead of capability questions.
What the episode covers:
The numbers behind the trade. Azure grew 43% and crossed $100 billion in annual revenue for the first time. The AI-specific business, bundling AI consumption with Copilot and GitHub, hit a $37 billion annual run rate, up 123% year over year. Total quarterly revenue reached roughly $90 billion, up 18%, with net income up 31% to nearly $36 billion.
The concentration question inside that number. Roughly $24 billion of the $37 billion AI run rate is hosted by OpenAI, about 7% of the company's total revenue. Peter frames the two ways to read it, as systemic customer concentration risk or as evidence that hundreds of thousands of workloads now run through Azure, and explains why he leans toward the second.
The toll booth position. Microsoft is the only cloud provider that hosts the OpenAI, Anthropic, and Mistral models alongside its own MAI family on the same platform. As model orchestration becomes a real buying criterion, that means an Azure customer never has to leave the platform to route traffic across model families, and gets one contract and one bill for all of it.
Can U.S. Frontier AI Labs Survive a Price War with Open-Weight Models?
mardi 18 août 2026 • Durée 32:27
When Kimi K3 landed, the headlines said Chinese open-weight models had caught the American labs, and the AI trade sold off from chipmakers to the labs themselves. Nobody stopped to ask whether cheaper and more profitable mean the same thing.
In this week's AI to ROI Big Story, Ray Rike and Peter Buchanan run the actual business math on both sides of the fight and find that neither side has the balance sheet to fight a sustained price war.
The setup is stark. Anthropic is projecting its first-ever quarterly operating profit of roughly $559 million in Q2, with an annualized revenue run rate near $47 billion, up from $9 billion at the end of last year. That profit disappears the moment it tries to match open-weight pricing. Gross margin is running around 40%, roughly 10 points below the internal forecast, against more than $350 billion in data center commitments coming due over the next three to five years. OpenAI's picture is even thinner: roughly $30 billion in ARR, a projected $14 billion operating loss this year, data center commitments approaching $1 trillion, and an advertising business off to a slow start that needs to reach $100 billion by the end of the decade to close the gap.
What the episode covers:
Why the Chinese open weight labs are not the subsidized price killers the coverage assumed, with Z.ai's gross margin falling from 41% to roughly 15%, DeepSeek near break-even on about $500 million of revenue and already back in market after a $7 billion round, and Moonshot and MiniMax raising at rising valuations rather than running toward profit
The open weight versus open source distinction that changes the entire economic model, since every new customer requires more chips, power, and data center capacity, and Z.ai's own numbers show roughly 49 to 50% gross margin on customer hosted deployments versus about 19% when they host and serve via API
The price war math itself: frontier models cost roughly $6 to $8 per million output tokens to serve, Anthropic's $25 per million on Opus produces about a 70% gross margin, and repricing down to the $4 to $6 range where Meta's Muse Spark sits flips that margin from positive 70% to negative 65%
Why DeepSeek cut prices on a low-end model and then, two weeks later, told customers to prepare for substantial increases across the line, particularly on API access
Two lessons learned. Operationally, he moved to second order process redesign before the organization understood first order automation, and had to reset to crawl, walk, run. On the product side, AI generated ten times more code and therefore more total defects, until he introduced adversarial review across different model families rather than same family self review.
If you are finding value in these conversations, please subscribe on your favorite podcast app, leave a five star rating, and connect with Ray Rike on LinkedIn to suggest future guests.
Retention is still experimental. Some AI native application companies report churn between 25 and 45 percent, roughly triple a mature SaaS benchmark, which reflects how little is deeply embedded yet and how low switching costs remain in the early departmental deployment.
The visibility math. Tool Radar tracked roughly 11,500 mentions across 387 tech media sources between February and August. Only seven percent of the more than 10,000 software products tracked received any coverage at all, and ChatGPT alone accounted for close to 22 percent of the sample.
What the breakouts have in common. Harvey built for legal depth, moved onto a purpose-built model to fix its cost structure, and staffs 40 to 50 percent of pre-sales with people from the legal industry. Abridge went narrow on clinical documentation. EvenUp prices against recovered damages rather than seats. Fieldguide earned the AIUC-1 certification and uses audit firms as a distribution channel. The five traits Peter pulls out of those cases include embedded domain context, an expensive workflow worth solving, proprietary context accumulated from customer interaction, expansion from single task to full system, and outcome-aligned pricing.
Ray closes with the position he will stake his reputation on. The AI native applications that win will own complex multi step workflows and the data around them, make the economic value obvious and tightly coupled to the pricing model, and make the underlying model the least interesting part of the value proposition.
If you are getting value from these episodes, please subscribe to the AI to ROI podcast, give us a five star rating, and reach out to Ray Rike on LinkedIn if you are an AI native application company or an enterprise executive with an AI success story to share.
Customer success as the ultimate AI to ROI metric. How ThoughtSpot renamed customer success to customer outcomes and FDEs to AI outcome managers, opens every staff meeting with a five-metric customer review, and assigns a C-suite owner to every AI initiative
Follow the AI to ROI podcast on your favorite podcast app, and connect with Ray Rike on LinkedIn to suggest future guests.
This is not OpenAI versus Anthropic. It is both of them defending against hyperscalers above and open-weight economics below while battling each other for the same enterprise budget.
The report closes with role-specific fixes for the CFO, CIO/CTO, FinOps, product leadership, and engineering, with the throughline being visibility first, economic value second, before committing to multi-million dollar AI bets.
For the comprehensive research report, click here, and to subscribe to the AI to ROI newsletter at ai2roi.substack.com.
A market with no agreed definition. Gartner sizes the pure play AI gateway market at roughly $250 million in 2025 growing to about $2.9 billion by 2031, a 51.6 percent CAGR. IDC frames it far more broadly as AI cost governance and orchestration at roughly $11.2 billion, potentially subsumed into a $54.8 billion API management market. The category is crowded fast: Portkey, Requesty, Kong, Eden AI, plus gateways shipped this year by Ramp, Snowflake, Databricks and Cursor, and a rumored Meta entry.
What boards and CFOs should watch. The open question is whether Stripe keeps OpenRouter operating as a standalone neutral entity for the next two to three years or bundles routing, metering, billing and payments into a single lock in system too early. Ray's guidance to enterprise buyers: resist the instinct to negotiate AWS style multi million dollar model commitments. Models are evolving too quickly to lock in on today's price performance. Preserve the flexibility to switch, and evaluate every routing provider on one question. Is it an honest broker?
Read the full analysis in the August 25th AI to ROI Big Story newsletter at ai2roi.substack.com, published every Tuesday alongside our weekly summary and analysis of the top 10 stories in AI.
Ray and Sundeep close with three rapid-fire questions on who should own AI ROI measurement, the key variables for driving it, and advice for early-career professionals navigating an AI-disrupted job market.
If you're finding value in the insights our guests share, subscribe to the AI to ROI podcast on your favorite platform and leave a five-star rating. Have a guest suggestion? Reach out to Ray Rike on LinkedIn.
The equity stakes that make Microsoft indifferent to who wins. A 27% position in OpenAI now valued around $220 to $240 billion, plus an Anthropic stake that produced a $3.2 billion unrealized gain last quarter and added 33% to earnings per share.
Seven years of OpenAI partnership history and how it changed. The original $1 billion investment in July 2019 in exchange for Azure exclusivity, the $13 billion follow-on in 2023 with exclusive commercial API rights, the mid-2025 strain when OpenAI signed separately with Oracle for compute, and the public benefit corporation restructuring that converted profit sharing rights into equity while Microsoft gave up hosting exclusivity, retaining a first mover window on new releases and non-exclusive licensing rights running through 2032.
The MAI model family and Satya Nadella's frontier diffusion strategy. Seven models announced at Build in June with Microsoft owning the IP. MAI Thinking One is a mixture-of-experts reasoning model with about a trillion parameters, only 35 billion of which are active at any time, making it cheap to run at scale. MAI Code One Flash, a 5 billion parameter coding model, became the default in GitHub Copilot. Image, voice, transcription, and cybersecurity models fill out the set.
What that routing actually saves. Microsoft reported an 84% reduction in GPU costs for image generation in PowerPoint and an 89% reduction in GPU costs for voice processing in the Dynamics 365 contact center by keeping high-volume, repetitive traffic on models it controls end-to-end rather than sending it to a frontier lab.
Two production examples. Dragon Copilot, the clinical assistant built into Dragon Medical One, and DAX Copilot that drafts clinical notes from a patient visit, and Project Perception, an agentic cybersecurity platform running continuous vulnerability scanning and patching through coordinated agents rather than generating alerts for human triage. Both are high-volume, low-novelty work, which is exactly the profile these models were built for.
Where the benchmark story does not match the marketing. On SWE Bench Pro, MAI Thinking scores around 53% against roughly 69% for the current Anthropic flagship on independent leaderboards. Microsoft's own AI chief has acknowledged a lead measured in months rather than weeks. Ray's point for buyers: MAI benchmark figures come from Microsoft's internal technical reports, and should be treated as vendor claims until independently validated.
The distribution advantage that does not show up in any benchmark table. 70% of the Fortune 500 using Microsoft 365 Copilot, more than half a million companies in the Microsoft AI Cloud Partner Program, and 30 million paying Copilot seats, up from 20 million the prior quarter, plus more than 20 million developers and over 140,000 organizations on GitHub Copilot.
The sovereignty play in Europe. A multi-billion dollar Mistral partnership funding European data center build out in exchange for prioritized capacity and distribution rights, which sidesteps the capital intensity of building EU capacity from scratch and answers the sovereignty question directly in a region where OpenAI and Anthropic are less established.
Frontier Company, and the forward-deployed engineer story revisited. A $2.5 billion subsidiary staffed largely from existing engineers rather than new hires, with more than 6,000 forward-deployed engineers implementing inside client environments. Ray connects it back to the FDE economics the show covered previously, and Peter explains why pulling the large consultancies into the tent matters more than the headcount itself.
What CFOs and GTM leaders should take away:
Route by workload, not by vendor loyalty. Frontier quality for the problems that need it, cheaper controlled models for Excel formulas, transcription, and support tickets. The difference in margin between those two decisions is the whole story.
Treat self-reported model benchmarks as vendor claims. Until independent evaluation catches up, keep MAI models off your most technically demanding workloads.
Don't skip the security diligence. Microsoft shipped real vulnerabilities this year, including a Copilot flaw that could have exposed customer files. The new Microsoft and NVIDIA open weights security association, which grew from 20 founding members to 125, is a good leading indicator but not a substitute for your own testing.
Recognize what a vendor with equity in multiple labs is actually incented to do. When the provider profits regardless of which model you select, the steering pressure toward a specific model family drops.
Weigh the balance sheet underneath the AI investment. Microsoft is funding all of this from a core business that grew 18% and saw net income up 31%, a cushion the pure-play labs do not have.
Ray's read: never bet against Microsoft, and the combination of a large profitable core business plus existing enterprise access is a genuine advantage. Peter's read: this is a price-performance play rather than a capability play, and whether it is a durable moat or a stopgap until the cost curve moves again remains an open question.
For the full analysis behind this week's big story, subscribe to the AI to ROI newsletter at ai2roi.substack.com.
Google as the structural outlier, with 83% growth in its cloud and AI segment, Gemini embedded across fifteen products with more than a billion users each, and the Apple Siri deal extending reach toward two billion devices
Where Kimi K3 actually fits, including the caveats nobody is pricing in: weights released only last week, no published large-scale production deployments, two to three times the token consumption on complex tasks, and infrastructure requirements around a 72 GPU rack that costs millions to install and millions a year to run
The geopolitical wildcard, with Washington weighing sanctions or outright bans on Chinese open-weight models and distillation-related IP exposure still unresolved
The metric that resets the argument: cost per completed task
Price per token is the easiest unit to measure and the wrong one to buy on. Ray and Peter walk through a frontier lab evaluation that assumed a fully loaded remediation cost of $17 per failed attempt, roughly 10 minutes of a human operator's time. Claude Opus completed the task about 90% of the time at roughly $2.56. Meta's Muse Spark, priced at a quarter of Opus on tokens, succeeded 75% of the time and landed above $5.80 per completed task. The list price was 75% lower, and the delivered cost was more than double. A fifteen-point reliability gap did all the work.
The formula Ray offers turns a squishy quality debate into something a CFO can actually evaluate: cost per attempt, plus failure probability times fully loaded remediation cost, divided by success rate.
What CFOs and GTM leaders should take away:
Build cost per completed task into vendor evaluation and make vendors compete on that number rather than on a token price list
Segment AI workloads by what a failure actually costs, since a wrong answer in a regulated process or a customer support interaction carries a very different price than the model delta suggests, and standardizing on one model across both workload types to save on price is the common mistake
Treat orchestration as a cost lever, not just the model choice, since Cursor's internal testing found coordination across multiple models delivered comparable code quality at a fraction of the cost of a single large model
Run your own evaluations instead of trusting public leaderboards, since LMSYS Chatbot Arena measures human preference rather than task completion and says nothing about your workload
Watch the switching trap, because a cheap model that attracts heavy traffic today can reprice two or three times higher in a quarter once the vendor needs margin
Do not overreact to the open weight scare by standardizing on the cheapest option, since compute, power, and people costs are accelerating, and vendor durability still belongs in the evaluation
Ray's read: on real-world performance and total cost of ownership, the closed-weight labs still hold the advantage, their prices keep coming down, and they have every incentive to avoid a price war. Peter's read: open weight economics favor whoever hosts the model more than whoever built it, which makes the hyperscalers the quiet winners regardless of how this resolves.
For the full analysis behind this week's big story, subscribe to the AI to ROI newsletter at: ai2roi.substack.com