<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
<channel>
    <title><![CDATA[Research - Creative Strategies]]></title>
    <link>https://ghost-development-2724.up.railway.app/research/</link>
    <description><![CDATA[Independent technology analysis since 1969.]]></description>
    <atom:link href="https://ghost-development-2724.up.railway.app/feed/research/" rel="self" type="application/rss+xml" />
    <lastBuildDate>Tue, 01 Sep 2026 20:34:51 +0000</lastBuildDate>
    <language>en</language>
    <ttl>60</ttl>
            <item>
                <title><![CDATA[Breaking the Memory Wall With CXL]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/breaking-the-memory-wall-with-cxl/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/breaking-the-memory-wall-with-cxl/</guid>
                <dc:creator><![CDATA[Ben Bajarin]]></dc:creator>
                <pubDate>Thu, 06 Aug 2026 20:00:57 +0000</pubDate>
                
                <description><![CDATA[We spent the last few days at FMS (Future of Memory and Storage) and had meetings with all the key players in the memory, storage, and now also interconnect/networking, ecosystem. A clear takeaway was better line of sight to the CXL standard to start to solidify. If you are not familiar with CXL it]]></description>
                <content:encoded><![CDATA[<p>We spent the last few days at FMS (Future of Memory and Storage) and had meetings with all the key players in the memory, storage, and now also interconnect/networking, ecosystem. A clear takeaway was better line of sight to the CXL standard to start to solidify. If you are not familiar with CXL it is a standard interface that gives the industry a common way to attach and share memory over a PCIe-based physical link. Up to this point the deployment model around that interface is still being developed, which is why the ecosystem can support several architectures, from a memory box inside a rack to pooled or optical memory systems.</p><p>From our conversations we believe deployments are more likely to start next year and build into 2028, but the ecosystem is maturing enough for CXL to move from a standard into something that can be deployed with enough customers to meaningfully start to deploy in their AI compute infrastructure. The first and most logical customer base is the hyperscalers, which can put memory and inference workloads into custom compute clusters and use the surrounding ecosystem to qualify the architecture.</p><p>It is our conviction that solving the memory wall will take many different shapes by many different players but the same problem statement remains. We need more memory, and designers are up against how much memory can go on CPU/XPU/GPU package or near the package. Having ways to expand the available memory pool while keeping latency low enough for selected near-memory workloads is the promise of CXL.</p><h3 id="memory-is-becoming-a-fleet-problem">Memory is becoming a fleet problem</h3><p>The old server model ties memory capacity to one processor and its local channels. AI workloads make that boundary more expensive because context, KV cache, and orchestration state can grow faster than the memory attached to one compute device. A fleet can have plenty of memory in total while individual CPUs or accelerators still run short.</p><p>The practical distinction is important: “hot” describes the role memory plays in the workload, while local, off-die, off-board, and rack-level describe where that memory sits. HBM can remain the hot tier in an attached appliance when the fabric preserves the bandwidth and latency the workload requires. DDR can be split by role in the same way. Local DDR can serve the CPU’s most active data, while DDR4 or DDR5 behind a CXL controller can provide a larger attached tier and eventually a shared pool. That remote memory will not have the same latency as on-package HBM, but it can still be the highest-performance tier available outside the package.</p><p>CXL gives the system those placement choices through a coherent interface. The first commercial use is likely to be a card, module, or box that adds memory to one host. Later designs can attach several hosts to a pool, with software deciding which data belongs close to compute and which data can move into the attached memory. The value comes from expanding the working set and using capacity more fully without buying a full server for every increment of memory.</p><p>The graphic below frames the opportunity by workload rather than by device. HBM remains closest to compute for the hottest accesses, while local DDR/SOCAMM and CXL-attached memory support larger working sets, overflow KV cache, retrieval buffers, and shared state. CXL is not a fixed warm tier; its role depends on the memory type, topology, and workload.</p><p><strong>Exhibit 1. Inference moves the memory problem from model weights to active state. CXL expands the set of places where that state can live.</strong></p><h3 id="deployment-flexibility">Deployment flexibility</h3><p>It is noteworthy how early we are in the rack scale compute era. By our estimates, using our accelerator installed base model, we believe rack scale accelerators in the range of 9-11% of the total AI accelerator installed base. The shift to rack scale solutions are the necessary catalyst that will enable CXL, and other solutions even if custom, to be designed by the customer. Knowing the data center customer is scaling their rack scale infrastructure helps as a catalyst for CXL whether that deployment is north south or east to west in its location. Enough of the ecosystem was present, through announcements and demonstrations at FMS, to make the deployment path more visible. Memory suppliers, interconnect companies, switch vendors, custom silicon providers, and networking ASIC companies were all discussing a path toward viability with much different language than earlier in the year. As of now, we expect the larger deployments are more likely to arrive in 2027 and build into 2028, yet the ecosystem is now moving from evaluation toward qualification, driven by a key set of customers.</p><p>We believe hyperscalers are the first catalysts for several reasons. They control rack specifications, work with ODMs, and can make system-level TCO decisions that are difficult in standardized OEM deployments. <strong>They also have access to a large legacy memory resource that is not always useful in its original server configuration, but could become valuable again in a CXL-based system.</strong> We explain the size of that resource and the qualification path in the full report. That flexibility gives hyperscalers room to qualify a CXL memory box, attach memory beside a CPU or custom ASIC, or build a dedicated memory system. The economic benefit is that they can add controller and system content around capacity they already own.</p><p><strong>Our key take here is:</strong> We have increased confidence that the CXL standard will emerge as a preferred additional approach to infrastructure build out. Its economic and TCO benefits, detailed in the full report, along with its flexibility in implementation give it many advantages over other solutions. While we understand the tradeoffs, we believe the biggest customers in the world are positioned to drive the adoption of CXL and have distinct advantages over others in the market. The competitive advantage available to those customers through CXL could become evident.</p><h2 id="inside-the-full-report">Inside the Full Report</h2><ul><li>Why hyperscalers are likely to become the first large CXL customers, and how their ODM model helps them qualify custom systems faster.</li><li>How much existing memory could become reusable, what that changes for deployment timing, and the potential TCO benefit.</li><li>Where CXL fits alongside HBM, local DDR and MRDIMM, HBF, proprietary memory attach, and networked-memory approaches.</li><li>Our CXL market model through 2030, including the path from single-host expansion to rack-level pooling.</li><li>The beneficiary map across controllers, memory suppliers, interconnect, systems, and software, plus the production signals needed to validate each opportunity.</li></ul><figure class="kg-card kg-embed-card"><iframe class="no-lazyload" data-no-lazy="true" style="width: 100%; height: 400px; border: 1px solid #EEE; background: white; overflow: hidden;" src="https://www.thediligencestack.com/embed"><!--kg-card-begin: html--><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><!--kg-card-end: html--></iframe></figure>]]></content:encoded>
            </item>
            <item>
                <title><![CDATA[Copilot Seats Are the Bridge for Agents]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/copilot-seats-are-the-bridge-for-agents/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/copilot-seats-are-the-bridge-for-agents/</guid>
                <dc:creator><![CDATA[Carolina Milanesi]]></dc:creator>
                <pubDate>Thu, 30 Jul 2026 20:56:56 +0000</pubDate>
                
                <description><![CDATA[Microsoft disclosed more than 30 million paid Microsoft 365 Copilot seats this week, and the number works best read as an installation count. Every paid seat puts Work IQ inside a customer’s tenant, the context layer over roles, projects, artifacts and institutional knowledge that Microsoft’s agents]]></description>
                <content:encoded><![CDATA[<p>Microsoft disclosed more than 30 million paid Microsoft 365 Copilot seats this week, and the number works best read as an installation count. Every paid seat puts Work IQ inside a customer’s tenant, the context layer over roles, projects, artifacts and institutional knowledge that Microsoft’s agents need in order to produce anything specific to a company. The seat buys the context, the context makes the agent worth running, and the agent consumes tokens on Azure, which grew 43 percent and passed $100 billion in annual revenue for the first time this year.</p><p>That sequence explains why <a href="https://www.microsoft.com/en-us/microsoft-365/blog/2026/07/30/the-next-measure-of-ai-momentum-is-work-transformed/?ref=ghost-development-2724.up.railway.app">The next measure of AI momentum is work transformed</a>, published by Jared Spataro on the Microsoft 365 blog the morning after FY26 Q4 earnings, spends its customer proof points on agent counts. Premera Blue Cross has built more than 900 agents, most of them by employees outside IT. Microsoft’s own supply chain team runs more than 70 across planning, sourcing, fulfillment and logistics. Jared Spataro closes on the growth and governance of agents as the measure that matters next, which is a long way from the hours-saved framing Copilot launched with.</p><h4 id="what-the-seat-actually-installs">What the seat actually installs</h4><p>On the January earnings call Satya Nadella described the data underneath Microsoft 365 as the most important database any company running Microsoft has, pointing to the tacit information it holds about people, relationships, projects and artifacts. Copilot Chat users on the bundled tier do not get Work IQ grounding. The paid seat is the thing that turns that database into something an agent can query with permissions attached.</p><p>This makes broad deployment a technical requirement in its own right. A partial rollout produces a partial context graph, and agents grounded in a partial graph produce generic output. Enterprises that bought a pilot allocation of 500 seats to test ROI were, without knowing it, testing the version of the product least likely to work. The economics that looked unjustifiable at $30 per user for time savings look different when the seat is the prerequisite for an agent that closes a workflow.</p><p>Agent 365 extends the same logic to identity. Agents carry their own credentials and permissions, Entra governs them, and Microsoft Scout, introduced in June, runs in the background under that model. The governance layer is the asset competitors cannot assemble quickly, because it depends on already sitting inside the enterprise directory.</p><h4 id="the-trajectory-with-two-qualifiers">The trajectory, with two qualifiers</h4><p>Microsoft had never published a Copilot paid seat count before January 2026. It disclosed 15 million that month, 20 million in April, and more than 30 million this week, with net adds going from 5 million to 10 million quarter over quarter. Customers running more than 50,000 seats grew more than sevenfold year over year. NHS England is deploying to 505,000 clinicians and staff, HSBC has committed to 200,000 seats, and EY is putting the E7 suite in front of 400,000 employees.</p><p>Q4 is Microsoft’s fiscal year end, when large deals cluster to capture year-end discounting on exactly the product the vendor wants to push, so the 50 percent sequential jump needs one more quarter before it reads as a run rate. And the attach rate stays small. Against a Microsoft 365 commercial base of roughly 464 million paid seats, derived from the 6 percent annual growth Microsoft reported against the 450 million figure it last published in January, 30 million works out to about 6.5 percent. The seat count doubled in six months and more than 93 percent of the installed base has not bought.</p><h4 id="comprehensive-stopped-being-the-differentiator">Comprehensive stopped being the differentiator</h4><p>Microsoft’s portfolio breadth is real. Microsoft 365 Copilot, GitHub Copilot, Security Copilot, Copilot Studio, Agent 365, Dynamics, Power Platform, Azure AI Foundry. No competitor matches that span from developer tooling through security to the knowledge worker.</p><p>Breadth stopped sorting the field sometime last year. Google has Gemini Enterprise, Workspace, the Gemini Enterprise Agent Platform that replaced Vertex branding at Cloud Next, Code Assist, and a security line that now includes Gemini 3.5 Flash Cyber. Andy Jassy has been describing AWS share gains in terms of a top to bottom stack since 2025, and Quick Suite put Amazon directly into the knowledge worker layer. Every hyperscaler claims full coverage now. What separates Microsoft is the installed base and the directory, which is a distribution argument.</p><h4 id="what-this-means-for-a-copilot-evaluation">What this means for a Copilot evaluation</h4><p>For anyone still running a Copilot pilot, the allocation size is the thing to reconsider. Testing 500 seats against a time-savings baseline measures the configuration least likely to produce a return, because the context graph underneath it is thin and the agents built on it stay generic. An evaluation that produces a usable answer covers a full function, targets one workflow that crosses systems, and measures completion rather than minutes saved.</p><p>Fiscal year end matters to the negotiation. The cohort above 50,000 seats grew more than sevenfold this year, which tells you volume pricing exists at thresholds Microsoft has an incentive to reach. Buyers sitting between 5,000 and 50,000 seats have more room than list price suggests, and the leverage is largest in the quarter Microsoft is trying to close.</p><h4 id="microsoft-is-competing-with-its-own-supplier">Microsoft is competing with its own supplier</h4><p>The most revealing line in the Copilot post sits in footnote 2. Microsoft ran 125 test runs across 12 prompts comparing Copilot Cowork against Claude Cowork with the Microsoft 365 connector, both running Opus 4.8, and reports its own product came in 30 to 40 percent cheaper. A post celebrating 30 million paid seats does not normally carry a competitor cost benchmark in its footnotes.</p><p>The threat that footnote names is a model provider reaching Microsoft’s own surfaces through a connector, without Microsoft’s seat and without Microsoft’s margin. Copilot Cowork runs on agentic technology Microsoft brought in from Anthropic, and Microsoft booked a $3.2 billion gain on its Anthropic investment in the same quarter it published the comparison.</p><p>Nadella made the architecture explicit on the same call. Answering UBS analyst Karl Keirstead, he said the platform design requires you to “keep your harness separate from the model,” with memory and context held externally so that any model is swappable at any time. He described models as an input to the enterprise, and the goal as a firm that controls its own human capital and token capital.</p><p>That is the grounding argument stated as vendor strategy, and it names the layer Microsoft intends to own. It also explains the benchmarking. Footnote 2 puts Copilot Cowork against Claude Cowork on cost. On the same call Nadella claimed MAI-Cyber-1-Flash outperforms a much larger Mythos model at half the cost when paired with Microsoft’s multi-agent security harness. Two comparisons against Anthropic in the same week, from a company holding a stake in it.</p><p>Cowork’s multi-model design lets a task pull whichever model fits, which is Microsoft declaring models a commodity input it intends to arbitrage. The position holds while the harness stays proprietary and the context stays inside Microsoft’s directory. Thirty million seats are what keep it there.</p>]]></content:encoded>
            </item>
            <item>
                <title><![CDATA[The Behind-the-Meter AI Buildout]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/the-behind-the-meter-ai-buildout/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/the-behind-the-meter-ai-buildout/</guid>
                <dc:creator><![CDATA[Ben Bajarin]]></dc:creator>
                <pubDate>Thu, 30 Jul 2026 20:00:48 +0000</pubDate>
                
                <description><![CDATA[In our prior report, The Hyperscaler Capacity Partner Hierarchy, we talked about the importance of partners for hyperscalers, particularly those with a specific type of business model that allowed them to keep a preferred margin profile when they could not secure infrastructure fast enough to conver]]></description>
                <content:encoded><![CDATA[<p>In our prior report, <a href="https://www.thediligencestack.com/p/the-hyperscaler-capacity-partner?ref=ghost-development-2724.up.railway.app">The Hyperscaler Capacity Partner Hierarchy</a>, we talked about the importance of partners for hyperscalers, particularly those with a specific type of business model that allowed them to keep a preferred margin profile when they could not secure infrastructure fast enough to convert their backlog. We think, as the market adopts behind-the-meter power and standards emerge, this can accelerate time to power and perhaps shift the speed at which hyperscalers can get land, shells, and other physical infrastructure in place (all things easier to secure and build), then work with a partner like Bloom, and others, to bring capacity online while they still work to get broader grid connectivity.</p><p>Grid connectivity is the hardest part of this equation. The hyperscalers that needed to monetize their compute were doing deals with those who had secured power via the grid, often on less favorable terms that impacted their margins, because they were desperate for power. As BTM scales and becomes more viable, it gives them another choice since it can operate without going through all the hoops needed to secure a grid connection. We are not saying they will stop pursuing grid connections. By adopting BTM, they can scale the easier part of the project, if anything in this process is actually easy: getting a shell built and all the surrounding infrastructure in place. That gives them first-party ownership instead of forcing them to go through partners simply because those partners have grid contracts done and ready to go. The maturation of BTM means the option now exists to add the grid later. That is why BTM can cost more and still be the better choice.</p><h3 id="inference-changes-the-equation-for-btm">Inference changes the equation for BTM</h3><p>From conversations a year ago, when BTM seemed more like a theoretical with potential but needed to be proved out, a few things have changed. First, the AI infrastructure mix moving from training toward more inference is a big driver. In training, the workloads are bursty and thus prone to massive power spikes versus more nominal and consistent power draw. Inference is less bursty and more manageable as a workload, meaning that having a range of redundancies in place to handle power spikes, which was a challenge for pure BTM, is less of an issue with inference. But even with training workloads now, the latest GPUs and rack-scale systems are putting more power control into their systems to help regulate power spikes and make them smoother. One data center operator we spoke with told us anything around 25% or lower for this kind of power spike was low enough that BTM can suffice, and that the latest rack-scale systems can keep those spikes at 25% or lower.</p><p>Public data supports the view from our conversations with those in the power industry. <a href="https://ghost-development-2724.up.railway.app/content/files/www-microsoft-com/en-us/research/wp-content/uploads/2024/03/gpu_power_asplos_24.pdf">Microsoft production data</a> measured a maximum two-second change of 37.5% of provisioned power for training, compared with 9% for interactive inference. <a href="https://docs.nvidia.com/multi-node-nvlink-systems/multi-node-tuning-guide/power-thermals.html?ref=ghost-development-2724.up.railway.app">NVIDIA’s GB200 documentation</a> describes programmable power smoothing, while its newer <a href="https://developer.nvidia.com/blog/inside-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=ghost-development-2724.up.railway.app">Vera Rubin rack architecture</a> adds more local energy buffering. We would not treat 25% as a universal engineering cutoff. It is one operator’s practical marker, and the broader point is that better rack-level power management gives the onsite generation system a smoother load to follow. Smoother workloads are one enabler of the BTM shift. Grid delays and the economics of bringing compute online earlier remain the larger forces.</p><h3 id="customers-have-moved-from-%E2%80%9Ccan-we-use-btm%E2%80%9D-to-%E2%80%9Chow-will-we-use-it%E2%80%9D">Customers have moved from “can we use BTM?” to “how will we use it?”</h3><p>We are not saying BTM becomes the standard or only source of power, only that BTM was not in a place where it was as viable as it appears now. That waas clear from Bloom’s call where they disclosed every major customer is now engaged to use BTM as a way to get to revenue faster. Customers have moved from BTM looks good on paper to all systems go to ramp BTM. We are in a cycle where the supply chain needs to scale and start to ramp to meet demand so customers can use BTM either to bridge the time to grid service or offset some grid needs and work toward more favorable power economics by using a hybrid architecture. BTM seems to be inflecting, for all the reasons we include, and now the main players need to start ramping their supply chains to meet demand.</p><p>From talking to operators in the field, we had consistently heard about the challenge that is electrical integration for the full site. While the equipment that goes into a BTM solution is one piece of the puzzle, the customer still needs to secure labor and testing, then bring together switchgear, transformers, UPS capacity, microgrid controls and a host of other things. The best-positioned companies are those that have already secured the fuel, have utility relationships and have sites ready to go.</p><p>Power is only part of the equation when analyzing who captures most of the value. This is why the ability to control the complete delivery schedule and stand behind the contract matters, because any delay has a direct impact on the customer’s time to revenue. The guarantee must cover an operating outcome. That shifts the risk of turning a power plan into operating compute away from the hyperscaler and onto the power or capacity partner.</p><h3 id="bloom-is-the-clearest-public-proof-point-for-now">Bloom is the clearest public proof point for now</h3><p>Bloom went from essentially entering this market nine months ago, per the call, with one customer, to management saying its technology is now validated and approved by all major U.S. hyperscalers and more than a dozen other AI infrastructure customers. Those relationships span live deployments, booked and shipped systems under construction, and definitive agreements. Customer engagement has clearly inflected. The next proof is whether those approvals become repeat deployments and accepted, revenue-producing critical IT MW.</p><p>The reported results now sit behind that claim. Revenue reached $1.065 billion in Q2, up 165.5% from a year earlier, while GAAP gross margin increased to 33.4%. The next proof is repeat orders and accepted, revenue-producing critical IT MW.</p><p>A handful of large projects can make the market appear broader than it is. Developers may reserve several generation options for the same campus, while different suppliers count the same prospective demand. Repeat deployments provide stronger evidence. They show that the first installation worked well enough for the customer to use the architecture again. Accepted, revenue-producing critical IT MW then confirms that the project cleared every gate: fuel, permits, generation, electrical integration, commissioning, and usable compute. That would establish BTM as a repeatable procurement model rather than a collection of emergency power projects.</p><h2 id="inside-the-full-report">Inside the Full Report</h2><ul><li>How BTM gives hyperscalers a faster path to owned compute</li><li>The six gates between a power plan and usable capacity</li><li>Where Bloom leads and how engines and turbines compete</li><li>What Bloom’s Q2 confirms and what still needs proof</li><li>Who captures value across the BTM delivery stack</li><li>How BTM can scale and what would weaken the thesis</li></ul><figure class="kg-card kg-embed-card"><iframe class="no-lazyload" data-no-lazy="true" style="width: 100%; height: 400px; border: 1px solid #EEE; background: white; overflow: hidden;" src="https://www.thediligencestack.com/embed"><!--kg-card-begin: html--><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><!--kg-card-end: html--></iframe></figure>]]></content:encoded>
            </item>
            <item>
                <title><![CDATA[The Hyperscaler Capacity Partner Hierarchy]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/the-hyperscaler-capacity-partner-hierarchy/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/the-hyperscaler-capacity-partner-hierarchy/</guid>
                <dc:creator><![CDATA[Ben Bajarin]]></dc:creator>
                <pubDate>Tue, 28 Jul 2026 20:00:03 +0000</pubDate>
                
                <description><![CDATA[Inside the Full Report For Clients and Subscribers This report is the third installment in our work on hyperscalers, neoclouds, and the economics of external AI capacity. We have spent a lot of time tracking power and utilities because we continue to believe power will set the pace of the AI infrast]]></description>
                <content:encoded><![CDATA[<h2 id="inside-the-full-report-for-clients-and-subscribers">Inside the Full Report For Clients and Subscribers</h2><p><em>This report is the third installment in our work on hyperscalers, neoclouds, and the economics of external AI capacity.</em></p><p>We have spent a lot of time tracking power and utilities because we continue to believe power will set the pace of the AI infrastructure buildout. Our recent conversations with industry sources keep coming back to the same problem. We have known about the power bottleneck for some time. What we are hearing now is that it has not gotten any better and, if anything, is getting worse. Data center demand continues to move faster than local utilities can support it. Grid studies can take six to twelve months, and those timelines are getting longer. A new substation can add another two to three years. A company can sometimes move faster if it pays for the substation work itself, but that makes the project more expensive. Even then, the equipment and skilled labor needed to finish the work are getting harder to secure on schedule. Permitting and local resistance can push the timeline out again. Based on what we are hearing, we think the power constraint is likely to stay with the industry through at least 2030 and probably longer.</p><p>As we have heard directly from hyperscalers, this timing problem has been shaping how they think about capacity for some time. Demand is coming in faster than utilities and the normal data center build cycle can support. That leaves them using outside partners that already have firm power or can bring a site online sooner.</p><p>Calling a company an power/shell/ landlord, neocloud or data center developer only tells us so much. What we care about is whether the site can actually get power and be delivered on time. Where we may differ from many consensus is our belief that the best fit is often a partner that can bring the power and the building online, then stop there. The hyperscaler still controls the compute and the customer, which keeps more of the economics inside its own business.</p><h3 id="amazon-and-microsoft-will-put-the-capacity-gap-back-in-focus">Amazon and Microsoft will put the capacity gap back in focus</h3><p>This report will publish during the same week <a href="https://news.microsoft.com/source/2026/07/08/microsoft-announces-quarterly-earnings-release-date-68/?ref=ghost-development-2724.up.railway.app">Microsoft reports fiscal Q4 results on July 29</a> and <a href="https://ir.aboutamazon.com/events/event-details/default.aspx?ref=ghost-development-2724.up.railway.app">Amazon reports Q2 results on July 30</a>. We expect both companies to raise the capital-spending bar and show another increase in contracted cloud demand. The labels differ between backlog and remaining performance obligations, but the economic signal is the same: Google (reported), AWS, and Azure are booking demand faster than internal infrastructure can be delivered.</p><p>The chart below puts the imbalance on one scale. Across the companies shown (full universe including meta and neoclouds), backlog and RPO have grown faster than cost-adjusted capex. That pulled the ratio from 44% in 2024 to about 37% in 2026. We would not treat this as a measure of physical capacity coverage because capex is an annual flow and backlog is a point-in-time balance. Still, the direction is clear. Contracted demand has grown faster than the capital response, which is why we do not expect capex to slow anytime soon. Key point to remember – revenue always lags capex.</p><p>That leaves the economics of the partner as the main question. Hyperscalers will keep using outside capacity while they build more of their own. On Alphabet’s latest earnings call, management called that capacity a bridge and said it would create modest near-term margin pressure. Google is willing to pay the premium because the revenue is available now. Waiting for its own data centers would mean leaving some of that demand unserved. As Google brings more controlled capacity online, part of the premium should go away. The type of partner, and how much of the stack that partner owns, will determine how expensive the bridge becomes. <a href="https://www.investing.com/news/transcripts/earnings-call-transcript-alphabet-beats-q2-2026-estimates-shares-fall-on-capex-surge-93CH-4807140?ref=ghost-development-2724.up.railway.app">Alphabet Q2 2026 call transcript</a></p><p>Our prior reports looked at the AI cloud from the supplier side. We mapped where hyperscalers and neoclouds compete, then separated neoclouds by the layers they own.</p><p>This report looks at the same issue from the buyer’s side. We think AWS and Azure face the same choice Google described. They can, and will, buy outside capacity now and accept some pressure on unit economics, then move more of the workload into controlled infrastructure as their own supply arrives. The question is who is best suited to the hyperscalers are a foundational capacity partner.</p><p>That is why we are spending more time on the partner type. A company that secures firm power and delivers the physical layer can remain an attractive partner after the scarcity premium fades. The hyperscaler still controls the server system and the customer. From here, we want to know which providers can deliver the same kind of project more than once.</p><h3 id="external-capacity-is-not-one-product">External capacity is not one product</h3><p>A complete compute service and a leased data center shell can both deliver capacity, but the buyer is paying for very different things. A managed compute provider supplies the building and the server system, then operates it. The price has to cover the cost of the hardware and the risk that the provider cannot keep it fully used. It also has to cover the chance that the equipment loses value faster than expected. All of that adds another supplier margin to the cost.</p><p>A third-party-owned shell stops at the physical layer. The landlord develops the site and delivers a building that can support the required density. The hyperscaler can still own the accelerators and decide how the network is designed. It also keeps the workload inside its own control plane, which is what the hyperscalers generally prefer. More of the cloud margin stays with the buyer, while someone else owns the slow, long-lived physical asset.</p><p>This is the capacity hierarchy we are trying to describe. When there is enough time, hyperscalers prefer to own and operate strategic capacity themselves. A leased shell can get them there sooner without giving up control of the compute. Managed third-party compute still makes sense when the value of serving demand now is greater than the premium being paid.</p><p><em>Exhibit 1. The Hyperscaler Capacity Hierarchy. Relative profiles are Creative Strategies analytical judgments, not reported provider margins. Source: Creative Strategies analysis of Alphabet public disclosures and public infrastructure agreements.</em></p><h3 id="the-best-landlord-owns-more-than-land">The best landlord owns more than land</h3><p>Power delivery is the known hook making these partners attractive. The latest disclosures help answer the next question: what turns that power into a valuable contract? Applied Digital delivered another 75 MW on schedule, while Core Scientific had 437 MW billing by mid-July against 1.1 GW leased. More important for our thesis, Applied Digital said direct hyperscaler leases should lower its financing cost over the full term. Core Scientific’s mix of direct AMD leases and neocloud leases with AMD protections shows how the risk can change by project. The value sits in converting power into accepted capacity under a contract the market can finance with a specific margin profile. <a href="https://ir.applieddigital.com/news-events/ir-calendar/detail/20260727-q4-2026-earnings-call?ref=ghost-development-2724.up.railway.app">Applied Digital Q4 2026 earnings call</a> <a href="https://investors.corescientific.com/news-events/press-releases/detail/139/core-scientific-announces-second-quarter-2026-results?ref=ghost-development-2724.up.railway.app">Core Scientific Q2 2026 results</a> <a href="https://investors.corescientific.com/news-events/press-releases/detail/138/core-scientific-and-amd-announce-infrastructure-partnership?ref=ghost-development-2724.up.railway.app">Core Scientific and AMD partnership</a></p><p>Hut 8 may add another version. The company has leased 704 MW at Beacon Point to an unnamed investment-grade tenant. The Financial Times reported that NVIDIA is the tenant and may sublease the capacity to neocloud partners, although neither company has confirmed it. If accurate, NVIDIA would be securing physical capacity itself rather than waiting for partners to find it. <a href="https://www.hut8.com/news-insights/press-releases/hut-8-fully-commercializes-1-gw-beacon-point-ai-data-center-campus-with-second-352-mw-it-lease?ref=ghost-development-2724.up.railway.app">Hut 8 Beacon Point announcement</a> <a href="https://www.brecorder.com/news/40432141/nvidia-behind-50-billion-lease-on-texas-data-center-ft-reports?ref=ghost-development-2724.up.railway.app">Reuters summary of the FT reporting</a></p><p>Chip vendors may be starting to secure physical capacity themselves. AMD is doing it directly with Core Scientific, while NVIDIA is reported to be leasing Hut 8 capacity that it may place with neocloud partners. The obvious reason is speed, but we think there may be a competitive angle as well. Hyperscalers are putting more custom ASICs into the infrastructure they control. Securing outside sites could help NVIDIA keep scarce power tied to NVIDIA systems before those sites are absorbed into hyperscaler builds. That is our interpretation, not something NVIDIA or Hut 8 has said, but it would expand the buyer pool for the physical layer.</p><p><strong>For now, our perhaps out of consensus view, is we still think the best long-term position may sit with the company that controls the physical asset and stops before the compute layer.</strong> Hyperscalers will use neoclouds when speed is worth the premium. Chip vendors may now compete for the same sites. All of them still need the landlord to deliver.</p><ul><li>An original model for comparing the cost of delay with the premium paid for managed compute.</li><li>A buyer-side margin and capital sensitivity across owned, leased-shell, and managed-compute structures.</li><li>Google and Meta case studies showing how the same buyer uses different structures for different time horizons.</li><li>A contract-quality ladder separating direct hyperscaler leases from backstopped neocloud tenancy and uncontracted pipeline.</li><li>A selected capacity map showing where supply sits, followed by public-company market, beneficiary, and partner-fit maps that separate economic role from announced MW.</li><li>A deployable-MW proof ladder and monitoring framework for testing whether announced capacity can become revenue-producing infrastructure.</li></ul><figure class="kg-card kg-embed-card"><iframe class="no-lazyload" data-no-lazy="true" style="width: 100%; height: 400px; border: 1px solid #EEE; background: white; overflow: hidden;" src="https://www.thediligencestack.com/embed"><!--kg-card-begin: html--><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><!--kg-card-end: html--></iframe></figure>]]></content:encoded>
            </item>
            <item>
                <title><![CDATA[In 2023-2024 We Weren’t Bullish Enough]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/in-2023-2024-we-werent-bullish-enough/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/in-2023-2024-we-werent-bullish-enough/</guid>
                <dc:creator><![CDATA[Ben Bajarin]]></dc:creator>
                <pubDate>Fri, 24 Jul 2026 20:00:58 +0000</pubDate>
                <category><![CDATA[AI]]></category>
                <description><![CDATA[We thought it would be useful to look back at a wide range of research from 2023 and 2024, just as it became clear that AI was about to change everything, and ask where those early forecasts ended up being right, wrong, or simply too conservative. This retrospective draws on a wide range of research]]></description>
                <content:encoded><![CDATA[<p>We thought it would be useful to look back at a wide range of research from 2023 and 2024, just as it became clear that AI was about to change everything, and ask where those early forecasts ended up being right, wrong, or simply too conservative.</p><p>This retrospective draws on a wide range of research from third party sources as well our own. Much of that work saw AI coming. The largest misses came from underestimating how quickly the pieces around the accelerator would have to scale together. Something we now refer to as the GPU Tsunami.</p><h3 id="how-much-the-market-moved">How much the market moved</h3><ul><li><strong>Four years early:</strong> A 2022 forecast put the semiconductor market at <strong>$1 trillion in 2030</strong> (<a href="https://www.mckinsey.com/industries/semiconductors/our-insights/the-semiconductor-decade-a-trillion-dollar-industry?ref=ghost-development-2724.up.railway.app">historical framework</a>). Our current model reaches <strong>$1.79 trillion in 2026</strong> and <strong>$3.49 trillion in 2030</strong>. Adding adjacent AI infrastructure brings the combined 2030 silicon-systems pool to <strong>$3.95 trillion</strong>.</li><li><strong>5.8 times:</strong> Our 2027 accelerator model is now <strong>$720 billion</strong>, compared with an old <strong>$125 billion</strong> AI-compute endpoint.</li><li><strong>60%:</strong> One quarter of NVIDIA Data Center revenue reached about 60% of the old full-year 2027 AI-compute endpoint.</li><li><strong>2.7 to 3.2 times:</strong> The 2025 CoWoS capacity estimate was revised from 25,000 to 30,000 wafers per month to more than 80,000.</li><li><strong>79%:</strong> The 2027 AI-switching endpoint was revised higher in less than a year.</li><li><strong>3.0 to 3.6 times:</strong> Current frontier rack power surpassed an old 2030 expectation four years early.</li><li><strong>2.4 times:</strong> Our seven-company 2027 capex model is already 2.4 times an old global data-center endpoint.</li></ul><p>We think its important to look back, as a post mortem, to understand where the analysis went wrong and what the collective industry (us analysts) underestimated, so we can recognize those patterns when they appear again. Some earlier calls were right, while others moved too quickly on adoption timing. We also missed how much the unit of analysis was changing as the market developed. We were particularly too conservative on the ASP expansion across the semiconductor supply chain. We fully modeled semiconductor output based on our foundry capacity model, but we did not fully appreciate the pricing leverage created by the AI infrastructure cycle. Our midpoint scenario implies the industry’s blended semiconductor ASP rises from approximately $0.75 today to more than $2.10 by 2028, nearly tripling in just three years. <strong>Between 2025 and 2028, we estimate semiconductor unit shipments increase only 25–30%, while the average semiconductor ASP increases nearly 185%, driving approximately 260% industry revenue growth</strong>. This represents a structural shift in semiconductor industry economics as value per device compounds substantially faster than industry unit shipments. We map every constraint in the below report leading to this ASP increase, in the era we now call the era of margin expansion.</p><h3 id="what-the-archive-saw-and-where-its-boundary-failed">What the archive saw, and where its boundary failed</h3><p>We reviewed a broad set of historical forecasts to understand how expectations changed over time. Licensed historical work is described as third-party forecasts or archive estimates, while current comparisons use public outcomes and Creative Strategies models.</p><p>The scorecard separates outcomes, run rates, and model revisions. A newer forecast shows expectations moving; only an outcome proves the old forecast wrong.</p><p>Each row uses a different type of evidence, but the pattern is consistent. The largest forecast revisions appeared once the market moved beyond the accelerator and began modeling the complete AI system. Compute only becomes usable capacity when the full rack can be commissioned. The early forecasts treated the infrastructure around the accelerator as a supporting input. It turned out to be part of the market itself.</p><h2 id="compute-demand-escaped-the-original-boundary">Compute demand escaped the original boundary</h2><p>In March 2023, a third-party forecast assumed that 40 to 60 large-model builds over the following 12 to 24 months would create approximately $10 billion to $15 billion of incremental GPU TAM. A June 2023 custom-silicon model put AI computing semiconductors at roughly $43 billion in 2023, growing to $125 billion in 2027, with custom ASICs reaching as much as 30% of cloud AI semiconductor spending.</p><p>Those estimates were aggressive at the time, but their boundary was still too conservative. The model counted likely builds and the accelerators required, then extended the curve as custom silicon gained share. It did not capture the wider inference market that would form around installed compute. Ultimately, what most got wrong was underestimating the compute intensity of the workload that is agentic AI. We think many still under appreciate this today as well.</p><p>NVIDIA reported $75.2 billion of Data Center revenue in Q1 FY2027. Data Center includes systems and networking, so it is not like-for-like with semiconductor TAM. Even so, that quarter reached 60% of the old 2027 AI-compute endpoint. In one quarter, NVIDIA generated revenue equal to 60% of what the earlier model expected from the entire AI-compute market for all of 2027. Looking forward, our current model puts merchant GPU and XPU logic plus custom AI accelerators at $770 billion in 2027 and $1.36 trillion in 2030.</p><p>The models underestimated how much demand each new accelerator would create. Training built the installed base, while better performance expanded the market for inference. That pulled custom silicon and rack-scale systems into what had started as a GPU market. As supply grew, the TAM grew with it.</p><h2 id="packaging-was-the-clearest-under-call">Packaging was the clearest under call</h2><p>In August 2023, a third-party advanced-packaging forecast put industry CoWoS capacity at roughly 15,000 wafers per month, rising to 20,000 to 25,000 in the second half of 2024 and 25,000 to 30,000 in 2025. The forecast saw little likelihood that CoWoS would remain a meaningful bottleneck beyond 2024.</p><p>By January 2025, a later archive model expected TSMC CoWoS capacity to exceed 80,000 wafers per month by the fourth quarter of 2025. <strong>That was 2.7 to 3.2 times the 2023 estimate for the same year. </strong>In July 2026, TSMC still described packaging as tight enough to limit customer growth. <a href="https://investor.tsmc.com/english/quarterly-results/2026/q2?ref=ghost-development-2724.up.railway.app">TSMC Q2 2026 materials</a>.</p><p>In retrospect, what was missed was the timing, pull-in, of chiplet designs as AI accelerators were the first evidence of the broad shift from monolithic to systems based chip design.</p><p>Packaging became part of the capacity boundary, which is the central idea in our report <a href="https://www.thediligencestack.com/p/the-gpu-tsunami-tsmc-intel-and-samsung?ref=ghost-development-2724.up.railway.app">The GPU Tsunami: TSMC, Intel, and Samsung Foundry</a>. A leading-edge wafer becomes useful AI capacity only after advanced packaging brings it together with HBM and the rest of the rack.</p><h2 id="memory-became-a-much-larger-market-than-the-models-allowed-and-structural">Memory became a much larger market than the models allowed (and structural)</h2><p>Memory was not missing from the early forecasts. It was treated as a cyclical recovery. In late 2023, the public market forecast called for roughly $130 billion of memory revenue in 2024. Actual sales reached $165 billion, 27% above that estimate.</p><p>The longer-range forecasts were even more conservative. One public model published in 2024 put combined DRAM and NAND revenue at approximately $206 billion in 2025 and just $214 billion in both 2026 and 2027. The assumption was that the recovery would level off once pricing normalized.</p><p>Our current model looks very different. We estimate combined DRAM and NAND revenue of $230–240 billion in 2025, $550–570 billion in 2026, and $800–850 billion in 2027. Against the old forecast, the 2026 market is now modeled at 2.6 times the prior estimate. By 2027, the difference reaches approximately four times.</p><p>This is a forecast revision, not yet a realized outcome, yet trending in that direction. But it shows what the earlier models missed. AI did not create demand only for HBM. It increased the amount of server DRAM and enterprise NAND required around each accelerator. At the same time, HBM consumed capacity that would otherwise have served conventional memory. The result was a much broader shortage and a much larger market than the original forecasts allowed.</p><h2 id="the-bottleneck-kept-moving">The bottleneck kept moving</h2><p>The early work saw the constraint moving beyond the accelerator, but underestimated how quickly the network and site would control the deployment schedule.</p><p>A June 2023 forecast put 2027 AI-switching revenue at $8.5 billion. Eleven months later, that endpoint had risen 79% to $15.2 billion. Our broader networking-silicon model now reaches $105 billion in 2027. The categories differ, but reported revenue confirms the direction: Broadcom’s $10.8 billion of quarterly AI semiconductor revenue reached 77% of its old full-year FY2025 forecast. <a href="https://investors.broadcom.com/news-releases/news-release-details/broadcom-inc-announces-second-quarter-fiscal-year-2026-financial?ref=ghost-development-2724.up.railway.app">Broadcom Q2 FY2026 results</a>.</p><p>Power moved faster still. A July 2024 note expected average rack density to reach 40 kilowatts by 2030. GB200 and GB300 systems reached 120 to 142 kilowatts four years early, or 3.0 to 3.6 times the old endpoint. These are frontier systems rather than fleet averages, but they set the next building standard. <a href="https://docs.nvidia.com/mission-control/docs/systems-administration-guide/2.1.0/prs/faq.html?ref=ghost-development-2724.up.railway.app">GB200 specification</a> and <a href="https://docs.nvidia.com/enterprise-reference-architectures/nvl72-ai-factory/latest/components.html?ref=ghost-development-2724.up.railway.app">GB300 reference architecture</a>.</p><p>Capex followed. A May 2024 model put global data-center capex at $500 billion in 2027. Our seven-company cost-adjusted base now reaches $1.18 trillion, or 2.4 times that endpoint.</p><p>The bottleneck moved from the chip into the network and then the site. Each fix exposed the next constraint. That is why <a href="https://www.thediligencestack.com/p/counting-real-ai-capacity?ref=ghost-development-2724.up.railway.app">Counting Real AI Capacity</a> and <a href="https://www.thediligencestack.com/p/gigawattonomics?ref=ghost-development-2724.up.railway.app">Gigawattonomics</a> focus on commissioned systems and economic output per watt. Until the full system is commissioned, announced capacity remains intent.</p><h2 id="the-top-line-hid-the-infrastructure-reallocation">The top line hid the infrastructure reallocation</h2><p>A January 2024 semiconductor model forecast sales of $645 billion in 2024 and $718 billion in 2025. WSTS reported $630.5 billion and $795.6 billion, respectively. The model was 2.3% high for 2024 and 10.8% low for 2025. <a href="https://www.wsts.org/76/103/Global-Semiconductor-Market-grows-26-in-2025-to-796B?ref=ghost-development-2724.up.railway.app">WSTS 2025 result</a>. It caught the recovery but missed the next acceleration in memory pricing and AI-infrastructure mix.</p><p>Public cloud was also broadly on track. A June 2024 model forecast spending rising from roughly $675 billion in 2024 to $1.38 trillion in 2028. The surprise occurred inside the total as AI infrastructure grew faster and carried more capital intensity than conventional workloads.</p><p>A top-line forecast can land within normal error while the internal demand map changes enough to redirect capital, capacity, and profit. The semiconductor and cloud totals did not reveal how much of the next dollar would be pulled toward accelerators and the infrastructure required to deploy them.</p><h2 id="enterprise-ai-had-two-opposite-forecast-errors">Enterprise AI had two opposite forecast errors</h2><p>Infrastructure estimates were generally too small. Enterprise adoption produced two errors in opposite directions. Spending and usage intensity ran above the early markers, while scaled-production timing was sometimes too aggressive.</p><p>Two 2023 third-party forecasts framed an $820 billion enterprise-software TAM and approximately $150 billion of GenAI software spending within three years.</p><p>Our current working range puts enterprise GenAI spending across software, services, and inference compute at $175 billion to $200 billion in 2026, rising to more than $600 billion by 2030. The boundary is broader than software alone, so the comparison is directional. Even with that caveat, the 2026 midpoint is about 25% above the old three-year marker, and the decade-end range is more than four times as large.</p><p>The survey archive explains the other error. A 2024 IT survey found 52% of respondents live with at least one AI use case, but only 10% in production at scale. Later CIO surveys measured production at 30% in 2024 and 39% in 2025, while an earlier CTO survey expected nearly universal use by year-end 2024.</p><p>The error was treating a company with one live use case as equivalent to a company that had reorganized a production workflow. An enterprise can pay for AI and increase token consumption while only a small number of workflows reach scaled production. In our report, <a href="https://www.thediligencestack.com/p/from-ai-usage-to-ai-earnings-power?ref=ghost-development-2724.up.railway.app">From AI Usage to AI Earnings Power</a> picks up there. The test is whether repeated use changes a measurable operating baseline enough to earn a durable budget.</p><h2 id="where-the-current-diligence-stack-work-goes-next">Where the current Diligence Stack work goes next</h2><p>The current research starts from this revised unit of analysis. <a href="https://www.thediligencestack.com/p/counting-real-ai-capacity?ref=ghost-development-2724.up.railway.app">Counting Real AI Capacity</a> establishes what has actually been commissioned. <a href="https://www.thediligencestack.com/p/gigawattonomics?ref=ghost-development-2724.up.railway.app">Gigawattonomics</a> then asks whether that capacity can earn an acceptable return from the power it consumes.</p><p>Below the site, <a href="https://www.thediligencestack.com/p/the-gpu-tsunami-tsmc-intel-and-samsung?ref=ghost-development-2724.up.railway.app">The GPU Tsunami</a> applies the same capacity discipline to foundries. <a href="https://www.thediligencestack.com/p/where-ai-constraints-become-pricing?ref=ghost-development-2724.up.railway.app">Where AI Constraints Become Pricing Leverage</a> asks which shortages can support durable economics rather than temporary scarcity pricing. The memory and enterprise reports carry the method into capacity allocation and workflow monetization.</p><h2 id="bottom-line">Bottom line</h2><p>As a firm that has been doing market models, forecasting, sizing, and more for over 40 years, we understand being conservative, and how forecasts are almost always wrong.  A forecast is only as good as the underlying assumptions we stay in a state of constant learning, observing, and looking for the right past and present patterns in order to continually strengthen our assumptions on every part of the industry we research.</p><p>We shared a lot of our internal forecasts and models in the look back, knowing full well they may be wrong and also pose the scenario, within the historical view, that what if even today we are not bullish enough.</p><figure class="kg-card kg-embed-card"><iframe class="no-lazyload" data-no-lazy="true" style="width: 100%; height: 400px; border: 1px solid #EEE; background: white; overflow: hidden;" src="https://www.thediligencestack.com/embed"><!--kg-card-begin: html--><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><!--kg-card-end: html--></iframe></figure>]]></content:encoded>
            </item>
            <item>
                <title><![CDATA[Cybersecurity and the Enterprise AI Control Layer]]></title>
                <link>https://ghost-development-2724.up.railway.app/research/cybersecurity-and-the-enterprise-ai-control-layer/</link>
                <guid isPermaLink="true">https://ghost-development-2724.up.railway.app/research/cybersecurity-and-the-enterprise-ai-control-layer/</guid>
                <dc:creator><![CDATA[Ben Bajarin]]></dc:creator>
                <pubDate>Thu, 23 Jul 2026 20:00:23 +0000</pubDate>
                <category><![CDATA[AI]]></category>
                <description><![CDATA[We have been thinking about cybersecurity as a much deeper enterprise function than the word security tends to imply. The same framework can be extended to sovereign nations. Closed and open models, including models developed outside a company’s or country’s borders, are becoming capable enough that]]></description>
                <content:encoded><![CDATA[<p>We have been thinking about cybersecurity as a much deeper enterprise function than the word security tends to imply. The same framework can be extended to sovereign nations. Closed and open models, including models developed outside a company’s or country’s borders, are becoming capable enough that persistent threats will be a fact of operating life. Those threats will reach a company through its products and systems. They will also reach the company’s data and, increasingly, information tied to its employees. The question is how enterprises and governments defend themselves when the tools available to attackers keep improving.</p><p>Over the last few months, we have spent considerable time speaking with enterprise decision-makers about agentic deployments, including through our own CIO/CTO survey and conversations with large corporate customers. One concern kept surfacing: security, and cybersecurity in particular. As capable open models continue to advance, the cost and technical barrier for bad actors will come down with them, increasing the threat surface corporations have to defend. For that reason, we think security will be one of the earliest and most durable sources of pull-through from enterprise AI adoption. Security attaches to an AI project as it moves into production, then returns through the existing cyber budget at renewal.</p><p>While the early evidence indicates spending is unlikely to arrive as one neat new AI-security category. Instead, AI increases the burden on security controls companies already have in place. In conversations on this with stakeholders one of the first challenges we hear brought up is data access. Before an enterprise can put a model into production, it needs to know what information the model can retrieve and whether that information should be available to the person making the request. Identity becomes more important when an agent begins to act, and runtime controls enter the picture once those actions reach production systems. Some of this need will create new products, but a meaningful share may appear as deeper use of platforms customers already own. That makes the demand easier to see than the eventual revenue pool. The need for greater control can become obvious well before investors can measure where the value is accruing. We also have outlined the need for a new class of compute, to go with a new class of models, or model variants specific to security, cyber security.</p><h3 id="open-models-raise-the-market-floor">Open models raise the market floor</h3><p>The external threat environment pushes that same budget in a secondary direction. Our companion research on <a href="https://www.thediligencestack.com/p/chinas-ai-arms-race-runs-through?ref=ghost-development-2724.up.railway.app">China’s AI stack</a> argues that China does not need semiconductor parity at every layer to keep model development moving. Adequate sovereign compute, paired with competitive open-weight models, is enough to widen access to capable AI. For cybersecurity, capability diffusion does not wait for chip parity. Any open source model, regardless of where it comes from, expands the pool of models available outside controlled services. That linkage has a clear limit: model access alone does not create a successful attacker.</p><p>Strategic competition gives nation-states a reason to keep investing in that capability. Enterprises absorb much of the operating cost because their systems are common targets. Models can package parts of reconnaissance or exploit adaptation into tools that are easier to use, which may raise attempted attack volume faster than a human-led defense process can scale. Defenders gain from the same models, but they still need a control layer that can operate at machine (agent) speed.</p><p><strong>Exhibit 1. AI Deployment Is Outrunning the Control Layer</strong></p><h3 id="the-budget-converts-through-existing-control-points">The budget converts through existing control points</h3><p>A direct survey in the source set points in the same direction: deployment is running WELL ahead of dedicated AI-security tooling. The gap is wider if the standard is the full control path around data access and agent behavior. Survey definitions differ, so we do not treat the size as a universal market estimate. The exact number is less insightful than the behavior it reveals. <strong>Enterprises are deploying AI before they can fully explain how it behaves and some of the unintended consequences of an unstructured deployment.</strong></p><p>Vendors that already own enterprise context or an enforcement point start with an advantage. A data-security platform can attach to a copilot rollout because it controls what the model can retrieve. Identity becomes relevant once agents receive permissions to act. Broader security platforms can spread AI-assisted workflows across an installed base, although the economics remain unclear until customers pay more or use more of the platform.</p><p>For stakeholders, a product announcement only shows that a vendor is participating. Continued use through renewal shows whether customers are willing to keep paying. The evidence so far suggests AI will expand cybersecurity spending because every production workload creates more systems and activity to secure. Much of the early spending is likely to flow through vendors already embedded in how companies protect their systems.</p><p>The rise in external threats may create a separate revenue opportunity at the model layer. As attackers gain access to more capable models, large enterprises and sovereign governments will need defensive models that can operate at the same speed. We think frontier labs such as OpenAI and Anthropic could build (may already be building) or license restricted models specifically for cyber defense. Some of these capabilities may be too sensitive for broad public access, giving the labs a direct path into enterprise and government security budgets.</p><h2 id="inside-the-full-report">Inside the Full Report</h2><ul><li>A two-pool analysis separating AI for Security from Security for AI, including the different maturity curves and budget sources.</li><li>The seven-layer Enterprise AI Control Layer map, showing where identity, data, runtime, and access controls sit around production AI.</li><li>A category conversion matrix distinguishing high-probability budget pull-through from capabilities likely to be bundled.</li><li>A vendor-positioning exhibit that maps established control-path platforms, specialists, emerging options, and exposed point tools.</li><li>A capability-diffusion framework connecting sovereign AI and Chinese open-weight releases to the enterprise security-spend ratchet, without assuming chip parity or proven attack causality.</li><li>A falsification and monitoring framework centered on paid attach, dedicated budgets, machine identity, platform consolidation, and realized SOC productivity.</li></ul><figure class="kg-card kg-embed-card"><iframe class="no-lazyload" data-no-lazy="true" style="width: 100%; height: 400px; border: 1px solid #EEE; background: white; overflow: hidden;" src="https://www.thediligencestack.com/embed"><!--kg-card-begin: html--><span data-mce-type="bookmark" style="display: inline-block; width: 0px; overflow: hidden; line-height: 0;" class="mce_SELRES_start">﻿</span><!--kg-card-end: html--></iframe></figure>]]></content:encoded>
            </item>
</channel>
</rss>
