Back to Daily Brief

Compute & Infrastructure

25 sources analyzed to give you today's brief

Top Line

Anthropic is building an in-house ASIC design team to co-develop custom AI inference chips with Samsung as manufacturing partner, a direct strategic move to reduce dependence on NVIDIA and lower long-run inference costs.

AMD acquired AI chip startup Taalas to deepen native inference capabilities across its silicon portfolio, signalling that the inference efficiency race is now driving M&A as well as internal R&D.

Amazon Web Services is rationing EC2 CPU capacity among its own engineers to prioritise external customers, a concrete signal that agentic AI workloads are creating compute scarcity even inside the hyperscaler building the infrastructure.

RAM prices have reverted to 2007 per-gigabyte levels in a matter of months, with AI-driven HBM and server DRAM demand effectively erasing two decades of consumer memory deflation.

JPMorgan has raised its 2026 tech bond issuance forecast to over $500 billion as hyperscalers accelerate AI capital deployment through debt markets, and data centre operator Switch has filed confidentially for an IPO, reflecting continued investor appetite for infrastructure exposure.

Key Developments

Anthropic Moves to Break NVIDIA Dependence with Custom Inference ASICs

Anthropic has announced it is assembling an in-house team to co-design custom ASIC accelerators specifically for inference workloads, with Samsung reported as the manufacturing partner. This is a significant structural shift: rather than simply buying more H100 or Blackwell capacity, Anthropic is pursuing a vertically integrated silicon strategy comparable to what Google achieved with TPUs and Amazon with Trainium and Inferentia. The inference layer is where Claude's operating costs are concentrated, so purpose-built silicon — optimised for the model's specific sparsity and attention patterns — can deliver step-change improvements in cost-per-token and throughput per watt. Tom's Hardware

The Samsung partnership is strategically notable. Samsung competes directly with TSMC in advanced logic and with SK Hynix in HBM, and it has been aggressively seeking anchor customers to validate its 2nm and 3nm nodes. For Anthropic, Samsung offers an alternative to TSMC's congested leading-edge capacity — though Samsung's yield track record at sub-3nm remains a risk factor that will require careful qualification. The move also adds Anthropic to a growing list of inference-focused custom silicon programmes alongside Meta's MTIA, Microsoft's Maia, and Amazon's Inferentia, collectively representing a structural demand shift away from merchant GPU suppliers.

Why it matters

Every major foundation model lab that builds custom inference silicon reduces its addressable spend at NVIDIA, collectively threatening the revenue concentration that currently supports NVIDIA's datacenter margins.

What to watch

Which Samsung node Anthropic targets for tape-out, and whether the programme remains fabless-style co-design or evolves toward a full custom chip team — the latter would be a multi-year, multi-billion dollar commitment.

AMD's Taalas Acquisition Embeds Inference Logic Natively Into Future Silicon

AMD has acquired AI chip startup Taalas, whose core technology focuses on inference optimisation at the microarchitecture level. Financial terms were not disclosed, nor was a completion timeline. According to analysis from Next Platform and Data Center Dynamics, the acquisition is intended to allow AMD to integrate inference-specific capabilities directly into its chiplet architecture rather than relying solely on software-layer optimisation via ROCm. This positions AMD to compete more credibly in the inference segment, where NVIDIA's dominance is increasingly challenged by custom ASICs and where AMD has historically underperformed relative to its training-side gains.

Separately, AMD's broader agentic AI positioning — described in Next Platform's analysis of its recent earnings — suggests the company is betting that agentic workloads, which combine frequent inference calls with task orchestration, will reward chips with lower latency and better memory bandwidth efficiency rather than raw FP8 throughput. Taalas appears to be a tactical acquisition to close that gap at the hardware level.

Why it matters

If AMD can natively embed inference efficiency into its Instinct line, it becomes a more credible alternative for hyperscalers diversifying away from NVIDIA, tightening competition in the segment that will dominate AI compute economics through the rest of the decade.

What to watch

Whether Taalas technology surfaces in AMD's next Instinct architecture announcement and how quickly AMD can demonstrate inference-per-watt benchmarks competitive with NVIDIA's Blackwell NVL variants.

Compute Scarcity Deepens: AWS Internal Rationing, RAM Prices at 2007 Levels, and Consumer GPU Supply Dysfunction

Three separate data points this week collectively paint a picture of AI-driven compute demand outpacing supply across every layer of the stack. Amazon Web Services is actively telling its own engineers to reduce EC2 CPU usage, citing pressure from external customer demand driven by agentic AI workloads. Tom's Hardware reports that low-utilisation instances have become a hot commodity internally — a notable inversion where AWS is prioritising paying customers over its own engineering teams. This is a confirmed operational policy change, not a speculative capacity warning.

On the memory side, a researcher's analysis cited by Tom's Hardware finds that per-gigabyte DRAM pricing has reverted to 2007 normalised levels in a matter of months, driven by AI server and HBM demand consuming the memory supply chain's capacity growth. SK Hynix's announced $38 billion new fab investment, noted in the Semiconductor Engineering week-in-review, is a direct response — but new fab capacity takes three to four years to come online, meaning memory price pressure is structural through at least 2028. At the consumer GPU tier, NVIDIA RTX 5090 cards are trading at nearly 2.5x MSRP on the marketplace, with NVIDIA cancelling orders mid-transaction due to upstream price increases from AIB partners — a supply chain coordination failure that reflects how thin margins and allocation constraints are across the board.

Why it matters

When the world's largest cloud provider is rationing CPU capacity internally and memory prices have structurally reversed two decades of deflation, the AI infrastructure buildout is operating beyond current supply limits — a condition that will sustain elevated capital expenditure and hardware pricing through the medium term.

What to watch

Whether SK Hynix's $38 billion fab commitment triggers parallel capacity expansions from Micron and Samsung, and how quickly AWS and other hyperscalers can convert announced datacenter capacity into live inference-serving infrastructure.

Hyperscaler Debt Financing and IPO Pipeline Signal Infrastructure Buildout Entering a New Capital Phase

JPMorgan has revised its 2026 tech bond issuance forecast above $500 billion, driven by hyperscaler AI capital programmes being financed through debt rather than cash flows alone. Bloomberg notes growing investor fatigue with the pace of issuance but confirms demand remains sufficient to absorb supply. Separately, data centre operator Switch has filed confidentially for a US IPO, joining a cohort of infrastructure platforms seeking public market capital to fund expansion. Bloomberg reports the filing is positioned to capture AI-themed investor appetite. These two developments together indicate the buildout is shifting from hyperscaler balance sheets into broader capital markets, which both accelerates deployment and distributes risk.

In the sovereign and regional infrastructure dimension, Indosat (part of the Ooredoo Group) has announced a 1GW data centre programme in Indonesia with NVIDIA and Nokia as partners, confirming Southeast Asia's emergence as a serious compute geography. Data Center Dynamics notes this is an announced plan, not confirmed construction. Australian neocloud Firmus raised $2 billion in equity from Blackstone, Coatue, NVIDIA, and Jane Street — a confirmed financing round that makes NVIDIA an equity stakeholder in its own customer's infrastructure, a structural alignment worth tracking. Data Center Dynamics reports this separately from a smaller Australian neocloud, Sharon AI, which has announced a $373 million five-year cloud contract — the commercial status of which should be treated as claimed rather than independently verified.

Why it matters

The shift from hyperscaler-funded to capital-markets-funded infrastructure buildout means AI compute capacity expansion is now partially decoupled from individual company earnings cycles, accelerating the pace of deployment but also creating leverage-dependent capacity that could stall if credit conditions tighten.

What to watch

Switch's IPO pricing and whether it achieves a premium valuation multiple consistent with AI infrastructure peers, and whether NVIDIA's equity stake in Firmus represents the beginning of a broader vendor-equity strategy to secure long-term GPU placement commitments.

SpaceX's Compute Revenue Signals a New Class of Vertically Integrated AI Infrastructure Competitor

SpaceX's first public quarterly earnings disclosure reveals that the company is, by revenue composition, primarily a telecom and compute rental business — with the AI division posting a $1.26 billion operating loss on substantial capital investment. The Verge notes that SpaceX's revenue structure, following the absorption of xAI assets, makes the 'space company' framing increasingly nominal. The infrastructure implication is that SpaceX is now a vertically integrated compute provider with its own satellite backhaul (Starlink), its own AI models, and its own datacenter ambitions — competing for the same hyperscaler customers and sovereign contracts as AWS, Azure, and Google Cloud, but with a differentiated connectivity stack for remote and austere environments.

The $1.26 billion operating loss in AI is a confirmed figure from SpaceX's own earnings statement, not an analyst projection. It indicates xAI's infrastructure buildout — Colossus and successor clusters — is consuming capital at a rate that has not yet been offset by inference revenue. Whether SpaceX's compute business achieves the scale to become self-sustaining remains an open question, but the disclosure confirms that Musk is deploying SpaceX's cash generation from launch and Starlink to subsidise AI compute infrastructure at a scale previously visible only to publicly traded hyperscalers.

Why it matters

SpaceX entering the compute rental market with Starlink-integrated connectivity creates a genuinely differentiated infrastructure competitor for edge, sovereign, and geographically constrained deployments that hyperscalers cannot easily replicate.

What to watch

Whether SpaceX discloses compute capacity figures — GPU count, data centre locations, power contracts — in subsequent quarters, which would allow direct comparison with hyperscaler infrastructure growth rates.

Signals & Trends

Custom Silicon is Becoming the Default Inference Strategy, Not the Exception

Within a single week, Anthropic announced a custom ASIC programme with Samsung and AMD acquired Taalas to embed inference logic into its chiplets. Combined with existing programmes at Google (TPU), Amazon (Inferentia/Trainium), Microsoft (Maia), and Meta (MTIA), the pattern is now clear: any organisation with sufficient inference volume and engineering resources is pursuing purpose-built silicon. The strategic consequence is a progressive narrowing of NVIDIA's total addressable inference market to workloads where flexibility, software ecosystem maturity, or rapid deployment speed outweigh unit economics. Analysts should track the timeline compression between foundation model lab announcements and first silicon tape-out — historically 18-24 months — as a leading indicator of when custom inference supply begins to meaningfully displace merchant GPU orders.

Land and Power Acquisition is Becoming a Binding Constraint on Datacenter Expansion

A Northern Kentucky farming family's refusal of a $26 million offer for datacenter land conversion illustrates a friction that is becoming systematic rather than anecdotal. Across the US, UK, and Europe, data centre operators face organised community resistance, agricultural land protection rules, water use restrictions, and grid interconnection queues that are extending site-to-operational timelines well beyond financial projections. The sponsored analysis from Data Center Dynamics on grid-to-chip infrastructure, and the separate piece on decarbonisation accountability Data Center Dynamics, reflect an industry increasingly aware that social licence and grid access — not capital availability — are the true bottlenecks. The shift of major buildout announcements toward Southeast Asia (Indosat's 1GW Indonesia programme), the Middle East, and Australia partly reflects operators routing around constrained Western grid markets. Infrastructure professionals should model site permitting and grid connection lead times as probability-weighted constraints on any capacity forecast, not as administrative formalities.

The Memory Supply Shock Has Structural Duration: Plan for Elevated Pricing Through 2028

The reversion of DRAM per-gigabyte pricing to 2007 levels is not a temporary demand spike — it reflects a fundamental reallocation of memory manufacturing capacity toward HBM and high-density server DRAM that has no short-term relief valve. SK Hynix's $38 billion new fab commitment, noted in the Semiconductor Engineering week-in-review, confirms that the industry recognises the structural gap, but fab construction and qualification timelines mean meaningful new supply enters before late 2028 at the earliest. For AI infrastructure operators, this has two practical consequences: server procurement costs are structurally elevated, compressing the unit economics of inference at scale; and the consumer and prosumer GPU market will continue to exhibit the dysfunction visible this week — grey market memory mods on seven-year-old cards, GPU orders cancelled mid-transaction, and MSRP pricing treated as a novelty event. Operators should model memory as a constrained input and evaluate HBM supply agreements as a strategic procurement priority equivalent to power purchase agreements.

Explore Other Categories

Read detailed analysis in other strategic domains