Compute & Infrastructure
Top Line
Envision Group has commissioned the initial phase of a gigawatt-scale AI data center in Inner Mongolia, marking a strategic pivot by a major Chinese renewable energy company into AI infrastructure and signalling that China's buildout is increasingly integrating collocated clean power with compute capacity.
Elon Musk confirmed SpaceX and xAI will exclusively deploy NVIDIA GPUs — including a space-optimised Vera Rubin NVL72 planned for orbit in 2027 — reinforcing NVIDIA's near-total lock on frontier AI compute demand across radically divergent deployment environments.
Frore Systems claims its LiquidJet direct-die cooling solution can reduce NVIDIA Rubin GPU temperatures by 10°C and improve token-generation efficiency by 15%, with hyperscalers actively evaluating delidded GPU configurations, signalling that thermal management is becoming a primary lever for data centre economic optimisation.
Multi-physics co-simulation bottlenecks — not toolchain gaps — are now the principal constraint on heterogeneous chiplet design iteration speed, with direct implications for the timeline of next-generation AI accelerator packaging.
NSF announced a $100 million AI Infrastructure Hubs programme to distribute compute access across underserved US regions, a modest but directionally significant sovereign infrastructure move to prevent geographic concentration of AI R&D capacity.
Key Developments
Envision's Inner Mongolia Gigawatt Data Center: Renewable-Integrated AI Infrastructure at Scale
Envision Group, one of China's largest wind and energy storage companies, has commissioned the first phase of a gigawatt-scale data center in Inner Mongolia, according to Bloomberg. The facility represents an unusual but strategically coherent move: a company that controls renewable generation assets building compute infrastructure adjacent to its own power supply, effectively short-circuiting the grid interconnection bottleneck that is throttling data center expansion globally. Inner Mongolia offers abundant wind resources, relatively low land costs, and existing high-voltage transmission infrastructure, making it one of China's preferred locations for power-intensive AI workloads.
The confirmed commissioning of the initial phase distinguishes this from the large volume of announced-but-unbuilt capacity in the market. That said, 'gigawatt-scale' describes the full build target — the commissioned phase is almost certainly a fraction of that figure. The strategic model here — energy company as data center operator — is one to watch as a template for bypassing utility constraints, particularly in markets where grid connection queues run to years.
NVIDIA Rubin Cooling Economics: Thermal Management as a Competitive Differentiator
Frore Systems has publicly claimed that its LiquidJet direct-die coldplate technology can lower NVIDIA Rubin GPU junction temperatures by 10°C and translate that thermal headroom into a 15% improvement in token-generation throughput, per Tom's Hardware. Crucially, hyperscalers are reportedly evaluating delidded GPU configurations in production — removing the integrated heat spreader to allow direct contact between the coldplate and the die. This is aggressive thermal engineering that trades hardware warranty and handling complexity for sustained frequency and power delivery gains.
The 15% token-generation efficiency claim requires scrutiny: it likely reflects a best-case scenario under sustained thermal throttling conditions rather than a universal baseline improvement. Nevertheless, the direction is analytically important. As GPU TDP continues to climb — Rubin-class hardware operates in the 1,000–1,200W range per GPU — the delta between thermally-constrained and thermally-unconstrained performance widens, and cooling vendors increasingly determine effective compute yield per rack. This is shifting cooling from a facilities afterthought to a first-order design parameter in AI factory architecture.
SpaceX and xAI's Exclusive NVIDIA Commitment: Vera Rubin NVL72 Headed to Orbit
Elon Musk confirmed via social media that both SpaceX and xAI will exclusively use NVIDIA AI accelerators for training and inference workloads, citing the Vera Rubin architecture as the best available compute platform, per Tom's Hardware. Musk further stated that a space-optimised variant of the Vera Rubin NVL72 is planned for launch in 2027, implying NVIDIA is developing or has already developed radiation-tolerant or thermally-adapted versions of its flagship AI accelerator for orbital deployment.
The exclusive commitment from xAI reinforces NVIDIA's demand position at the frontier, though it also carries strategic concentration risk for both parties. For NVIDIA, dependence on a small set of very large customers creates revenue volatility; for xAI, exclusive platform lock-in limits negotiating leverage as AMD, Intel, and custom silicon alternatives mature. The orbital deployment angle is analytically distinct: space-based AI compute is nascent but directionally significant for defence, Earth observation, and autonomous satellite operations. If NVIDIA's silicon is qualifying for radiation environments, it expands the addressable market and creates a new certification moat.
Chiplet Co-Simulation Bottleneck: Design Velocity Constrained by Compute, Not Tooling
A Semiconductor Engineering analysis confirms that multi-physics co-simulation across 2.5D and 3D chiplet stacks has become the primary constraint on heterogeneous design iteration — not toolchain capability gaps, per Semiconductor Engineering. Simulating thermal, electrical, mechanical, and signal-integrity interactions simultaneously across a multi-die package is computationally expensive to the point where full-fidelity simulation runs take days or weeks, forcing engineers to work with degraded models or serialise what should be parallel design loops. This is a direct tax on the development velocity of next-generation AI accelerators, which almost universally rely on chiplet-based packaging — HBM stacking, I/O chiplets, compute tiles — to circumvent monolithic die scaling limits.
The implication for the AI infrastructure supply chain is that simulation compute infrastructure — high-core-count workstations, EDA-optimised cloud instances, and simulation acceleration hardware — is itself a bottleneck in the path from architecture to production silicon. As AI accelerator design cycles compress under competitive pressure, the firms that invest in simulation infrastructure gain a compounding velocity advantage.
NSF's $100M AI Infrastructure Hubs: Sovereign Compute Distribution, Modestly Scaled
The US National Science Foundation has announced a $100 million AI Infrastructure Hubs programme designed to extend AI compute access to research institutions in underserved regions of the US, per Data Center Dynamics. The programme is confirmed as announced — it is not yet confirmed as funded and deployed. At $100 million, the investment is modest relative to the scale of private hyperscaler buildout (Microsoft, Google, and Amazon each announced over $50 billion in 2026 capex commitments), but the policy objective is distinct: distributing compute access geographically rather than maximising aggregate capacity. This addresses a real structural problem — AI R&D capability is concentrating in a handful of coastal metro areas, limiting the diversity of institutions that can participate in frontier research.
Signals & Trends
Thermal Architecture is Becoming an AI Compute Capacity Variable, Not a Facilities Variable
The Frore LiquidJet development and the broader hyperscaler interest in delidded GPU configurations point to a structural shift in how data center operators think about compute yield. Historically, cooling was a facilities and mechanical engineering problem solved after hardware selection. As GPU TDP climbs toward and beyond 1,000W per device, and as rack densities push into the 100–200kW range, the gap between thermally-throttled and thermally-unconstrained GPU performance is becoming large enough to materially affect AI economics. This means cooling vendors — liquid cooling suppliers, direct-die thermal interface specialists, and immersion cooling operators — are increasingly competing on the same value metric as GPU vendors: tokens per dollar. Operators who integrate thermal optimisation into their GPU procurement and deployment strategy, rather than treating it as a downstream facilities problem, will extract structurally better economics from the same hardware. This dynamic will intensify as Rubin and its successors push power envelopes further.
China's AI Infrastructure Model: Vertically Integrated Energy and Compute
Envision's Inner Mongolia facility is a data point in an emerging pattern: Chinese AI infrastructure buildout is increasingly structured around collocated or adjacent renewable energy assets, rather than relying on grid interconnection. This is partly a response to China's grid congestion in coastal demand centres, and partly a deliberate strategy by energy companies — which control large land banks and generation assets in resource-rich regions — to capture value from the AI compute boom. Western hyperscalers are pursuing similar strategies (Microsoft and Google have both signed power purchase agreements with adjacent generation facilities), but the Chinese model appears to be moving faster toward full vertical integration, with the energy company operating the data center rather than simply supplying power under contract. If this model proves economically superior — by eliminating interconnection queues and providing guaranteed power pricing — it could accelerate Chinese AI training capacity relative to Western counterparts in ways that aggregate GPU import data does not capture.
Hybrid Digital-Analogue Computing Research is Maturing Toward System-Level Credibility
The joint paper from Nottingham Trent, Imperial College London, and Aston University arguing for hybrid digital-analogue and neuromorphic computing as a credible energy-efficiency pathway — moving beyond peak TOPS/W as a metric — reflects a maturing of the research agenda in this space. The key analytical shift is from device-level efficiency claims to system-level efficiency accounting, which has historically been where analogue and neuromorphic approaches underperform once interconnect, memory, and precision costs are included. If the system-level framing holds up under scrutiny, it could accelerate industry interest in mixed-substrate AI accelerators at a time when the energy cost of AI inference is under intense pressure from regulators and corporate sustainability commitments. This is a signal to track over an 18–36 month horizon rather than an immediate infrastructure implication, but the convergence of academic credibility and commercial energy pressure makes it worth monitoring.
Explore Other Categories
Read detailed analysis in other strategic domains