The Infrastructure Behind AI

AI Infrastructure: Scale Up or Scale Out?

To meet the growing need for AI inference, enterprises must scale the right infrastructure in all the right locations

Glenn Dekhayser
AI Infrastructure: Scale Up or Scale Out?

TL:DR

  • As AI inference becomes the dominant enterprise workload, scaling out distributed digital infrastructure across locations is now as critical as scaling up compute capacity.
  • Hybrid multicloud architecture with seamless interconnection empowers enterprises to place workloads in the right environments while maintaining control, security & cost-efficiency.
  • Equinix enables enterprises to scale AI inference globally by combining colocation across strategic markets, cloud adjacency, interconnection & managed infrastructure services.

The AI ecosystem is diverse, and different players have different priorities when it comes to scaling their infrastructure:

  • Neoclouds are primarily concerned with scaling up, because they need to meet massive demand for compute capacity from enterprise customers.
  • The enterprise customers themselves are typically more concerned with scaling out to support distributed workloads such as agentic AI across many different locations.

Many enterprises find scaling out AI infrastructure challenging because it’s so different from what they’ve done in the past. They typically started their AI journeys in the public cloud, and they were content to let someone else worry about scaling on their behalf. But as we’ve established, the scaling priorities of enterprises don’t always align with those of their service providers.

This distinction matters because AI infrastructure decisions increasingly shape business outcomes. As enterprises move AI from experimentation into production, decisions about where workloads run, where data resides and how systems connect can influence latency, operating costs, governance, and the speed at which new AI capabilities are deployed.

Enterprises may also be looking for more cost-effective ways to scale their infrastructure, due to a convergence of different factors. For one thing, many AI service providers are ramping up their monetization efforts. For instance, some prominent AI services have recently switched from a subscription model to usage-based billing, increasing prices by up to 100x for some users.[1]

At the same time, the shift from AI experimentation to production means that enterprises need to support large AI workloads consistently, making a pay-per-use pricing model particularly concerning.

To build scalable, future-proof AI infrastructure, enterprises need control. They need control over what they run, where they run it, and who they connect with to help them. They can’t always count on cloud providers to give them this level of control. However, that’s not to say that they should avoid the cloud altogether. What they need is a true hybrid multicloud architecture backed up by seamless connectivity.

With hybrid multicloud, they’re empowered to scale infrastructure by placing the right workloads in the right environments. Then, they can interconnect those workloads to maintain control, security and cost-efficiency across their scale-out architecture. This ability is especially important in the era of AI inference. Today’s enterprises are increasingly concerned with how to keep latency low for inference workloads, and this new era requires a new approach to AI infrastructure.

The shift to inference underscores the importance of scale-out infrastructure

Just like service providers and enterprises have different priorities around scaling, different varieties of AI workloads must scale in different ways. One of the basic principles of distributed AI is that training workloads and inference workloads have different infrastructure requirements, and thus need to run in different places:

  • Training workloads require a lot of compute capacity, meaning that organizations will likely need to scale up large GPU clusters in a few core locations.
  • Inference workloads tend to be more spread out in proximity to data sources, because the primary concern is keeping latency low. Therefore, organizations would likely scale out their inference hardware: smaller deployments, but more of them in more places.

In the early days of generative AI, much of the industry commentary focused on training workloads. We heard a lot about demand for GPUs and investments in massive new data centers to host those GPUs. Now that the AI market has matured, many organizations are less concerned with getting the right models and more concerned about how to apply the models they already have. This shift puts inference in the spotlight.

For enterprise leaders, it’s as much a strategic shift as it is a technical shift. As inference becomes the dominant AI workload, making the right infrastructure choices increasingly determines the quality of customer experiences, the economics of AI deployment and the ability to scale new AI-powered services globally.

The advent of agentic AI has further underscored the importance of scaling out inference infrastructure. According to one prediction from IDC, the number of actively deployed AI agents will exceed 1 billion worldwide by 2029, which is 40 times more than in 2025. More significantly, these agents will execute more than 217 billion actions a day.[2]

Every single action an agent performs comes with the potential for increased latency. And since agents often operate autonomously, enterprises may not have direct control over where agents run and which tools they connect with. Thus, they may find it difficult to optimize proximity and ensure low latency.

As AI workloads become more distributed, performance increasingly depends on how quickly data moves between clouds, models, enterprise environments and end users. This means that the network layer has become just as important as the compute layer. Recognizing this shift, the network service provider Zayo recently deployed 400G connectivity across Equinix’s global footprint to support high-bandwidth, low-latency AI and data-intensive workloads for our shared customers.

For all the reasons mentioned above, scaling out will become more important than ever before. We know that agents will live everywhere, and we know they’ll need to connect with data everywhere, so it only makes sense that we’ll need AI infrastructure everywhere.

How enterprise leaders can start scaling out AI infrastructure today

While much of the industry remains focused on scaling up AI capacity, enterprise leaders must tune out the noise. They need to focus on where their AI workloads run and how distributed AI systems connect. They know that limited data center capacity could become a bottleneck for their AI strategy, but they also know they don’t need to wait for anyone’s help to build their way out of this problem.

Instead, they can take a more opportunistic approach. Because they’re scaling out, they don’t need to worry about finding a massive data center footprint in any one location. They can take advantage of capacity that’s already available across many different locations. Even if their average deployment is small, that aligns better with their distributed infrastructure strategy anyway. And with the right connectivity, their distributed hardware can still function as one integrated computing environment.

Hybrid multicloud connectivity solutions can help enterprises:

  • Link their private infrastructure with their multicloud environments to optimize performance, cost-efficiency and security benefits
  • Apply distributed security and governance policies across the network to scale globally without introducing risk
  • Get the network agility and intelligence needed to support tomorrow’s demanding AI agents

Find the right partner to scale out your AI infrastructure

To scale AI successfully, enterprises need the following foundational capabilities: proximity to users and data, freedom to choose the right cloud and AI service providers, secure connectivity between distributed environments, and consistent operations at global scale.

Equinix brings these four capabilities together in a single platform:

  • Private environments on a global scale: Equinix colocation data centers are available in 77 different strategic markets throughout the world. Our customers can quickly deploy wherever they need to be to ensure proximity to their users and data sources.
  • Cloud adjacency: Equinix is the market leader in native cloud on-ramps to all major cloud providers. This means our customers can quickly connect to the clouds of their choice, while still maintaining full control over data inside their private environments.
  • Global connectivity solutions: No matter where our customers deploy their infrastructure or who they partner with, they can find the connectivity solutions needed to tie all the different pieces together. This includes Equinix Fabric Intelligence™ for intelligent network automation and Equinix Network Edge for on-demand network services without hardware constraints.
  • Managed infrastructure services: With Equinix Managed Solutions and Enablement Services, our customers can let someone else handle the nuts and bolts of their AI-ready data center environments, freeing up IT resources to focus on more high-value work.

This combination of capabilities is unique to Equinix. We are the only digital infrastructure provider that offers global colocation, cloud adjacency, interconnection and managed services under one roof. Therefore, we’re the ideal partner to help enterprises enable AI inference by scaling the right infrastructure in the right places.

Customer story: Cognitiv pairs low-latency inference with full control over hardware

Equinix customers are already tapping into our offerings to scale their global inference infrastructure. For example, the leading ad tech company Cognitiv deployed their AI inference workloads in Equinix data centers across multiple global regions.

By hosting inference at the edge and colocating with many customers that are also deployed inside Equinix data centers, the company can now evaluate millions of bid requests per second. They also maintain full control over their hardware, support workloads that can’t run in the public cloud, and keep end-to-end latency in the low single milliseconds.

To learn more, read the full Cognitiv case study.

 

[1] Bruno Ferreira, Github Copilot customers report up to 100-fold price hikes — AI sticker shock bites as Microsoft switches to usage-based pricing, Tom’s Hardware, June 3, 2026.

[2] IDC Blog, Agent Adoption: The IT Industry’s Next Great Inflection Point, December 10, 2025.

Avatar photo
Glenn Dekhayser Global Principal, Global Solutions Architects
Subscribe to the Equinix Blog