TL:DR
- For agentic AI inference workloads, proximity to data sources is as critical as hardware selection, since latency compounds quickly when agents send requests at machine speed.
- Metro edge colocation delivers the low latency, security, private interconnection and high-density power that distributed agentic AI infrastructure requires.
- Enterprises building edge AI strategies must plan for agents that dynamically discover, reach and connect to the nearest available data sources across multiple metros.
Many enterprise leaders have spent the past few years laser-focused on their AI strategies. They’ve obsessed over every detail of which hardware they should acquire and which models they should run. But now, they’re realizing that where they run AI workloads is just as important as what they run.
In particular, inference workloads can be highly sensitive to latency, so it’s important to place infrastructure in proximity to the data sources needed for useful results. Enterprises that don’t account for this could end up with AI applications that look great on paper, but ultimately fail to meet expectations.
The emergence of AI agents that operate at machine speed further underscores the need for proximity and governed interconnection. As they pursue their goals, AI agents need to connect with one another, access tools, move data, and adapt to changing conditions. This is all happening constantly, with actions completed in milliseconds, so there’s no room for delays or downtime.
Today’s agents already depend on distributed infrastructure, but this need will grow stronger as agents proliferate. According to IDC, enterprise agent use will increase 10x by 2027, while token and API call loads will increase 1000x.[1] This surge will drive enterprises toward distributed and federated data management, while still necessitating centralized data lakes. This creates a new paradox where data wants to be centralized and distributed at the same time. For agents to get the timely, context-rich data they need, new infrastructure challenges need to be met.
The future of AI will be agents running everywhere, all the time, all at once. It only makes sense that AI hardware will also need to be everywhere. Enterprises and neoclouds will need to deploy infrastructure at many metro edge locations to get the ideal mix of performance and security for this new reality.
Why agentic AI workloads belong at the metro edge
Latency is the inevitable byproduct of distance, regardless of which networking solutions you use. If data has to cover long distances, there will be latency, and the user experience will suffer. This is true for all different kinds of users:
- Human users may have to wait several seconds to get a response from a chatbot, which may lead them to get frustrated and try something else.
- Agent “users” won’t be able to access timely, accurate data when and where they need it, making it more difficult to achieve their goals. And since agents make far more requests than the average human user would, the impact of latency will compound much quicker. This could also result in underutilized GPU hardware, which has a powerful negative impact on the monetization of capital investments.
Enterprise leaders increasingly recognize the need for distributed infrastructure at the edge to support their AI inference workloads. For that matter, deploying one large inference cluster wouldn’t be practical anyway, since organizations would likely struggle to find enough data center capacity in a single location. For all these reasons, spreading AI workloads across multiple interconnected data centers has become the new standard. As one Google executive recently put it, the continent is the new data center.
Enterprises are aligning their investments to this new reality: According to IDC, worldwide spending on edge computing is expected to reach $450 billion by 2029. This is nearly double the edge spending we saw in 2025, reflecting a pivotal expansion fueled by rapid edge AI advancements.[2]
Edge computing is not merely a technical option; it is a fundamental pillar of AI's future. It directly enables the real-time responsiveness, data privacy, and operational resilience that advanced AI agents demand.”Dave McCarthy, Research Vice President, IDC[3]
It’s important to note that “the edge” isn’t a specific location. The edge is simply anywhere that’s closer to end users and data sources than the centralized “primary” data center is. Each organization defines this in their own way. They’ll typically have a hierarchy of different edge locations, including:
- The far edge: Where organizations deploy hardware in the field. One example would be a manufacturer placing a server rack on their own factory floor to service industrial IoT workloads.
- The metro edge: Where organizations deploy processing hardware in a dedicated facility located in the same metro area as the workloads they’re servicing.
The metro edge represents the sweet spot for AI inference workloads. It provides lower latency than remote cloud data centers, but also better security and reliability than running workloads at the far edge. It also provides more interconnection possibilities to multiple NSPs and neoclouds, which helps enable resilience, performance, and cost control.
This is especially true because organizations can work with a colocation partner to deploy their metro edge infrastructure, or to privately connect to a local GPU provider within that metro for a low-latency experience. The colocation provider handles physical security and power redundancy to protect sensitive assets and keep them online. They may also provide liquid cooling to support high-density GPU hardware for the enterprise or their chosen neocloud provider. It would be essentially impossible for organizations to get these capabilities for themselves at the far edge.
In the future, agentic AI will be happening everywhere, all the time
Placing AI inference at the metro edge doesn’t necessarily mean deploying in one specific metro. This is especially true for AI agents, which often need to work across different metros to ensure the best results.
Looking toward the future, enterprise leaders must plan for a time when agents will be able to decide for themselves where they should operate. We can safely assume that future agents will be equipped with service discovery capabilities that allow them to understand how latent they are to a particular data source, and then automatically position themselves in the nearest available metro area.
This is why enabling low latency for agentic workloads isn’t just about having infrastructure in all the right locations; it’s also about connecting all the different locations using direct, private networking solutions. This allows agents to avoid sending requests or moving data over the public internet, which both exacerbates latency caused by distance and creates privacy, security and sovereignty concerns.
These networking solutions must be agile and intelligent, particularly as agents continue to grow more self-directed. Imagine an agent that needs to access a data source inside a particular cloud region. Not only should it be able to reach the metro area closest to that cloud region, but it should also be able to set up its own connection to that cloud. Instead of the organization paying for an always-on private cloud connection, even when it’s not needed, the agent should be able to spin up that connection on-demand and take it down once it’s finished using it.
What enterprise leaders need to know about the future of agentic AI
We’ve established why it’s so important for enterprises to consider location in their AI strategies. Now, let’s consider a few questions enterprise leaders should ask to determine what that might look like:
- Which data sources or services will our agents need to connect with? Creating a detailed account of different partners, devices and end users provides an initial understanding of where data’s coming from, which in turn helps determine where agents might need to go.
- Are we prepared for our infrastructure to follow the data? Again, it’s not just about where the agents need to be right now; in the future, you’ll need to ensure agents can go wherever their goals might take them.
- How will we bring our different metro edges together? A distributed infrastructure strategy for agentic AI depends on connectivity to link all the different data sources, and that doesn’t mean using the public internet.
- What governance constraints do I need to be aware of? The question of “What do I want to do?” must come with the follow-up questions “How do I need to do it?” and “What can’t I do?” All these questions must be codified in the automation that agents execute.
Equinix offers digital infrastructure solutions that are ready-made to support agentic workloads running at the metro edge. Our colocation data centers are available in 77 different metros worldwide. This means you don’t have to worry about finding data center capacity where you need it; there’s a very good chance that capacity already exists.
With Equinix Fabric Intelligence™, we’re pioneering the kind of agile, intelligent networking solutions that future generations of AI agents will depend on. This AI-native solution allows the network to observe real-time conditions, make goal-driven decisions, and adapt accordingly. In short, it makes the network just as agentic as the agents themselves.
Since Equinix solutions are vendor-neutral by design, they make all data sources accessible, including those distributed across different cloud environments.
To learn how agentic AI is changing the way enterprises approach distributed infrastructure, read the IDC analyst brief “Growth in AI Agents Will Require an Edge Inferencing Strategy.”
Also, watch the video “AI Inference at the Edge: How Distributed AI Architecture Reduces Latency” for more information:
[1] IDC, Agent Adoption: The IT Industry’s Next Great Inflection Point, December 10, 2025.
[2] IDC press release, Edge Computing Global Spending to Grow at 15%, reaching $450 Billion by 2029, February 25, 2026.
[3] IDC, Growth in AI Agents Will Require an Edge Inferencing Strategy, May 2025.