TL:DR
- AI infrastructure requires observability across both compute clusters and facility systems, as the gap between these environments is where outages most commonly originate.
- Integrating Equinix Smart View’s facility telemetry with NVIDIA Mission Control’s compute orchestration creates a unified operational view across power, cooling & hardware systems.
- Shared observability across compute and facility infrastructure enables faster recovery, more efficient utilization & greater resilience for production-scale AI workloads.
Modern AI workloads are growing ever more powerful. Enterprises are adopting AI agents that can integrate with each other and pursue common goals without human intervention. According to one IDC prediction, agentic AI spending will reach $1.3 trillion by 2029, accounting for more than 26% of total IT budgets worldwide.[1]
As enterprises scale AI, they also scale their dependencies. The more capabilities and permissions they grant to agents, the greater the risk when those agents can’t function properly. Ensuring resilient AI infrastructure has become a top priority for enterprise leaders, because when their AI infrastructure goes down, key automated business processes go with it. The organizations that scale AI most effectively will be those that can detect, isolate, and recover from outages before they impact business operations.
Fortunately, we already have the blueprint for AI resiliency: It’s about detecting interruptions quickly and automating recovery to minimize downtime. When Meta trained their Llama 3 model using NVIDIA GPUs, they logged 419 unexpected interruptions in 54 days, but still achieved 90% effective training time, thanks to automated cluster maintenance.[2]
This kind of automated response requires observability. The challenge is that AI infrastructure doesn’t operate as a single system. Compute clusters and facility operations both generate critical signals, yet those signals are typically managed separately. As AI workloads become more distributed and power-intensive, that separation becomes an operational risk.
These signals live in the orchestration software that monitors the AI compute cluster and in the data center infrastructure management (DCIM) software that monitors power and cooling systems. Traditionally, these solutions haven’t been able to integrate with one another. This gap between the hardware and the data center is exactly where AI outages tend to originate.
To ensure reliable AI infrastructure, teams need an integrated view that spans the compute and the facility. As we like to put it, they need observability from the cluster to the circuit. This shift represents more than better monitoring. It’s a new operating model for AI infrastructure, where compute orchestration and data center operations work from the same picture rather than responding independently to the same event.
Resilient AI infrastructure depends on data center observability
AI workloads are inherently more dense than conventional IT. Today’s GPUs may use as much as 200 kilowatts per rack, compared to the 10-20 kW that traditional data centers were designed to deliver. This kind of density demands liquid cooling, and that introduces new risks not found in air-cooled systems.
For instance, organizations must be aware of issues with coolant temperature, pressure and flow, as well as the potential for leaks. And since a single coolant distribution unit (CDU) can serve an entire compute cluster, it only takes one small problem to cause serious disruption.
Power systems are another common source of failure for AI infrastructure. In fact, the most recent Uptime Institute annual outage analysis report found that power issues were once again the leading cause of impactful data center outages.[3] These issues could include both grid constraints at the utility level and problems with UPS systems, transfer switches and generators at the facility level.
The increasing density of today’s AI workloads will only exacerbate these power issues. And just like with cooling, even relatively small problems with power systems could be extremely impactful. This is especially true of synchronous AI training jobs. When even a single node is knocked offline, thousands of others will sit idle while the job recovers.
Data center operators should provide redundancy for cooling and power systems, but this is only helpful if the compute environment is designed to take advantage of that redundancy.
Many of the warning signs that indicate a potential outage come from power and cooling systems. That’s why it’s problematic when IT teams that manage AI clusters can’t access facility telemetry. When there’s an issue—like a power circuit approaching its limit or coolant drifting outside standard temperature ranges—they won’t know about it until the facility team manually notifies them.
This delay inevitably leads to idle GPUs. On the other hand, when the facility team and the IT team have shared observability, they’ll both learn about the problem at the same time, and the delay disappears.
Equinix Smart View and NVIDIA Mission Control close the observability gap
NVIDIA Mission Control is the software intelligence solution that operates an NVIDIA DGX SuperPOD. It’s responsible for scheduling jobs, monitoring GPUs, networks and storage, and automating recovery.
Equinix Smart View is our DCIM solution. It gathers operational data from data centers, including:
- Power draw at the circuit
- Temperature and humidity across cabinets, cages and sensors
- The status of mechanical and electrical assets
These two solutions are integrated to provide a complete view of both sides of the equation: the hardware, as well as the power and cooling systems that keep the hardware running.
The Smart View facility telemetry enters Mission Control via NVIDIA’s building management system (BMS) integration capabilities.[4] This means that IT teams can now find visibility into power and cooling systems in the same place they already find information about compute, networking and storage.
Equinix Smart View carries facility telemetry into NVIDIA Mission Control
Integrated observability enables action, not just visibility
The value of integrated observability isn’t simply seeing more data; it’s enabling predictive infrastructure maintenance. When compute orchestration and facility operations share the same operational context, organizations can detect issues earlier, automate responses and reduce cascading failures that can interrupt large-scale AI workloads.
The integration between Smart View and Mission Control enables both automation capabilities for compute recovery and manual intervention from skilled operations professionals.
On the compute side, Mission Control’s Autonomous Recovery Engine can restart failed jobs from the most recent checkpoint, with no human-in-the-loop required. According to NVIDIA, this can be up to 10x faster than recovering jobs manually.[5]
Mission Control also includes advanced power optimization capabilities. When paired with Smart View’s facility telemetry, it can match GPU power draw to the actual amount of power the circuit can provide. This allows organizations to make the most of the capacity they’re already paying for. In contrast, setting a static cap based on what the circuit might be able to provide can result in underutilization and inefficiency.
Facility telemetry also helps identify the circuits that are best suited to support recovery jobs. Otherwise, the system might move a recovery workload to a circuit that’s already at risk, thus setting up a repeat of the issue that caused the failure in the first place.
Automation accelerates recovery, but resilient AI operations still depend on coordinated action across software and physical infrastructure. That‘s because automation contains the problem and protects the workload for the moment, but it doesn’t address the root cause. That can only be done on the facility side. Whether it’s an overloaded circuit or a coolant leak, physical events in the data center require a human response.
A data center operations team responds to facility signals by sending a technician to the floor to repair or replace the impacted system. Integration with compute observability helps them respond to issues faster. They already know exactly which circuit or cooling loop caused the problem, because they could see the cluster react to it in near-real time.
When the automation capabilities on the compute side and the operations team on the facility side have the same view from the cluster to the circuit, they can work together to keep an AI factory running at full potential over the long run.
What integrated observability means for enterprise leaders
The future of AI operations depends on treating compute infrastructure and facility infrastructure as one integrated operational system. As AI workloads become denser, more autonomous and more business-critical, organizations that anticipate failures—not just respond to them—will achieve higher utilization, greater resilience and better returns on their AI investments.
With the integration between Equinix Smart View and NVIDIA Mission Control, organizations gain a unified operational view across compute and facility systems. IT teams no longer have to respond to hardware failures after the fact. Now, they know when a failure is likely to occur, and more importantly, they know why. The result is faster recovery, more efficient infrastructure utilization and greater confidence running AI workloads at production scale.
This integration is the latest development in the Equinix AI Factory accelerated by NVIDIA. With this turnkey solution, organizations can deploy privately owned NVIDIA AI infrastructure that’s installed, monitored and maintained by our managed services teams. Thanks to our portfolio of 280+ colocation data centers spread across 77 global metros, they can deploy that infrastructure wherever it delivers the most business value and maintain operational resilience.
In the future, integrated observability will be a foundational capability for resilient, production-scale AI. By connecting visibility from the cluster to the circuit, Equinix helps enterprises build AI infrastructure designed to scale with confidence.
Learn more about joint Equinix and NVIDIA solutions for enterprise AI: Read our solution brief.
[1] IDC Blog, Agentic AI to Dominate IT Budget Expansion Over Next Five Years, Exceeding 26% of Worldwide IT Spending, and $1.3 Trillion in 2029, According to IDC, August 26, 2025.
[2] Llama Team, AI @ Meta, The Llama 3 Herd of Models, July 23, 2024.
[3] Douglas Donnellan, Andy Lawrence, Rose Weinschenk, Annual outage analysis 2026, Uptime Intelligence, May 11, 2026.
[4] NVIDIA Mission Control Integration with Building Management System

