Retrofitting In-Building Data Centers

May 5, 2025 | Doug Makishima, CEO

Retrofitting existing data centers for AI is becoming increasingly common as organizations aim to harness the power of artificial intelligence without the high cost and complexity of building new facilities. Here’s why companies choose to upgrade their current infrastructure:

Cost Efficiency
By avoiding new builds, businesses can save on the capital expenditures associated with constructing a new data center and arranging for power, connectivity and cooling.

Time to Market
The process of constructing a new data center is very time consuming and can significantly delay the availability of the newly planned AI infrastructure. It may also take time for utilities to build out the power grid and network infrastructure.

High Cost of Public Clouds
Costs to rent GPUs in the cloud have skyrocketed with demand, meaning that there is either limited availability at the price point businesses are willing to pay, or they will be paying a premium for GPU reservation and utilization.

Data Privacy and AI Sovereignty
Data privacy and AI sovereignty is essential to ensure both compliance and control in an increasingly data-driven landscape. Ensuring these principles during a retrofit not only secures sensitive data but also enables greater strategic autonomy, resilience, and trust in the AI systems being deployed.

Maximizing Utilization

When replacing servers with ones that offer more GPU/CPU power for AI/HPC applications, the objective should be to maximize utilization. By using optimized infrastructure management, businesses can take advantage of advanced schedulers and other tools that keep GPU workloads running as often as possible. This is where DCIM (Data Center Infrastructure Management) comes in.

ECOBLOX offers solutions for both AI/HPC in-building data centers and modular data centers, as well as retrofitting existing data centers. Our proprietary DCIM platform is helping data centers run more efficiently across all resources giving operators a comprehensive view of operations and providing opportunities for energy savings, resource usage optimization, and managing cooling systems. In fact, these efforts can offer energy savings up to 25% and can reduce the number of GPU nodes required

The Role of DCIM

Let’s look more closely to see why data centers are choosing ECOBLOX to help optimize their existing resources through our DCIM platform. In a sense, a DCIM platform combines infrastructure management, resource scheduling, monitoring, automation, and reporting tools to maximize ROI.

Infrastructure Management
From power distribution units (PDUs), uninterruptible power supply (UPS) systems to generators and power conditioners such as surge protectors and voltage regulators, operators can understand power consumption and failover mechanisms. Infrastructure management covers the data center’s cooling and environmental controls such as HVAC systems that provide precise temperature and humidity controls and cooling systems to manage heat dissipation from servers and other equipment.

From computing and storage systems to virtual and software services, racks, cages and cabling, all of these assets must be managed and done so in harmony with facilities management and automation systems.

In addition to hardware configuration settings of each physical and virtual device and maintaining firmware versions, there are network configurations and software (both operating system and application settings) that must be considered. Not only should these configurations be documented including change management and maintenance, but they must also include data backups and replication processes. Maintaining an accurate asset inventory is helpful when setting standards for procuring, deploying, and when equipment has reached its end of life, decommissioning them.

The management platform also includes the physical security of the data center: access control, surveillance systems and fire suppression systems. The connectivity and network infrastructure (cabling, network devices) as well as views of the physical space and layout of the data center also fall under facilities infrastructure.

Resource Scheduling
By using the ECOBLOX DCIM platform to schedule GPU tasks, operators can maximize GPU usage. Resource scheduling in a data center optimizes the use of existing hardware by ensuring that workloads are distributed efficiently across available resources. This can reduce the need for additional hardware by maximizing the capacity and performance of the current infrastructure. This includes running jobs at off-peak hours when GPUs would otherwise be idle, and assigning additional resources when jobs need to be completed immediately.

Monitoring
Real-time monitoring is a key aspect of ECOBLOX’s DCIM’s automation capabilities. We monitor compute and GPU usage, storage, network throughput, power usage, temperature, humidity. The DCIM is configured to generate alerts if normal thresholds are exceeded, and alarm forwarding allows these alarms to be funneled into an existing altering system. In addition to alerting, the DCIM can also regulate resources to keep temperature and humidity within desired ranges to prevent damage and to maximize efficiency. Monitoring can also be linked with automation to provide automatic remediation actions when a fault is detected and diagnosed.

Automation
ECOBLOX’s proprietary DCIM platform uses AI to analyze resource usage and historical data. In this way, the DCIM platform can predict future needs and recommend upgrades or adjustments. It can also automate the deployment of both hardware and software – this includes provisioning machines, configuring network settings and deploying applications.

Routine maintenance tasks, work order management, and reporting/analysis are areas that automation can play a role. The ECOBLOX DCIM platform can also integrate with other IT management tools facilitating organization-specific workflows, custom processes and centralized reporting.

Historical Reporting
In an AI/High-Performance Computing (HPC) data center, historical reporting is essential for monitoring performance, optimizing resource usage, ensuring uptime, and managing costs. Various types of reports provide insights across hardware, software, and operations over longer periods of time. Here are the main categories and types of reports available:

  • Node Utilization Reports – Showing operators details of CPU, GPU, and memory across compute nodes
  • Job Performance Reports – Tracking execution time, resource consumption and identifying any bottlenecks in the job queue
  • Storage I/O Performance Reports – This measures input/output rates across storage systems to pinpoint times when the storage network was at capacity, and which jobs were running
  • Network Performance Reports – Evaluate data transfer speeds and latency between nodes to ensure fast communication

With a DCIM platform that manages all that has been described above, it is easy to understand how one, fully integrated AI/HPC platform, is essential for all data centers – both modular and in-building.

With the high demand for AI/HPC data centers, we have seen demand for the ECOBLOX DCIM platform rise significantly. While new in-building data centers may be getting all the attention and investment, many businesses should consider retrofitting existing facilities and using DCIM to get the most out of the space and power they already have. Contact ECOBLOX sales to find out how.