Evaluate the server against a named job, not a GPU model name.
This decision matrix separates evidence about the workload from the storage and network connections it needs, the facility conditions it requires, and the responsibilities of whoever will operate it.
Data center deep-dive series
What must the workload prove?
Start by naming the task—such as training or inference—and defining what a successful run must achieve. Job completion time and token throughput are examples of end-to-end measures discussed in Cisco’s benchmarking approach, not performance promises for a proposed server. A GPU model name cannot establish whether the complete deployment will pass your test.
Workload result
- Enter or request
- Named task, inputs, acceptance measure and a result from the proposed configuration
- Decision
- Pending until tested; consider excluding a configuration that fails a required measure
Accelerator
- Enter or request
- Accelerator type, installed configuration and its role in the task
- Decision
- Pending if only a model name or category is given
Working memory
- Enter or request
- Memory required while the job runs and memory available in the proposed configuration
- Decision
- Pending until capacity and workload fit are verified
Accelerators are not interchangeable simply because they are described as AI hardware. IBM describes GPUs as one type and also discusses NPUs for real-time AI processing and TPUs for tensor computations; those category descriptions give no model-specific capacity or benchmark result. Keep working memory separate from persistent storage.
IBM describes high-bandwidth memory in GPUs, other accelerators and some SSDs, but that fact does not establish how much usable memory a particular job has or whether its model fits.
U.S. Department of Energy — Best Practices Guide for Energy-Efficient Data Center Design
IBM — What Is an AI Data Center? | IBM
IBM — What Is a Data Center? | IBM
How will data and network traffic reach the server?
Record where frequently accessed data resides and how it reaches the proposed server. Local direct-attached storage (DAS) can keep frequently used data near the CPU. Network-attached storage (NAS) gives multiple servers access over standard Ethernet, while a storage area network (SAN) provides shared storage over a separate network.
A site can use more than one arrangement. None of those labels specifies the capacity or throughput available to this job.
Data location and connection
- Enter or request
- Location of frequently accessed data; actual DAS, NAS or SAN path
- Decision
- Pending if the path is unspecified
Online capacity and copies
- Enter or request
- Data that must remain readily accessible, required copies and proposed capacity
- Decision
- Consider exclusion if a required capacity or copy arrangement cannot be provided
Cluster traffic
- Enter or request
- Paths for requests, logging and data ingestion; separately, any GPU-to-GPU collective traffic
- Decision
- If collective traffic is required, pending until its path and tests are identified
Cross-site connection
- Enter or request
- Whether clusters span sites; if so, proposed path and assessment of distance, traffic, oversubscription, buffering and optics
- Decision
- Not applicable for a single-site deployment
DOE advises right-sizing storage redundancy because additional storage modules increase power use. Its discussion of moving data out of the production environment applies to data that need not be readily accessed—not to every dataset placed on NAS or SAN.
For a proposed GPU cluster, distinguish application-facing traffic from communication among GPUs. Cisco calls these its frontend and backend fabrics. Its lossless, non-blocking backend is an architecture example, not a requirement for every AI server.
Cisco also distinguishes raw RDMA session throughput, GPU-to-GPU collective-communication tests and end-to-end job measures: a result at one level does not by itself prove another. If clusters span sites, Cisco cautions that extending the backend is sensitive to loss and delay and that the design may require network tuning or workload changes. These are questions to test on the proposed path, not universal sizing rules.
Can the facility power and cool that configuration?
Ask for the proposed configuration’s initial load, expected operating and part-load conditions, and future load. Compare them with the site’s electrical distribution path and decide which equipment actually requires UPS protection. DOE notes that a scientific-computing facility and a financial institution can have different UPS needs.
Check power-supply efficiency at the loads where the server will normally run, rather than relying only on its rated maximum; redundant UPS arrangements can also operate at low load factors.
Power and continuity
- Enter or request
- Configuration loads, distribution path, equipment needing UPS protection and efficiency at expected loads
- Decision
- Pending without configuration and site evidence; consider exclusion for an unmet mandatory requirement
Air-cooled installation
- Enter or request
- Equipment class, intake and exhaust direction, rack airflow, cable obstructions and inlet-monitoring plan
- Decision
- Pending if the required equipment-inlet conditions cannot yet be demonstrated
Direct liquid cooling
- Enter or request
- Components cooled by liquid, residual air-cooling load, and cooling distribution unit (CDU) and facility-loop requirements
- Decision
- Not applicable if liquid cooling is not proposed; assess missing required connections before acceptance
Cooling capacity
- Enter or request
- Proposed heat load and site response at initial, part and future loads
- Decision
- Pending if only a cooling-method name is supplied
For air-cooled equipment, check conditions at the equipment inlet, not an unspecified room temperature. DOE reproduces the 2021 ASHRAE A1–A4 recommended inlet-air temperature range of 18–27 °C and different allowable ranges by class.
The recommended range is an operating target intended to maintain reliability; allowable boundaries concern equipment functionality. Neither is a liquid-supply specification or a substitute for the proposed equipment’s requirements. Intake and exhaust direction, separation of cool supply air from hot exhaust, cable obstructions, and temperature and humidity monitoring at inlets all affect the installation check.
A proposal for direct liquid cooling needs its own interface check. DOE notes that some implementations still leave heat for room-air cooling and that a CDU commonly connects the IT cooling circuit to the facility loop while supplying liquid at appropriate temperature, pressure and chemistry. Do not assume that every GPU server needs liquid cooling—or that naming a cooling method proves the site can handle changing loads.
Where will it run, and who controls access?
The existing rack is not the only deployment option. Record where the workload runs and its data resides, who owns the hardware, who operates it and who responds to faults. On-premises, managed, colocation and cloud arrangements divide space, hardware and operational responsibilities differently; the arrangement’s name alone does not settle suitability or cost.
Location and operations
- Enter or request
- Workload and data locations; hardware ownership, operating duties and access rights
- Decision
- Pending if a required responsibility has no assigned owner
Security and access
- Enter or request
- Required controls for hardware, storage, administrators, connections and applications
- Decision
- Consider exclusion if a mandatory control cannot be provided
IBM’s data-center security description spans physical equipment and storage, administrative access, applications and organizational policies. Network virtualization may change how resources are allocated, but it does not replace a review of those controls at the chosen location.
What counts as enough evidence to decide?
Mark a requirement met only when the proposed configuration, workload result or site check demonstrates it. Mark missing evidence pending; mark a condition not applicable with a reason, such as no cross-site connection; and consider excluding a configuration when it cannot satisfy a mandatory requirement.
Architecture descriptions cannot fill in an unknown application result, storage capacity, electrical load or cooling demand.
Keep facility and job measures separate. DOE defines power usage effectiveness (PUE) as annual facility energy divided by annual IT-equipment energy; it describes supporting-infrastructure efficiency, not this job’s completion time or throughput. The useful next step is to request evidence for each pending row rather than treating a GPU specification—or a facility-wide metric—as proof that the deployment is ready.
Sources
- blogs.cisco.com — Scale-across: Why the future of distributed AI isn’t in one data center - Cisco Blogs Search
- IBM — What Is a Data Center? | IBM
- IBM — What Is an AI Data Center? | IBM
- U.S. Department of Energy — Best Practices Guide for Energy-Efficient Data Center Design
No comments:
Post a Comment