Summarize

By Chris Kawalek, Director of Product Marketing, NVIDIA
Enterprises are putting AI to work in products and services, software development and business operations. According to a Gartner® survey, “75% of survey respondents said they were piloting, were deploying or had already deployed some form of AI agents into their organization.”1 AI has become a day-to-day resource that companies must manage effectively.
Many are turning to AI factories: specialized production environments for intelligence. Like other factories, they transform raw materials—energy and data—into useful output. Their success depends on five interdependent layers: energy, chips, infrastructure, models and applications.
Build the AI Factory as One Efficient Solution

Success depends not simply on how much computing capacity an organization deploys, but on how productively the complete system operates. Keeping infrastructure utilized, data available and applications responsive improves both sides of AI factory economics: Throughput expands the capacity to generate revenue and business value, while cost per token measures how economically that output is produced. The key to strong AI factory economics is to coordinate design and operations across the stack.
1. Operate and measure the AI factory as a system.
Infrastructure economics are ultimately determined in production, where efficiency, reliability and utilization decide how much useful work an organization receives from its investment.
Like traditional factories, AI factories need specialized machinery and management systems. Accelerated computing, networking, storage, software, power and cooling must work together as one production system. Otherwise, capacity can lose value through poor utilization, failed jobs, downtime and fragmented management.
GPU underutilization has a direct and material impact on AI economics, infrastructure scalability and I&O credibility. Without intervention, organizations will continue to absorb an 80%-to-85% waste tax on some of their most expensive assets, limiting AI adoption while increasing scrutiny on infrastructure spend. Gartner2
Priorities should include visibility, AI workload orchestration, system health monitoring, rapid recovery and matching resources to business needs. Power and cooling are key factors because facility decisions affect how much available energy becomes productive computing. At data-center scale, NVIDIA DGX SuperPOD™provides a blueprint for deploying AI factories, while NVIDIA Mission Control demonstrates how scheduling, monitoring and recovery can be unified in an operations layer.
2. Make data and storage part of the efficiency equation.
The NVIDIA Rubin platform is designed to deliver dramatically more tokens per watt than prior generations. NVIDIA DGX Rubin NVL8 brings that architecture into a purpose-built AI system combining NVIDIA Rubin GPUs and Intel® XeonⓇCPUs. Realizing its full potential, however, depends on keeping the system fed with data.

Earlier AI storage strategies focused on capacity and throughput for training. Agentic AI adds additional challenges: providing fast, efficient access to growing context, and preparing and retrieving enterprise data quickly enough to ground agents in accurate, relevant information. Slow data paths increase latency and reduce utilization, while unreliable retrieval can require repeated attempts that consume more tokens and energy. NVIDIA AI Data Platform addresses the enterprise data challenge by bringing accelerated computing into the data path to make enterprise data AI-ready and improve retrieval efficiency.
Enterprise data center storage will nearly triple to 23.3 ZB by 2030. Enterprises’ share of the installed base will rise from 78% in 2025 to nearly 90% in 2030 as organizations scale storage to support core AI infrastructure spending. IDC Global StorageSphere Forecast, 2026-20303
Executives should plan the data path alongside accelerated computing from the beginning, so enterprise data remains AI-ready and valuable infrastructure stays productive.
3. Place compute in the right location.
AI workload placement has two dimensions. Within the data center, the core AI factory provides centralized scale, while specialized computing closer to storage and data pipelines can reduce data movement and prepare information for models more efficiently. Beyond the data center, latency-sensitive inference may need to run at the edge, closer to users, devices and points of action.
Leaders should coordinate these locations as one environment, placing compute and models according to function, latency, data gravity, security and economics. Open models, such as NVIDIA Nemotron, provide options across a range of sizes and capabilities. Intelligent routing can direct work to the model and infrastructure best suited to each task, using larger models centrally when greater capability is required and smaller models near the edge when responsiveness matters more. NVIDIA DGX systems can anchor the core AI factory, while a common software foundation helps organizations develop, deploy and run these open models consistently across locations.
4. Use the ecosystem to accelerate deployment.
Even as AI systems improve performance per watt, rising rack power density is changing facility requirements. High-density systems increasingly use liquid cooling to remove heat effectively, support greater compute density and reduce cooling overhead. Because many enterprise data centers were designed for lower-density, air-cooled equipment, deploying these systems may require facility upgrades or alternative deployment approaches.
An established ecosystem gives enterprises multiple paths forward. The NVIDIA DGX-Ready Colocation Data Center program connects enterprises with certified providers offering liquid-cooling-ready capacity, while mechanical, electrical and plumbing specialists; system integrators; and deployment partners can adapt facilities and coordinate installation and commissioning. Bringing these partners into planning early gives decision-makers the flexibility to modernize an existing facility, use colocation capacity or combine both approaches.
Treating facility readiness as part of the AI factory architecture from the beginning can reduce redesign, accelerate deployment and help the infrastructure operate at its intended efficiency sooner.
Putting the Four Pillars to Work
Enterprise AI factories will span core systems, data infrastructure, edge deployments and partner-operated capacity. What should not vary is the operating discipline: measure useful output, keep data moving, place workloads deliberately and plan facilities early.
The next era of enterprise AI will not be defined by computing capacity alone, but by how organizations productively operate their AI factories, converting energy and data into useful intelligence reliably, efficiently and at scale.
Sources:
- Gartner® Press Release, Gartner Survey Finds Just 15% of IT Application Leaders Are Considering, Piloting, or Deploying Fully Autonomous AI Agents, Sept 30, 2025. GARTNER is a trademark of Gartner, Inc. and/or its affiliates. https://www.gartner.com/en/newsroom/press-releases/2025-09-30-gartner-survey-finds-just-15-percent-of-it-application-leaders-are-considering-piloting-or-deploying-fully-autonomous-ai-agents
- Gartner®, The 2026 GPU Optimization Playbook, May 15, 2026.
- IDC, Market Forecast: IDC Global StorageSphere Forecast, 2026-2030 (Doc #US53425526, June 2026)
