Dr. Sanjay Basu Explores Why the Inference Era Will Be Won with the Cheapest Millisecond

Drawing on his Chicago keynote, Dr. Basu examines how networking, power, and proximity to users are redefining the design and economics of AI inference infrastructure.

In a new article reflecting on his keynote at The Connected World LIVE! in Chicago, Dr. Sanjay Basu examines how the infrastructure requirements of artificial intelligence are changing as the industry moves from training increasingly large models to delivering inference at scale.

The central question, he argues, is no longer simply, “How large a model can we build?” It is now, “What does the next token cost, and how far away is the person waiting for it?”

This shift fundamentally changes how AI infrastructure must be designed, operated and measured. While the training era prioritized the size and performance of compute clusters, inference introduces a different set of demands: continuous availability, millisecond-level utilization, unpredictable global traffic and stringent latency expectations that directly affect the customer experience.

Dr. Basu explains that the economics of inference cannot be determined by processing power alone. They are shaped by the interaction of compute, network architecture, power availability, caching, workload scheduling and geography. As prefill and decode workloads become increasingly disaggregated, the movement of the KV cache places the network at the heart of AI performance—transforming what may appear to be a memory challenge into a networking challenge.

The article also distinguishes between the three fabrics supporting modern AI infrastructure: scale-up within the rack, scale-out across the data hall and scale-across between metropolitan markets. Together, these layers determine not only how efficiently a model can operate, but also how quickly its output reaches the person waiting for it.

At the heart of Dr. Basu’s argument is the distinction between throughput and goodput. Tokens only create value when they are delivered within the required service-level objective. For infrastructure providers and operators, this means that tokens delivered within latency targets—not theoretical peak performance—must become the true measure of capacity and economic yield.

He identifies five principles for infrastructure leaders navigating this transition:

  • Plan in megawatts, but measure in tokens.
  • Report goodput rather than throughput.
  • Engineer for tail latency.
  • Treat the KV cache as infrastructure.
  • Design for continuous, 24x7 operation.

Drawing on the discussion in Chicago, which brought together perspectives from Oracle, NVIDIA, EdgeCore and InterGlobix, Dr. Basu demonstrates why silicon, data centers and interconnection can no longer be optimized in isolation. Each controls a different part of the same latency budget—and it is the combined performance of the system that ultimately determines what the customer experiences.

As Dr. Basu concludes: “The training era was won with the largest cluster. The inference era will be won with the cheapest millisecond.”

Read Dr. Sanjay Basu’s complete article and insights here.

READ THE FULL ARTICLE

Amplified by InterGlobix News Embargo Digital Marketing

This announcement is being amplified by InterGlobix as part of its News Embargo Digital Marketing Amplification Services, supporting strategic visibility for major developments across the global digital infrastructure, AI, cloud, connectivity, and data center ecosystem.

For more information:
InterGlobix News Embargo Services

Go Top