John Ousterhout, Stanford Professor and author of “A Philosophy of Software Design”, turns our attention to the evolving nature of AI networking workloads and why traditional protocols like TCP and RDMA are becoming bottlenecks in modern data center environments!
The Shift in Workloads
Historical context:
AI traffic was dominated by massive, long-running transfers (gigabytes of gradients), where throughput was the primary metric.
Modern AI:
Workloads, especially inference and agentic applications, now rely on frequent, small coordination messages (e.g., KV cache lookups, barrier synchronization). These small messages are highly sensitive to latency. The Bottleneck: When small synchronization messages are mixed with large traffic, they get trapped in queues (caused by incast), significantly increasing 99th percentile (tail) latency. This causes GPUs to sit idle, wasting expensive compute resources.
Why Legacy Protocols Struggle:
Sender-Driven Congestion Control: TCP and RDMA rely on the sender to detect congestion, often via packet drops or delayed signals from switches. This process is inherently reactive and oscillates, leading to unstable performance.
Byte Stream Model:
These protocols view data as an opaque stream of bytes rather than discrete messages, making it difficult to prioritize short, critical tasks.
The Homa Solution
John introduces Homa, a clean-slate transport protocol designed for data centers:
Message-Based:
Unlike byte streams, Homa understands message boundaries, allowing it to predict traffic and prioritize short messages using Shortest Remaining Processing Time (SRPT).
Receiver-Driven:
The receiver controls the flow by issuing grants to senders, effectively managing congestion before it occurs at the switch.
Priority Queues:
Homa leverages the multiple hardware queues already present in modern switches to bypass long, queued traffic with low-latency short messages.
Performance Results:
Benchmarks show Homa can reduce tail latency for short messages by over 10x compared to TCP, while simultaneously improving performance for large messages.
Learn More about HOMA Technology?

