1 | ©2023 SNIA. All Rights Reserved.
Virtual Conference
September 28-29, 2021
How Bad is TCP?
(And What Are the Alternatives?)
Endric Schubert, Ph.D. endric@mle.biz
2 | ©2023 SNIA. All Rights Reserved.
MLE Mission: “From Software to Silicon!”
High-Performance (Embedded) Compute & Connected Systems-of-Systems need “Offload Engines” for
better performance, lower and deterministic latency and improved energy efficiency.
Focus on standards such as PCIe, NVMe, Ethernet, TCP/UDP/IP, TSN.
3 | ©2023 SNIA. All Rights Reserved.
Multi-Gigabit Real-Time Networking
Market & Technology Forces
6G Radio Integrated
Communication and
Sensing (ICAS)
Zone-Based In-Vehicle
Networking (Auto/TSN)
100G Real-time Backbone
for Virtualized PLC
(Robo/TSN)
4 | ©2023 SNIA. All Rights Reserved.
Work Motivation
Systems-of-systems
● Tightly-coupled: i.e. distributed processing with microservices
● Loosely-connected via networks (which continuously are the bottleneck)
Need to optimize
● for power / energy efficiency
● for throughput
● for (deterministic) latency and real-time delivery
Domain-Specific Architectures:
● “Offload” (protocol) processing software onto silicon
but adhere to (defacto) standards and APIs
● Make networks more deterministic and Time-Sensitive
5 | ©2023 SNIA. All Rights Reserved.
Presentation Outline
▪ How Bad is TCP?
▪ Computational Burden
▪ Tail-end Latencies
▪ What Are the Alternatives?
▪ Homa from John Ousterhout’s team at Stanford University
▪ QRP (Quad R P) - a Hardware Accelerated Version of Homa
6 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
7 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
8 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
9 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
10 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
11 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
12 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
13 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
14 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
15 | ©2023 SNIA. All Rights Reserved.
The Computational Burden of TCP
Benchmark TCP Point-to-Point with Netperf 2.6
● TCP_STREAM (and TCP_MAERTS) for
throughput in Gbps
● CPU Load on Tx and Rx side
● TCP_RTT for Round-Trip-Time
● Efficiency = Throughput / CPU Load
System Setup
● AMD FPGA w/ NPAP-25G
(Fraunhofer HHI’s TCP/UDP/IP Full Accelerator)
● 25 GigE NIC: Mellanox ConnectX-4 LX
● CPU: Intel(R) Xeon(R) e5-1620 v0 @ 3.60GHz
1 socket, 4 cores, 1 thread per core
RAM: 32G (4x8G) (Samsung, DDR3, 1600 MT/s,
syn reg)
Effects of different Linux kernels
● Vanilla 4.19 vs vanilla 5.15. vs vanilla 6.3.6
● Vanilla 4.19 vs Centos 4.18
Experimental Setup
16 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
17 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
Tx Side Rx Side
18 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
Tx Side Rx Side
19 | ©2023 SNIA. All Rights Reserved.
TCP_RR Results for Different Linux kernels
Average RTT
of FPGA Full
Accelerator clearly
outperforms (Linux)
software
20 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
21 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
22 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
23 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
24 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
25 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
26 | ©2023 SNIA. All Rights Reserved.
Wireshark Dissector for Homa
27 | ©2023 SNIA. All Rights Reserved.
NPAP - Network Protocol Accelerator Platform
● Full Accelerator = no CPUs, no software
● Standard IEEE Ethernet PHYs with RMII,
GMII, XGMII, etc via PCS/PMA via
ASIC/FPGA Ethernet Subsystem
● Ethernet, ARP, IPv4, ICMPv4, IGMPv4, UDP
& TCP, DHCP
● Optional TSN, optional TLS
● Datapath via AXI4-Stream 128-bit
● Complete stack uses generic VHDL code
● In production use for automotive, aero &
defense, industrial test & measurement, telco
applications
28 | ©2023 SNIA. All Rights Reserved.
QuadRP - Reliable, Rapid Request-Response Protocol
● Based on Homa
● Implemented within NPAP
Tested and proven Ethernet and IPv4
● Complements TCP/IP and UDP/IP
⇒ Best of both worlds:
● No CPU load
● Very low, deterministic latency
● Option for handling messages in
○ Programmable Logic, or
○ in Linux software
QRP
29 | ©2023 SNIA. All Rights Reserved.
Homa’s Benefits for NVMe-over-Fabric
LAN is the bottleneck already, with now additional burden from heavy SAN traffic.
Homa latencies can be 100x faster that TCP and promises to put less load on the network.
NVMe eliminated the legacy software overhead and uses fast PCIe Posted Writes for better response times and IOPS.
So, is TCP then the proper foundation for NVMe-over-Fabric?
NVMe-over-Homa can be a
drop-in replacement or an add-on,
achieving storage latencies close
to DAS performance,
with less overhead on network
and servers.
30 | ©2023 SNIA. All Rights Reserved.
HOMA References
[1] John Ousterhout, Stanford University: https://web.stanford.edu/~ouster/cgi-bin/papers/replaceTcp.pdf
[2] John Ousterhout’s presentation at USENIX ATC’21 (15 minutes)
https://www.usenix.org/conference/atc21/presentation/ousterhout
[3] Montazeri’s presentation at SIGCOMM18 (starts at 1:22)
https://www.youtube.com/watch?v=o_sg1nnN2bQ&t=4927s
https://conferences.sigcomm.org/sigcomm/2018/files/slides/paper_4.4.pptx
[4] Homa Linux kernel module implementation
https://github.com/PlatformLab/HomaModule
[5] Montazeri’s PhD dissertation
http://purl.stanford.edu/sp122ms2496
31 | ©2023 SNIA. All Rights Reserved.
Please take a moment to rate this session.
Your feedback is important to us.