1 | ©2023 SNIA. All Rights Reserved.
Virtual Conference
September 28-29, 2021
How Bad is TCP?
(And What Are the Alternatives?)
Endric Schubert, Ph.D. endric@mle.biz
2 | ©2023 SNIA. All Rights Reserved.
MLE Mission: “From Software to Silicon!”
High-Performance (Embedded) Compute & Connected Systems-of-Systems need “Offload Engines” for
better performance, lower and deterministic latency and improved energy efficiency.
Focus on standards such as PCIe, NVMe, Ethernet, TCP/UDP/IP, TSN.
3 | ©2023 SNIA. All Rights Reserved.
Multi-Gigabit Real-Time Networking
Market & Technology Forces
6G Radio Integrated
Communication and
Sensing (ICAS)
Zone-Based In-Vehicle
Networking (Auto/TSN)
100G Real-time Backbone
for Virtualized PLC
(Robo/TSN)
4 | ©2023 SNIA. All Rights Reserved.
Work Motivation
Systems-of-systems
Tightly-coupled: i.e. distributed processing with microservices
Loosely-connected via networks (which continuously are the bottleneck)
Need to optimize
for power / energy efficiency
for throughput
for (deterministic) latency and real-time delivery
Domain-Specific Architectures:
“Offload” (protocol) processing software onto silicon
but adhere to (defacto) standards and APIs
Make networks more deterministic and Time-Sensitive
5 | ©2023 SNIA. All Rights Reserved.
Presentation Outline
How Bad is TCP?
Computational Burden
Tail-end Latencies
What Are the Alternatives?
Homa from John Ousterhout’s team at Stanford University
QRP (Quad R P) - a Hardware Accelerated Version of Homa
6 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
7 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
8 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
9 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
10 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
11 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
12 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
13 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
14 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
15 | ©2023 SNIA. All Rights Reserved.
The Computational Burden of TCP
Benchmark TCP Point-to-Point with Netperf 2.6
TCP_STREAM (and TCP_MAERTS) for
throughput in Gbps
CPU Load on Tx and Rx side
TCP_RTT for Round-Trip-Time
Efficiency = Throughput / CPU Load
System Setup
AMD FPGA w/ NPAP-25G
(Fraunhofer HHI’s TCP/UDP/IP Full Accelerator)
25 GigE NIC: Mellanox ConnectX-4 LX
CPU: Intel(R) Xeon(R) e5-1620 v0 @ 3.60GHz
1 socket, 4 cores, 1 thread per core
RAM: 32G (4x8G) (Samsung, DDR3, 1600 MT/s,
syn reg)
Effects of different Linux kernels
Vanilla 4.19 vs vanilla 5.15. vs vanilla 6.3.6
Vanilla 4.19 vs Centos 4.18
Experimental Setup
16 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
17 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
Tx Side Rx Side
18 | ©2023 SNIA. All Rights Reserved.
TCP_STREAM Results for Different Linux kernels
Tx Side Rx Side
19 | ©2023 SNIA. All Rights Reserved.
TCP_RR Results for Different Linux kernels
Average RTT
of FPGA Full
Accelerator clearly
outperforms (Linux)
software
20 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
21 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
22 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
23 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
24 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
25 | ©2023 SNIA. All Rights Reserved.
Courtesy of John Ousterhout, Stanford University
26 | ©2023 SNIA. All Rights Reserved.
Wireshark Dissector for Homa
27 | ©2023 SNIA. All Rights Reserved.
NPAP - Network Protocol Accelerator Platform
Full Accelerator = no CPUs, no software
Standard IEEE Ethernet PHYs with RMII,
GMII, XGMII, etc via PCS/PMA via
ASIC/FPGA Ethernet Subsystem
Ethernet, ARP, IPv4, ICMPv4, IGMPv4, UDP
& TCP, DHCP
Optional TSN, optional TLS
Datapath via AXI4-Stream 128-bit
Complete stack uses generic VHDL code
In production use for automotive, aero &
defense, industrial test & measurement, telco
applications
28 | ©2023 SNIA. All Rights Reserved.
QuadRP - Reliable, Rapid Request-Response Protocol
Based on Homa
Implemented within NPAP
Tested and proven Ethernet and IPv4
Complements TCP/IP and UDP/IP
Best of both worlds:
No CPU load
Very low, deterministic latency
Option for handling messages in
Programmable Logic, or
in Linux software
QRP
29 | ©2023 SNIA. All Rights Reserved.
Homa’s Benefits for NVMe-over-Fabric
LAN is the bottleneck already, with now additional burden from heavy SAN traffic.
Homa latencies can be 100x faster that TCP and promises to put less load on the network.
NVMe eliminated the legacy software overhead and uses fast PCIe Posted Writes for better response times and IOPS.
So, is TCP then the proper foundation for NVMe-over-Fabric?
NVMe-over-Homa can be a
drop-in replacement or an add-on,
achieving storage latencies close
to DAS performance,
with less overhead on network
and servers.
30 | ©2023 SNIA. All Rights Reserved.
HOMA References
[1] John Ousterhout, Stanford University: https://web.stanford.edu/~ouster/cgi-bin/papers/replaceTcp.pdf
[2] John Ousterhout’s presentation at USENIX ATC’21 (15 minutes)
https://www.usenix.org/conference/atc21/presentation/ousterhout
[3] Montazeri’s presentation at SIGCOMM18 (starts at 1:22)
https://www.youtube.com/watch?v=o_sg1nnN2bQ&t=4927s
https://conferences.sigcomm.org/sigcomm/2018/files/slides/paper_4.4.pptx
[4] Homa Linux kernel module implementation
https://github.com/PlatformLab/HomaModule
[5] Montazeri’s PhD dissertation
http://purl.stanford.edu/sp122ms2496
31 | ©2023 SNIA. All Rights Reserved.
Please take a moment to rate this session.
Your feedback is important to us.