2022-07-07
1
A 10 Gigabit Ethernet TCP/IP Stack
Implementation on MicroChip PolarFire
for High-Speed Camera Image Transport
Missing Link Electronics
Ulrich Langenbach, Andreas Schuler
2022-07-07
2
MLE - Experts for Domain-Specific Compute Architectures
Our Mission is
to support customer projects with deep
expertise and hands-on design services
Offering pre-validate FPGA subsystems of
FPGA IP blocks and open-source software
Applying novel FPGA design
methodologies for increased productivity
Partners to
Headquartered in Silicon Valley with Design
Offices in Germany
Founded 2010, employee owned
17+ Certified FPGA Designers
50+ Presentations at Technology
Conferences, 5 Patents awarded
2022-07-07
3
Our Design Services Expertise
RTL and High-Level Synthesis using Intel or Xilinx Toolflows
Zynq-7000 SoC in designs since Q1/2012
Zynq Ultrascale+ MPSoC in designs since Q4/2015
Zynq UltraScale+ RFSoC in designs since Q2/2018
Arria-10, Cyclone V SoC PCIe subsystems
PetaLinux / Vanilla Linux and Yocto-based SW development
Multigigabit transceiver configurations
PCIe Gen2/3/4/5, SATA 3/6G, SAS 6/12G, NVMe,
CAPI, JESD204B, DP/HDMI, MIPI CSI-2 D-PHY
10/25/4050/100G Ethernet, Low Latency Ethernet
Radar & Lidar for civil, mil/aero, automotive, industrial
Image processing for HDMI, Displayport, SDI
Time Sensitive Networking, Detnet, Layer-2 Switching
Functional Safety Design Flows ISO 26262 (ASIL), IEC 61508 (SIL)
Security & Trust (PUF, Crypto, OP-TEE)
2022-07-07
4
Agenda
1) Application - Camera Image Transportation
=> Why TCP/IP?
2) Microchip Polarfire Overview
3) Protocol Overview
4) TCP/IP
a) Why TCP/IP?
b) How TCP/IP Works
5) NPAP
a) Overview Stack
b) Overview ERD
c) Latency
6) NPAP Applications
2022-07-07
5
Camera Image Transport
Cameras getting more demanding in regards of Bandwidth
Preprocessing is not always possible - raw data is required
Long distance between camera and server/operator
Server / Storage
Pre-
processing
Image
Sensor
Connectivity
2022-07-07
6
Zone-Based 10 GigE Automotive Backbone
2022-07-07
7
MicroChip - PolarFire
2022-07-07
8
PolarFire - What’s inside?
2022-07-07
9
Why Polarfire?
Non-volatile FPGA fabric
Low Power
Low device static power
Low inrush current
Low power transceivers
Reliability Features
Configuration cells single event upset (SEU) immune
Security Features
Differential Power Analysis protection
Physical Unclonable Function
Secure Non-volatile Memory
2022-07-07
10
Overview
2022-07-07
11
High-Speed Transceivers
https://www.microsemi.com/blog/2018/04/10/polarfire-fpga-transceivers/
10 GbE SFI 1 GbE SFI
2022-07-07
12
Market Development of Image Transport Techn.
https://www.get-cameras.com/How-to-select-a-machine-vision-camera-interface-USB3-GigE-5GigE-10GigE-Vision
2022-07-07
13
Protocol Overview
SDI
TCP/IP
A-PHY
FPDL-III
GMSL
2022-07-07
14
Protocol Overview - Wide Area > 50 m
SDI
TCP/IP
2022-07-07
15
Protocol Overview - Interoperable with IT Equ.
TCP/IP
2022-07-07
16
Protocol Overview - Interoperable with IT Equ.
TCP/IP
We do
2022-07-07
17
Why TCP/IP with cameras?
Mature Protocol - it is around for more than 40 years
De facto standard of the Internet
Guaranteed delivery, back pressure capability -> it’s a big, distributed FIFO!
Widely available commercial off-the-shelf (cots) hardware
Options to add features through additional Layers, on top or below:
Time Sensitive Network (TSN)
Media Access Control Security (MACsec)
Transport Layer Security (TLS)
2022-07-07
18
TCP Facts
Layered architecture
• “Packet”-based with data segmented
into Protocol Data Units (PDU)
TCP message – PDU at TCP layer
Datagram – PDU at IP layer
Frame – PDU at link-layer
Communication is
Reliable
Ordered
Error-checked
2022-07-07
19
TCP/IP Header
https://packetpushers.net/radiuid/
2022-07-07
20
TCP/IP Header
Options
https://packetpushers.net/radiuid/
Various web application
driven additions available,
e.g. via TCP Options, such
as session cookies reducing
the number of 3 way
handshakes required to
load a single web page
2022-07-07
21
How TCP works - The Handshake(s)
No, not this one
https://www.freepik.com/vectors/corona-virus-cartoon
created by brgfx - www.freepik.com</a>
2022-07-07
22
The 3 Way Handshake (establish connection)
https://afteracademy.com/blog/what-is-a-tcp-3-way-handshake-process
2022-07-07
23
The 3 Way Handshake (teardown connection)
https://afteracademy.com/blog/what-is-a-tcp-3-way-handshake-process
2022-07-07
24
The 3 Way Handshakes
1. Make sure both sides are on the same page
2. Enable both sides to detect if something got wrong
(a packet was lost)
3. Respective flags are handled as if they were a Byte of payload
=> This actually provides integrity and consistency
Data may already or still be
passed during partially
established / teared down
connections
2022-07-07
25
TCP Implementation, usually a software domain!
BUT
Dataflow processing fits best to the power, compute and space
requirements!
2022-07-07
26
NPAP
Network Protocol Acceleration Platform
2022-07-07
27
NPAP - Network Protocol Acceleration Platform
Key features:
IPv4 with ICMP and IGMP
TCP/UDP with AXI-S interfaces
DHCP client
Different speeds available (10/25/40/50/100 GE)
Jumbo frame support
Low latency and deterministic
Configurable buffers for each session and direction
What is NPAP?
NPAP is a TCP/UDP/IP Full accelerator and is
operated processor independent
2022-07-07
28
NPAP - Why Platform and not IP
It is delivered as an Evaluation Reference Design
(ERD) and consists of:
MAC (depending on speed/ FPGA technology -
eval required)
TCP/UDP/IP full accelerator
Control Flow
Examples for handling
TCP Sessions
UDP
Stack control
Netperf
Open Source Network Bandwidth Measurement tool
2022-07-07
29
NPAP - Evaluation Reference Design (ERD)
2022-07-07
30
NPAP Control Application
AXI4-Lite Interface
NPAP Control (IP, MAC, …)
DHCP Control and Status
NPAP Reset
One IP instance per NPAP instance
2022-07-07
31
TCP Command Application
AXI4-Lite Interface
Example TCP Command Interface
implementation
One IP instance per TCP session
Controlled by TCP Demo Application
or standalone usage
2022-07-07
32
TCP Demo Application
AXI4-Lite Interface
Data stream control (loopback, discard, external)
TCP Session Reset
One IP instance per TCP session
Controls TCP Command Application
2022-07-07
33
UDP Demo Application
AXI4-Lite Interface
Data stream control (loopback, discard, external)
TUSER setting (per Datagram meta-data, e.g. source + destination ports)
One IP instance per UDP port
2022-07-07
34
Design Flows
Vivado Block Diagram Flow
Based on IPXACT packages IP cores
Allows for quick design generation
Classic RTL Based Flow
File inclusion into project
Library assignment for files
Traditional Verilog or VHDL module instantiations
Tool specific IP integration Flow(s)
Usually TCL script based
Adds sources and may provide interface bundles
2022-07-07
35
NPAP - Performance & Metrics
TCP Payload Size [Byte] Latency [ns]
1 462,8
10 457,4
16 485,8
64 520,0
160 656,0
448 1092,1
720 1502,7
960 1868,9
1216 2251,3
1456 2622,7
Simulation Latency Results
Testbench
DUT Wrapper 0 DUT Wrapper 1
156.25 MHz
175 MHz
175 MHz
NPAP 0 MAC 0
APPS
MAC 1 NPAP 1
APPS
XGMII
TX RX
2022-07-07
36
NPAP - Performance & Metrics
Round Trip Time
and
Throughput
2022-07-07
37
The Bandwidth-Delay-Product
Is a metric for network system performance
Provides an estimate for buffer sizing
=> Let’s have a closer look!
NPAP - Performance & Metrics
2022-07-07
38
Bandwidth
Node 1 Node 2
Send data 1
Node 1 Node 2
Send data
t0
t1
(Equally sized packets = 1 unit [Bit])
Send data 2
Send data
T
t0
t1
T
Bandwidth B = data quantity per time interval
= data quantity / (t1 - t0) [Bit/s]
Transfer 1: B1 = 3u / T
Transfer 2: B2 = 9u / T = 3 * B1
Transfer 1
Transfer 2
2022-07-07
39
Process
data
Delay a.k.a. RTT
Node 1 Node 2
Send data
Ack data
Node 1 Node 2
Process data
Send data
Ack data
RTT
t0
t1
RTT = t1 - t0 [s]
RTT = 2 * Latency (for symmetric systems)
2022-07-07
40
Bandwidth-Delay-Product - Low Bandwidth
Node 1 Node 2
Send data
Ack data
Process data
Node 1 Node 2
Send data
t0
t1
(Equally sized packets = 1 unit [Bit/s])
T
Transfer 1
Bandwidth-Delay-Product = Bandwidth * Delay
= B [Bit/s] * T [s]
= BDP [Bit]
2022-07-07
41
Bandwidth-Delay-Product - High Bandwidth
Node 1 Node 2
Ack data
Process data
Node 1 Node 2
(Equally sized packets = 1 unit)
Send data
t0
t1
T
Transfer 2
Send data
More data in flight during the RTT -> larger buffer
required to cover for potentially missed packets (re-
transmission buffer)
2022-07-07
42
NPAP on PolarFire
MLE NPAP Application
2022-07-07
43
Resource Utilisation
The following table shows resources synthesized for MicroSemi PolarFire
MPF300TS-1FCG1152I using Libero 2021.1 - instantiating the following design
features:
Ethernet
IPv4
UDP
3 instances of TCP
2022-07-07
44
Challenges migrating to Microchip Polarfire FPGAs
Microchip Polarfire FPGA
Registers cannot be initialised
during SRAM cell initialisation
(bitstream load)
To provide a defined POR state a
Power-on-Reset is a must on this
platform!
FPGAs of other vendors
Registers are initialised during
FPGA SRAM cell initialisation
(bitstream load) to a specific
value
a. An initial value is chosen by the tool
b. An initial value is provided by the
developer
Power-on-Reset is nice-to-have,
but a good design practice
2022-07-07
45
Challenges migrating to Microchip Polarfire FPGAs
RTL descriptions written for other vendors’ devices / families must not
necessarily perform similar on Microchip Polarfire devices
Code must be carefully reviewed and re-written to implement
POR for all registers that define the circuit state,
including re-writing potentially present initial values into a POR
2022-07-07
46
NPAP Applications
2022-07-07
47
Distributed PCIe NTB
NTB: Non-transparent
Bridge
Prototypical
implementation based on
Xilinx ZU+ devices
2022-07-07
48
PCIe Non-Transparent Bridge
2022-07-07
49
NTB: Multi-CPU Interconnect via a Daisy-Chain
2022-07-07
50
NTB: Multi-CPU Interconnect via a Daisy-Chain
NPAP
Network
2022-07-07
51
PCIe Range Extension via TCP/IP
Presented at PCI-SIG
Developers Conference
2018
Results of a prototypical
implementation based on
Xilinx Z7000
Since than a new
generation of prototypes
is available based on
Xilinx ZU+ devices
2022-07-07
52
PCIe Transport via TCP/IP
Fully transparent to network equipment
Just a bunch of TCP sessions
No special traffic handling required
Fully transparent to PCIe
Reliable transport via TCP
Congestion control via TCP
A “distributed” PCIe Switch
In accordance to PCIe Spec
Scalable via TCP session count
Supports latency requirements for “sideband” signals
Special care needed to avoid deadlocks
Independent of lower network layers, e.g. physical layer
Could be used as PCIe range extension
2022-07-07
53
A Prototypical Implementation
Adds visibility to the
PCIe path
Traditional Network
tools now
applicable to PCIe,
e.g. Wireshark
2022-07-07
54
Concept of PCIe-over-TCP (1)
Network Protocol Stack
Encapsulation of PCIe
Transaction Layer
Packets (TLPs) into TCP
2022-07-07
55
Concept of PCIe-over-TCP (2)
NPAP
2022-07-07
56
Key-Value-Store Hardware-Acceleration
Joint work of MLE and
Xilinx Research Ireland
Presented at SNIA SDC
2016 and SNIA SDC 2017
Single-Chip Solution
2022-07-07
57
Key-Value-Store Hardware-Acceleration
Ethernet Attached
KVS Storage Node
Fully Pipelined
Local NVMe memory
Local DRAM memory
Multi-hierarchy
storage
Small and Fast
Large and “Slow”
NPAP
2022-07-07
58
Key-Value-Store Hardware-Acceleration
Data-Flow Processing enables real on-the-fly meta-data extraction
No additional server load
Provides very low latency
2022-07-07
59
Linux Server w/ NVMe SSDs and Xilinx Alveo
for data storage and processing
“Edge” FPGA w/ sensors for
Data acquisition and pre-processing
TCP/UDP/IP over
10/25/50/100 GigE
LAN with
10/25/50/100 GigE
High-speed ADC / DAC
MIPI CSI-2
GMSL
FPD-III
LVDS
FPGA-Based “Edge” - Server Connectivity
2022-07-07
60
NPAC - Network Protocol Accelerator Card
2022-07-07
61
NPAC - Features
First MLE NPAC PCIe card, namely NPAC-KETCH, will be available soon:
Targeted to Intel Stratix 10 GX 400
Netperf and TCP-/UDP-Loopback example instances
4x SFP+ for 4x 10 GigE via Twinax or Fibre
Supports Quartus design flow with High-Level Synthesis design option
Runs on MLE NPAC-40G Cost-Optimized SmartNIC
Other device vendors on the Roadmap: Microchip, Xilinx
2022-07-07
62
Contact Information
Email contact: sales-web@missinglinkelectronics.com
Missing Link Electronics, Inc.
+1 (408) 475-1490
2880 Zanker Road, Suite 203
San Jose, CA 95134
United States
Missing Link Electronics GmbH
+49 (731) 141149-0
Industriestraße 10
89231 Neu-Ulm
Germany