2022-07-07
1
A 10 Gigabit Ethernet TCP/IP Stack
Implementation on MicroChip PolarFire
for High-Speed Camera Image Transport
Missing Link Electronics
Ulrich Langenbach, Andreas Schuler
2022-07-07
2
MLE - Experts for Domain-Specific Compute Architectures
Our Mission is
● to support customer projects with deep
expertise and hands-on design services
● Offering pre-validate FPGA subsystems of
FPGA IP blocks and open-source software
● Applying novel FPGA design
methodologies for increased productivity
● Partners to
Headquartered in Silicon Valley with Design
Offices in Germany
● Founded 2010, employee owned
● 17+ Certified FPGA Designers
● 50+ Presentations at Technology
Conferences, 5 Patents awarded
2022-07-07
3
Our Design Services Expertise
● RTL and High-Level Synthesis using Intel or Xilinx Toolflows
● Zynq-7000 SoC in designs since Q1/2012
● Zynq Ultrascale+ MPSoC in designs since Q4/2015
● Zynq UltraScale+ RFSoC in designs since Q2/2018
● Arria-10, Cyclone V SoC PCIe subsystems
● PetaLinux / Vanilla Linux and Yocto-based SW development
● Multigigabit transceiver configurations
○ PCIe Gen2/3/4/5, SATA 3/6G, SAS 6/12G, NVMe,
○ CAPI, JESD204B, DP/HDMI, MIPI CSI-2 D-PHY
○ 10/25/4050/100G Ethernet, Low Latency Ethernet
● Radar & Lidar for civil, mil/aero, automotive, industrial
● Image processing for HDMI, Displayport, SDI
● Time Sensitive Networking, Detnet, Layer-2 Switching
● Functional Safety Design Flows ISO 26262 (ASIL), IEC 61508 (SIL)
● Security & Trust (PUF, Crypto, OP-TEE)
2022-07-07
4
Agenda
1) Application - Camera Image Transportation
=> Why TCP/IP?
2) Microchip Polarfire Overview
3) Protocol Overview
4) TCP/IP
a) Why TCP/IP?
b) How TCP/IP Works
5) NPAP
a) Overview Stack
b) Overview ERD
c) Latency
6) NPAP Applications
2022-07-07
5
Camera Image Transport
Cameras getting more demanding in regards of Bandwidth
Preprocessing is not always possible - raw data is required
Long distance between camera and server/operator
Server / Storage
Pre-
processing
Image
Sensor
Connectivity
2022-07-07
6
Zone-Based 10 GigE Automotive Backbone
2022-07-07
7
MicroChip - PolarFire
2022-07-07
8
PolarFire - What’s inside?
2022-07-07
9
Why Polarfire?
Non-volatile FPGA fabric
Low Power
● Low device static power
● Low inrush current
● Low power transceivers
Reliability Features
● Configuration cells single event upset (SEU) immune
Security Features
● Differential Power Analysis protection
● Physical Unclonable Function
● Secure Non-volatile Memory
2022-07-07
10
Overview
2022-07-07
11
High-Speed Transceivers
https://www.microsemi.com/blog/2018/04/10/polarfire-fpga-transceivers/
10 GbE SFI 1 GbE SFI
2022-07-07
12
Market Development of Image Transport Techn.
https://www.get-cameras.com/How-to-select-a-machine-vision-camera-interface-USB3-GigE-5GigE-10GigE-Vision
2022-07-07
13
Protocol Overview
SDI
TCP/IP
A-PHY
FPDL-III
GMSL
2022-07-07
14
Protocol Overview - Wide Area > 50 m
SDI
TCP/IP
2022-07-07
15
Protocol Overview - Interoperable with IT Equ.
TCP/IP
2022-07-07
16
Protocol Overview - Interoperable with IT Equ.
TCP/IP
We do
2022-07-07
17
Why TCP/IP with cameras?
Mature Protocol - it is around for more than 40 years
De facto standard of the Internet
Guaranteed delivery, back pressure capability -> it’s a big, distributed FIFO!
Widely available commercial off-the-shelf (cots) hardware
Options to add features through additional Layers, on top or below:
● Time Sensitive Network (TSN)
● Media Access Control Security (MACsec)
● Transport Layer Security (TLS)
2022-07-07
18
TCP Facts
Layered architecture
• “Packet”-based with data segmented
into Protocol Data Units (PDU)
● TCP message – PDU at TCP layer
● Datagram – PDU at IP layer
● Frame – PDU at link-layer
● Communication is
● Reliable
● Ordered
● Error-checked
2022-07-07
19
TCP/IP Header
https://packetpushers.net/radiuid/
2022-07-07
20
TCP/IP Header
Options
https://packetpushers.net/radiuid/
Various web application
driven additions available,
e.g. via TCP Options, such
as session cookies reducing
the number of 3 way
handshakes required to
load a single web page
2022-07-07
21
How TCP works - The Handshake(s)
No, not this one
https://www.freepik.com/vectors/corona-virus-cartoon
created by brgfx - www.freepik.com</a>
2022-07-07
22
The 3 Way Handshake (establish connection)
https://afteracademy.com/blog/what-is-a-tcp-3-way-handshake-process
2022-07-07
23
The 3 Way Handshake (teardown connection)
https://afteracademy.com/blog/what-is-a-tcp-3-way-handshake-process
2022-07-07
24
The 3 Way Handshakes
1. Make sure both sides are on the same page
2. Enable both sides to detect if something got wrong
(a packet was lost)
3. Respective flags are handled as if they were a Byte of payload
=> This actually provides integrity and consistency
Data may already or still be
passed during partially
established / teared down
connections
2022-07-07
25
TCP Implementation, usually a software domain!
BUT
Dataflow processing fits best to the power, compute and space
requirements!
2022-07-07
26
NPAP
Network Protocol Acceleration Platform
2022-07-07
27
NPAP - Network Protocol Acceleration Platform
Key features:
● IPv4 with ICMP and IGMP
● TCP/UDP with AXI-S interfaces
● DHCP client
● Different speeds available (10/25/40/50/100 GE)
● Jumbo frame support
● Low latency and deterministic
● Configurable buffers for each session and direction
What is NPAP?
NPAP is a TCP/UDP/IP Full accelerator and is
operated processor independent
2022-07-07
28
NPAP - Why Platform and not IP
It is delivered as an Evaluation Reference Design
(ERD) and consists of:
● MAC (depending on speed/ FPGA technology -
eval required)
● TCP/UDP/IP full accelerator
● Control Flow
● Examples for handling
○ TCP Sessions
○ UDP
○ Stack control
● Netperf
○ Open Source Network Bandwidth Measurement tool
2022-07-07
29
NPAP - Evaluation Reference Design (ERD)
2022-07-07
30
NPAP Control Application
● AXI4-Lite Interface
● NPAP Control (IP, MAC, …)
● DHCP Control and Status
● NPAP Reset
● One IP instance per NPAP instance
2022-07-07
31
TCP Command Application
● AXI4-Lite Interface
● Example TCP Command Interface
implementation
● One IP instance per TCP session
● Controlled by TCP Demo Application
or standalone usage
2022-07-07
32
TCP Demo Application
● AXI4-Lite Interface
● Data stream control (loopback, discard, external)
● TCP Session Reset
● One IP instance per TCP session
● Controls TCP Command Application
2022-07-07
33
UDP Demo Application
● AXI4-Lite Interface
● Data stream control (loopback, discard, external)
● TUSER setting (per Datagram meta-data, e.g. source + destination ports)
● One IP instance per UDP port
2022-07-07
34
Design Flows
● Vivado Block Diagram Flow
○ Based on IPXACT packages IP cores
○ Allows for quick design generation
● Classic RTL Based Flow
○ File inclusion into project
○ Library assignment for files
○ Traditional Verilog or VHDL module instantiations
● Tool specific IP integration Flow(s)
○ Usually TCL script based
○ Adds sources and may provide interface bundles
2022-07-07
35
NPAP - Performance & Metrics
TCP Payload Size [Byte] Latency [ns]
1 462,8
10 457,4
16 485,8
64 520,0
160 656,0
448 1092,1
720 1502,7
960 1868,9
1216 2251,3
1456 2622,7
Simulation Latency Results
Testbench
DUT Wrapper 0 DUT Wrapper 1
156.25 MHz
175 MHz
175 MHz
NPAP 0 MAC 0
APPS
MAC 1 NPAP 1
APPS
XGMII
TX RX
2022-07-07
36
NPAP - Performance & Metrics
Round Trip Time
and
Throughput
2022-07-07
37
The Bandwidth-Delay-Product
● Is a metric for network system performance
● Provides an estimate for buffer sizing
=> Let’s have a closer look!
NPAP - Performance & Metrics
2022-07-07
38
Bandwidth
Node 1 Node 2
Send data 1
Node 1 Node 2
Send data
t0
t1
(Equally sized packets = 1 unit [Bit])
Send data 2
Send data
T
t0
t1
T
Bandwidth B = data quantity per time interval
= data quantity / (t1 - t0) [Bit/s]
Transfer 1: B1 = 3u / T
Transfer 2: B2 = 9u / T = 3 * B1
Transfer 1
Transfer 2
2022-07-07
39
Process
data
Delay a.k.a. RTT
Node 1 Node 2
Send data
Ack data
Node 1 Node 2
Process data
Send data
Ack data
RTT
t0
t1
RTT = t1 - t0 [s]
RTT = 2 * Latency (for symmetric systems)
2022-07-07
40
Bandwidth-Delay-Product - Low Bandwidth
Node 1 Node 2
Send data
Ack data
Process data
Node 1 Node 2
Send data
t0
t1
(Equally sized packets = 1 unit [Bit/s])
T
Transfer 1
Bandwidth-Delay-Product = Bandwidth * Delay
= B [Bit/s] * T [s]
= BDP [Bit]
2022-07-07
41
Bandwidth-Delay-Product - High Bandwidth
Node 1 Node 2
Ack data
Process data
Node 1 Node 2
(Equally sized packets = 1 unit)
Send data
t0
t1
T
Transfer 2
Send data
More data in flight during the RTT -> larger buffer
required to cover for potentially missed packets (re-
transmission buffer)
2022-07-07
42
NPAP on PolarFire
MLE NPAP Application
2022-07-07
43
Resource Utilisation
The following table shows resources synthesized for MicroSemi PolarFire
MPF300TS-1FCG1152I using Libero 2021.1 - instantiating the following design
features:
● Ethernet
● IPv4
● UDP
● 3 instances of TCP
2022-07-07
44
Challenges migrating to Microchip Polarfire FPGAs
Microchip Polarfire FPGA
● Registers cannot be initialised
during SRAM cell initialisation
(bitstream load)
● To provide a defined POR state a
Power-on-Reset is a must on this
platform!
FPGAs of other vendors
● Registers are initialised during
FPGA SRAM cell initialisation
(bitstream load) to a specific
value
a. An initial value is chosen by the tool
b. An initial value is provided by the
developer
● Power-on-Reset is nice-to-have,
but a good design practice
2022-07-07
45
Challenges migrating to Microchip Polarfire FPGAs
● RTL descriptions written for other vendors’ devices / families must not
necessarily perform similar on Microchip Polarfire devices
● Code must be carefully reviewed and re-written to implement
○ POR for all registers that define the circuit state,
○ including re-writing potentially present initial values into a POR
2022-07-07
46
NPAP Applications
2022-07-07
47
Distributed PCIe NTB
● NTB: Non-transparent
Bridge
● Prototypical
implementation based on
Xilinx ZU+ devices
2022-07-07
48
PCIe Non-Transparent Bridge
2022-07-07
49
NTB: Multi-CPU Interconnect via a Daisy-Chain
2022-07-07
50
NTB: Multi-CPU Interconnect via a Daisy-Chain
NPAP
Network
2022-07-07
51
PCIe Range Extension via TCP/IP
● Presented at PCI-SIG
Developers Conference
2018
● Results of a prototypical
implementation based on
Xilinx Z7000
● Since than a new
generation of prototypes
is available based on
Xilinx ZU+ devices
2022-07-07
52
PCIe Transport via TCP/IP
● Fully transparent to network equipment
○ Just a bunch of TCP sessions
○ No special traffic handling required
● Fully transparent to PCIe
○ Reliable transport via TCP
○ Congestion control via TCP
● A “distributed” PCIe Switch
○ In accordance to PCIe Spec
○ Scalable via TCP session count
○ Supports latency requirements for “sideband” signals
○ Special care needed to avoid deadlocks
● Independent of lower network layers, e.g. physical layer
● Could be used as PCIe range extension
2022-07-07
53
A Prototypical Implementation
● Adds visibility to the
PCIe path
● Traditional Network
tools now
applicable to PCIe,
e.g. Wireshark
2022-07-07
54
Concept of PCIe-over-TCP (1)
● Network Protocol Stack
● Encapsulation of PCIe
Transaction Layer
Packets (TLPs) into TCP
2022-07-07
55
Concept of PCIe-over-TCP (2)
NPAP
2022-07-07
56
Key-Value-Store Hardware-Acceleration
● Joint work of MLE and
Xilinx Research Ireland
● Presented at SNIA SDC
2016 and SNIA SDC 2017
● Single-Chip Solution
2022-07-07
57
Key-Value-Store Hardware-Acceleration
● Ethernet Attached
KVS Storage Node
● Fully Pipelined
● Local NVMe memory
● Local DRAM memory
● Multi-hierarchy
storage
○ Small and Fast
○ Large and “Slow”
NPAP
2022-07-07
58
Key-Value-Store Hardware-Acceleration
● Data-Flow Processing enables real on-the-fly meta-data extraction
● No additional server load
● Provides very low latency
2022-07-07
59
Linux Server w/ NVMe SSDs and Xilinx Alveo
for data storage and processing
“Edge” FPGA w/ sensors for
Data acquisition and pre-processing
TCP/UDP/IP over
10/25/50/100 GigE
LAN with
10/25/50/100 GigE
High-speed ADC / DAC
MIPI CSI-2
GMSL
FPD-III
LVDS
FPGA-Based “Edge” - Server Connectivity
2022-07-07
60
NPAC - Network Protocol Accelerator Card
2022-07-07
61
NPAC - Features
First MLE NPAC PCIe card, namely NPAC-KETCH, will be available soon:
● Targeted to Intel Stratix 10 GX 400
● Netperf and TCP-/UDP-Loopback example instances
● 4x SFP+ for 4x 10 GigE via Twinax or Fibre
● Supports Quartus design flow with High-Level Synthesis design option
● Runs on MLE NPAC-40G Cost-Optimized SmartNIC
Other device vendors on the Roadmap: Microchip, Xilinx
2022-07-07
62
Contact Information
Email contact: sales-web@missinglinkelectronics.com
Missing Link Electronics, Inc.
+1 (408) 475-1490
2880 Zanker Road, Suite 203
San Jose, CA 95134
United States
Missing Link Electronics GmbH
+49 (731) 141149-0
Industriestraße 10
89231 Neu-Ulm
Germany