05-07-26: Today, I caught up with an old friend, Ron Renwick, and the team from Xelera Felix Winterstein, the CEO, and Andrea Suardi, Head of Acceleration, to discuss their new Silva offering and how they are bringing AI inference to the network edge. Recently, the Xelera team completed the STAC-ML™ Markets (Inference) benchmark audit on a stack that includes a STAC-ML™ Pack for Xelera Silva with AMD Alveo™ V80 on an HPE Proliant DL385 Gen10 Plus v2 server. For context, here are some highlights from this report: For the small (GBT_A) and medium (GBT_B) models, 99th percentile latencies were 560K inferences per second at the highest NMIs tested For the large (GBT_C) model, the 99th percentile latency was 2.88µs, with worst-case instance throughput of 379K inferences per second The maximum latency was = 12.3µs across all models and NMI tested Table of Contents 00:00 Hello & Welcome 00:09 Ron Background 01:04 Felix Background 01:40 Andrea Background 02:30 Xelera Origin Story 06:00 FPGAs are very sticky once you start working with them 07:33 What is Xelera, what do they do? 07:55 It’s not about FPGAs 08:30 It’s network acceleration for the data center 08:40 We are seeing a massive buildout in data center capacity 09:05 Number one KPI (Key Performance Indicator) is compute capacity 10:31 All that processing is starting to reach its limits 10:50 In cybersecurity and networking, we see this lack of computing becoming painful 11:10 The thing every infrastructure buyer or architect needs to consider 11:24 What are the products that Xelera offers, and who do they address? 11:45 First is a classic DPU softNIC 12:05 Strong footprint in Cyber Security vertical 12:26 Second product is AI Acceleration 12:48 We bring Andrea back in to talk about his STAC Research Event Talk 13:40 When you attend STAC events, you’re in a room with really deeply technical people 14:00 All these people are here to solve one problem: optimize their execution stack 14:57 Tail latency spikes are discussed, and the impact on trading 15:17 Ron drills into this tail latency issue a bit more 16:12 The cost of latency varies from firm to firm 16:45 Too slow, and you are a price taker, not a price maker 17:13 It’s important to win more than 50.001% of the time 17:24 What we’re selling today is the ability to trade both faster and smarter 19:40 Gradient boosting 50 us running in the CPU, went down to 5us running on the FPGA 20:10 This became Silvia, an agent who runs fast. 20:30 What other markets can Silva be applied to, for example, security 22:00 Silva can look at each packet at the edge, for Ransomware or DDoS, pattern detection 23:00 Silva is the engine, the model can vary depending on the use case, time 23:30 What numbers can you provide? 24:00 Two main Silva modes, the first is Offload, inference via API with one 2M nodes 25:10 Now 30us to run gradient boosting on CPU, if you use Silva, it’s 1us on FPGA 25:40 LSTN with a 1M parameter model is 1ms. If you offload to an FPGA, you get 3us 26:10 There is a cost using the PCIe bus, 500-700ns depending on packet size 26:40 The second Silva mode is Inline mode, and this runs entirely on the FPGA 27:10 Are there specific use cases or environments driving deployment environment 28:15 The lowest latency requires optimize the stack; cloud won’t work 29:00 We rewrote the CPU kernel to tune the cache and have it work with Gradient Boosting Chapters (00:00:03) - Scott Schweitzer vs Ron Renwick(00:01:00) - Zelera CEO and COO on TechCrunch Disrupt(00:01:44) - Xelera's origin story(00:08:05) - Nvidia Network Accelerator: Data Center, Network(00:11:21) - Celera provides two product lines for network acceleration and AI(00:12:50) - Determining ultra-low latency Inferring for Trading(00:15:19) - Inference and the Latency challenge(00:20:24) - How SYLVIA works in cybersecurity & DDoS(00:27:13) - Inclination for CPUs and FPGAs(00:30:29) - What are the advantages of SmartNICs?(00:33:17) - Silva Networks: Who Do You Talk To About Network Security?(00:35:53) - SmartNIC and Silver: Roadmap(00:39:25) - A Taste of Accelera's Future