Why Hardware Still Matters - CTO Advisor Virtual Seminar

50:15 · Watch on YouTube ↗

Transcript 7,855 words · about 52 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

hey it's keith townsend principal of the cto visor and welcome to the cto advisor virtual seminar what makes this a virtual seminar versus a webinar well one there's a chat next to the screen next to this video try to convert with me and unlike a webinar where you don't know if there's really other participants of your peers there's a bunch of crosstalk i'm there rebecca weekly is there we're all having the conversation a while why hybrid matters and is that pure engagement that makes this different than a webinar is truly a seminar so let me introduce rebecca weekly to you just in case you didn't register specifically to hear her speak this is a woman with chops unlike most vps rebecca started out or didn't start out she was a principal engineer at intel she moved up into the vp ranks at intel in 2020 she became the chairperson of the open compute foundation which has members such as meta et cetera she's talked a lot about the advancements and the shared knowledge around openshift compute and hardware and she's now the vp of hardware systems at cloud flare she has operational responsibility for delivering some of the most advanced networking and cloud services on the web she was going to tell us why the cloud matters but also why hardware matters please welcome rebecca in this session make sure to ask plenty of questions she's armed with additional data points and references for you don't be shy to engage in the conversation again thanks for attending hello my name is rebecca weekley and i'm the vp of hardware systems engineering at cloudflare this session today is on why hardware matters i'm going to try and convince you that it actually is something that very much matters in a cloud-first world so first of all we're going to talk about the cloud and the fact that the cloud in is a solution that solves every problem of course right we'll go through some of the great learnings and key lessons of the cloud where it's working really well and where we have challenges as consumers and entrepreneurs then i'll go through within cloudflare's experiences what we're doing why hardware matters to us and why it creates unique experiences and finally how we build that for others so let's talk about the cloud first things first we're going to talk a little bit about the sort of evolution of the enterprise stack to the cloud this is like 2015 so probably you know all this but i'll just start here because it gives us an orientation point so the enterprise stack of yesteryear was you know you bought from an oem your systems they ran an operating system that was fully controlled whether it was sun or windows or something that you were running with oracle or some other option you were running that operating system and a series of applications that were embedded on those systems managed within your it environment using some sort of network providers mainframe style in the new world in the cloud world you're running on maybe your own private cloud you're running and using the public cloud these are all your compute and storage mechanisms then you're going to be running application stacks you might be running applications for your hr department uh you might be running ones that help you with debug triage you know bug reports etc and then you also will be running something that is doing your global dissemination your networking that could be sdn within your own environment or that could be something like a cloudflare so this is sort of the new world that we're a part of and this new world has interesting challenges so to talk a little bit about sort of why and what we have why we're seeing such significant cloud adoption let's first base in in what we're actually seeing in the ecosystem so step one we're definitely seeing 91 of businesses are use the public cloud as of 2019.

t tech telecoms you know they have a lot of on-prem footprint and they will continue to because of the nature of their business either ip protection and concerns in that domain or you know their telecom and they actually are in the business of selling bandwidth to others so they need to have significant capital expenditure in that domain um where you see people who have embraced i think much more significantly is in this retail distribution transport and obviously you know various sectors have different reasons for why they're looking at these distribution and solutions and how their footprint is distributed so every company is going to make this decision with different criteria and they should but fundamentally cloud as an architecture is very efficient make sure you're getting those efficiencies and not just passing them onto your cloud service provider by ensuring you have the opportunity to diversify across multiple clouds or your on-prem cloud you know having visibility and telemetry okay so hopefully i have convinced you that the cloud is amazing for many many many reasons but it is nuanced and that you want to take an intelligent approach toward leveraging cloud services now i'm going to kind of shift gears to why the hardware matters and why the cloud is actually enabling innovative services through new things within hardware it's actually a major area for both us at cloudflare as well as without all of the different hyperscale ecosystem partners so let's talk about hardware first off you might have known if you're alive that we are seeing massive workload growth in the domain of ai ml machine learning of all varieties deep learning uh and certainly transcode right especially with hybrid work through kovid especially with all of us shut in and into our homes we were watching netflix we were watching and ordering everything online the cloud grew immensely we'll see if all of that continues in the next coming years but fundamentally this workload growth is extreme we are seeing so so much come up so at cloudflare we've been doing a lot of different things in the domains of domain-specific accelerators what are dsas dsas also known as xpus literally for any kind of processing unit are really domain-specific accelerators asic chips that have been optimized for a specific workload so unlike general purpose processors cpus gpus which run a whole bunch of different workloads on that particular asic these ones focus on specific workloads so we have done a bunch of work and i put some of the uh you know blog posts on the right to look at where we can leverage gpus and specifically the more inference oriented gpus on the edge this also we also leverage aspects of standard general purpose processors for this as well you know there's lots of great work that's been done by various x86 both intel and amd to actually improve inference on general purpose processors but we do still see accelerators as very interesting for certain workloads so that is one domain video transcode we have a product called stream which really is an important part of how we deliver live streams out to the world those actually can run very very well leveraging for the transcode aspects accelerators why do we care about these things why are they awesome for us well in general a domain specific accelerator for a workload its workload whatever workload it's operating upon can get about 10x performance for every power what what invested in terms of power so this is why it's interesting if we can get 10 times more performance that becomes an economically compelling situation the barriers the challenges we've had are in the software layers often the sdks that come with a lot of these accelerators don't support versions that we utilize in terms of kernel versions many of them don't have compliance as a globally distributed network provider we really care about making sure we are compliant to everybody in every region so there are real challenges to using these things at scale but where we've been able to leverage them where they've had great software support they've really been able to allow us to do greater processing the last category of domain specific accelerator we've been playing with and i can't say we've fully deployed this yet but it is actually an actively a project my team is working on is infrastructure accelerators what are infrastructure accelerators you've heard people call them ipu's dpus there's so many different names that they go by but fundamentally this is usually an asic that is in line to your network interface card that is doing some aspect of infrastructure workload overhead offload yes i just put all of those together what i mean by that is let's say i am you know a user has come in they are getting served a website you know for doing some aspect of security on the front end and then they decide to use our workers platform and they need to go access storage that access of storage is happening likely on another node so i am going to go through the network interface card maybe over you know the towards the spine to another node that has the data that that person wants to process if i can do my or maybe that person wants to store that data and it needs to be geo-replicated i can do deduplication hashing aspects of the security protocol and aspects of the storage stack in line on the nic as the data transfers so i reduce my network bandwidth requirements and i reduce my my overhead on the primary cpu the reason why that matters is that the primary cpu is doing hopefully customer oriented work that you guys all want to use at scale so the more i can move my overhead my operations my management onto a more optimal architecture again as measured by you know throughput per watt the better off i will be in serving end users having great experiences but also being sustainable in my ecosystem and footprint for my data center so these are key areas where we're looking to innovate within the domain specif specific accelerator space and we bring those innovations right to our customers through solutions like workers so another key area that i think is really um probably the most important part of how we are addressing challenges at scale right now is reducing the cost of storage and memory so there's a whole bunch of different techniques and literally you cannot go a week without seeing a press release in this domain from one of the hyperscalers i highly recommend you start looking at it this is the domain space of memory disaggregation memory compression and then faster storage media for specific data tiers i could talk for 45 minutes just on this topic you can tell by the number of words on this slide how much i wanted to say so really what we're talking about here is the cost of memory and storage so just to use to orient it in my own data over the last three product generations that we have built the cost of memory and storage has nearly doubled now this is large as a percentage of our bomb right this is largely because core counts have been doubling generation on generation and if you haven't gotten significantly more efficient in writing your code you're going to see that you generally want to scale your memory capacity with the course and that means that you're going to be paying more and more and more for each and every one of those dimms and each and every one of those ssds so unfortunately that has really created some cost structures pain points uh reliability challenges i mean name your favorite flavor of what comes as you start scaling up these devices which you know were not the most reliable part of the bomb to begin with so we have been looking at a lot of different solutions in this space but so has the ecosystem at large so there's a bunch of different technologies that have been upstreamed into the linux kernel that really are looking to do hot and cold memory that are doing transparent compression of memory to be able to get better performance for your overall solutions with lower memory footprint there's a bunch of different ones we could talk about uh tmo which facebook transparent memory operations which facebook meta upstreamed into the linux kernel google has written papers on this netflix has written papers and open sourced technology in this domain to work on memory compression this is like the no-brainer you don't have to change anything about your hardware systems compress your memory if you're not touching it too often compress your memory it's in that same category of like hot caches so if you're doing a bunch of cloud workloads and you're going to do small file accesses repeatedly don't swap out to a hard drive use a fast ssd right if you're at all familiar with the latencies of hard drives they're very cheap of course but you know there's spindles there's read right arms there's a physical entity that has several physical entities that have to move for you to get data off that spindle that is or plate technically that is going to cause latency if you can use a storage an ssd solid state drive you're going to have fast access to flash there's absolutely no movement uh it's no vibration you know like none of the factors that can make hard drives difficult from reliability to manage you're not going to see that with ssds you get fast access for that hot cache allows you to do swap again many of us had to turn off swap because we just scaled up memory to get more performance as hard drives got slower and slower and lower relative to the speeds of memory so you get to have swap back but you get to use a near cache that's very smart so hot caches very smart choice we're starting to see this whole striation of the domain space of memory and storage this you know pyramid that was in the current paper from meta on the cxl tier of memory which is a disaggregated tier of memory um through the coherent express link that's what cxl is available with pcie with certain vendors pcie gen 5.

this really is enabling end users to be able to leverage transparent paging what's called what's hot between the fast ddr bus which is more than 10x faster than you know what is attached off of a pcie bus and then your typical storage media behind that so same model same software compact being used for writing to memory but it's on this disaggregated tier which allows for cheaper memories to be stored behind them or other options more creative options maybe over time so i think this is an amazing area of innovation again compression is like just do it everything there is just low hanging fruit you know many hyperscalers have been working none more than meta platforms has been working to you know upstream these things into linux it's available for people to use that now the last point i wanted to make within this memory and storage domain is that data's only growing right we saw in that previous section that end users have kept most of their data fifty percent of their data on prem to date but they see it growing and they're expecting to use the public cloud for that storage media we're creating new and new tiers this all requires innovation and contributions and investment when innovation is used generally that means investment of various vendors so you know when you see something like qlc nand right um quad level bit writing bitsell nand this is really allowing people to write more you get more effective density you're reducing your costs relatively about 33 versus tlc and you're doing that with great performance versus you know a hard drive um and pretty good performance versus tlc in certain scenarios but if you have sort of dynamic loads you're going to see very different very different performance so i don't see qlc replacing tlc anytime soon and that's just my summary of why it's still complicated to have all these different tiers of storage you want to look and optimize your tiers looking at what is archival you're not going to access it very often and that you should absolutely look at hard drive i mean even tape is having a resurgence then database you know what you want to look at for databases is really are there opportunities to use your right ahead logging on a faster medium whether that's an optane style drive or tlc style ssd maybe even qlc although again for performance reasons i would probably go towards the tlc for that and then leverage you know again cheaper media on the back end and then all the way to i ai and ml where the more data most of these high-end processors gpus for ai for deep learning specifically or asics in that domain they're actually the memory the effective memory bandwidth is what limits the t flaps so everyone will advertise you know teraflops of performance that you want you should really look for is the effective memory bandwidth to the device because if you can't feed it it's not going to give you any insights and so this is again a domain where very fast tiered approaches to the memory and storage for your gpus can be interesting ways to get more out of what you're investing from a hardware perspective so data is growing it's going to grow faster than we can really store it i mean that's what the global projections are looking at and the more that we can process without moving that data the more we can save ourselves egress fees the more we can stay aligned to whatever regulatory and compliance concerns we might have um and that increasingly means we want to take advantage of memory compression interesting techniques for memory expansion and faster storage where we're logical to get more out of it okay so hopefully i've convinced you that innovating in storage and memory is a completely necessary and b there's never been a more interesting time for us as an ecosystem so let's talk a little bit about another domain that is i would argue probably the most interesting to me to us as cloudflare uh and you know i think probably interesting to the overall ecosystem so open networking this is a buzzword two buzz words there's two words there uh that many people throw out there but i really think of open networking as having its seminal moments when james hamilton and 22 in 2010 uh did a paper that said data center networks are in my way and he absolutely tore apart with respect and data the way to do it uh the entire ecosystem of current network design it is beautiful it is recorded it is a paper you should read it nobody should work in a data center without ever having been exposed to this it is amazing what he really equated the data center network world of 2010 to was the mainframe business model and it is a pretty incredible piece when you basically look at it what the tl dr is is that software defined networks are critical because you spend most of your cost from a capex perspective on your servers in his paper but if your network doesn't scale dynamically so that you can get data to those servers they're less efficient so the network may only be eight percent of your power budget that's his data but it actually is the great enabler of your servers so we have to do it in a way that's more scalable more software defined more capable so this was the original sort of teardown of traditional network architectures i would argue the hybrid not in the sense of hybrid cloud but the hybrid in the sense of hybrid work conditions we've all been in for the last two years as covid wreaked havoc with lives actually is making it even more critical that we all embrace software-defined networks now you've probably heard you know secure access uh security edge um model sassy models for security this is another reason why software-defined networking is so important because if you are engaged in a traditional network design where every upgrade is by chassis and you have a traditional model of the user has to be on-prem at my data center at my location to be secure and to have access to their email their data their you know remote work so in the fourth domain of areas where we've been working to improve our hardware systems to deliver better performance at scale we really are seeing the overall power optimized architecture having an interesting resurgence so one of these is obviously the work that's happened within the arm domain whether you're watching what has happened with graviton or ampere or any of the others by the way qualcomm became before that and there's been many many offerings within the arm ecosystem for server over the years but we absolutely see the opportunity of risk-based architectures to deliver more optimal performance per watt it's not highest performance it is optimal performance per watt at good enough performance and starting to be much more compelling especially with neoverse as an option from an architecture perspective so we're seeing a lot of interesting things we've been very public about what we've done in that domain to try and really give a power optimized compute footprint in places where power is incredibly constrained or in places where power is incredibly expensive just for regulation reasons or truly accessibility reasons so that's one aspect of power optimized architectures that is interesting i think the other one is sort of the combination of domain-specific accelerators and asics as the world moves towards chiplets so there's always been value in optimizing functionality onto the mesh of a processor integration is you know whether we were talking about the early aughts and how intel processors were able to bring on vector-based avx-based you know acceleration so that everything could run faster in terms of floating-point math there's so much that general purpose processors over the years have brought in in terms of acceleration and integration into the core processor to make everything a little bit faster and enough faster that the value of being able to run all the different kinds of workloads is worth it and we've seen this within the gpu space as well you know tensor processing unit came out in 2016 and gosh darn it gpus the next generation of gpus had support for tensor math directly built in so this sort of integration play has been something that the general purpose vendors have always worked on and ultimately delivered i think in incredible ways what's interesting about what's happening right now is that the world of chiplets has exploded there's so many different startups in the domain of hardware and asic development whether infrastructure accelerators or ai you know training or inference chips or vcus things in this domain and what you're seeing in the ecosystem now is this desire to integrate so amin veda wrote an amazing piece i think uh in the an amazing blog on the soc as the new motherboard and what you've seen in the ecosystem is a huge embrace of this so uci the universal chip interconnect express was open sourced and as a consortium this year the slide shows that picture on the end of all the different semiconductor vendors foundry vendors ip vendors that came together to define this and it had many predecessors and this is not the first space obviously i'm very involved in the open compute project being the chairperson of the foundation and you know we've had odsa the open domain specific accelerator work group for many years working in this domain of open chiplets and having a chiplet marketplace why is this interesting i fundamentally believe it is for power so we've seen many general purpose processor companies and and different ecosystem providers try to find that optimal mix of cores plus accelerators plus a little fpga over here so you can define it yourself and try and get the right mix to be a commercial success and almost inevitably those projects all get cancelled because what is one person's you know perfect solution is not perfect for anybody else and so it's just very hard to build a general purpose business merchant on that model but if we can get to a place where the world is more composable and we still have security and reliability and visibility across that ecosystem that is where we can start to see 2d stacking and 3d stacking and the best of all these worlds come together which will give us better supply chain reassurance which will give us better optionality in the products that we build in ways that feels like legos i mean that's the vision that's the hope that's the prayer so we'll keep watching this i can't say we have a chiplet project going on at home uh there's a little bit more going on but you know this is absolutely an area where we continue to look and invest in staying really current through projects like odsa okay so that's kind of a high level of why i think hardware matters why it's the renaissance of hardware in my personal opinion and why i can't imagine a better job than to be in hardware systems engineering right now because the intersection of software and hardware especially at a company like cloudflare where we are you know running all of our own software effectively at scale we get to do really cool things where we can innovate and build it break it and blog about it pretty darn quickly so we do a lot of that so maybe i'll just give a little context of who cloudflare is and why i'm talking about all these things and why we build hardware systems so cloudflare we run a globally distributed network and all of our hardware runs all of our services so whether you're talking about ddos or waff or any of the other solutions they're running everywhere all the time now that globally distributed network we're talking about let's go to the next um we're talking about you know over 140 terabits per second 270 countries we're running 17 of the world's internet traffic it is a huge distributed network of services that we run for our customers and our entire goal is to build a better internet you have your own private cloud you have your own development you're running in public cloud great no problem we distribute that effectively across the world for you so that your end users have a very highly performant experience and a secure experience that's what we do that's what we focus on we focus on the network at scale and that is you know really what has allowed us to think through how to grow and build systems and innovate faster this is cloudflare's stack of services our overall suite and i would say the part of it that i find so interesting is not just our services layers for security and reliability and performance but it's also the compute innovation that's happening with workers at our platform level so there's so much that we're working on including r2 which is an active project my team is working on to help actually globally distribute storage everywhere across the world to bring the data to where the compute is and to hopefully help you stay more compliant to all the different you know data sovereignty issues that occur and can be difficult with observability and reliability so the last thing i would say is we also work to enable hybrid hybrid in the real sense with cloudflare one with a legacy environment helping that scale into this global network of secure sassy ready-to-go cloudflare one services so for us we're taking an approach of legacy lift and shift like assuming that somebody will start and move to a green field deployment whether cloud or otherwise is illogical helping people meeting them where they are and helping them improve that posture and leverage what we've built on top that is actually how we think about the systems that's how we think about using them and how we help our users come on board so that's just to give you a sense of who we are what we do and why i spend a heck of a lot of time on hardware systems so final section for why hardware matters why the cloud matters fundamentally i would say the cloud is amazing amazing set of technologies whether private or public that are enabling you to get the most efficient use of hardware that is fundamentally what we all are doing we're trying to get insights from data we're trying to create transformative experiences for our end users we need systems to do that and the cloud has driven more innovation both in the systems themselves as well as the efficiency of utilization than any other technology in the entire world i could be wrong about that but i believe it is true fundamentally if you are a real user if you are a cio a cto who has to operate real systems at scale you have to look at private clouds you have to think about your deterministic footprint your deterministic workloads and identify where you can be cost efficient with those within your own ecosystem and leverage the public cloud for non-deterministic workloads for something that you're researching in the ai domain that you don't yet know what you need to do and certainly for disaster recovery this is and business continuity planning obviously the last year has taught us that this is really critical i think there's a lot of work in hybrid apis i'm very hopeful that we will see a continued investment in solutions that work across regardless of of whatever acquisitions are in play because it is critical that we are able to move workloads encapsulate workloads well across different clouds and solutions to take advantage of what power users really need when they come to the cloud and finally there's global dissemination that you can leverage to make applications work better faster stronger and reliably so that you can improve user experiences and we'll keep looking at all those hardware solutions to help drive them and make them better and we will absolutely blog about them actively cloudflare's blog is probably the nerdiest place i've ever been so i'm happy to help in any way thank you so much