Web3 - Is it just all Blockchain and Crypto?

In this podcast, Keith and David Aronchick ( @aronchick ) Co-director of Research Development at Protocol Labs talk about WEB3 and the theory of decentralizing trust in storage solutions. Learn more about Interplanetary File System at IPFS.io The CTO Advisor Web3 - Is it just all Blockchain and Crypto? Play Episode Pause Episode 1x 00:00 / Subscribe Share Apple Podcasts Spotify RSS Feed Share Link Embed <blockquote class="wp-embedded-content" data-secret="PBVOhxDqhz"><a href="http://thectoadvisor.com/web3-is-it-just-all-blockchain-and-crypto/">Web3 &#8211; Is it just all Blockchain and Crypto?</a></blockquote><iframe sandbox="allow-scripts" security="restricted" src="http://thectoadvisor.com/web3-is-it-just-all-blockchain-and-crypto/embed/#?secret=PBVOhxDqhz" width="500" height="350" title="&#8220;Web3 &#8211; Is it just all Blockchain and Crypto?&#8221; &#8212; The CTO Advisor" data-secret="PBVOhxDqhz" frameborder="0" marginwidth="0" marginheight="0" scrolling="no" class="wp-embedded-content"></iframe><script> /*! This file is auto-generated */ !function(d,l){"use strict";l.querySelector&&d.addEventListener&&"undefined"!=typeof URL&&(d.wp=d.wp||{},d.wp.receiveEmbedMessage||(d.wp.receiveEmbedMessage=function(e){var t=e.data;if((t||t.secret||t.message||t.value)&&!/[^a-zA-Z0-9]/.test(t.secret)){for(var s,r,n,a=l.querySelectorAll('iframe[data-secret="'+t.secret+'"]'),o=l.querySelectorAll('blockquote[data-secret="'+t.secret+'"]'),c=new RegExp("^https?:$","i"),i=0;i<o.length;i++)o

Transcript 3,928 words · about 26 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Hey, you're listening to yet another episode of the CTO advisor. We're getting into kind of our stride for the new year. I'm breathing deeply 2021 is over 2022 is in and we are full diving into web three on this podcast. I have a web three guy from the Twitters because that seems to be where most of the conversation is going that and clubhouse. If you still do clubhouse clubhouse, I have will be David Aronchick day. What do you work? Where do you work?

Other than you know, a Twitter personality that just know you from an introduction from a common mutual associate on on neutral friend on on Twitter Alex Ellis who we had on the podcast. Where do you work? Yeah, Alex is amazing, by the way. So I am co-director of research development at protocol labs, protocol labs, you probably haven't heard of we we try and stay a little bit in the background, but many of the folks who are here at protocol labs worked on IPFS and file coin IPFS is the interplanetary file system.

It's been around for about seven years, and I think I think it was 2013. So now we're at nine years, I guess, and it's the idea is to have your data be shared everywhere in the world in a much more distributed, friendly way, and I co-lead research development here. So it's really interesting what I guess my first question, like right out the gate, like I absolutely get the utility of IPFS. But when I think of web three, my first thoughts go to crypto, because that's the main use case that I see on basically mass media and even some of the more technical trade platforms and by trade platforms, I mean, IT platforms.

Break down to me what let's let's first define web three. What is web three? Well, if you can define web three, you're smarter than than most folks out there, because just about everyone has a different definition. Let me give you a little bit of my background and how I kind of landed here, because that'll give you the basis for what web three, my interpretation of web three. So before I was doing this, I worked at Google and I led product management for Kubernetes, the global scale distributed platform for taking a whole bunch of resources and hiding them behind a very clean API.

So I can say, oh, I just want to deploy this thing and I want to do it across whatever 17 different machines, Kubernetes takes care of that for you. Very, very powerful stuff. And there were other ones out there at the time and they all have their pluses and minuses. For better or worse, Kubernetes did pretty well and we had a lot of success. From there, I went and co-founded the Kubeflow project, which is kind of same thing, but for machine learning.

So it's taking Kubernetes and allowing you to distribute across all of these various resources, disks and VMs and so on behind a single API, making it very easy to, if you want to go out and you want to use machine learning, you want to train machine learning, things like that, you can take this very large number of distributed machines, merge them together behind a single API, and then do what you needed to get done. Those two together really gave me a sense for the power behind this distributed world out there.

In the olden days, you would go out and you'd spend whatever, $250,000 with Sun Microsystems on a single machine. And if that single machine's power supply went out or disk went out or memory went bad or the motherboard went bad, you were in bad shape. The idea, it's certainly not new, but behind these systems is how do you distribute the load between all of these various machines, giving you additional resilience, giving you the ability to target what you need and make specific resource requests and things like that and pack a whole bunch of things together rather than having everything contention in a single box sitting there.

When I came to that understanding and started looking at various jobs out there, I did talk to the protocol folks because IPFS really leaned into that kind of space. It's the idea that, boy, there's a lot of machines out there and each of those machines has a disk attached. Wouldn't it be great if we could treat them all like a single machine, like a single giant disk that you could write stuff to and pull it out? And if a single one of those boxes went away, the system automatically repaired and figured out and redistributed that load to other machines and so on and so forth.

So a lot of real genius there. That is really the core of Web3. Certainly there's a lot of air sucked out of the room by people doing decentralized finance and speculating on the price of this and the coin of that and whatever. At the end of the day, it's this idea that you can decentralize across everything, identity, trust, compute, resources, so on and so forth. How do you spread that out? Now, the problem is that as you think about spreading this thing out, the number one thing you're going to run into is trust.

How do you prevent someone coming along and looking at your data or lying to you saying, I did run this compute, except they didn't actually run the compute, they lied to you, or there's some malicious actor or so on and so forth. And so the core of this decentralized system is this idea of how do you decentralize trust? And the way to do that today is with things like blockchain, where you can say without question that this thing happened in this way, you prove it with math, with private keys and so on, and then once that's complete, you can now spread that awareness to everyone and say, hey, everyone, this thing happened at this time, this person signed in at this time, this person exchanged value with that person over there, so on and so forth.

And so that kind of decentralized trust is really, really powerful. That's also, by the way, how you get to the value of coins, right? Because what you want to do is you want to have a way to capture the value exchange. So for example, if I have a disk and I offer up a gigabyte of disk to the network, I should get paid. Ideally, there's not some central authority paying me, it's the network that pays me. And whoever wanted to store data on there, I can prove to them that I stored this for an hour and have that come back, and that's something that we have in Filecoin called proof of space time.

At the end of it, they put in a certain amount of Filecoin to the system, the person who was storing it paid me in Filecoin, no transaction fees, no anything, it just happens automatically through math. So it's not an easy answer, there's certainly a lot of subtlety to it, but when you hear Web3, really think about decentralization, and in particular, decentralized trust that is validated with the system. And that's really what we're trying to get to here. So some of my nascent thoughts on this back in 2013, 2014, when Google first did Google Fiber, I kind of thought about, well, this is obviously pre-pandemic and pre-work from anywhere, but one of my initial thoughts was, man, you know what, it'd be really cool if somebody came up with a distributed cloud, Web3 basically, this idea that there's an awful lot of idle compute connected to this highly, maybe there's high latency relative to a data center or even corporate network, but the sheer amount of distributed compute, distributed storage, those resources that are idle.

If I took a gig from 100 users or 10,000 users across a pretty resilient network, I can create some really interesting services. I thought about it more in the centralized traditional web, that two type of concept that if a centralized agency or a centralized entity coordinated all of that compute, the trust would be with the centralized agency. But, you know, IPFS and blockchain and all of these efforts have looked to solve the trust issue, distributed trust issue with technology. It's a completely valid look.

I think where we're at now is the usability of it. When I think of my initial set of desire that if I have a bunch of storage that I need to archive, let's say a petabyte of storage, and I don't want to manage that petabyte of storage in my data center, I'm not going to need it. It just needs to be archived and it needs to be spread, and I don't care if it's spread across 10,000 nodes or not. I just need to make sure that when I need it, I can get it, and that when I'm getting it, what I requested is what I'm getting back in the time frame that I want to get it back.

And I think IPFS is a good thing to pick on because it's been out there for a little bit now. What are some of the solutions or what are some of the problems something like a distributed file system like IPFS solves? Well, so you bring up a really excellent question. First off, I will be the first to say that IPFS is nowhere near as usable as we would like it to be. We want it to be as usable as any POSIX file system out there in the world.

You should be able to write to it as easily as you write to your local machine. There is a core element here, though, and I really don't want to light paper over it. There are a lot of folks out there who are like, Web 3 is going to replace Web 2. I cannot disagree with that strongly enough, right? You need to pick and choose the things that work in your new world and use them where it makes sense. So take your petabyte example.

There's a dashboard out there that you can go and check. I haven't checked recently, but storing a petabyte of data on IPFS, I think, is going to cost you about $10 a year or something like that. Like, it's absurdly cheap. We have 14 exabytes of space. That's not an exaggeration. Exabytes of space available on the network, ready for anyone, and it is unbelievably cheap. However, it is way worse performance than anything out there. You could go and measure it against a, I don't know, 1200 RPM drive, and it's going to be slower than that.

That's not bad. It just does a different function, right? And so you start to get smart about how to build caching layers and things like that. And there are many, many people in the Filecoin ecosystem who actually do this, right? So you can use a third-party service, and we already have deals and partnerships with folks like Cloudflare and so on to make those caching layers and performance much better. So I'm a huge fan of like, look, pick your Web 1 technology, use that.

DNS is still great. You know, if you have a big database, you'll probably want to run that on a single machine. There's lots of reasons to use Web 1. There's lots of reasons to use Web 2, things like Kubernetes and Kubeflow and so on. And there's lots of reasons to use Web 3. Just pick and choose the things that are good fits for those particular situations. I don't want to paper over your comment. It's funny, because the performance of IPFS, I'm not that concerned about it, right?

Like, yes, it's slow. It'll get faster. It'll get much faster. We actually have a lot of innovation coming out this year that's going to make it much faster. But I don't ever want to compete with your local disk, let alone something, you know, sitting on NVMe or some super high-performance thing. It just serves a different function. What I want, more than anything, is, you know, brain-dead easy ways to submit your storage to it, to run processing over this enormous amount of data that's out there at extremely low cost, and to get results, and to make those APIs so easy that, you know, you kind of are scratching your head, even like, wow, maybe I should write to that and just figure out how to use really slow disk, because it's so cheap, and it goes to exactly what you just described.

I don't have to think about managing those thousands of hard drives, physical hard drives, to store that petabyte of data, because IPFS does that for me. So that's how I see. The biggest failures for me are just continuing to make it easier to use. Yeah. My mind is kind of racing with, like, HPC type of use cases, that once you get the data out onto the network, you have this amazing amount of compute. You need a way to coordinate all of that compute, et cetera, and I'm sure there's probably projects and solutions out there that's helping to coordinate the running of compute, the isolation of instances, the validation.

Again, the validation that the compute that I wanted to run is actually the compute that I've run. I'm not getting a result that's different. You know, the fact that I can run the same thing slightly differently across 10,000 nodes has some type of utility to me, naturally. I haven't really sat down and thought about it, but the fact that I can run the same thing the fact that I can get either the same result or the same expected result when I make slight changes across a set of data or a different set of, I can version date.

I can, you know, think about earlier in your career and think about Q flow and the maturing of this. There's like an unlimited number of use cases. So I buy the concept that Web 3 is real. I buy the concept that there are utilities across Web 1 and Web 2 that Web 3 will either enhance or enable. One of the things that we haven't cracked a nut around, I remember maybe it was about five years ago when I was looking at kind of blockchain and beyond the trust models, I think IBM played around, I haven't looked at it in about four or five years, but IBM played around with other models that weren't token based for paying for work done on the blockchain or proof of work done on the blockchain.

There were other rewards, like whatever it is. What I haven't kind of cracked my head around or wrapped my head around is the level of effort that it takes, the amount of compute that it takes to do this distributed compute. I did not expect, when I thought about this circa 2013, 2014, I didn't think about the cost of the compute. I think we pick on GPUs as the biggest problem. Is there a model in which we're not costing more than the compute, than it was just idle?

The computers are on. It is a really fascinating question. I remember when I first got to Google and talked to a lot of the folks in the MapReduce space. For those that don't know, MapReduce was one of the innovations that allowed Google in very early days to index the entire web, which is a really hard problem. You have all the storage, you have all these machines. How do you have each machine do a processing on a single thing and then merge all that together?

So you map it all out, and then you work on it, and you reduce it by bringing it back. It was really fascinating because it's one of these problems where you're kind of like, well, that seems pretty easy, right? Couldn't I just write a, I don't know, bash script and just execute it? Like, how hard could that be? And it turns out that it is easy to do on a single machine. But the moment you start spinning it out to 1,000 different machines, how do you know that, exactly like you said, machine number 770 finished properly, that it didn't crash halfway through, that it submitted the result?

And, oh, by the way, this is duplicate with machine number 322, right? Like, there's all these weird problems that you have to solve for. And so having a structure around that, internally at Google, they built their own, and then they published papers around it. You can go read it, folks that are listening, if you haven't read it already, where they talk about all these kind of things. Today, you could argue that the vast majority of money in data is spent on machine learning.

The vast majority of money in data is being made on this very problem, right? When you go out and look at the big data companies, the ones that are getting the most funding, we're talking about hundreds of billions of dollars of valuation. They are, you know, being paid simply for storing a bunch of data and then letting you run compute against them. Now, you know, it's not, certainly not their only thing. Some of the scenarios that you're going to have there are very high performance ones, but lots and lots of them are very, very slow.

It's like, all right, hey, by tomorrow, I need this data processed in a way that is usable for me to do a training job against, or that lets me, you know, figure out what the average number of people in whatever, Peoria, you know, spent on Starbucks, right? Like, whatever it is, I don't need that right now. But if you can get that to me by tomorrow, that'd be great. And those are the kind of scenarios where if you had a appropriate structure over, you know, low performing disk that has lots of spare capacity, you just submit the job.

And by tomorrow, you have the answer and you have it at, you know, one one thousandth of the cost or less, right? Because the disk is just sitting there and the space is just sitting there. So why not use it? So I'm completely with you. Like, there's no question the needs are out there. It's just a matter of how we bridge from this world where we are today into a distributed and trustless and, you know, infinite capacity world tomorrow. Yeah.

So the drive is home for much of my audience. I have audience that's kind of in the cloud native world. I have audience that's in traditional enterprise IT. If you're in traditional enterprise IT, think of SAP, SAP query that takes three weeks to run and it fails. I mean, that the problem isn't that it takes three weeks to run. The problem is that it fails. Yeah. And it normally fails because of some type of capacity, timeout, et cetera.

What happens when so when we go to SAP with that problem, the solution is SAP HANA. SAP HANA in-memory database, I can do the same query in a day or less or in a half an hour, 45 minutes, doesn't matter the gain of time. The benefit isn't the reduced time, but the certainty that I will get the result in a time that the business needs the result. What happens when there's a in-between? What happens when there's a solution that will, let's say that it takes three and a half weeks or two and a half weeks.

The SLAs are the same, but the difference is the reliability. I can reliably put something on to, I can reliably put a question onto a network and get the answer within a pre, within a dependable amount of time. And with some certainty, I think this is the potential of web three and augmenting enterprise IT capabilities. So without spinning up a whole staff to deploy SAP and the SAP HANA and the resources that it requires, there's kind of this middle step that appropriately provides the solution for the problem.

Yeah, that's such a perfect example. And, you know, we would, you know, we in the web three world, this is my opinion, would be idiotic to attempt to replace SAP. Right. You know, we talk about something that has global adoption and is great APIs and great all these kind of things. But imagine that a component of SAP instead of being, you know, SAP can deploy to a number of different databases, can deploy to a number of different like solutions out there.

Like you could say, hey, you know what, this is a slow running job. SAP, can you communicate with, for example, in this case, IPFS to submit this long running job? We're going to give you, you know, seven nines, eight nines of reliability because we're going to figure out how to get this thing done no matter what. But you and user keep using SAP. It's great. It's going to get you your answer in, you know, defined amount of time, whatever your business requirements are.

We're going to get you that advice or that answer. But we're going to take on being reliable for that. And that's a perfect example. Like where you're not replacing your stuff that you're using today. You're using this stuff, you know, and love just behind the scenes. It becomes far more reliable and far more cost effective. So, David, we're a little bit over. I appreciate the time. Where can folks find out more about what you're doing? So your best thing is just go to, you know, the Protocol Labs website.

We have everything up there. ai. In addition to that, by the way, if you if you're interested in distributed compute in any way, please reach out to me. I'm on Twitter. My last name, Aronchik, that's A-R-O-N-C-H-I-C-K, Aronchik, on Twitter. I'm sure that you'll get links or whatever in the show notes. You know, email me directly. We're actually actively doing a lot of development around this right now. And if you have particular use cases or scenarios or you'd like to whatever, help us code it, we can do that too.

But there's a lot of activity around this. And then, of course, IPFS. io to read all about it. All the insanely cool and advanced math necessary to make everything you see here cryptographically secure and validated and so on is all up there. And you can read white papers to your eyes bleed. All right. com. I'm always available on Twitter. DMs are open at CTO Advisor. Rate the podcast. Share it with your aunt. She will love talking about crypto related IPFS in the penitentiary file system.

I guarantee it. Talk to you next CTO Advisor podcast.