The Evolution and Impact of Cloud Data Solutions: A Chat with Miles Ward

Host: Keith Townsend, Futurum Global Advisor Guest: Miles Ward, CTO of Sada Overview: This episode of the CTO Advisor podcast features a return guest, Miles Ward, to discuss the technological and strategic advancements in data management and cloud migrations, focusing on the challenges of data sovereignty and the innovative solutions developed to address these complex [...]

Transcript 4,317 words · about 29 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, so I'm coming out of Google Cloud Next a couple of weeks ago, and I cannot tell you how many people referenced a CTO Advisor podcast episode. And I felt a little bad because I haven't been producing the podcast as consistently. You know, I'm with this much bigger company, Futurum, I have a big title, I'm a global advisor. I've been upgraded from the CTO Advisor to a global advisor. What does that mean? I don't know, and we'll get into that later.

But the opportunity to talk to one of our previous guests, one of my, I think, more entertaining guests, Myles Ward, CTO of SADA. Welcome back to the podcast. Thank you, Keith. I appreciate that. I aim to be entertaining. I was super excited to reach out to you about this one because they killed my baby. The snowmobile is getting shut down, and I know too much. So I'm looking forward to talk to you. So I'm looking forward to this conversation because I know this, you know, as I looked at the announcement, I kind of said, hmm, I don't know the details of why they didn't share why, but I could infer to why they would.

Some reasons why snowmobile may be a little bit past its prime. But let's talk about, you know, kind of the origin of snowmobile. Break it down to us. What is it? What was the use case for it? And, you know, why does any of this even matter? Sure. So I was in the first hundred employees at AWS, and one of the most incredible folks there built out the whole patent program. He said, if you if you come up with something you think hasn't been done in cloud, I want to hear about it.

It's great. So I sat on a plane next to a movie producer who described the complexity of on-site shoots with 3D cameras, 120 hertz, all the color resolution that they need. And they were bringing out half racks of storage equipment, and they would have to call stop on set because they ran out of disk space to be able to keep the cameras rolling. Now, all that sounds like a bad problem. In the meantime, I've been working, selling Apple S3 and helping everybody consume petabytes on the S3 side.

And I'm looking at what you can actually get in a half rack locally deployed going, well, of course, that sucks. Like, that's just going to be terrible. So I'm sitting with this producer who's he's a pretty technical guy, too. He's bought a bunch of gear. And we start ginning up this experiment we want to run. Like if I put a bunch of storage gear in a Prius, then I'd have a generator on site. That's kind of cheating. I can bring gas cans with me.

I can carry way more than you can put in like a shippable storage container because I can just drive it there. And if I showed up with wire, I could run it out to the, you know, run 100 gig Ethernet or whatever, you know, Mellanox or whatever cable connected to you want. And then I could drive it back to one of the Amazon regions and plug in and voila, it would all just be objects in S3. I laid out all this crazy technical shit about how to do the replication on Amazon's at the S3 system is built on this thing called a tree topology, whatever.

So I laid this whole design out and I explained this to Eric, the crazy inventor lead. And he's like, well, that's nuts. We should we should file a patent for that. So we do that. And then I bail. I go. I end up going to Google. I'm I'm working. Eric Schmidt pulls me over to work as the first solutions architect for Google Cloud. And they drive that snowmobile on stage. So so I send a chat message to Werner, I'm like, Werner, is that the Prius?

That's an awful big Prius. And and he's like, oh, if you were still here, we'd let you drive it on stage. You made terrible choices. So that's just like a knife in the chest where he was digging me around in there. But the use case, the core of that is like. If I can put a hundred petabytes local to a big data center facility that's got hopefully more than that storage and I and I'm going to truck it back to any one of the AWS facilities at the time, there were only three major regions in the US.

Now there's more. , you were doing so at six point four gigabits a second, which is pretty fucking fast. So I'm not giving a sorry terabits a second, six point four terabits. So that's like actually. So. That, you know, there are plenty of data center facilities that piled up storage, but just did not have the connectivity to be able to effectively migrate to cloud. And that would make the migration plans look like two year, three year, four year, five year things where Amazon and all the cloud providers, frankly, are trying to do them way faster.

So they saw it more of an enabler for non cloud to cloud migration and dragging all that data up to the farm. I always thought of it as kind of more of an edge enabler where you'd be able to tear around. It's pretty hard to drive a 40 foot container truck to your average movie set. But but that was that's that's kind of how it all got started. So that's a great origin story. And, you know, snowmobiles, no balls, no snow ball edge, I think was the kind of HCI type solution to build out.

I think obviously this especially with sovereign cloud and door in the in the in Europe, people are trying to figure out how to make use of local. Data stores to one being compliance with regulatory concerns and then to even with the challenges around availability, around accelerators, the fastest data transfer is still the data transfer you don't have to make. So there is a little bit of we're going a little bit of a tangent and we'll get back to like the demise of this noble noble.

What are you seeing when it comes to customers and their challenges around data and data sovereignty? Sure. I helped design really early on a bunch of Google's push into what now they call Google Distributed Cloud hosted, which is a system that allows them to run Anthos and the underlying container management infrastructure and the management infrastructure that's required to stand up effectively a private region of Google Cloud. So today, Tallis, the effectively like infrastructure and training company in France, T-Systems in joint with T-Mobile and the other components in Germany, those two, as well as a couple of others that are in planning, are already deployed functional sovereign clouds in Google's ecosystem and Amazon and Azure have similar building blocks.

I think Oracle at one point was advertising if you just sort of stroke the pen efficiently enough, they would deploy an entire region wherever you wanted one. So everybody has gotten to a place where at least some flavor of their offering is deployable in a way that is sovereign. The complexity for all that, in most cases, is that in order to really do that, not only does the gear have to be in your location sovereign and the software has to be licensed and owned in a way that that is sovereign to the legal jurisdiction that you're talking about, but the people, the staff have to be citizens or, you know, legally authorized to do that kind of work.

So like I remember taking a long phone call with the folks from Tallis when we started out with them and it was all about what is SRE? How do we train people to run this thing they're going to hand us called a region because Google won't run it for you. They wouldn't be allowed to. They don't have. Now I got to put hundreds of French citizens in the French region of sort of operated on your behalf. They expect if you're going to deploy this thing that you as a provider want to do that work.

That's why you're reaching out about asking and taking it on. I think the last thing Google wants to do is send a bunch of staff to other companies. So that that whole line of questioning and the complexity there, I mean, like you think it's hard to get certifications on AI stuff or solutions, architecture work or the other kinds of enablement. There's no good training courses today on what does it take to run the Google region or what does it take to run an Oracle region or an Azure region or any of the rest of that.

It's very specialized stuff. So I think as much as countries want it. I remember mentioning to Urs at one point that this seemed like the most tractable problem in computer science. Right. There's only one hundred ninety two or one hundred ninety six countries just deploy a region and all of them. And then you're done. Like you're just actually done. You have one in every jurisdiction success. None of the companies are even close to that. I think the biggest country coverage right now is is Azure with something like 16.

So it shows goes to show you how much there is left to do to get this into everybody's legal jurisdiction. Yeah. And, you know, coupled with that security concerns, data gravity, which we'll get back to, you know, we I think I came up with this concept. I didn't go out as much as Dave McCurry did, who came up with the data gravity concept in theory. I've had him on the podcast a couple of times to talk about it. There is you know, there is a actual theory, there's algorithm behind it.

But I challenged him a couple of years ago and say, you know what, services are starting to have gravity and he calls it more inertia. I don't care what you call it, but I can't get Vertex inside of my data center as easily as I don't know that, you know, I need Vertex. So in order to get Vertex, I need to ship my data somewhere for that to happen, because that service is that important to me to be able to take advantage of it.

So, you know, we're having these extremely unique challenges where we can't necessarily build and you stated this. We can't build the services cloud providers can build. But yet our data is on site, hence the snowball bill. You know, OK, let's let's get the the value is worth the risk. Let's get the data up there. We need to get it up there and and, you know, at six gigabits per seconds, terabits, terabits. Because I've had connections, I've had 100 gig connections and in the data center before and they're just not big enough.

So, you know, you're you're at the practitioner level, you're doing massive projects. What do you think of the demise of something like a snowmobile? So I mean, I led the migration for Twitter before it was called X from their private data centers into Google Cloud's facilities, and they had three hundred and sixty petabytes active in what is inarguably the largest Hadoop cluster in the world. And they didn't like it anymore because Hadoop is terrible and BigQuery is a lot better. Well, I was just about to say, I'm sorry, but that's that's a lot.

Yeah, that's a that's a heartbreaking state, right? Like this is the lovely part of my career, right? Like Hadoop, inarguably, absolutely saved my startup, like reduced our cost by three orders of magnitude to build the indexes that we need to in comparison to SQL Server. And now just 15 years later, we sort of giggle about it. And of course, you got to tear that shit out of your data center. So when we did the deployment for for Twitter, you have exactly this option, right?

Like somebody could bring a snowmobile out, somebody could bring a large format. There's a couple of competitive services from Equinix and Megaport that would do the same kind of thing. And we looked at their existing physical footprint and built eight one hundred gigabit a second GBIC connected links over Juniper devices into Google's facilities and transferred the data over 43 days. And if you think about what kind of a pain in the neck it is to find a parking spot for the truck and get the wire through the security doors into the racks, and then we make sure the security people are eyeballing it all the correct times of day and night, all that stuff might take more than 43 days to figure out.

So I think at the end of the day, and we were moving three times, three and a half times what would fit in a snowmobile if even that was available as an option. So I think what what Amazon certainly found and Google found quite a bit earlier is they have a network that's some seven times the throughput of the public Internet. So they are way leaned into grabbing fiber everywhere, is that you have to have a lot of data in some place really weird where it's hard to get big fiber in order to be able to in order to be able to have a solution like a snowmobile actually make sense.

And so they ended up not getting used all that often. I think, you know, the original one that I proposed that was a lot smaller probably would get used a lot more often because the latency would be commensurately lower, right? We had all the time customers would show up with, you know, two suitcases full of external drives to a Google office and say, I want to move to the cloud. You guys have been with like, God, like, can you can you only plug all these things in?

We're like, this is not actually what we do. But yeah, sure. Shit. Yeah, let's let's roll. Right. And we would set everybody up and we'd move a bunch of data that way. So I think, you know, the metro area city latency of driving around is probably viable like semi size was maybe better for the press release than really the actual use the practical pieces of it. And then I think you make a huge point that when it first came out, there was three regions.

Now there's regions, there's zones, there's with Equinix, other data center providers. So if you can get your data relatively close to one of these providers, you can ingest the data much quicker than you can. Look at the the size of hard drive of SSDs now held in my hand on a GTC, a solid dime has a sixty one point forty four terabyte, which for the most part, you're listening to this. You promise you get me one. A sixty one terabyte SSD, low power device.

You know, you have a suitcase of those and you get it into a co-location that has the direct connects directly into all the cloud providers. The cross connects into the cloud providers. You have essentially near cloud speeds, capabilities that you can either ingest the data or leave the data is there. You know, Google has been here. You can do amazing things with BigQuery across platforms. So I think the idea was cool, like, you know, the truck I've been trying to get folks to sponsor putting a data center in the end, the CTO advisor in cloud, which I'm I'm recording from now.

Cool idea. It's niche. It gets people. It gets eyeballs coming. People are going to come look at it. But at the end of the day, I think we're finding that there's just more practical solutions. Yeah, I to two things there, right, that you were talking earlier about, you know, about the kind of distribution of workloads, the things you can build in cloud, you can't build in data center. And it'd be nice if the software showed up there. Google is clearly pushing in that direction.

The first real offering there is this. They set it up so that the Vision AI model could be run in Anthos on prem. So that's them taking a Google managed, Google supplied, Google certified, all the various SLAs and contractual obligation bits. And they put that thing in in a facility on your side. Now, I think most customers were like, well, that's one API, I want the other 806 of them. And and would have preferred the entire regional footprint. The last time I talked to Google about the actual total regional services count, because remember, Google Cloud is made out of stuff, some of which was built for Google Cloud, like, say, GKE, which was designed in concert with the Google Cloud team, and some products which predate Google Cloud by quite a lot.

So like BigQuery is Dremel is a pre org hosted runs on the internal Google three data systems that are distinct from the stuff that's built above that layer in Google. So any of the really shiny shit that Google built for itself before they even thought about cloud is really difficult to pry out and cram someplace else. Even given that difficulty, they figured out how to do it on the other clouds. So BigQuery omni allows you to run the BigQuery codebase on Amazon or on Azure, you can write a query that federates data from s3 and GCS and get results in the BigQuery terminal.

And it does the processing Amazon site, it's it's an incredible architecture, they don't really built a crazy thing. So are they also have Dataflow omni and a couple others. So I think over time, with the omni initiative, they're going to try to get as many of these services running in as many facilities and hopefully more of the bare metal facilities and local customer facilities as possible. The other toy if you want toys, Western Digital just came out with this thing. It's called an ultra star transporter.

And it's kind of looks kind of like one of those Halliburton suitcases, right? It's like a little tiny attache case, but it's 384 terabytes. So like, that's my kind of size. It's perfect, right? Like, it fits in a 19 inch, if you want to slide it into a rack row, it doesn't have any rails off hanging off, you got to clip that stuff on comes a little carry handle, and you can put, you know, a third of a petabyte in your carry on, like, that's, that's perfect.

That's just great. Yeah, and I think these types of solutions, I don't know what the lead time it was for to get a snow bill out to market, but you know, the things change fast. Like, yeah, man, you know, you have to adapt in the world. Last time we talked, we talked about, you know, the difference in cloud services from when you ran infrastructure for the political campaign you worked on to now, and how much the cost structure would be different, then, you know, how quickly AI, the models, I mean, llama three was just announced, I don't know how many parameters of llama three, I have a hard time keeping up with all this stuff.

But just priorities change. And even around this much data, the priority changes around this much data faster than we can, frankly, faster than we can move it, even in a snowmobile snowmobile. Yeah, no, I think there's a lot of situations where I remember from the very, very beginning in public cloud, that the concept was always to be flexible to the specifics of customers needs to be able to pay only for what you use, and use only the parts of things that you need and be able to programmatically manage all of that.

So you could be as nimble as your software is, and that pattern, you know, data centers have gotten a lot cheaper, and they've gotten a lot denser. And they I think they've gotten to a place where they are a lot more reliable, right? I think a lot of early people that I spoke with when they came to cloud went because the uptime was great. And their data center would go offline for like days and shit, right? Like I, you know, I remember working on those kinds of problems all the time.

Now I think, you know, a local data center probably caught up, like, the gap isn't nearly as big on reliability for base, you know, virtual machines or something like that. But then all of a sudden, you know, things change. And like VMware comes out and, you know, multiplies the cost of your data center by a factor of 10. And then you go right back to like, hold on a minute, how do I get out of this thing and get over into into GCP, I don't care how much the data cost to move, I'm the compute just got costed gazillion times more.

So somebody helped me figure out how to get out of here. So they, you know, I'm sure next year, Nvidia will change the pricing model on all these devices. And you'll keep swinging the pendulum back and forth between shit, I got to get everything to cloud or shit, I got to get everything over into into into my data center. But wherever it is that that stuff goes, I think it's critical for our customers and our users to recognize that what we see every day is just this incredible diaspora.

Their data is everywhere, their control layers are everywhere, their access and privilege, their logging is everywhere. The complexity we ask most CIOs and CTOs to take on this year is totally daunting in comparison to the way these kinds of roles worked even five, 10 years ago. So we have to give them a lot of grace because it's it's not an easy job. Yeah, I'm not the dust off my data infrastructure framework that I was working on a few years ago that turned out to be repetitive, but the research has already been done.

But the premise of it has been that it is much more important to put governance around your data than to worry about where it lands technically, because that's going to always change. Whether we're talking about SAS, IAAS, on prem really doesn't matter. You need to have a framework for how you're going to share data in between systems and you're going to need to have a framework around how you're going to secure that data, because where the economies of scale to put that data or where technically that data is going to be the fastest to services, all of that is going to change.

What's not going to change is your policies around the data itself. And you need to understand the interfaces around how you're going to access the data and how you're going to protect that data. So, Miles, it's always great to talk to you. Where can people find out more about Sodom? com is super straightforward. A bunch of solutions materials and 200 plus customer case studies. You got details on the 1500 successful migrations in the last couple of years. So we're seven times in a row, back to back to back to back to back to back to back partner of the year on the Google side.

And we do not intend to screw up the streak. I think if we get to 10, they'll like they'll call it a lifetime achievement award. And, you know, they'll invite us to sort of anoint to other winners. And they'll give us a spot on the. Yeah, no, you're just kind of like emeritus that you're your team. Yeah, yeah, you have the team. Yeah, and the email is not hard. It's Miles. It's not it's not like it's complicated.

I'm happy to talk shop with folks. Your last bit on like where the data goes. There's a great project back allow named after it's like, you know, code over data there. They're a remote execution framework. That's a really interesting way to try to put the computation wherever the data happens to be as a as a model approach. So cool, cool project from David there. That's cool. I have to try and check out check out the project and reach out to David to have a conversation.

All right. If you want to find out more about the CTO advisor, you can follow us on the web, the CTO advisor dot com and our parent company future group dot com, where we're doing all kinds of stuff like acquiring Textron. We have the intelligence platform, which I was opposed to drop a intelligence dime about, you know, AI. I'll make sure to do that the next time. But check out at that future group dot com. Talk to you next CTO advisor podcast.

Thanks a lot, Miles. Hey, thank you. Cheers.