AWS Snowmobile is No More: Chat with Miles Ward

25:07 · Watch on YouTube ↗

Transcript 4,478 words · about 30 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

all right so I'm coming out of Google Cloud next a couple of weeks ago and I cannot tell you how many people referenced a CTO advisor podcast episode and I felt a little bad because I haven't been producing the podcast as consistently you know I'm with this much bigger company future I had a I have a big title I'm a global advisor I've been upgraded from the CTO advisor to a global advisor what does that mean uh I don't know and we'll get into that later

but the opportunity to talk to one of our previous guests one of my uh uh I think more entertaining guests mes Ward CTO of Sada welcome back to the podcast thank you Keith I appreciate that I I I aim to be entertaining I I was I was super excited to reach out to you about this one because uh because my they killed my baby uh the the Mobil is getting shut down and I I know too much so I'm looking forward to talk shop so I'm

looking forward to this conversation because I know this uh you know as I looked at the announcement I kind of said H I don't know the details of why they didn't share why but I could infer to why they would some reasons why snow bill may be a little bit past its prime but let's talk about you know kind of the origin of snowall Bill break it down to us what is it what was the use case for it and you know why does any of

this even matter sure so I I was in the first 100 employees at AWS and one of the most incredible folks there uh built out the whole uh patent program he said if you if you come up with something you think hasn't been done in Cloud I want to hear about it is it great so I sat on a plane next to a movie producer who described the complexity of onsite shoots with 3D cameras 120 htz all the color resolution that they need and they were

bringing out half racks of storage equipment and they would have to call stop on set because they ran out of disc space to be able to keep the cameras rolling now that all that sounds like a bad problem in the meantime I've been working uh selling Apple S3 and helping everybody consume pedabytes on the S3 side and and I'm looking at what you can actually get in half rack locally deployed going that well of course that sucks like that's just going to be terrible so I'm

I'm sitting with this producer who's he's a pretty technical guy too he's bought a bunch of gear and we start jinning up this experiment we want to run like if I put a bunch of storage gear in a Prius then I'd have a generator on site that's kind of cheating I can bring gas cans with me I can carry way more than you can put in like a shippable storage container because I can just drive it there and if I showed up with wire I could

run it out to the you know Run 100 Gig Ethernet or whatever you know melanox or whatever cable connectivity you want uh and then I could drive it back to one of the Amazon regions and plug in and VOA it would all just be objects in S3 I laid out all this crazy technical [ __ ] about how to do the replication on Amazon's the S3 system is built on this thing called a tree topology whatever so I I lay this whole design out and I

explained this to Eric the uh the the crazy inventor lead and he's like well that's nuts we should we should file a patent for that so we do that and then I bail I go I end up going to Google I'm I'm working Eric Schmidt pulls me over to work as the first Solutions architect for Google Cloud uh and they drive that snowmobile on stage like what the [ __ ] so so I sent a chat message to wner I'm like Werner is that the Prius

that's an awful big Prius and uh and he's like oh if you were still here we'd have let you drive it on stage you made terrible choices so that's just like a knife in the chest wers digging me around in there but the use case the core of that is like if I can put a 100 pedabytes local to a big data center facility that's got hopefully more than that storage and I and I'm going to truck it back to any one of the AWS facilities

at the time there were only three major regions in the US now there's more but uh we calculated if you filled the box which is just a full 40ft shipping container and drove it from New York to LA you were doing so at six 6.4 gigabits a second which is pretty [ __ ] fast so or not not g sorry terabits a second 6.4 terabits a second so that's like actually actually so that uh you know there are plenty of data center facilities that piled up

storage but just did not have the connectivity to be able to effectively migrate to cloud and that would make the migration plans look like two year three year four year five year things where Amazon and all the cloud providers frankly are trying to do them way faster so they saw it more of an enabler for non-cloud to Cloud migration and dragging all that data up to the farm I always thought of it as kind of more of an edge enabler where you'd be able to tear

around of course it's kind of hard to drive a 40 foot container truck to your average movie set but uh but that was that's that's kind of how it all got started so that's a great origin story and you know snowmobile snowball snow snowball Edge I think was the Y kind of HCI type solution to build out I think obviously this especially with Sovereign cloud and Dora and the in the in Europe people are trying to figure out how to make use of local data stores

to one being compliance with regulatory concerns and then two even with the uh challenges around availability around accelerators the fastest data transfer still the data transfer you don't have to make so talk to me a little bit about we'll go on a little bit of a tangent and we'll get back to like the demise of the snowm snowm what are you seeing when it comes to customers and their challenges around data and data sovereignty sure I I helped design really early on uh a bunch of

Google's push into what now they call uh Google distributed Cloud hosted which is a system that allows them to run anthos and the underlying container management infrastructure and the VM management infrastructure that's required to stand up effectively a private region of Google Cloud so today uh Talis the um the effectively like infrastructure and train company in France uh t- systems in joint with uh T-Mobile and the other components in Germany uh those two as well as a couple of others that are in planning are already

deployed functional Sovereign clouds in Google's ecosystem and Amazon and Azure have similar building blocks I think Oracle at one point was advertising if you just sort of stroke the pen efficiently enough they would deploy an entire region wherever you wanted one so everybody has gotten to a place where at least some flavor of their offering is Deployable in a way that is s the complexity for all that in most cases is that in order to really do that not only does the Gear have to be

in your location Sovereign and the software has to be licensed and owned in a way that that is Sovereign to the legal jurisdiction that you're talking about but the people the staff have to be citizens or uh you know legally authorized to do that kind of work so like I remember taking a long phone call with the folks from TSS when we started out with them and it was all about what is Sr how do we train people to run this thing they're going to hand

us call the Region because Google won't run it for you they wouldn't be allowed to they they don't have they're not going to put hundreds of French citizens in the French region to sort of operate it on your behalf they expect if you're going to deploy this thing that you as a provider want to do that work that's why you're reaching out about asking and taking it on I think the last thing Google wants to do is vend a bunch of staff to other companies so

that uh that whole line of questioning and the complexity there I mean like you think it's hard to get certifications on AI stuff or Solutions architecture work or the other kind of enablement there's there's no good training courses today on what does it take to run a Google Fusion or what does it take to run an moral region or a n region or any of the rest of that it's very specialized stuff so I think as much as countries want it I remember mentioning to S

at one point that this seemed like the most tractable problem in computer science right there's only 192 or 196 countries just deploy a region in all of them and then you're done like you're just actually done you have one in every jurisdiction success uh but uh but there there are none of the companies are even close to that I think the biggest country coverage right now is is azure with something like 16 so it shows goes to show you how much there is left to do

to to get this into everybody's legal jurisdiction yeah and you know couple with that security concerns data gravity which we'll get back to uh you know we I think I came up with this concept I didn't go out as much as Dave mccu did who came up with the data uh gravity concept and Theory I've had him on the podcast a couple of times to talk about it there is you know there is a actual Theory there is algorithm behind it but I challenged him a

couple of years ago and said you know what services are starting to have gravity uh uh and he calls it more inertia I don't care what you call it but I can't get vertex inside of my data center as easily as I don't know the the uh you know I need vertex so I in in in order to get vertex I need to ship my data somewhere for that to happen because that service is that important to me to be able to take advantage of it

so you know we're having these extremely unique challenges where we can't necessarily build and you stated this we can't build the services Cloud providers can build but yet our data is on S hence the snowall Bild you know okay let's let's get the the value is worth the risk let's get the data up there we need to get it up there and and you know at six gigabits per seconds uh terabits second yeah terabits I'm sorry terabits because I've had gig connections I've had 100 Gig

Connections in in the data center before and they're just not big enough so you know you're you're at the practitioner level you're doing massive Pro projects what do you think is demise of a of something like a snowmobile so I I mean I led the migration for Twitter before it was called X uh from their private data centers into Google Cloud's facilities and they had 360 pedabytes active in what is inarguably the largest Hado cluster in the world and they didn't like it anymore because hadoop's

terrible uh and big query is a lot better I was just about to say I'm sorry that's that's a l yeah that's that's a heartbreaking state right like this is the lovely part of my career like like hop inarguably absolutely saved my startup like reduced our cost by three orders of magnitude to build the indexes that we need to in comparison to SQL server and now just 15 years later we sort of giggle about it and of course you got to tear that [ __ ]

out of your data center so the when we did the deployment for uh for Twitter that you have exactly this option right like somebody could bring a snowmobile out somebody could bring a large format there's a couple of competitive services from equinex and megaport that would do the same kind of thing and we looked at their existing physical footprint and built uh 100 gbit a second GB connected links over Juniper devices into Google's facilities and transfer the data over 43 days um and if you think

about what kind of a pain in the neck it is to find a parking spot for the truck and get the wire through the security doors into the racks and then make sure the security people are eyeballing at all the correct times a day and night all that stuff might take more than 43 days to figure out so I I think at the end of the day and we were moving three times three and a half times what would fit in a snowmobile if even that

was available as an option so I think what what Amazon certainly found and and Google found quite a bit earlier is they have a network that's some seven times the throughput of the public internet so they are way leaned into to grabbing fiber everywhere uh is that the uh you have to have a lot of data in someplace really weird where it's hard to get big fiber in order to be able to uh in order to be able to have a solution like a snowmobile actually

makes sense and so that it ended up not getting used all that often um I think you know the original one that I proposed that was a lot smaller probably would get used a lot more often because the latency would be commensurately lower right we had all the time customers would show up with you know two suitcases full of external drives to a Google office and say I want to move to the cloud uh you guys band with like God like can you can you help

me plug all these things in we're like this is not actually what we do for but yeah sure [ __ ] let's let's roll right and we would set everybody up and we'd move a bunch of Daya that way so I I think uh you know the metro area City latency of driving around is probably viable like semis siize was maybe better for the press release than really the actual use the Practical pieces of it and then you I think you make a a huge point

that when it first came out there was three regions now there's regions there's zones there's Parts with equinex other data center providers so if you can get your data relatively close to one of these providers you can ingest the data much quicker than you can you look at the the size of hard drive of of ssds now I held in my hand doing uh a GTC uh solid has a 61 1.44 terabyte which for the not solid you're listening to this you promise you get me

one a 61 terabyte SSD low power devices you know you you have a suitcase of those and you get it into a uh collocation that has the direct connects directly into uh all the cloud providers the cross connects into the cloud providers you have essentially near Cloud speeds capabilities that can either ingest the data or leave the data is there you know Google has big cury you can do amazing things with big cury across platforms so I think the the idea was cool like you know

the truck I've been trying to get folks to sponsor putting the data center in the in the CTO advisor uh yeah Cloud which I'm I'm I'm recording from now cool idea it's niche it gets people it gets eyeballs coming people are going to come look at it but at the end of the day I think we're finding that there's just more practical Solutions yeah I two two things there right that you were talking earlier about uh you know about the the kind of distribution of workloads

the things you can build in Cloud you can't build in Data Center and it'd be nice if the software showed up there Google is clearly pushing in that direction the the first real offering there is this um they they set it up so that the vision AI model could be run in anthos on PR so that's them taking a Google managed Google supplied Google Certified all the various slas and contractual obligation bits and they put that thing in in a facility on your side now I

think most customers were like well that's one API I want the other 86 of them uh and and would have preferred the entire Regional footprint the last time I talked to Google about the actual total Regional Services count because remember Google cloud is made out of stuff some of which was built for Google cloud like say gke which was designed in in in in concert with the Google Cloud team and some products which predate Google Cloud by quite a lot so like big query is Dremel

is a pre dorg hosted runs on the internal Google 3 data systems that are distinct from the stuff that's built above that layer in Google so any of the really shiny [ __ ] that Google built for itself before they even thought about cloud is really difficult to pry out and cram someplace else even given that difficulty they figured out how to do it on the other clouds so big query Omni allows you to run the big query code base on Amazon or on Azure you

can write a query that federates data from S3 and GCS and get results in the big query terminal and it does the processing Amazon side it's it's a incredible architecture they built really built a crazy thing so are they also have data flow Omni and a couple others so I think over time with the Omni initiative they're going to try to get as many of these Services running in as many facilities and hopefully more of the bare metal facilities and local customer facilities as possible uh

the other toy if you want toys uh Western Digital just came out with this thing it's called an ultra star transporter uh and it's kind of looks kind of like one of those halberton suitcases right like a little tiny attache case but it's 38 84 terabytes so like that's my kind of size it's perfect right like it fits in a 19inch if you want to slide it into a rack rail it doesn't have many rails off hanging off you got to clip that stuff on comes

with a little carry handle and you can put you know a third of a petabyte in in your carry-on like that's that's perfect that's just great yeah and and I think these types of solutions uh I don't know what the the lead time was for to get a sow bill out to Market but you know the things change fast like yeah man you know you you have to adapt in the world last time we talked we talked about you know the difference in cloud services from

when you ran uh infrastructure for the uh political campaign you worked on to now and how much the cost structure would be different oh man then you know radically different quickly AI the models I mean llama 3 was just announced I I don't know how many parameters llama 3 I I have a hard time keeping up all the stuff but just priorities change and even around this much data the priority changes around this much data faster than we can frankly faster than we can move it

even in a snowmobile snowmobile yeah no I I I think there's a lot of situ ations where um I remember from the very very beginning in public Cloud that the concept was always to be flexible to the specifics of customers needs to be able to pay only for what you use and use only the parts of things that you need and and be able to programmatically manage all of that so you could be as Nimble as your software is and that pattern uh you know data

centers have gotten a lot cheaper and they've gotten a lot denser and they I think they've gotten to a place where they are a lot more reliable right I think a lot of early people that I spoke with when they came to Cloud went because the uptime was great and their data center would go offline for like days and [ __ ] right like I you know I remember working on that those kinds of problems all the time now I think you know a local data

center probably caught up like The Gap isn't nearly as big on reliability for base you know virtual machines or something like that but then all of a sudden you know things change and like vmw comes out and uh you know multiplies the cost of your data center by a factor of 10 so we're right back to like hold on a minute how do I get out of this thing and get over in into into gcp I don't care how much the data cost to move I

the compute just got costed gazillion times more so somebody helped me figure out how to get out of here so you know I'm sure next year Nvidia will change the pricing model on all these devices and you'll keep swinging the Pendle them back and forth between [ __ ] I got to get everything to cloud or [ __ ] I got to get everything over into into into my data center but uh wherever it is that that stuff goes I think it's critical for our customers

and our users to recognize the what we see every day is just this incredible diaspora their data is everywhere their control layers are everywhere their access and privilege their logging is everywhere the complexity we ask most cios and CTO to take on this year is totally daunting in comparison to the way these kinds of roles worked even five 10 years ago so uh we have to give them a lot of Grace because it's it's not an easy job yeah I'm have to dust off my data

infrastructure uh uh framework that I was working on a few years ago that turned out to be repetitive but the uh of research that's already been done but the premise of it has been that it is much more important to put governance around your data than a worry about where it lands technically because that's going to always change whether we're talking about SAS inra uh IAS on Prem really doesn't matter you need to have a framework for how you're going to share data in between systems

and you're going to need to have a framework around how you're going to secure that data because where the economies of scale to put that data or where technically that data is going to be the fastest two Services all of that is going to change what's not going to change is your policies around the data itself and you need to understand uh the interfaces around how you're going to access the data and how you're going to protect that data so Mouse it's always great to talk

to you where can people find out more about s sure.com is super straightforward uh bunch of solutions materials and 200 Plus customer case studies you got details on the 1,500 successful migrations in the last couple of years so we're uh seven times uh in a row back to back to back to back to back to back to back partner of the year on the Google side uh and we do not intend to screw up the streak I think if we get to 10 they'll like they'll

call it a Lifetime Achievement Award and you know they'll invite us to sort of anoint other winners and they'll give us a spot on the yeah you just you just kind of like emirat that you you your team yeah yeah you you have the team yeah and yeah and the email is not hard it's miles it's s it's not like as complicated I'm happy to talk shop with folks your last bit on like where the data goes there's a great project uh back alow named after

it's like you know code over data they're they're a remote execution framework that's a a really interesting way to try to put the computation wherever the data happens to be as a as a a model approach so cool cool project from David ear there there cool stuff try and check out check out the uh project and reach out to David to have a conversation with them all right if you want to find out more about the CTO advisor you can follow us on the web the

CTO advisor.com and our parent company fugm group.com where we're doing all kinds of stuff like acquiring Tech strong we have the uh intelligence platform which I was supposed to drop a intelligence dime about you know AI I'll make sure to do that the next time but check out app futurum group.com talk to you next CTO advisor podcast thanks a lot miles hey thank you cheers