Data Gravity - Reaching Escape Velocity
Transcript
all right i hope you've enjoyed my session rebecca session and if you haven't seen those you missed those at the end of the day beginning tomorrow you'll be able to watch all of the sessions live or on replay on this platform for the next 30 days joe honesty is up next talking about data gravity joe is the co-founder of transformational continuum where they help sales teams and organizations make that transformation into selling in this new environment this new digital environment check them out transformationalcontinuum.com uh joe
is about to speak on date of gravity this was a popular request amongst you i visited joe at his ranch drawing the cto advisor road show and we did a quick blurb on data repatriation and it was extremely popular a lot of conversation around that so joe agreed to continue the conversation and give this dedicated session i know he's known as a internet curmudgeon but he is a wonderful person presenter i think you'll enjoy it make sure you at him in the chat session he's there
hanging out listening this may be your only opportunity to talk to joel online until he goes back into his hobbit hole there in on the border of wyoming in montana let's get started take it away joe hello and welcome to this session data gravity reaching escape velocity in this session we will cover what data gravity is why it matters and how you can architect around it like any problem or challenge in technology if we know where it exists and we know what it can cause we
can typically architect around it and this session is designed to give you a baseline for how to think about data gravity and how to architect the way in which you deliver data around those challenges created by data gravity my name is joe onisick i'm a principal and co-founder at transformation continuum i've got 20 plus years operating architecting and deploying data center and cloud infrastructure at scale i'm a former javelin missile system expert so working on javelin anti-tank missile systems and i am a first generation cattleman
completely cluelessly diving into cattle ranching as an aspiration to becoming a bit of a cowboy i'm also a high school dropout and so i put all this together to say that it's not where you learn something that matters it's what you know and more importantly your willingness to learn something new and i hope to teach you something new in this session lastly i'm a big believer that there's no problem you can't solve if you can find the big enough wrench and so in this session we're
going to cover two big wrenches that you can use as tools to architect around data gravity problems let's start off with framing our discussion and i'm going to frame our discussion with a brilliant quote that i love from a friend dave mccrory dave mccrory is a hell of a technology thought leader he's been around for quite a while uh dave was a big big thought leader in the early days of cloud and still to this day looking forward at the way in which we architect information
systems years ago in a conversation with dave he said to me we are in the information technology industry not the data technology industry information has value data doesn't i'm going to pause there because that's a really important statement said in a very simple fashion what is the actual difference between information and data information is data plus the context to make that data relevant so when we look at information versus data we can take a single data point as an example let's use a temperature and let's
call that temperature 80 degrees fahrenheit that's a single data point that single data point on its own has no value to me i'm going to need some context to make that data valuable so it's 80 degrees fahrenheit is that measuring the outdoor temperature now this is becoming more valuable based on context is that the outdoor temperature in the arctic circle or in the sahara desert those very different contexts will change the way i interpret that data as useful information in our world of information technology the
way in which we turn data into information is via the applications and services that utilize that data to service the people things and machines that need it so we use apps and services to provide context to turn data into information and that marriage between the apps and services that provide the context using the data to create information is what we'll be talking about for the rest of this session that leads us directly to the concept of data gravity so let's start off by understanding what gravity
is we all know its effects but the definition is very simple it's the natural attraction between physical bodies so one physical body will have a natural attraction on another physical body and the larger the mass of the first physical body the larger that attraction will be as mass grows the gravitational effect will grow with it so when we look at that from a data perspective we start to look at placing a data set in a given location that data set is going to have its own
gravitational effect that gravitational effect will be applied to your applications and services what you'll find is that your applications and services gravitate towards the data that they rely upon so if we take data and place it in a data center the tendency will be to place our applications and services in that same data center we may replicate that data to a different data center and then place new apps and services in that second data center for production requirements or disaster recovery business continuity requirements whatever it
may be the same effect occurs if we place that data in the cloud so if we place this data in our cloud provider of choice the odds are that our applications and services will gravitate towards that same cloud provider what they'll get by being local to that data set is typically a series of benefits the first being a lower latency connection to the data by being closer they will have faster access to that data which typically relates directly to application or service performance and possibly user
experience they will also tend to benefit from higher bandwidth or larger connections and lower costs for those connections so these are the things that create that data gravity effect for your apps and services pulling them closer to that data set our cloud providers know this and they utilize it to create more revenue more value more profit so if we look at these examples which were pulled on july 4th 2022 utilizing one of the major cloud providers online pricing and pricing estimation tools we can see a
few very interesting things first general purpose storage in the cloud tends to be very very cheap it's hard to beat the pricing of general purpose storage in the cloud when you look at storage and isolation we're looking here at about two plus cents per gigabyte for storage in a general purpose fashion getting my data into the cloud is most often free and if it's not free it's the lowest cost of most of the services you'll look at within the cloud so i'm allowed to put my
data into the cloud for free and store it there very low cost but that data doesn't live in isolation apps and services have to access it users have to access it through those apps and services so that's where the costs start to truly come in if we look at getting data out of the cloud or the egress charges for the cloud that's one of your larger costs for the cloud here we're looking at 9 cents per gigabyte to access that data from the internet that is
four times as much as storing it it's a 4x cost difference to access that data from the internet whereas internally within that cloud that access may be no cost or low cost for your applications and services that utilize it then we start to look at the value add services that come on top of that if i have data in one location and something happens to that location either the data itself or the location itself now i've lost all of my data so i'm probably going to
want to do things like replicate that data for the purposes of backup and recovery i may also want that data distributed across the globe and accessed via the optimal path so i may need advanced services like multi-region routing if my data stored in a private data center or co-location those are my responsibility if my data stored in the cloud the chances are i'll need to or want to rely on their resources and services to do that routing to the data and again at an additional cost
so if we look at the basing basic pricing example on the right what we'll see is that storing the data tends to be about half the cost of using the data or accessing it once it's stored so the cloud provider is incentivizing you to get your data into the cloud because they're going to charge you to use that data and this is an important factor to understand it's neither bad nor good it's just important to know that this is how it's working and one of the
ways that they drive additional revenue additional costs within the cloud now why does all of this matter to you why data gravity matters first it can affect my operational flexibility when i look at the effects of data gravity as my data set grows in a given environment private or public the tendency to place applications and services close to it will grow with the size of that data much like gravity grows with the size of the mass so my operational flexibility is affected because now my decision
of where to place an app or service is going to be heavily dictated by where my data exists so if i have one app and some data and i'm choosing where to put that app and data i have all the options on the table but if i've already built 500 apps and terabytes of data in a given location the tendency will be to need to place the next app service or data in that same location so i'm reducing my choice by increasing the pro and con
of where i place that data minimizing my flexibility or at least reducing my operational flexibility my operational choice i also have to worry about feature and service availability in the given location if there's a feature or service i require to utilize on my data then i'm going to need it available where that data is so if i start to build large data sets in location a and location a doesn't have the new feature or service that i require now i have to figure out am i
going to have that feature or service access it externally and incur those additional costs or am i going to figure out a way to build or replicate that within the given provider i'm using and then finally cost optimization wherever my data is it will typically be cheaper to access that data locally so i will optimize my costs by having those apps and services next to that data but i may not always want those apps and services next to that data or i may not always be
able to place those apps and services locally with that data what we're trying to avoid as we think through data gravity is the creation of data gravity black holes i work with clients around the world who are already dealing with data gravity black holes that they've created they started to use cloud provider a because it was the cheapest per cost terabyte of storage they could get and so they decided to move from on-premises storage to cloud storage for a cost savings and as that data set
grew in a given cloud their applications and services naturally gravitated towards that cloud and so now that data has created its own form of lock-in to that cloud or its own data black hole now maybe it happened in the opposite way maybe they began using some cloud instances to roll out new applications and services well as those applications and services rolled out they created data and more data was migrated to that cloud gaining in mass and growing its gravitational effect again creating these data black holes
so how do we avoid data black holes we need to consider the way in which we design our data the way in which we design where our data will be located and how our data will be accessed when we consider this let's use the example of cloud storage with that example getting my data into the cloud the ingress charges to the cloud are typically minimal or no cost i can get my data to the cloud in many different fashions at typically no cost to me or
my company now when i want to start accessing that data either via another region in the same cloud or via the internet i'm going to start to incur egress costs or costs for accessing that data externally two things increase as i do this the first will be cost how much i'm paying to access that data and the second will be latency the speed at which i can access that data the further away from the data i get the higher the latency or lower the speed of
access will be and again this typically will have a direct effect on things like performance and user experience very important when we start to think about things like real time data or near real time data that cost and latency become your two major considerations or two major factors to design for when you're looking at where to place your data and which services to use for that data there are several other factors that need to be assessed but these will sit typically at the top of that
list so how do we alleviate the challenges that come with data gravity how do we escape that data gravitational pull how do we reach escape velocity for our data escape velocity is the minimum velocity that a body must attain to escape a gravitational field completely the easy example is a space shuttle when a space shuttle wants to leave earth's orbit it has to reach escape velocity it has to reach a certain speed to be able to break free of that gravitational pull or gravitational effect so
how do we do this when we're talking about our data the first big wrench we'll talk about is data locality data distributed based on its requirements by utilizing data locality i can place the data in the right public or private environment where it's local to the applications and services that will rely upon it most by designing where i place my data i can decrease my costs and decrease my latency therefore theoretically improving my performance and user experience by designing my data to sit resident with its
apps and services i can improve the way in which i operate with that data so for instance i have some applications on private data center that are going to stay on private data center and rely on specific data in that data center let's leave that data there locally where it's going to have fast low cost access i may need services from cloud provider a that provide some functionality that i require and they rely on specific sets of data that i can place there in cloud provider
a the same on with sas providers and other cloud infrastructure providers wherever i need services and applications i can design my data to primarily sit there locally and while there will be exceptions and data sharing across of this the majority of my data can be placed locally to the apps and services that utilize it and the less locations i end up requiring the easier data locality will be if we look at this what we're creating is locally sourced hormone-free artisan data i've got my prehistoric legacy
applications that i'm never going to re-architect even though i kid myself into thinking i will they'll run in permanence and i'm going to keep those here on premises and so on premises i have the bulk of the data that those legacy applications will utilize i'm also going to start building new applications in the cloud and so either utilizing multiple cloud providers or multiple regions within given cloud providers for redundancy and resiliency purposes i'm going to place the data for those new modern applications resident next to
those apps and services in the cloud and then maybe i'll have some other data that's being used for artificial intelligence machine learning advanced operations that gets placed in that cloud and that data may be non-persistent meaning it's used while processing and then disappears when not processing so that the answers or results can be fed back to other systems or maybe it is persistent and it's maintained there for the purposes of analytics on that data in either event what i'm looking to do is place the bulk
of my data as close to the apps and services that require it wherever possible what i get from this is data that's local to that app and service which will tend to mean lowest cost access highest performance access although caveats of course exist it means that i can have replication where required for instance if some of the data being used by my prehistoric apps needs to be accessed by cloud a or region b i can replicate that data to the local data store and have it
accessed there locally by those applications if it's a cloud provider where that data store exists moving that data to the cloud provider will typically cost me nothing as long as that data isn't modified in such a way where it needs to be egressed back to my local data store my data transfer fees will stay low to no cost and then next i get those minimal egress access charges because most of my processing is being done locally in any given location i'm minimizing the times where data
has to be moved out of a location and incur any additional costs that may come with that like all decisions there are pros and cons to data locality let's start with the pros represented here by my professional who's accessing data locally there on that construction site what he's experiencing is optimal performance the data is sitting locally to the applications and services that turn that data into the usable information he needs to make the decisions on site they're also experiencing an optimized accessibility cost by reducing or
removing egress costs the overall costs are coming down theoretically now we also have some cons represented by my convict here at the bottom we have increased complexity i have to design my data for where it will be most utilized and this is easier at a small scale than it is at a large scale but it does increase the complexity across the board in a previous life i was a storage architect and i are architected the storage and network solutions for large-scale data systems and in those
days we primarily use spinning disk when using spinning disk we had varying costs of disk and so one of the things we would often try and do is use different disk tiers for the cost and performance benefits of those specific tiers and what you would try to do is make sure your data was always existing on the given set of disk that provided the optimal performance for that data at the lowest cost possible and this was easy when you got started with a new data set
but became far more complex over time same thing applies here as we start to look at data locality across a cloud multi-cloud or private infrastructure the other thing we end up with is data fragmentation and so data fragmentation on its own is not necessarily a good or a bad thing but it's definitely a consideration to think through oftentimes we'll need access to that data for multiple purposes security scanning logging analytics you name it so the more fragmented that data becomes the more complex difficult and possibly
costly it will become to utilize those tools so the data fragmentation created by driving data locality can become a problem so what other tools do we have in our tool bag let's talk about the next big wrench the next big wrench is data neutrality data maintained in mutually accessible locations a neutral location or demilitarized zone dmz for data represented here by the dmz on the border of south and north korea when we look at data neutrality what we're trying to do is place the bulk of
our data in a location that has the performance requirements for our data access at the lowest possible cost regardless of where this is being accessed from so for instance if this data is being stored on private infrastructure typically my bandwidth to the internet is already paid for i'm paying for access to a pipe of a certain size at a certain latency so the data ingress and egress charges aren't specific to the amount of data i'm moving they're already included in the costs i'm looking at so
therefore we're negating a lot of the egress charges or reducing the egress charges in some cases by creating data dmz's we can create a single data store that allows us to access that data from anywhere without additional or excessive costs what this becomes is a concentrated animal feeding operation for our data or a factory farm for our data so what we can see here is i'm utilizing private infrastructure either a data center i owned or a co-location i lease to hold the bulk of my data
that data is then made accessible to the various cloud infrastructure and cloud software that i utilize to run my it operation so one local data store or set of data stores accessible by all of the different applications and services wherever they live now what i'll need to do is ensure that i'm getting the optimal performance and latency requirements to be able to do this and that's possible both in private infrastructure and in co-location infrastructure this isn't necessarily going to mean all of my data is going
to sit locally or sit in that neutral zone there may be instances where i also have some data somewhere else and in this graphic we're representing that with some non-persistent data being used for artificial intelligence or machine learning processing in a cloud that's providing those services what i get out of this is centralized data store in a neutral location or data dmz one location free of fragments for the most part i get high performance connections if i properly design for this so the ability to do
that is going to be dictated both on the data center i have the service provider i use and the physical locality i have to the resources that need to access this for instance if i'm using cloud provider a and the data center i'm utilizing whether my own or a co-location lease is physically located near a cloud on-ramp for that provider i'm going to get latency performance and bandwidth that's comparable to being directly inside that cloud so that's an important thing to design for and consider as
you're looking at this if your data center sits in the middle of nowhere and you're utilizing this concept of a data dmz in that data center all of your connections will experience a higher level of latency and the other advantage i get here is minimal egress access which is similar to the data locality i'm minimizing the overall egress charges because what i'm typically doing is using data ingress to get the data from my local neutral data store to the given apps and services residing in other
data centers or cloud environments when we look at data neutrality we're also looking at some pros and cons my pros are represented by my professional here at the top she's experiencing normalized performance because i've used a data dmz or utilize data neutrality to provide consistent access from a central data store to all of the various environments where my apps and services that rely upon it reside we've also eliminated the access costs that come from egress charges and the various services that assist in data egress or
at least minimize those overall costs but again we have some cons and those are represented here by my convict we have some performance considerations will probably get the best performance by having the data accessed locally so if i'm using a data neutral dmz now i may have some slightly increased performance for all of my access but that will again be normalized performance which is easier to predict and easier to architect for so it's kind of a two-sided coin and i also have some major design requirements
where can i place this data where will i get the appropriate cost and performance metrics to deliver this data to my applications and services wherever they may reside so this can create some lock-in represented by these handcuffs locking me to that local data store very similar to the login i may get if i place everything in a given cloud so again that data gravity effect will be created as i use data neutrality but that data gravity effect will be created in a location where i'm not
incurring additional costs and charges for accessing that data from a separate location as i would in a public cloud environment the important thing to remember here is that data locality and data neutrality are not mutually exclusive options we don't have to look at this as either or we can look at this as and i'm not a fan of compromise in fact i hate compromise i'm an and guy why choose between cake and pie when i can simply have both so how do we look at data
and locality and data neutrality in a both system the and approach the and approach starts with a foundation of private data center or co-location utilized as a data dmz so what we can see in the bottom of this graphic is three connected data centers being utilized for data replication providing business continuity and disaster recovery for that data as a central data store in a neutral location or data dmz this will be the bulk of my data in the and approach in fact my recommendation when using
the and approach is that everything goes to that data dmz unless there are specific requirements that force it to live somewhere else for performance cost or other reasons those specific requirements i recommend you pre-define define those up front it goes into the data dmz unless x y or z this means it will never be a tough decision when you're trying to decide where to place data apps or services and the point of this is that you always want the bulk of your data to exist in
these neutral locations or data dmz's which can be private data center that you own or private data center co-location that you lease from here we utilize data locality on top of this data neutrality foundation that we've built so we can then have specific data sets smaller data sets that exist within the given environments both private public infrastructure as a service and sas that we utilize to deliver technology so i've got the bulk of my data in my co-location or private dc data dmz and then where
i need specific data for specific application and service processing i can have that data reside locally that data may be non-persistent meaning it changes and is never really stored or persistent where it's kept and replicated and backed up like we traditionally think of as data but in either event what i'm trying to do is minimize the fragmentation where i'm putting data in different locations but still utilizing that data locality to get the performance and cost metrics i need for specific applications and services this becomes extremely
applicable when we think about two things that are very common in the it world today the first is going to be real real time or near real time processing anytime i need real time or close processing i'm going to need that data very accessible and we need to remember that even the fastest links on earth are still bound by the limits of physics the fastest we can move data is going to be the speed of light plus the serialization and deserialization time to get that light
off the wire and into the system so the further away we are from our data the more latency we induce or the slower that access becomes meaning we quickly negate the ability to really do real time or neil real time so in those instances i will want that data very close that same concept applies to the field of edge computing where we're looking to do things like have agricultural iot monitoring a large scale farm or ranch analyzing that data in real or near real time there
locally and then providing back out results feedback or computation to larger scale systems on the back end in those locations in those instances i will probably want some locally stored data for that access maybe the the data coming directly off those iot devices which gets processed to some level and then fed back up to a more centralized system so as we look at that and approach again we haven't created a mutually exclusive relationship between data locality and data neutrality we've combined the two to get the
best of both worlds where we're reducing cost and reducing latency while providing a robust and flexible operational model for the way in which we engage with and access our data so let's take this down one level in strategy to a more tactical view of how do we approach this getting started into how to look at data locality data neutrality or even ignoring data gravity starts with a big question and that big question may actually be the easiest question you answer when assessing and architecting for data
can you standardize on one cloud or one set of private infrastructure in some cases the answer is yes and there are big name companies who have made the decision to do this because while they may incur some costs and some reduction in flexibility from their operational perspective what they get is an overall efficiency of operations that allows them to use one provider for all of the data apps and services that they need so in some instances this will work or make sense for your business and
in some instances this will not work or make sense for your business or mission so it's the first question to ask if you can meet your performance and cost metrics utilizing a single provider that may be the correct choice there may be other factors that negate that choice or other factors that support that choice but it's the first thing to think through if you can utilize a single cloud or single private infrastructure for all of your data and compute needs then the question is which which
cloud which private data center and the primary questions that will factor into that are going to be cost metrics the available feature set in that given location so for instance cloud provider a has a different feature set than cloud provider c versus cloud provider b but also what features can you replicate in a private data center or co-location for instance and then finally supportability your staff's ability to support that environment if you have a staff that has 20 years of private data center infrastructure experience and
zero years of cloud experience that is going to factor in to where you're going to place that data and the apps and services that rely upon it if one specific cloud for all or most of your needs is not a solution that works for your organization now the next question is do you own data centers if you do own data centers are they usable for what you're looking at now this question becomes far more complicated usable is going to come down to the power and cooling
density for modern applications that's required it's going to come down to the cost of those facilities the cost of that maintenance it's going to come down to the longevity of those brick and mortar facilities were they built 20 years ago and ready to be phased out or were they built last year and still depreciating over time and then finally are they in a location that can provide the connectivity requirements you have from a cost performance and latency perspective so if your data centers in herndon or
reston virginia where there are plenty of high-speed fiber backbone connections and there are large cloud data centers locally resident you may have a yes if your cloud if your data center is in the outskirts of kansas city maybe that's a no it's very much going to depend on what you have available to that data center remember you're going to want to have the lowest latency and lowest cost connections to all of the apps and services that are going to rely upon that data if you don't
own data centers then the question is should i lease co-location space and the amount of space you'll need to lease can be as little as a single rack with some storage and connectivity requirements in that rack to be able to provide your data dmz or data neutrality all the way out to entire cages of data center computing and application environments what you typically can get from the larger tier one co-location environments is rapid low-cost or no-cost cloud on-ramps often those tier one co-location environments sit locally
very close to the cloud provider data centers and often on the same fiber backbones what this means is that your access speeds and access costs for co-location to cloud are comparable or similar to your access speeds within a given cloud and if you're not being up charged for that cloud on-ramp access then you're reducing a lot of your overall cloud connectivity costs so will you lease does the co-location you're choosing to lease have the appropriate connectivity options to create this and model of data neutrality plus
data locality and are they close enough to the locations where you will do the primary bulk of your data processing or application and service the primary takeaway i want you to bring from this is that regardless of whether you choose to accept data gravity as it is to utilize data neutrality to alleviate some data gravity concerns or you use data locality to alleviate some of those concerns or some other method what you want to avoid when you're thinking about distributing applications services and access for a
global i.t delivery is free range free range is good for cattle free range is bad for data and i used my little memorial day calf who came as one of two calves this memorial day to represent my my free range cattle so there you can see him with his mom happily free-ranging the reason this calf and his mother are free-ranging happily is because they're not data do not let your data roam wherever you wherever it gets built or wherever it gets put as you start to
scale that out as you start to develop a little bit in this environment and a little bit in that environment reining it back in is extremely complex and when you start to think about the security implications the performance implications and the cost implications this is something you very much want to think through as early as possible in your design phases let's recap we are in the information technology industry not the data technology industry data itself has no value the data needs the context to create information
that is usable by our businesses the way in which we get the context to turn data into information is through the applications and services that rely upon that data to serve the users and devices that require it this relationship between data and the apps and services utilizing that data creates an effect that dave mccrory coined the term data gravity for the data itself creates a gravitational pull bringing the apps and services close as the data set grows the pull or gravitational effect grows much like as
mass grows in the real world the gravitational effect of that mass grows with it this gravitational effect known as data gravity is not bad or good in and of itself it's simply a consideration that you should be thinking through when you're deciding where to place your data and or where to place your applications and services if utilizing a single provider public cloud or private infrastructure provider for all of your data and application service needs or at least most of your data and application and services needs
will work for your business or mission then you don't need to worry much about data gravity simply be aware of its existence for most of us that will not be an option as we live in a very multi-cloud world that's being further distributed by the idea of edge computing we discussed two tools or big wrenches as i described them the first was data locality placing the right data in the right location based on the requirements of the apps and services acting upon that data and the
second was data neutrality utilizing a neutral location or data demilitarized zone dmz to place the bulk of your data therefore normalizing access and reducing costs by the apps and services regardless of where they reside finally we discussed the and method utilizing data neutrality as a foundational element to store the majority of your data reducing costs through things like egress charges from the cloud and using data locality to further benefit that by placing specific data near the applications and services that may utilize it an example of
that may be local storage for 90 and utilizing cloud storage next to cloud services that are providing ai and ml use cases and lastly we walked through a simple set of decisions that can get you started along the path of understanding where you can place your data for your organization's specific requirements while that is very much not a comprehensive list it helps to get the thought process started as the first step in looking at where you place your data again my name is joe onisick i
am a principal with transformation continuum you can contact me at joe at transformationcontinuum.com if i've done my job right you won't find me on social media but i'd be happy to answer any questions comments concerns or general hate via email thank you very much i hope this was a great use of your time