The Data Gravity Index with Dave McCrory

In this podcast, Keith and Data Researcher Dave McCrory( @mccrory ) touch on the concept of Data Gravity. Dave explains how it’s measured and the related benefits of tracking this science. The CTO Advisor The Data Gravity Index with Dave McCrory Play Episode Pause Episode 1x 00:00 / Subscribe Share Apple Podcasts Spotify RSS Feed Share Link Embed <blockquote class="wp-embedded-content" data-secret="yRrwqkIIWK"><a href="http://thectoadvisor.com/dave-mccrory/">The Data Gravity Index with Dave McCrory</a></blockquote><iframe sandbox="allow-scripts" security="restricted" src="http://thectoadvisor.com/dave-mccrory/embed/#?secret=yRrwqkIIWK" width="500" height="350" title="&#8220;The Data Gravity Index with Dave McCrory&#8221; &#8212; The CTO Advisor" data-secret="yRrwqkIIWK" frameborder="0" marginwidth="0" marginheight="0" scrolling="no" class="wp-embedded-content"></iframe><script> /*! This file is auto-generated */ !function(d,l){"use strict";l.querySelector&&d.addEventListener&&"undefined"!=typeof URL&&(d.wp=d.wp||{},d.wp.receiveEmbedMessage||(d.wp.receiveEmbedMessage=function(e){var t=e.data;if((t||t.secret||t.message||t.value)&&!/[^a-zA-Z0-9]/.test(t.secret)){for(var s,r,n,a=l.querySelectorAll('iframe[data-secret="'+t.secret+'"]'),o=l.querySelectorAll('blockquote[data-secret="'+t.secret+'"]'),c=new RegExp("^https?:$","i"),i=0;i<o.length;i++)o.style.display="none";for(i=0;i<a.length;i++)s=a,e.source===s.contentWindow&&(s.removeAttribute("style"),"height"===t.message?(1e3<(r=pa

Transcript 3,161 words · about 21 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, you're listening to another episode of the CTO advisor podcast. This is the second podcast we recorded today. So I wanted to kind of like, as if you were listening to the conversation we just had with magic, as if, you know, you, you were walking with me, but this is a really great episode that we have with Dave McCrory, global head of insight, research, insight, and analytics and digital analytics. Okay. I'm not research. I'm like, I wanted to add research because you, you do an awful lot of research.

I do do a lot of research. That's part of my professional side and part of my, you know, personal project side is doing lots of research. Absolutely. So, uh, I've known Dave for a few years now, right? Dave, we've, we've rolled in the same circles. We've ate brisket together. We so that, that, that means that we're officially, we can officially call ourselves friends because we've ate a brisket in Texas. I don't think, although I don't, Texas, I think the bar to eat brisket together is probably fairly low.

It is, but you know, we've met in a few other places as well, uh, outside of Texas and, uh, that's for you and I, it might not be that high of a bar, but for most people that would actually be a really high bar. Yes. Texas is a really big state. The, uh, we, uh, visited, uh, Texas as part of the CTO advisor road trip, uh, mutual friend, was kind enough to host an event. Dave came to it and we, uh, kind of said to ourselves, why haven't we done content together?

Then we, uh, ran into each other again. Did we run into each other? Re-invent again? Yes, that's right. Another re-invent follow-up conversation where Dave talked about his continued research. For those of you who don't know, if you don't know Dave, you probably know, probably one of his most famous pieces of work, which is this concept that data has gravity. You might've heard the term data has gravity. Dave actually coined the term, not just coined the term, but you have a whole framework and data and analytics to show that data has gravity.

Is that correct? Yes. Yes. I've spent, I've spent over a decade now, um, looking at data in several ways. And one of those ways, uh, was exploring data gravity and why does it exist? How, uh, how does it manifest itself? How does it work? And when I went to work for digital realty, we decided to build a data gravity index. And so for those unfamiliar with the concept of data gravity, data gravity is the idea that data has an attractive force.

And so that attractive force manifests itself in two ways. The first way is as data grows, let's say in the database, it creates more data and it attracts applications and services to it. So if you had a, we'll say you had a 10 petabyte database with data in it, um, the likelihood that many applications and services are going to want to consume that data is fairly high. Um, so that's one, one side of the data gravity coin. The other side of the data gravity coin is activity.

So much like, um, in the world of physics where you can have gravity from a mass or gravity from a high amount of energy. Uh, the same thing exists with data. You can have a large mass, like I explained before, or you can have a large amount of activity that ends up having an attractive effect. And so that might be, uh, you post, uh, you post some incredible meme or something else and a very large number of people come and interact to view it.

And maybe it's on Twitter and it gets retweeted and it gets sent again and all of that activity. And it can be compute activity or user interactions. Those are both activities, uh, generate data gravity. And they do that because every time you interact with data, you create more data. Um, so if you think about that, uh, anytime you use a service or an application, you're creating more data, you're creating more data through the interaction itself, through log files, security checks, uh, audit trails.

When you save things, all of that adds to the data. So it accumulates faster and faster. The more interactions that occur with the data. Um, so that's from a high level, the overall view of data gravity itself. So thanks for that. Uh, recaching of data gravity. Let's talk about like your current work. So we, we know that data has gravity, but over the past few years, since you've come out with a data gravity theory, the landscape for services has changed.

I think the reality of data hasn't changed. Data still has gravity, but I think it's the, you know, the, me and you have talked about kind of, uh, service inertia, the, what happens when gravity meets inertia and we start to depend on, uh, centralized services provided by cloud providers and they have both our data in our service, or we have services in a data center that required that accesses data that's not close to it. So as, as the landscape has changed around, uh, data and services, what new findings are you discovering and via what, uh, vehicle we we've talked about kind of the data index that you're doing with, uh, digital realty.

What's the update on your work? So in, uh, in 2020, uh, I released two versions of something called the data gravity index with digital realty. And, uh, the data gravity index is focused on understanding the largest metros that are, um, generating data gravity. So we, uh, we developed with a fairly large amount of data, a, uh, a formula and a methodology to measure data gravity in metro areas. 0. 5, which expanded to 53 metros globally, um, looking at data gravity.

And so those, some of them are what I would call the usual suspects, uh, the, uh, the New York's, uh, Hong Kong, um, London, uh, very large cities, large metropolitan areas, LA. Uh, and then when we expanded to 53, uh, metros, we saw some, I would say some more interesting cities like, uh, Zurich or Madrid or Mexico city. Um, uh, so more of an expansion, uh, in looking more detailed at, uh, at the measurement. And what we did to measure it was we looked at the amount of data being generated, uh, in a metro area.

We looked at the amount of processing taking place and we looked at the overall, uh, data storage occurring in these metro areas in aggregate. And we, um, effectively combine those. And we look at what the, uh, what the latency of that metro area is, uh, what the average network latency is. And we come up with an index score. And then that index score determines, uh, the level of data gravity for a particular metro. com and, uh, easy link. There's no form to fill out.

You can just click download the PDF and you can look at the report. Uh, so that's very easy to do. 0. Uh, I've not spoken a lot publicly about it, uh, but it is, uh, well underway and, uh, it, uh, goes far more in depth, providing even more information across a greater number of metros, uh, and focuses even more on, uh, the different elements of data gravity that kind of make up the index. And so I'm very excited about that coming out, uh, later this year.

And, uh, and then in addition to that, uh, there will be, uh, an interactive, uh, tool of some type that we will also release that, uh, that will allow people to look at data gravity in, uh, in more of an analysis, uh, on their own kind of thing, um, than just purely looking at a report. So there are so many different ways I want to go with that, but we'll, uh, cause one, I, I just, I'm thinking about the scale of the actual work and how to collect that data and the data science around the actual work.

And then the results of the work, like what does that mean for enterprise? So we'll take the second route. And if we have time, we'll circle back to kind of the, the, the, the science behind the science. Sure. Uh, so now that I have an idea of the gravity or the impact of data gravity in a specific region, how does that help me help me inform to make decisions? You know, I live in Chicago. Uh, I would imagine I haven't looked at the index for Chicago, but I would imagine that, uh, because we have pretty good bandwidth here.

We have pretty, uh, centralized services. We, I can build applications that are, uh, latency sensitive. Uh, how does that interaction of the, the gravity index and how I decide on services and application architectures, how does those two kind of correlate? Sure. Sure. So what's, what's really happening is you're, you would be using the data gravity index to provide guidance on planning and such. And you would do that by, uh, by looking at your customers, uh, or what your intent of whatever this application is that you're building.

Those customers could be internal to your enterprise. Maybe you have, uh, maybe you have internal customers that are located in, uh, five or six different, uh, cities or metro areas. Maybe it's a hundred or you could have end users that consume your application that are more just, uh, general customers. And so you would obviously have information about those things. What you don't have is an understanding necessarily of where those, uh, where those customers are, what is their latency like in their city?

What is the amount of, uh, consumption or generation of data that is occurring in their city? Or perhaps, uh, there's not a lot in their city, but there is a large metro that's nearby that they can effectively connect to and leverage. And so this gives you guidance on that for today. And then it also in the report projects out what's going to happen over the next five years. So you can look at planning and see, hmm, is this metro going to be growing at a faster pace or is it going to be growing more slowly?

And should you make an investment or should you potentially locate physical hardware, um, for whatever your purposes are? Uh, maybe you, uh, maybe you deploy specialized equipment or something else. So does it make sense to expand in one place or another? And, uh, we found quite a few customers who have, uh, done workshops and things, uh, with us where they realized that they needed to be investing in, uh, one or more markets that they did not even consider looking at until they looked at the index and then went back and said, well, if this is true, then we should be seeing a lot of customers or traffic from this area.

And then they went and looked and they found out that they were having a lot of customers and traffic for that area. And that then drove them to do an entirely new deployment in, you know, Indonesia. And they did not expect to be deploying in Indonesia, but because of this, it kind of sparked them to do the necessary, uh, work on their end to figure out that that was actually going to be required. Yeah. I'm thinking through kind of the alternative, uh, approach to getting these insights and data.

And most of it is reactive. So the idea that I can put in probes, uh, client size probes, or I can use services that, uh, tell me about the user experience after I've built the application. You know, the, the, I can do synthetic transactions against my SAS offering. Uh, once my SAS offering is up and I have clients in the locations that I want to build or basically after I built the application, that may be too late. Like the, I've built the thing and now I have to be reactive and fix these, these foundational data gravity related issues.

Uh, this is something that I experienced a lot in the mid 2000s, uh, where I was running the IT shop and we had a very distributed business line and they would adopt the SAS offerings that would not perform well. And they had no way of knowing whether or not, you know, the thing that they signed up for was going to, uh, perform well in North Carolina. Most of this was US based and North Carolina versus Texas versus where it was bought, which was New York city, which has plenty of bandwidth.

So I can absolutely see how the index has utility for someone designing a product. 0, how quickly, there's a couple of questions. How quickly does the various indexes change based on time? So, uh, an easy way to look at it is, uh, if we look at it as year over year growth, or if we look at a compound annual growth, uh, for example, you'll see, uh, in the report anywhere from the one twenties to the one 50, uh, 150% plus in overall growth.

So it's incredible growth. And there's so many factors driving this from digital transformation to, uh, requirements. Uh, and this is globally government requirements around audit or data sovereignty. Um, there are, you know, the, the pandemic and the increased use of work from home and people interacting, uh, people doing digital entertainment more. So more and more, all of that is online, whether it's, you know, movies and things or gaming. And all of this takes up more data. We have IOT and smart sensors, all of this again, generating more data, that data is generated, it's processed, it's analyzed and consumed.

And all of these things all mean more data, more frequently, and that drives more data gravity. And it also means somewhere someone's buying more data storage. Someone's buying more data processing. And so that's, what's driving this continued growth in cloud, in hybrid. Uh, and so the growth rate continues to just be incredibly high, uh, year over year. Dave, I have so many questions that won't fit into a 20 minute presentation. So I'm going to start with you. Dave, I have so many questions that won't fit into a 20 minute podcast.

Like I have questions around methodology. I have a, uh, questions around, uh, the size of the data set. I have questions about how do you, uh, collect all this data? How do you munch it? How do you find, you know, how do you, uh, peer review it? How do you validate it? Like the, you can have a theory, you can run, uh, algorithm and, uh, uh, against the data and then come out for answer. How do you backtrack and get challenged on whether or not that is correct?

There's just too many questions to, uh, uh, answer in a, in a 20 minute podcast. So we're going to have you back on to, to kind of tackle some of these practical problems, because I think there's a lot of, of insight for my audience on how do they approach big data problems and how to, what are some of the problems you ran into and how did you solve that? You're both, uh, you know, you're, you're, you're, you have on two hats. You're, you're the person who, uh, correlates the data and you're also kind of the person who arranges your infrastructure to analyze the data.

So you have a unique hat on. Yes. Uh, it's, uh, I'm kind of the army of one when it comes to, uh, the data gravity index and the work behind it. So I do have to solve, as you said, the data science side and the challenges of architecting, creating, analyzing, uh, and compiling. Luckily, I don't have to actually build the report itself. Uh, there's a team that, that does that. Uh, that would be, uh, you know, far too much for me and it would take years for me to, uh, to get out.

Uh, and then they'll report exactly, exactly. Right. So, uh, from that side, yes. But, you know, at the same time, um, it means I understand the data a lot more. It means that, uh, I have a much deeper level of ownership. Um, and, uh, it means I get to learn from solving the problems, whether it's, uh, you know, dealing with some very large files that, uh, you know, you can't load in an Excel and you can't easily work with, uh, in a database.

And so how do you deal with that? And, uh, so that introduces some interesting coding challenges and other things. So, uh, I get to do a lot of things that I enjoy, uh, with this, which I think is, uh, one of the best things about my job is I'm constantly having fun and solving new and interesting problems. All right, Dave. For those that do the whole social media thing, do you have a social media presence? I do. I do.

Um, on Twitter, it's my last name, which is M C C R O R Y. And so it's just at McCrory on Twitter. Very easy to find. All right. And for those of you who want to follow us, we're the CTO advisor on the web. com for the data gravity index. That that's pretty easy, uh, to follow us on Twitter. It's at CTO advisor. If you have a question for that, you're not brave enough to ask Dave in public, which he is pretty interactive on Twitter.

You can DM me. I will X on your behalf. Uh, DMS are open. Make sure to rate the podcast and iTunes. It helps with discovery. The more people that listen to the podcast, the more we're able to get guests like Dave on the line. Talk to you next CTO advisor podcast.