Data Repatriation with Joe Onisick
Keith and Joe Onisick, ( @JoeOnisick ) Principal at Transformation Continuum talk about Data Repatriation. The CTO Advisor Data Repatriation with Joe Onisick Play Episode Pause Episode 1x 00:00 / Subscribe Share Apple Podcasts Spotify RSS Feed Share Link Embed <blockquote class="wp-embedded-content" data-secret="kypbWePB3m"><a href="http://thectoadvisor.com/joe-onisick/">Data Repatriation with Joe Onisick</a></blockquote><iframe sandbox="allow-scripts" security="restricted" src="http://thectoadvisor.com/joe-onisick/embed/#?secret=kypbWePB3m" width="500" height="350" title="“Data Repatriation with Joe Onisick” — The CTO Advisor" data-secret="kypbWePB3m" frameborder="0" marginwidth="0" marginheight="0" scrolling="no" class="wp-embedded-content"></iframe><script> /*! This file is auto-generated */ !function(d,l){"use strict";l.querySelector&&d.addEventListener&&"undefined"!=typeof URL&&(d.wp=d.wp||{},d.wp.receiveEmbedMessage||(d.wp.receiveEmbedMessage=function(e){var t=e.data;if((t||t.secret||t.message||t.value)&&!/[^a-zA-Z0-9]/.test(t.secret)){for(var s,r,n,a=l.querySelectorAll('iframe[data-secret="'+t.secret+'"]'),o=l.querySelectorAll('blockquote[data-secret="'+t.secret+'"]'),c=new RegExp("^https?:$","i"),i=0;i<o.length;i++)o.style.display="none";for(i=0;i<a.length;i++)s=a,e.source===s.contentWindow&&(s.removeAttribute("style"),"height"===t.message?(1e3<(r=parseInt(t.value,10))?r=1e3:~~r<200&&(r=200),s.height=r):"link"===t.message&&(r=new URL(s.getAttribute("
Transcript
All right, you're listening to another episode of the CTO Advisor podcast. For the first time, you know, it's odd, Joe, me and you interact on Twitter. I think if I was to do one of those like closest circle things, you'll probably be like absolutely in my first circle. We interact a lot, but never had you on the podcast. First time caller, and I'm very happy to be here. It's great to be on, and you know, I'm a big fan of everything you do and especially the work there at CTO Advisor.
So Joe, I want to say, what's your co-founder, what is your title at the Transformational Continuum? So I am a principal with Transformation Continuum. I am a co-founder of the company, but I don't go by that. Principal's nice because it lets me be the principal of whatever my co-founder and CEO needs me to be the principal of in any given moment. We're going to talk about data repatriation and continue the conversation we had when I visited you on the ranch way back in 2021, it seems like forever ago.
And we just had a quick chat about data repatriation. So let's set the framework. How would you define data repatriation first? So I guess there's a little bit of context around it. I'd want to start with, and the context being that idea that we talked about on that last episode was data gravity, right? So if I'm putting my data in Amazon or Azure or Google, I'm going to end up putting my applications and workloads there to be close to the data for the lower cost, lower latency, higher bandwidth I can get.
So that kind of traps me from a gravitational perspective to that cloud provider. So the idea of data repatriation is, A, if I'm moving to cloud, deciding the right location for my data so that I have more flexibility with my workloads, applications, and services, or B, repatriating the data I have in the cloud already into a more neutral location, demilitarized zone, if you will. But those can be an on-premises facility that I own that has good cloud access, cloud on ramps. Or in most cases, it's going to be a tier one data center provider that has those extremely low latency connections to the cloud with better cloud contracts for those egress charges.
So I'll reference the recent podcast we did with Dave McCrory, who invented this term or coined the term data gravity and doing additional work around it. This data gravity index and looking through and measuring data gravity or the effects of data gravity based on whether or not you're in a colo or if you're in a pop that's far away from your data and how you should consider the impacts of data gravity on your workloads. But as we think through like every major vendor, whether it's VMware with their cross-cloud strategy, Google with Anthos and their multi-cloud strategy, or even now AWS talking more about the edge, they're all addressing, I think, the central problem of data gravity, which means that I have to accommodate my data via placing workloads closest to that data.
If I'm in an oil field and I'm taking in a petabyte of data a day or a terabyte of data a day, I can't afford the time it takes to transfer that data from that remote location into where my compute is that I have to position the compute. But how do I take my cloud control plane and move it closer to the data that cloud control plane is heavy? How should companies begin to think about, I think, ownership of data placement? That's a great question.
And I think the first thing I'd caveat with is I'm a big fan of simplification and not over-engineering a solution. So if your organization has made a conscious decision to be all AWS or all Azure or all Oracle Cloud, then this conversation isn't for you because you've already made that decision, you understand the lock-in and the pros and cons that come with that, and you've assessed those. And that's a fair decision. But if you are going to use multiple clouds or on-premises infrastructure with cloud, or you think that you may want to do that in the future and have some freedom from your cloud provider in case they change prices or decrease the service level or whatever effect would make you leave, then you're really looking at this idea of data ownership.
And the data is the big thing that's going to dictate where your apps services live, how well they run, how well they perform. So you need to take a serious look at who owns your data and where are the costs with that data. The cloud providers are all going to give you very cheap storage in comparison to most other methods of storing your data. And they do that because where they're going to make their money is on the compute, the higher level services, and then as we all know, those egress charges.
When you want that data to leave their cloud, you're going to pay for it. So those are the considerations that really come in, I think, as the major ones when you're looking at data ownership. As you were speaking, I think there's a second part of data ownership that I wanted to take a look at. 0. 0 kind of overview, this idea of building data infrastructure that spans not just physical location of data, but data formats, et cetera. As we look at SaaS services, PaaS services, infrastructure as a service, various different formats for storing flat data and object data, et cetera, the challenges of data don't change across that data infrastructure.
I need to still understand access control, data retention policy, et cetera. Up to this point, and I know the technology isn't there, but up to this point, customers have kind of waved their hands in the air and said, you know what? My data that lives in Salesforce, while it's data that I do ETLs to bring into an Oracle database, et cetera, it's too hard to have a data policy across the platform. I have opinions about that approach. How should customers, from your perspective, as they go through data transformations and operational transformations, how should they think about ownership of their data infrastructure and their approach to making sure their policies are applied across kind of data platforms?
That's an interesting question and a big one, right? I think we all want these silver bullet solutions, which is one of the reasons Gartner makes so much money is like, just tell me what to do because it's all too much and too confusing. But when it comes to data, it's obviously going to come down to what kind of data do you use? How does that data get stored? How does that data get accessed and what all needs access to it? So I think you really want to think through all of these.
I'm a big fan of standardization, right? Anywhere I can use two tools instead of three, I'm going to use two. If I can use one instead of two, I'm going to use one. If that creates a level of lock-in, sure, but it also creates a lower cost, lower complexity solution all the time. So when I don't need to get out of that solution, I'm getting the best bang for my buck. So I think it really needs to come down to carefully looking at this, and if you're an organization who's going to use your data across multiple different platforms, so take your average enterprise who isn't getting rid of their legacy anytime soon, who isn't going to put everything into one cloud provider.
Now you really want to be looking at where can I get the best cost and performance for my data that doesn't end up locking the data access to a single provider, because I don't know that that's the provider I'm going to want to use. Yeah, and it's really interesting as you say that, I can't tell you how many times I've talked to not a tech person, but a business user, actually a non-for-profit. Really tiny non-for-profit, say, you know what, I have this, I have all of my contacts in this fundraising, this fundraising past database, and I want to send out ladders for events that we held with a subset of our sponsors, of our contributors.
Why do I have to pay somebody to by hand take the data out of one system, put it in another? Isn't there kind of just a master database or something that I can just sync to and do it? He didn't understand the backend technology, he just knew from a business process, he put all of his contributions in one system, all the information and data was in there. The forms that he needed to do, the platform he used to do the form letters and the data mining was in a different system, and he had to recreate the data sets, and now he's like, well, then I have to pay two different people, or I have to pay for two different processes to maintain and keep that data in sync.
This should all be easier. So people are natively coming to these conclusions, and I think my opinion about this is that if you don't have a data policy in general, like whether it's technically where we can enforce that with technology, if you don't have the policy to begin with, then you run into these types of problems that can't be solved because there isn't a technical solution. There's a technical solution to his problem, but as it scales, if you don't think about kind of how you want to manage your data, there's no technical solution to a problem that you really haven't well defined.
I think you hit on a brilliant point there, Keith. I think it's all about having a plan. I think in the moment, once you hit a problem, it's really hard to look at how you're going to get through that problem. But if you already have a plan for it at X scale, I need to do X with my data. For every platform I use, that data needs to be externally accessible in a programmatic fashion. Whether that's an API or some other tool, I want access to it.
These types of things going into a plan help make all of those decisions down the road, especially when those decisions have to be made fast or in an emergency. A really good example of this. I was reading this morning about this week's Okta breach. I was reading Cloudflare's version of it from how they were affected. One of the interesting things they said is that they import all of their logs from Okta, their access logs from Okta, and they store them on their own in their own SIEM, in their security system.
The reason they do that is because it allows them to access logs further back than Okta's going to store it. It's their data. They want it, and they want to do something that their SaaS provider doesn't, so they needed that access, and that's something Okta does provide. I think that's a critical thing. If you don't know going in to a decision to use something like a SaaS identity provider, that you're going to want your own access to the raw data, then you might pick a provider that does everything you need except that and end up with a problem down the road.
Oh, wow. I really love that example because, again, another conversation I get into consistently is my collaboration platform keeps eight versions of a document for three months, whatever the retention policy is. Therefore, I don't need backup. Again, I think people are looking at the technical ability versus what their requirement is. You need to sit down with a legal team or a business continuality team and understand what is it you want to do and what are the basic building blocks to do that so that you can supplement something as simple as log examination, a SIEM, and supplement it with simply exporting the data.
Even if you're not putting it into a SIEM, if you've exported the data and you have historically five years of log data, I can at some point take that, throw it into an analytics platform like Splunk and use it if it comes into play. I don't need to necessarily create the overhead of having a separate analysis platform than what my SAS provider gives me, but if I don't anticipate the policy and what I would like to do with the data, then I've just taken away the option.
I think the point is not necessarily to recreate what your provider has given you, but to make sure that you have the data in a format in which you can utilize within reason, given something like, who would ever thought that Okta, of all things, would get compromised? Yeah. Yeah. And there's a flip side there, right? Don't store the data, and this goes into your data policy, but don't store the data you don't need. So for instance, on your website, if you don't need to be logging people's IP address and system information and stuff that your website can do automatically, then don't.
If that doesn't have any particular future purpose that you can see, why collect that data? A, you're creating a possible breach hazard of data that you don't need that could be released through your company, which at best is a PR hit, but B, within certain regulated industries, any data collected might now hit retention and compliance policies. That means you now have to store it. And storage is cheap, so we think, okay, why is that a problem? It's not. You can store as much data as you want.
The problem is going to be when you get hit with a breach and it gets taken, or you get hit with a lawsuit and now you have to analyze it to show what was there, or something to that effect. So the idea of all of this is twofold, right? It's collect and have access to all the data you need where you need it, but it's also figure out what data you don't need, because we create data every time we look at data. We create a new set of metadata that becomes its own data about how we searched and researched that data.
So wherever you're not needing to store that ephemeral stuff, don't. Yeah, so the conversation started out with a repatriation conversation, but I think we've expanded the idea of not just repatriation, moving your physical data, but the idea of the repatriation and ownership of your plan around your data. We have a lot of data. It's growing exponentially. I'd love your metadata example, because one of the things that I've talked to data protection companies about is what happens when we start using our backups as opportunities to create QA and test environments.
And we actually increased the amount of metadata around that. And now we have to secure that metadata and we create yet another problem. We might have solved the problem, how do we activate a system quickly in case of a DR or test case, but we've now created this metadata problem that now that we know where the ability to be able to do this means that we need to have all this data about our data and how do we store and secure that? So it is a never ending battle to understand and own your data.
Joe, if people want to know about not just data repatriation, but the mindset in which you approach constant transformation, how can they get ahold of you? com. And that continuum word is a little weird. com. You can find me on Twitter, but I always kind of recommend against that because I'm a bit of a prickly bear on there. And I can definitely co-sign with that. One last question. How is the apple pie recipe coming? The apple pie recipe is coming along.
I think I've been perfecting that a bit, but I've also, I branched out into cheesecake. So next time you're here, if you're a cheesecake fan, I'll have to make a couple of those too. We might have to make a special stop to the border of Wyoming and Montana to get some cheesecake, some homemade cheese. That sounds delicious. For those who want to know more about the CTO Advisor, you can follow us on the web. com is the website. I'm a little less prickly on Twitter than Joe, but I do retweet him a lot.
So you might catch, you know, a little, the prickle, you know, uh, secondhand. It's at CTO Advisor on Twitter, DMs are open. You want to learn more or ask questions about this or anything else. I'm more than happy to have the conversation, talk to you next.