Philosophy of Data and Data People
Transcript
hello this is Karen Lopez you may know me from my snarky tweets on Twitter as data check or is my witty blog post at data model calm or perhaps you know me from appearances at other CTO advisor events I'm a Microsoft MVP in Data Platform avi expert I also enjoy being a NASA data not and by the way I'm a data chick because I was born that way you can follow me on Twitter a friend me on Facebook Linked with me or LinkedIn but no matter
what you do make sure that you love your data one note about this presentation since this is about some of my thoughts on data it's going to be a little bit eclectic but not weird I promise you there'll be no weird so let's talk about thinking about data why this topic data is perceived as being an old school technology I mean it's not about sexy software stuff and by old I just mean it's very experienced like me people get caught up on the fact that technology
is sexy and data might be less so but I told you I was born this way as a data lover but I think data is why we do the most of IT data is everything and everything is data what do I mean by this as the advent of software-defined services as virtualization becomes the norm of how we do things it seems like data is involved with most of it and I'm a data chick so I'm a bit biased but I'm here to tell you that all
data is suffering let's talk about why I think that if I'm such a fan of data well one of the noble truths about data being suffering is that we suffer when we want reality to be different than it is that's what causes suffering and so many times we expect our data to be there to be of good quality to have integrity to be available when we want it to and yet it isn't always that way the reason it's not that way is that we haven't put
in place the things to make sure our data has integrity and quality and protection and availability all data is suffering also because the world is much more complex than any one person sees all data is suffering also because we don't always do the right things or take the right path in order to make sure that it delivers to us what we need from it in this highly complex data model which by the way only covers in-store retail processes you can see the complexity there and how
everything is linked to every other thing or at least it seems that way my job as a data architect is to help decide what goes in what box how things should be related how they shouldn't be related what constraints we should have on that data how do we know whether the data is right what standards are we using to measure against it all of those things are part of my day-to-day job I'm here to tell you that I don't believe things are software-defined I think it's
more like their data to find think of those things that you think of was software-defined well there's a configuration there's desired States for them there's relationships between all the components and there's permissions and a bunch of other complexity that we need to define but as software defining that or do we have a whole slew of XML and JSON files and Jupiter notebooks and all of those things holding the configurations the desired States the relationships the security and the logic that ties them all together I'd like
to see you all using the term data defined but I know you really like the software-defined thing because most people think software is the reason we have information technology another thought about data is that one person's data is another person's metadata so the typical definition of metadata is data about data but we've expanded that in the world to mean properties or configurations or characteristics about something and yes that's exposed as data and it might be also some software meaning some scripts or some logic in there
in the data architecture world we talk about a lookup table we say it's just cup table who cares how it's designed well that one little lookup table let's say a list of countries and states is another person's full-blown data so you're just a lookup table or just a configuration or just a parameters file that's somebody's data another key aspect in thinking about data is data last longer than code as a very experienced data architect there are systems that I know the data existed before I did
and it'll exist long before I'm gone from this planet the other thing especially with the advent of big data is more data doesn't always mean better data sometimes more data is just more data how do you spot a date of professional in the wild well there are the traditional data rules data architect database administrator data scientist data miner data modeler data steward data governance professional all of those you can kind of tell their data professions because they've got that word in there but some of the
other data professionals there are the c-level people the CIOs the chief data officer z' the chief chief knowledge officer z' the data center dai the data moonlighter but what about these folks our DevOps analysts and DevOps engineers data people are developers data people how about the network people or the data center people or the storage people of course they are because they're all managing data they may not realize their database administrators as they produce all those JSON files and check them into github but their data
people and github is their database this is what we need to know just because the data stored in a non-traditional format doesn't mean it's not data let's also look at how these people fall on a matrix that goes from supports data to use as data you'd be shocked maybe to know that a data modeler supports data by creating data models but they don't really interact with the data that much themselves and DBAs especially like data to a DBA is really a database or maybe a database
instance or server the data center storage and infrastructure people their data the way they talk about it is sometimes hardware so they talk about data protection maybe that just means backups to them but to me data protection also means security and data quality and data integrity and data availability so you might think about where your title or where your role fits among supporting data or data technologies and using actual data the other thing that makes a big difference in data engineers and professionals is this kind
of conflicting viewpoints actual and analytical data users these are very different worlds in fact a traditional relational transactional data modeler often finds a hard time making a transition to being a data warehouse data modeler I'm one of those people these people have different points of view one is about preserving the integrity of the data that's why transactional systems are highly normalized so that we don't have update anomalies that's why analytical systems are highly denormalized because they're optimized for performance because the data integrity was in theory
already done at the transaction or the capture time these two groups of people have different needs they have different rewards and different skill sets much like the difference between someone who builds rockets and someone who flies on them there's a lot of hyper buzzwords in the data world now all these are legitimate data approaches or methods or tools but they're being abused all over the place artificial intelligence machine learning data science data virtualization and the all-important algorithm I think it's our job as IT professionals to
make sure that when someone says our product uses artificial intelligence our next question should be hey tell me about that artificial intelligence tell me how you do it what algorithms are you using what tools are you using same for mission learning it's a common understanding that some people think a bunch of if statements categorize machine learning and data science besides the fact that data science is the sexiest job on the planet and pays really well a lot of people think if you work with data you're
a data scientist some other ones big data that one's already sort of falling off the hype cycle but I see big data as a reason to not treat data well and if you're building a transactional system and it's not big data high volume high velocity then perhaps you shouldn't be using big data technologies for them no sequel has been through all this I think the fight between relational and non-relational data solutions has sort of fallen out - we'll just use the best tool for our data
story that makes me happy scheme alice is also a funny one that comes from the no sequel world because most of the no sequel solutions actually have a schema they might implement it differently but they have one they have a structure to the data there's an understanding of the data the difference is that structure might be applied after the data is read not when it's being written and then we have the whole code driven data things there have been along my 30-plus years of experience lots
of technologies where people are saying that we no longer need to design data we'll just let the code generate it I'm here to tell you that that means that the developer writing the code is also being the data architect and they are being measured rewarded and promoted based on how fast they get code written not how well they protect the data or preserve its integrity this is a natural conflict of interest that comes about through a reward mechanisms that come down from management so a code
driven data structure is one that says this data only exists for this piece of code that may not be the cost benefit and risk trade-off that you want for all of your data some data words that are newish in the overall scheme of data management data controllers and processes which arose out of the privacy legislation these are people that either steward the data in the term data controller or they're a third party that processes the data data governance is a huge issue right now and it's
going through a lot of growing pains is this some large multi departmental policing of data properties or is it a way of identifying the processes that make data better data stewards are people who curate data standards and curate data policies they often work with people like me data architects data sovereignty is the location of the data and what legislation or compliance or policies are required for data that either resides in a particular location or passes through it data models been around forever what's different about data
models now is they're not just relational data models entity relationship diagrams there's graph data models there's document data models a data model is still important and I'm here to tell you an ER D isn't always the most appropriate one metadata we already talked about data lakes swamps oceans all of these terms for bringing together a whole lot of data not worrying about overly restricting or curating it and providing it to make it available to all kinds of other people within the organization data estate is a
very buzzy word that has been used to talk about anything from hardware to software to data centers to databases to data models and data gravity data gravity says that the larger data becomes the more mass it has the more likely you are to bring compute to the data than you are to carry data to the compute some data truths let's start with the positive data people can be a bit snarky I'm one of those people and one of the reasons why I'm snarky is I've been
doing this data stuff for more than thirty years I'm very experienced at it I'm also very old and that makes me kind of occur about all of this but I also find that it means that I've been through a lot of innovation around data technologies and I definitely know the benefits of innovating in the data space but I also want to make sure that we don't forget about the fundamental truths of data I don't see data as a byproduct of software yet so many people do
we fund projects where data is a byproduct of an application project or a software out cuisine project data that's treated this way often is biased toward the application and when that application or solution is replaced the expense of migrating to a new solution is significant I see data as the reason we even have software now this is definitely a conflicting point of view with a lot of software professionals I'm just acknowledging this difference I also think that all data has value even data just for now
so data just for now that could be logs that could be configurations in a JSON file that are just going to be used for migrating this one project once I still want to apply the data integrity rules to make sure that we get the results we want I believe that data quality is a measure not a grade high data quality to me means that we've measured the data we have against a previous standard that we've published low data quality would be data that doesn't meet the
standard we expected but if the standard we posted was it doesn't have to be consistent on read we don't really care whether data is missing we don't really care if there are invalid values then the data has the quality we want and that's high quality on the other hand if the data is for my bank account and I have standards both Accounting Standards Bank standards and my own standards for its integrity and quality and it doesn't meet that I'm going to call that low quality so
we need to measure data against a standard report the quality of the data by measuring it to that standard and the implication for this is if you don't have a standard your data quality can be anything you say it is I mentioned this before data last longer than code sure there's plenty of systems where the data comes in it's processed it's analyzed and it's deleted but I've worked with data especially things like address and location data that's more than a hundred years old sometimes more than
200 years old that data has been collected has been curated for all of those years across dozens and dozens of professionals who were responsible for its integrity and quality and I can use it because it was treated well even COBOL isn't as old as a lot of the data that's still being used data deserves our protection for future users because of this so if we overly optimize our data to meet one application developers priorities we're definitely going to sacrifice the benefit we can get for it
for future users of that data so let's look at not just the truths well these aren't lies I'm going to talk about lies but the first thing I want you to know is that data is weird because people are weird I know this is a shock to you and systems are weird I know also a shock people lie and our data ends up impacted by this so let's look at this number in 1980s seven million children disappeared off the face of the earth just boom on
one day so let you ponder that for a second but I'll give you a hint this happened on April 15th in 1987 when the IRS started requiring Social Security numbers for children age five or over on April 15th of that year when the taxes were filed seven million children dependents really disappeared off the face of the earth and that was because prior to 1987 all you had to do was write down names of children let's just say that people weren't lying that they had extra children
it probably just felt that way to the parents but we know in 87 they all disappeared and it never even made the news that cost Americans 2.8 billion dollars in what appears to be people fudging data in order to benefit from it in this example I'll let you look at this number and let you guess what you think it is if you're in the US you'll probably guess right away this is a number I use when us-based companies ask me for a piece of information that
doesn't apply to me namely a zip code why don't I give something more popular like 902 100 from the TV series well that's because this is it code is from hell Michigan because the way I figure it if I'm gonna have to give bad data in order to buy something from you you might as well know what I think of it here's a Canadian example in this particular postal code h0h 0h0 this is a postal code for someone really important in Canada but people tend to
give it when they don't want to give their truthful postal code and that's because this postal code belongs to Santa Claus who lives in Canada I'm here to tell you Santa Claus lives in Canada what these examples show us is that if you want your data to be simple you got to go out and make the world simple then come back to me I get lots of comments from time-to-time from developers and DBAs that I've got too many tables in a database design that's because they
have some sort of fear of having to do a join in their code it's not really a fear it's really an irrational phobia but this is because they believe that all normalization makes code perform slower but it doesn't because normalization also decreases the size of your data and we all know decreasing the size of your data means you can return more of it faster the underlying truth that I want you to know about these problems with data is that they happen because people are weird and
systems are weird and people lie let's talk about some weird data so this news came about in 2017 that a teenager have found a flaw in NASA's rocket science data that particular flaw was that he found that when a certain sensor was being read when nothing hit it it returned a negative number but that energy level can't be negative thus he contacted US Space Agency and it turned out that no one else had noticed anyone in the data world knows that the reason a non reading
from a sensor returned negative one is because they were using that instead of a null they were using negative one as what I call a phono a phono puts the is put in place of an unknown value the problem with foe nulls is everyone has to be in on this weird little trick to fix the database this particular 17 year old wasn't in on that trick neither was all the media that just loved the story of a teenager telling NASA experts how to deal with their
data here's another example of weird data the envelope on the left is one sent to me by the US government specifically the State Department because it was a passport related correspondence you can't really tell from the photo but there's a sticker above my name that just says care of Canada which I found very endearing the key behind it is though that the third party mailing facility that the US government uses perfectly the State Department doesn't support countries in mailing addresses and therefore a person or a
machine had to go to the extra expense and time and error-prone Ness of putting a sticker on top of the little window that's kind of weird given that this particular governmental department is also responsible for overseeing all of our foreign embassies another example on the right same issue they were unable to add a country to their mailing address software so they wrote it in pen on the waxy window and it still got to me data is weird US citizens sometimes live in other countries customers sometimes
move to other countries just because people are weird here's an example of my name now formerly my name has an accent over the oh but this causes all kinds of problems with all kinds of computing systems and sometimes with people as you can see from my friend who Postal mailed me a card didn't know where the accent went so he just put it over all the letters this is the most infamous example of a blogger who also shares the same last name with me who ended
up with a package from ups that went through so many transformations he wrote a poem about it because it came up with Lopez spelled with all those extra escapes going on specifically there in Our Mutual name let's talk about why this issue is more than just a funny poem written about a really unfortunate series of events I think people's names are inherently the most personal pieces of information that we record about someone that means that we have designed computing systems with constraints that might have been
true decades ago we continue to carry along the wrong sorts of data types the wrong sorts of code that ignores the fact that all kinds of people have extra spaces in their names they have characters that they think are just regular characters that change how their names are pronounced that change the meaning of their name we force these people to lie about what their names are in our systems I never use the accented oh when I buy something online when I order something when I register
for something I didn't use it when I get my ID because we've all learned those lessons that eventually some system isn't going to be able to find our record or our license isn't legit it's more than just you know supporting extended characters this is about supporting people's names as they really are so we think about how people are weird having a special character or an accent or a different spelling or even new characters that go outside the North American data set those are important things to
people and we're not supporting them so I'd like you to consider the next time you're thinking about data maybe we should spell people's names correctly because data is weird if your data and your designs and your specifications and your standards don't account for data being weird or people lying or systems being weird then you're going to end up with chaotic outcomes here's one of my favorite ones for years the integration between several different travels sites and Airlines upended an a - my first name my name
is Karen but all of a sudden on Air Canada I became Kareena which I called my spy name and I would even tell TSA they didn't care that my name wasn't spelled the same they just cared that I had a passport with a name and that they knew that boarding passes often showed incorrect name information and you can see here from examples with my name being appended with salutations or not having a middle initial or just having an initial for our first but you can call
me Karina the next time we get together systems can be good the data can be weird and people can lie and systems can be good but we all know having to deal with systems that sometimes systems are just as good as the last person to work on them but they can lie for instance someone ship something to me they knew me from Twitter so they just use my Twitter name and it still showed up this is one of my favorites my Starbucks name is kitty and
that's also a nickname I have but you can see as I give that information to bar Reese's they have many different interpretations of this word kitty kitty spelling it both ways with two DS and two T's spelling it Kenny which is just ironic because that's my twin brother's name in this example I'm going to talk about how another way systems can lie so several years ago I went down to my Bureau of Motor Vehicles to get a new drivers license and I was most concerned with
my driver's license photo because that's how people are so I stood in line got my photo taken hoped for the best and got my driver's license I checked out the photo it was great so I was happy with it when I got home and showed my mother my new driver's license the first thing she noticed was my date of birth was wrong on it and that was a big problem so I went back to the Department of Motor Vehicles to get my license reprinted with the
correct date of birth on it and I had my birth certificate I had all my other ID but unfortunately their computer systems can't change the date of birth on anybody's driver's license because if you think about it it's really odd for the date of your birth to change I mean in real life what would that mean your date of birth changed you were born again well not quite that way but in my case the data was wrong someone had typed it in wrong so instead of
the 27th it was the 7 and that was wrong now you know some people might say what is this matter of course it matters this is your ID I said we have to be able to change this so they called their help desk and the help desk says well no one's date of birth can change and with that said to me was their system had been designed to focus on sort of we'll call it a real life truth but your data in a system isn't real
life it's an abstraction of who you are your name your address your height and your date of birth they weren't going to let me have that changed fortunately I knew some of the people that worked in this department and they knew I had a twin brother that's us there that's me on the left and him on the right and they knew that we were twins and therefore probably had either the same date of birth or very similar dates of birth I was also fortunate that he
was a law enforcement officer in this location and I got him to come down and talk to people and say there has to be a way he's very good at doing that type of stuff and eventually someone had to override the system and go in and manually update the data which if you know how data integrity works that's a very awful way of updating data but all's well that ends well I got my driver's license with the correct date well it was almost good enough because
they made me take my picture again and of course that one was horrible what I learned from this is that systems sometimes are really good at doing bad things we see this all the time when people implement constraints because they think it will make the data better only to find out it makes the data worse the typical example is on e-commerce systems on the web where they do sell to people outside the US they need to rename the label from zip code to postal code but
they put a constraint on it that it could only have numbers in it and in many locations such as mine our postal codes have letters and numbers so maybe the system was lying about my date of birth when someone typed it in wrong but the fact that they had no provision for updating that data except for a manual error-prone safety concern process was also a problem to data quality I know you might have some questions now that I've gone through these and in fact when I
give this in front of a live audience I often get these questions well artificial intelligence machine learning blockchain quantum computing or whatever solve all these issues you see lots of promise of these things sure ai and ml can help people do their jobs more efficiently by taking on sort of the low-hanging fruit or the 80% of things and activities that computes good at doing but the answer is no people being weird systems being bad systems lying people lying I don't think that modern technologies or even
traditional technologies are going to fix these problems these are part of being in the real world and we need to make sure that we have processes in place to deal with these anomalies the constant pushback I get from DBAs developers DevOps people and storage guys is what if I don't have time to worry about all those data things well a lot of us don't have time to worry about a whole lot of things and that's especially true as we're being asked to do so much the
tools and the technologies and the innovation that's happening around us is really hard to keep up with but as my grandmother would say obviously you have time to fix all the problems afterwards thinking about the data we're going to be supporting designing or using does require some time but the payoff for those is really huge the people in the agile world and some of the DevOps worlds the scrum worlds are going to say no no we can't do thinking first it's important we just start doing
sprints a lot of infrastructure is data as I talked about previously and a lot of DevOps is data I think it's important that we understand the more data we're producing as part of managing infrastructure defining infrastructure deploying infrastructure the more time we're going to have to spend to make sure that the data that we're using to do these things is properly managed that it has the right data quality that we version it just like we version other things that we control it as much as we
control other things another question I get is what about cloud technologies well cloud technologies fix any of these issues and the simple answer that is no yes software as-a-service can help us work on data quality can help us secure data help us monitor it but in the long run we're still going to struggle with these data issues that have to do with people lying systems lying and people being weird ah the big pushback isn't this just a problem with relational databases I hear this a lot
from the no sequel crowd well some of the issues with relational databases are the reason we have non-relational data storage solutions and they're solving a specific problem but again the people being weird that data being weird the real world being complex adding more complexity to our solution only makes sense when those solutions are a better fit for our data stories usually people get around to this one so I need a data professional I'm biased and I think you do but that data professional doesn't have to
be a data architect like me perhaps what you need is all IT professionals learning to love their data to think about their data to think about their data as a separate component of all the other things they're managing to treat it just as software means you have data that's biased towards that application we need to remember that people are weird and that people lie and that systems lie why do we need to do this well it's because I want you to be on team data team
data isn't necessarily just people who work with data all the time but they're also people who want to make sure that the data they use is treated with respect if we know that systems lie then we need to engineer out the lies if we know that people lie we need a way of validating the data they gave to us so it could be that the system is messed with the data for instance the integration errors I had while flying it could be that your data constraints
didn't match the real world such as when I have to give the zip code for hell Michigan it could be that I'm just lying about it because I don't want to give you my real email address there's all kinds of reasons why our data quality is poor if you understand that the world is complex you will allocate enough time to understand those complexities if you understand that people lie you'll allocate the time and resources to validate data if you know that systems lie you'll allocate the
budget to make sure that they stop lying the data you support and the data that you use deserves your love that includes security we all know for instance for data protection we need to make sure we have backups but as I'm known for saying we don't actually need backups what we need are restores if you're not retesting your restore systems both your files and your infrastructure and your networking it doesn't really matter if your backups were going really fast we also tend to deal with performance
as an issue where people propose that data quality and performance are conflicting points of view but the opposite of data quality isn't slow performance and the opposite of high performance isn't low data quality we need to work together on teams to ensure that we're not just moving bad data around faster and also make sure that the cost of having good data isn't preventing people from mix from executing their business processes our data also needs integrity integrity which is different than quality says that when we retrieve
a piece of data it's the value it's supposed to be so this works together with security constraints validations data profiling a lot of data activities to ensure that we can trust that data and then there's data quality which as I said is a measure against standards for those things I want to thank you so much for spending this time with me while I went through my rants and examples of people lying systems lying and the truths about data thank you [Music]