CTO Chat - Cancer research
Transcript
hiya Steve Townsend from the CT advisor calm with an old periscope coffee chat morning to the folks out there so we're going to talk a little bit about Big Data and specifically challenges that it brings to the infrastructure I had the pleasure of attending V Intel data center groups influence or day this past week and one of the things that they did it took us to ohsu which is a provider or a research facility uh researching cancer amazing the team does amazing work i highly suggest
that you guys go check out their website and some of the work that they're doing but a lot of the work they're doing in about the technology that they're being enabled with brought up some pretty salient challenges that's face when you're working with big data there are doing human genomics and therefore our human genome genomics creates a great great deal of data massive amounts of data so i think the numbers that i saw a couple of numbers that i saw oh they had a really cool
i think i know i'm going to get this name wrong it was a cry of thermal electronic Magnus cool that pretty much took the size of maybe two or three 42 you rats tall set on a energy absorbing floor and was amazingly complex piece of machinery that creates about 27 terabytes of data a day uh that data is creates imagery data which then can be analyzed by researchers then on the other end of peace when you're talking about creating a scene actual genomic data this data
if you were to sequence all the genomes or all the you're the sequence all the genomic data of all the cancer patients annually it will be something around four exabytes of data for exabytes of data that's an awful lot of data beyond the fact of how do you so sloshed store all of that data in a single location the question becomes how do you collaborate on that data when you're a large facility like ohsu who couldn't even collect all the data by themselves uh you still
need to make the data available to other researchers a lot of a lot of business challenges kind of come up with that as well so it's data and data is valuable so whether you're a non-profit for-profit research providers such as ohsu or a for-profit company like the one that I work for there's a challenge that you want to protect your IP you want to share the data with your community but you also want to share the parts of the data that you only want to control
that so how do you share that data this is where I think cloud solves a challenge that traditional infrastructure cannot solve one you can't as a single entity unless you're like the US government or something you can't afford to store that much data for exabytes of data is just incomprehensible I think to give it some perspective the numbers that they compared it to is that five exabytes of data is every word spoken ever mentioned in human language so in daily communication that's the enormity of the
amount of data that would be generated even if they could solve the problem how do you get all of this data into a system secondly how do you share the data I mean from a practical perspective if I need if I'm a researcher and I need to run if I have the CPU capacity to run queries against that data the data has to be somewhat local to the machines running the data so I could in theory put all this stuff in AWS and have AWS work
clothes from different regions running on it but AWS doesn't solve the business challenge for the technical challenge of every solution this is why the industry needs what a primary data the company of that linked into the title of this coffee talk you need this concept of data virtualization where you can virtualize where the data exists versus how the data is connected to the end system you know in systems consume this data in various ways one way is to present the data as NFS share so you
know we're getting into the technology a little bit you know when we talk about storage virtualization where I can connect the data I do wonder how much I exit by the data I don't know how much extra by the data would cost and physical storage let alone on s3 but that's a good question going back to the physical challenges of presenting the data you know different systems present the data in different ways some of it is by NFS s in NFS and I could create a
NFS share and store the data that or share the data that way even though for exabytes would be a crazy amount of NFS share and i don't know if that would work block storage doesn't work at that level how do you present and go through that data so i don't have solutions to any of this i think it's a great conversation to talk about of how do you virtualize that data and then from a physical infrastructure perspective how do you physically share the data across virtually
the globe right now we're just talking about us-based companies and research facilities but how do you make that data physically possible and then how do you control essex control meaning the metadata how do I share the valley you will metadata that I want to share how do I share the actual data and how do i compartmentalize the components of the metadata in the physical data that I want to share across universities and how do i get a crawl how do I get past the problem of
duplicate data changes to the data data sovereignty just a lot of a lot of questions that came up as part of the discussions with intel and hsu this is actually really in the medical world this is some old new problems i think the new problem is just before the past couple of years two to three years the price of genomics and genetic sequencing has come down by multitudes from the millions of dollars into the thousands of dollars or thousand dollars range per patient so the technology
is more available if the technology is more available that means that it's more likely be used so just within the past couple of years we've seen an explosion of data then coming from the side of the imagery when we're taking the physical image of something slicing it down they showed me another microsoft scope which is a light i can't remember the name but it's a light scanning slicing microscope that will slice a piece of genetic material down to a down to as low as for nano
meters of thickness to give some perspective the Intel the latest Intel manufacturing processes 40 nanometers so you can take a chance a intel transition transistor and a 40 nanometers transistor and carve it into for three to four sections of layers and do that across the entire let's say one millimeter by one middle of genetic piece of material and that creates an incredible amount of data how do you analyze that data and how do you share that data are relatively new challenges in the community so I
was really excited to see I'm digging a little bit more to the to the science behind cancer research and specifically the challenges when translating to the data that's created through all the vehicles that they have now to perform this recent research expect a lot of riding for me over the next couple of months on around the work that they're doing at OHS ohsu this important work in and of itself but I think the big data and IT community in general have a lot to learn about
the challenges that these researchers are in counting and how to encountering and how two weeks all those challenges ourselves that's it for this particular episode of coffee periscope chats talk to you guys maybe in the next couple hours as I as I think it's some new material talk to you guys later