Intro Deep Reinforcement Learning

12:15 · Watch on YouTube ↗

Transcript 1,701 words · about 11 min to read

Auto-generated captions from YouTube, not hand-corrected, so names and technical terms may be imperfect. The video is authoritative.

Hey welcome again to the CTO advisors coverage of sa P sapphire 2019 we have a treat for you we're going to go into a topic that we don't normally go deeply into email machine learning with GP GP you're a data scientist at Accenture right yes really appreciate you coming onto the show you have upcoming book on email AI give us the overall impression this show kind of what you've what your research is band and how it relates to the conference sure deep reinforcement learning is the

most advanced machine learning technique that this particular type of algorithm has completely off the charts it's it's on the top of the technology right now and like the brain Institute in the past ticket has invested more than a billion dollars to understand how the neural networks operate in the brain and how the consciousness powers the coordination of all the neurons in the human brain because there are around 87 billion neurons in the human brain and when a decision emerges out of the brain is it's so

synchronous and it's so simultaneous and it it just goes through all the neural paths and goes through the synopsis and since the signals out and in connects that and it provides admission the reinforcement learning is definitely based on the same architecture to write basically what it does for a robot to basically for example to come out of a maze there could be several paths and there could be several states and the robot has to perform a specific action to come out of the maze and the

reinforcement learning does the coaster multiple episodic train of the algorithms and it produces the action that is required in a most optimal path to come out of amis they could be several environments like that like open a gym once we understand how all of these neurons operate together and we can actually compute that into a neural network then we can find out the paths for these neurons and how does that decision emerge the neural paths and that's what this leads to the HDI and other than

the opening I team I was talking about I liver as lot of open airy open AI gym environments for the elimination of this reinforcement learning algorithms also like deep mind has come up with their own lab of environments for a reinforcement learning and there are a few other companies who are also implementing these reinforcement learning for the artificial general intelligence so let's talk about like the current state of the technology so there's the neural networks and leveraging neural networks in this reinforcement learning technique how can

companies apply that to a IML today in some of their decision processes so most of the artificial intelligence machine learning techniques that the corporation supply are based on the historical data sets right so what they do is they take a lot of historical data and then the machine learning algorithms simply learn from the data and then they produce the outcomes in terms of like predictive analytics or advance financial forecasting those type of applications corporations are building but so today so that yeah so today you know

until gave us that a couple of weeks ago or a few weeks ago at their data centric event and said that only 2% of the world's data has been analyzed that has been generated in the past five years so that's tremendous amount of data but the way that we're doing email and today we take that data see patterns in that data right and then we can make decisions or predictions based on that data today that's a pretty common use of AI however what your sounds like

what you're talking about for reinforced reinforcement a is there's a little bit more free world exactly right yeah so because it does not require any data at all right it's just like human decision because like a human makes a decision a dynamic decision like they said a car has to go onto a mountain and then there it's challenging because of the uphill task and either the person has to turtle or go in the right direction or live direction to balance or the GAR and reach to

the top of the mountain and that doesn't have any you've not seen the mountain before you don't you can have as a human I cannot intuitively know exactly how to navigate the mount exactly so teaching neural networks how to do that same thing right is the goal what's what's the key indicators that that we're starting to know that we're making ground in these areas so we have seen some significant improvements and achievements like on that chest mm-hmm where like the machine has totally has beaten the

humans in terms of right the chest and we also have seen few other that was alphago and a few other instances where we've seen some of the gaming environments also have shown significant improvements in terms of operating it through artificial intelligence applying the reinforcement learning like mojo environment is a environment in the gaming environment and so there are a few other environments where we have seen that kind of a pattern where the reinforcement learning completely operates on its own does not have any historical data said

it basically goes through episodic iterations for each iteration it gets some rewards for doing it right right and then there will be a discount factor and the reinforcement learning how much has been the reward versus if there is any loss there are also value approximation functions that can calculate the reward and and then finally when we get the output it's basically an action to be performed and that's what is needed to accomplish the AGI because most of the time we see like autonomous vehicles operating on

the road and then again there are pedestrians crossing the road and it does not know whether to save the passenger or to save the pedestrian right that is an action that it has to make instantly right on the spot it's not like something it has to have some data for it it's a dynamic decision and humans struggle with some of these decisions ourselves so the ideal of the the concept of rewarding during the learning phase the concept of rewarding rewarding the algorithm for making certain decisions

you know you have to imagine a lot of the difficulties and to your point is it is the AI making the decision based on the reward or because some other factors kind of getting those knobs right yes so that's what the reinforcement learning does it doesn't have any data it just has to perform a dynamic action in a number of states and there could be so many states of it like for example when it's navigating through a mountain or it's navigating through a road there are

several states for it and again some of the objectives have been fulfilled through like sensors like lead R for example right laser light range detection technique or radar sound technique where it senses the movement of people around the car and then basically it finds that these objects are there and then it kind of navigates through that this was actually implemented on Mars by NASA between for the Mars rover vehicles and and like Mars rover and so the way they calculate it is because of finding the

motion of an object and is detected by the vehicle and then it can apply some slam algorithm and then it can navigate through that so these type of decisions are like for example for any like a shuttle to land on a Mars it doesn't have a prior its Mars it's right right can I predict like how it's going to land there and because of latency we can't put a human remotely to navigate so there has to be some type of AI to to navigate that navigate

you make decisions exactly and that's what I'm trying to do in this book basically taking multiple environments it could be a gaming environment it could be like a mountain car it could be a Krabat different types of gaming environments and how it or even if it could be ahead simple traffic problem so as I was explaining about different algorithms for our reinforcement learning like trust policy region optimization algorithm and policy gradients so there are several Riis a piece that I'm developing for the book and the

book will be called pie tart 1.0 deeply enforcement learning okay cookbook and it will cover it will be there will be like hundred recipes in it and each chapter will explain a snippet of code how to implement these algorithms and how to implement each methodology for a most optimal path and optimal strategy to resolve and solve the environment and when you expect the book to be released well we were actually support it was supposed to be released in May that that's how books dole yeah right

I could only complete like three chapters so far we are expected to complete like ten chapters okay and so I mean I haven't completed all the rest of the seven chapters yet that's okay are you do you participate on social media Twitter or what what's your Twitter handle my twitter handle is a GP underscore pul IP aka and I share lots and lots of content on reinforcement learning and different machine learning techniques how these have been deployed by different companies and the production landscapes and how

it benefits and what are the most optimized ways of doing it so GP I really appreciate you coming on to the CTO dose if you've made it this far you are a ml hey I geek and we really appreciate you guys turning into the CTO vice we have other Enterprise related content you can find that on the web at the CTL visor calm you can follow me on twitter at CTO advisor talk to you next CTO DOS