Ransomware Protection Your Last Line of Defense
Transcript
hi my name is Kim Delgado and I'm a senior staff solution architect at VMware I'm also half of a Duo known as team candy thank you for joining me today as I share with you a little bit about ransomware recovery and how it is your last line of defense in the event of a ransomware attacks have been making headlines targeting companies and government agencies of all sizes causing larger negative impacts than other types of disasters to these organizations this costs these organizations millions of dollars in
Ransom sensitive data breaches damage to reputation and other sorts of related costs ransomware sets itself apart with its ability to live in systems for weeks or months before the attack happens or before you come become aware that an attack is underway that increases the risk of restoring your data from unreliable or compromised backups as a result this kind of attack requires a different disaster preparation and recovery process than your typical storm or power outage knowing how to be prepared for a ransomware infection and testing your
recovery process frequently can literally save your organization millions of dollars so let's have a look at what is required Beyond prevention and detection to be the I.T hero that your organization needs well today's the day your cell phone starts ringing in the middle of the night all of your company's critical applications are down business has come to a screeching halt now what today could be a career altering moment one thing is for sure now is not the time to say this isn't your job doesn't matter
what your role in it is you're going to probably play a part in helping your organization recover from this event so how prepared are you a lot of cios that we talk to really feel like unlike traditional Disaster Recovery events ransomware is really a when not an if scenario most companies plan for that if and maybe de-prioritize disaster recovery and business continuity as a result but ransomware attacks are becoming increasingly sophisticated there are literally entire business models that focus on developing ransomware as a service software
and then they partner with operators who execute attacks and exploit vulnerabilities then these organizations share in the profits your defenses are often like layers of Swiss cheese where different layers May cover some holes better than others but every so often those holes align and those Bad actors then are able to find a way in even if you have those multiple layers of protection and detection Solutions Solutions implemented Bad actors will penetrate your network through user error or stolen credentials and then leverage this remote access malware
once they gain access they focus on critical systems like your backups your active directory DNS servers storage all of those sorts of things and they also go focus on stealing admin credentials or they find ways if they can get into your active directory to create new accounts and then add themselves to admin groups in active directory if those ad if those groups are then given admin privileges in your uh in your sddc software Solutions they can then gain access to things like your backup admin console
and perhaps do do something like change your retention policies or disable your backups altogether they also managed to find their way into your network management tools then they can adjust your firewall rules and other protections that you have in place but you know you say well it's not going to happen to our organization we are running a tight ship according to Gartner by 2025 that's less than two years from now at least three quarters of it organizations are likely to face one or more ransomware attacks
when you talk to folks in it who manage the infrastructure and you ask them what their plan is for ransomware most of the time they say well that's the job of the security team not my problem when it comes to detection and protection that's probably true but what about when it comes to recovery then that comes back to the infrastructure team and the operations teams having to work together with the security team to recover as much data as possible as I mentioned earlier these organizations have
become incredibly sophisticated they understand the tools that are used in your data center they know how vsphere works and how to leverage access to network management tools and active directory at some point an attacker will likely find a way in again whether it's through human error or misconfiguration and they'll buy their time while they find ways to make their attack as impactful as possible this is what we generally call dwell time increasingly their goal isn't just to lock you out or encrypt your data it's also
to steal your data so that they can release that sensitive information if you choose not to write pay the ransom the problem is that a lot of people are choosing less and less to pay those ransoms and even if you do pay the ransom oftentimes they can't completely give you all of your data back or unencrypt all of your data it depends on what software they're using to do things for example like uh encrypting your VMS if they encrypt running VMS without questing them and shutting
them down properly chances are those VMS aren't going to be recovered even if the the disks themselves become unencrypted sometimes they're nice and actually shut down the VMS before they encrypt but not always nice I use that term very loosely the detection and protection tools that the security team have implemented are probably not going to be enough to protect your data and your workloads in all cases so you you need to have a tested recovery plan and Recovery Solution in place detection can be incredibly difficult
file based attacks with known hashes are becoming less common it's really easy to change a known hash by simply changing one bit in a malware file but more often than not once access is gained the malware will run in memory or as a library file and remove its trail of files that may be detected in file and backup scans the Bad actors are even finding ways to clear up logs or remove other traces of their activity these Bad actors may take weeks or months to spread
around the environment and hunt for vulnerable access points to Key Systems as a result different systems are infected at different points in time before the Bad actors have enough penetration to trigger the encryption or the lockout this makes it nearly impossible to pinpoint a single point in time to use for your recovery efforts for all of the different impacted workloads maintaining quickly recoverable backups for a minimum of 90 days is considered best practice but that may mean that you need to actually recover your base VM
from prior to the actual malware infection and then separately recover your data from a more recent backup ideally the most recent one prior to the Locking or encryption with the right types of tools this can be considerably faster than rebuilding the application environment from scratch but being able to rapidly iterate between the different snapshots or backups in order to verify the infection will be imperative so just keep that in mind identifying recovery points can be incredibly challenging as we just discussed especially when different applications or
workloads can be infected at different points in time one of the things that can help with this is having visual guides to see change rates and compression rates for each of your different snapshots or VMS that can at least help guide you to an initial point to starting point once you've selected a candidate based on that information that candidate needs to be verified again there's a dependency on security tools but having but fileless attacks can make make malware difficult to detect as we talked about so
this verification needs to be done in such a way that you don't end up reinfecting other VMS that are running in your recovery environment which means you need to have control over that East-West Network traffic easily in your recovery environment too and you rarely get it right the first time which means that you're going to be iterating between these first two steps until you find a stable candidate mixing and matching your full VM or Os level restore with file and folder level restores while keeping that
VM isolated can be nearly impossible this is why recovery often takes days or weeks to get critical systems back online the iterative process needs to be done as quickly as possible and preferably without the need to move bits around between different networks and different systems or between different environments if again having that isolated environment and having those backups easily available to that isolated environment is really critical so how prepared are you to recover do you have three to six months of backups held in an air-gapped
immutable storage location do you have that infrastructure to support an isolated recovery environment instantaneously trust me now is not the time to be cobbling together outdated hardware and building an environment from scratch mistakes will happen and reinfection will happen if you do that does your solution provide the ability to launch quarantined VMS for inspection security analysis and quick recovery does it offer Behavior Analysis and next-gen Antivirus scanning on live workloads in that isolated environment through integrated tools and can it continue to monitor workloads while you
restore data for more recent backups to prevent reintroduction of malware can it provide can it also provide you easy tools to promote uninfected or recovered VMS back to production or to another whole new clean environment if your solution is made up of many different services or Point Solutions Point products that are sort of stitched together you might want to consider starting to look for other options so what is missing in most common Solutions let's take a look backup and disaster recovery software often lack some critical
features that make them ideal for ransomware recovery often they lack effective air gapping although they may claim to have this feature effective air gapping means no direct connection to your storage is made from your production environment you can't connect for an hour a day to copy your backup files over to storage and be truly air gapped infections still find a way that that when that happens most Disaster Recovery Solutions require that you roll your own recovery environment whether that's in a different data center or in
the cloud it's typically not baked into the solution which can be a limiting factor and if you can't boot those VMS in the recovery environment directly from your air gap storage that means that you need to move bits around or restore from backups in order to start the process of evaluating any given backup from a point in time again remembering that this is going to be an iterative process and you're likely going to have to do this multiple times for each one of your workloads it
can really add a lot of time to that process if you're needing to move files around Cloud Disaster Recovery Services can help address the shortcomings of traditional Dr Solutions but they cannot but it can also lack critical requirements such as that built-in integration with next-gen antivirus and behavior analysis capabilities and some of these Solutions require you to change the format of your virtual machines in order to run them in your recovery environment that makes restoring those VMS back to production even more difficult my suggestion is
don't take a vendor's word on it for any of these key features you need to ask to do a POC and see it and test it and touch it for yourself to really evaluate if the solution is going to work for you so what can you expect when the inevitable happens well you can expect that's going to be brutally challenging to recover most organizations are full stop out of business for days or even weeks and you can accept the fact that you're likely going to lose
some amount of data but mostly you can expect chaos at least initially as you come to terms with the situation and assess the damage also keep in mind that attackers are still going to have access to your systems until you're able to figure out how they got in and block that and they will continue to attempt command and control access to prevent recovery so let's break this down a bit into the big steps now I've broken these down by days but these really are just groups
of priorities that you need to address first step is going to be finding a way as I said to cut off that command and control access that's very critical you also need to balance that with the access needs of your IT staff especially if they're remote are they able to come into the data center is there a possibility is it possible to cut off outside access without completely cutting off your IT staff that's something you need to consider business continuity planning is critical in circumstances like
this you need to understand what the businesses top priorities are and execute accordingly Central Communications are going to be a top Target for the Bad actors they live to create more chaos so be prepared to have alternate comms plans in case email or other centralized tools are compromised how will you coordinate and communicate plans without email or teams or confluence your P0 systems need to start with I.T services like domain controllers active directory backup systems and evaluating accounts to see if any Bad actors have gained
access you know have gained admin access to Key Systems once you've gotten things locked down a bit and have your have some priorities set further decisions need to be made on how to proceed with your data center are you going to be able to recover back to production saving your sddc layer and fix and restore VMS are the underlying components in your sddc potentially compromised did the Bad actors gain access to your network management tools to vcenter to other critical components those are things you need
to think about log files and copies of infected VMS and backups are going to need to be preserved for forensics and law enforcement log files can be invaluable in helping track progression of the attack if admin accounts were created or accessed all of those sorts of things your isolated recovery environment hopefully doesn't need to be built but it may need to if you weren't prepared but if any other prep work needs to be done and for that isolated environment before you can begin your recovery efforts
you need to get that done as quickly as possible and then begin with again those P0 components like your active directory your domain controllers and corporate comms platforms to get those re-established first then prioritize your business critical applications from this point it's going to be more of the same the security team and the infrastructure teams will be working in lockstep to continue to iterate on both the investigation and the recovery efforts but be prepared for a second round of attacks if something was missed from an
infrastructure and workload recovery perspective the very first step really like we talked about earlier is um the first phase is going to be getting to that most recent good-known state of a virtual machine the infrastructure teams don't usually have to deal with this sort of thing so having the tools in place to help them understand what a good state is is going to be critical to prevent re-infection knowing where to start can honestly be the hard part trying to understand when that infection started try to
make that first calculated guess of where to begin can sometimes be guided by looking at change rates or enter and entropy or compressibility of those snapshots or backups there's no real automated way to determine this you can't make this determination on a backup or a non-running version of a virtual machine so regardless of what solution you're using it's going to be an iterative process and it's going to be a cyclical process of picking and restoring and retrying different from different points in time so you need
those immutable backups to restore VMS into that pre-built recovery environment quickly so that you can test those running VMS and of course you need to be able to quickly iterate to be able to find that most ideal candidate and you need to do that before you can start restoring your additional files and data from that most recent available backup and you need to do all of that while that VM is running in an isolated state but in a state where you can be continuously monitoring it
with your next gen antivirus and behavior analysis tools so how can you be the it hero that's the question here right so first of course be prepared have your plan and your tools in place second be aware don't ignore odd behaviors or strange files investigate all suspicious behavior and don't assume that you're safe and third test your plan and your tools frequently consider this continuous training for your IT staff so that they will feel prepared and confident that they know how to handle the situation when
it when the when the time actually comes it will still be chaotic and stressful but you should be able to recover more quickly and thoroughly than most organizations if you follow the guidance that I've provided today thank you for your time today please feel free to reach out and connect with me on any of these uh with any of these options and I hope you have a great rest of the day and rest of this conference thank you as promised wasn't that a great session if
this is your first session of the day and you're wondering what's next well you have options you can hang out in the chat room now and you can see there's a conversation going on in the chat room in the session that you're in this will be open until the end of this hour and the top of the next hour the next session will start if you want to go into the breakout room there's a general chat session you leave here go into the lobby and then
enter generous chat you can chat with your regular attendees I'll be floating between the two chat environments love your feedback if you have any questions please you can DM me within the platform or on Twitter at CTO advisor enjoy the rest of the conference