Recuperar la identidad tras un ataque de ransomware es una tarea complicada. Si la restauración se lleva a cabo con demasiada lentitud, la empresa se paraliza; si se hace con demasiada rapidez, se corre el riesgo de que el atacante vuelva. En esta sesión se resumen las lecciones aprendidas de casos reales de respuesta a incidentes cibernéticos para crear un enfoque reproducible que permita una recuperación rápida y limpia de la identidad.
Aprenderás a:
- Diseñar patrones de copia de seguridad y restauración de la identidad que permitan una recuperación que tenga en cuenta el malware
- Evita los errores habituales en la recuperación que prolongan las interrupciones del servicio o permiten que los atacantes vuelvan a introducirse en el sistema.
- Verifica los bosques y dominios recuperados antes de volver a conectar los sistemas críticos
Welcome, everyone. I’m Allison Parrott, and I’m thrilled to welcome you to today’s event entitled Rapid Identity Recovery, lessons from the Frontline Incident Response sponsored by Semperis. Before we begin, I want to cover a few housekeeping details. If you have any questions throughout the presentation, please make sure to type those into the q and a box, and we’ll make sure to get those answered for you. Semperis is, has provided some resources which correspond with today’s event, so please make take a moment to check those out. They’re located to the right of your audience console. Today’s webcast is being recorded, so keep an eye out for a link in your email to rewatch the presentation or share with a colleague. And now I’m thrilled to announce our speaker for today. We have a pleasure of hearing from a member of the Semperis team. With us today, we have Tim Beazley, lead incident response at Semperis. So we are in for a great event. And with that, I’ll pass the time over to Tim to get us started. Thanks, Allison. Welcome, everybody. My name is Tim Beasley, like she was saying. I’ve been around in IT since when I was really sixteen, building computers and stuff, doing lots of professional things. Started Identity around NT four days, been with Active Directory ever since, along with specializing some PKI. My real true niche, though, comes with compromised recovery and incident response. I’ve been doing that for quite a while now. Did that at Microsoft. Was on the compromised recovery practice team as well as DART, for those of you in the industry who knows that acronym, whether it be fortunate or unfortunate. I’ve been able to participate or lead over three hundred different engagements between my time at Microsoft and now it’s in Paris. And now I run all things IR here at Semperis. So, wanted to share with you guys a lot of information and real world experience, basically talking about, the hardest, probably the most highest pressure jobs in the cybersecurity industry. And one of those things is making sure identity gets back on track after some sort of cyber event, whether it be ransomware, malware, whatever. And this the entire talk is drawn from, all that experience that I mentioned before, real world IR engagements, not theory. Those of you who are watching this who actually, own the work of, you know, managing identity, infrastructure, SecOps, incident response plans, CSOs, all you guys. So, the idea here is to give you guys a repeatable approach for recovering hybrid, Active Directory, Intra, Okta, Ping, all the identity platforms that are out there to state that you can trust at the end of an event. Not only do you wanna do that, trustworthy, and over time, right, but as soon as fast as you can. So fast enough really to save the business, clean enough to hopefully keep the attacker out, or at least prevent them from coming back and taking over again. And so we’re gonna move through three parts, of this session, how to design backups, do restorations, different patterns, how execute while avoiding different pitfalls, and then making sure everything’s validated before you reconnect into production. So as a preface, I just wanna mention this. Right? The goal of this talk is is not a sales pitch, even though I work for a company that provides products and solutions that covers everything in this. At the end, you know, at at the end of the day, it’s your your products that you guys pick. So, hopefully, you guys will be able to understand the concepts that I’m gonna be talking about and being able to implement that across the board. So as long as there’s something in place that’s designed well, to get you where you need to believe to achieve that cyber resiliency that we like to talk about a lot, that’s the goal. So to kick things off, we are gonna talk about what is actually hybrid. What does it mean to become hybrid? So first thing is quick. Couple questions. Right? Does your organization rely heavily on Active Directory? Most of you are gonna say yes. Second question, does your organization use or leverage o three sixty five? Probably. Right? And do you guys use single sign on for both Active Directory and o three sixty five? Meaning, the identity is the same. Do you log in with the same username and password in both on prem and up in the cloud? The answer is yes. You’re hybrid. So a lot of times, these solutions that we’re gonna be discussing focuses heavily around hybrid attacks. Most of the cybersecurity attacks and incidents that I’ve worked, the vast majority have all been hybrid in some sort of fashion, especially in recent years. So pay attention to that. Okay? And it doesn’t doesn’t matter what industry you’re in. I’ve been a part of engagements, whether it be health care, financial, government, anything public sector, private sector, all industries, trucking, farm equipment, critical infrastructure, you name it, banking systems, casinos, big names, health care stuff that’s been in the news that you guys have read about, you know, is probably behind the scenes in some of those too. So point being is hybrid attacks can occur anywhere and everywhere, and it’s good to be mindful of them. So what’s it kinda look like? Right? There’s this game board that we like to talk about. And with the the purpose of this is, like, when every hour that passes where identity is down, whatever it be, Active Directory, Entra, you name it, the entire business is down. Right? People can’t log in. Basically, everything comes through a screeching halt. Leadership gets amounts tremendous amounts of pressure. Maybe you’ve been forced to, like, try to be pushed to pay pay some sort of ransom, that kind of thing. And the the human nature, the instinct of being under that amount of pressure is to rush these things and go as fast as you can to get stuff back online. Problem is if you restore from an infected backup or if you reconnect before you convicted or even evicted a threat actor, you’re essentially giving the environment right back. You wanna avoid that. Right? Recovering identity really, after ransomware, honestly, any really attack can be really complicated, and it’s not just some easy, let’s restore from backup. And so the rest of my talk is gonna be resolving a lot of the tensions and, with this repeatable method So you don’t have to choose between going fast, versus safety. Right? So based on this game board, right, we’ve got this outage. You’ve been attacked. Right? Things are going down. So as we move through, we gotta find a team. We gotta gather team. We gotta get comms set up. Gotta remember what the plan was. Hopefully, you have a response plan. We’re gonna send those communications to different stakeholders, and that’s ongoing. Right? We’re gonna evaluate and hopefully pick a backup to recover from. Here at Semperis, we choose to use the most recent one. And then we establish this clean room kind of thing, like this level of isolation. We’re gonna We’re gonna restore that isolation. We’re gonna make sure everything is good and tested and sanitized. We run some forensic tools within that isolation. We do plan for the cutover. We do the restoration. We cut back over to prod. Hopefully, all that goes really well. We don’t have any real real craziness that happens in production. If we do, we have to go back. It’s a lot of fun going forward and then jumping back a few steps if something goes wrong. Right? Flipping DNS, establishes communications again, resyncing with Azure. There’s a whole slew of stuff. Right? And then once it’s all said and done, the for that Active Directory is back up and running in production, now it’s time for the real recovery. Right? And so once we get to that success place, hopefully, you’re in a trustworthy state. Identity is not I mean, it’s the core foundation. I look at it, as the backbone of any organization that’s out there. It’s not just one of many systems. Right? It’s the systems that most attackers are really should really targeting. Roughly, what is it, ninety percent of organizations worldwide run AD as their core identity service. Hello. And it’s implicated about well over ninety percent of of of nine out ten cyber attacks. Right? Yeah. Concept is and it’s well known throughout the industry. Attackers don’t break in. They log in. They’ll either phish, steal credentials, maybe they’ll buy them off the dark web. They’ll log in. They’ll get in. They’ll move laterally. They’ll bounce around from machine to machine, harvesting credentials as they go. Usually, in the traditional attack path, they’ll find some escalated or privileged credentials and bounce up, and then it’s essentially game over. And they don’t actually have to take over the domain controllers at that point. They can take over, say, the hypervisor that runs the domain controllers. Right? I don’t have to be a domain admin to be considered domain admin equivalent. Right? So if I own the backbone, if I own the identity plane, it means I own everything downstream. This visual shows that almost every attack path eventually converges on the identity player. K? So the most I’d say the most important thing you can look and target to recover from is identity. Right? And, unfortunately, it’s the one most organizations are really least prepared for. Nobody really recovers a single clean platform. They recover some hybrid web of all kinds of different things like AD, Ping, Okta, Entra. Right? A lot of attacks that I see, will start, from the cloud side and pivot to on prem or vice versa. That interdependency is exactly why you can’t recover simply one of those identity platforms in isolation and be like, okay. We’re done. That IR environment, it’s it’s extremely chaotic and, very unpredictable. Every one of those hundreds of engagements I’ve been in, it’s been unique in some sort of fashion, which is in one hand, it’s fun with my ADHD. On the other hand, not so fun because sometimes you have to think on your toes. But, yeah, at, at this point, most you know, a lot of our customers agree with this simple statement. You know, if identity is not secure, nothing is. So every user workload, now every AI agent relies on identity providers for different authentication, authorization, permissions, etcetera. So it’s very easy for a threat actor if one of those, identities get compromised. Eventually, it could lead to keys in the kingdom. He’s in control over different applications, data, infrastructure, all those different AI agents that people are throwing into their environments, all those ones that are designed to help, those are gone. Right? So when you’re dealing with hybrid cross platform attacks, it makes resilience a lot more complex. So lots of things to think about here. And in this particular slide, I like this one because AD usually is the source of authority. Right? You gotta start thinking about this from an identity perspective, especially when it comes to synchronizing IDs from on prem to the cloud. Right? This transitions identity to, like, SaaS and business applications that almost everybody uses. Right? So an active directory is compromised. That chain breaks from top to bottom. Syncs can fail, or they synchronize a bunch of bad stuff. Cloud SSO goes away, access to email, Teams, file shares, SharePoint, whatever, insert application here. As part of the outage. So that’s why it’s critical that the clock on identity recovery is really the clock on getting the company back. It’s that first thing. Right? And so, Forrester conducted an independent total impact of Semperis customers, and one exec really summed it up well, and it’s AD attacks are an essential threat to our entire organization. AD supports everything in our company. You can basically take that quote and apply it essentially almost everywhere. Right? They describe Active Directory attacks as business ending events, which you can’t recover fast from. It’s not just downtime. It’s the risk of losing core business operations either for days or even weeks, and I’ve seen it. I’ve seen incidents where threat actors have been in the environment for years. They, when when they finally were drawn up to the surface, so to speak, all insert word here broke loose. It’s it’s a mess. Right? And that’s where I wanna get into. Is not only making sure you guys understand the the complexity of things that go into this and recovering from these hybrid attacks, but the impact. So, like, when we talk about, you know, we have these different blind spots, especially with the the whole thing with AI and all this that’s coming. People are struggling and moving very, very quickly to adopt AI. They think that AI is going to help. They believe that. In a lot of cases, it does in in helping secure their infrastructure, etcetera. Maybe they’ll boost production, operations, administration, blah blah blah blah blah. Right? But they forget when it comes to protecting the identity that controls the AI or what the AI is actually using. It could leave some serious gaps. Like, when it comes to your ransomware, for example, different breaches. Right? A lot of these are successful attacks because identity is mismanaged from the get go. So what do we do to move forward? Well, we gotta prepare. K? Ninety percent of organizations out there roughly hit some major blockers during crisis response. K? I’ve lived through most of, if not all of, the blockers that I’m gonna mention here, but also a lot more. Right? And in either ransomware negotiations or ransomware events where, you know, customers are under that tremendous pressure that I mentioned before about paying that ransom. Teams get stuck. They start to freak out. The outage drags along and long and long. The paying paying that ransom suddenly starts to look like the easy way out. That same chaos repeats over and over again whenever identity is compromised, and that there’s all this stuff, you know, Zoom. You’ve got all these collaboration things. You’ve got different forms and files and you’ve got Seam. You’ve got all kinds of Word, office applications. It’s it becomes an ad hoc scramble. So with a tested playbook, hopefully, you guys have some, you can get a more controlled, somewhat faster restoration. Everything, right, you don’t wanna be improv improving a lot of the stuff that happens during these incidents. You wanna be able to have some sort of something that you can rehearse and plan before, a really bad day happens, and then people are calling me. So most crisis management plans assume that production systems, you know, including identity, especially identity, and those collaboration tools that you’re talking about, you know, where where we’re having these communications across, those things, the assumption is those are gonna remain on. But in reality, those are the first that are targeted. Have you ever been on the phone in a Teams meeting where the threat actor joined the call? I have. Right? In addition to that, we also have, you know, siloed comms. You got different teams doing other things. They’re not talking to these other ones. You’ve got applications that are trying to be brought up before the identity plane is brought up. Nothing’s working, and it’s it’s chaos. It’s mad, mad chaos. Nobody’s quite sure on maybe potentially who’s on email chains, who’s on conference bridges, who can approve what. There’s no racy charts in place, for example, how to proceed if the standard tools that are inherently, good for communications are bad, nothing’s been defined, vulnerabilities that may exist, what are the dependency maps, blah blah blah. Right? There’s a whole slew of stuff to consider when it comes to crisis management. And what’s kinda scary to me is that, you know, less than half, I’d say, you know, forty to forty five percent of organizations have different procedures to remediate different vulnerabilities. In my experience, I’ve seen just a slight over tick, maybe sixty ish percent of companies out there that have an actual AD recovery plan, less than that have an Intra recovery plan, even less than that have Okta. Right? When you think of threat actor groups scattered spider, they started out hitting Okta, which that’s why they were called octopus in the very beginning. So it’s like, you need all this time to prepare, especially when it comes to things like you need a lot of you’re gonna need more than, like, for example, a day to recover from a ransomware event. Right? So there’s three things that I want you to remember, coming out of this, and that these three words, design, execute, and validate. So we’re gonna design an identity backup solution and restore patterns that support malware aware recovery. We’ll get into what that means in a little bit. It’s not just making sure that you recover fast, but also clean. Right? This is the foundation of everything. When we execute, there’s lots of different pitfalls that people run into, either prolongs a different outage. You open the doors for threat actors to come right back in. And so I’m gonna talk about how to kind of sidestep each one of those, and then validate how to prove that redecovered forests and domains are actually trustworthy before you reconnect those prediction those production systems. So, again, design, execute, validate. So as we get in design, this is where most of you can start immediately. Right? You want backups and restoration patterns basically created or engineered so that when you do restore, you bring back identity without bringing back any sort of malicious code, potentially malware, etcetera. So we’re getting into some of how to do that. The first one is treating Active Directory as simply a server role. Right? So most teams assume that we’ve got backups. We’re good to go. We’re fine. Yeah. No. The problem is that conventional and bare metal backup tools treat Active Directory just as another server role and capture the whole OS with it. Right? You got your operating system. You got your registry. You got files. So, like, if the DCs were compromised when those backups were taken or even prior, if ransomware were staged at some point, your policies, etcetera, that backup pretty much contains the same bad stuff. So if you were to do a bare metal backup, you fully you were just brought back, like, rootkits, the malicious code, potential backdoors, etcetera. So those traditional backups, while it’s good to start from, in a lot of cases, there’s ways to enhance it. Right? Taking you guys to this next level of thinking. So, I mean, I’ve seen there was this one gig, man. I did a we restored the forest literally eight times because of different stuff the customer kept doing, and it it just turned into a nightmare. And a lot of it was restoring malicious code. They thought they had a a good backup, so they kept going back and back and back and back. Yeah. That was not fun. So the whole business stayed down while we played that game. But, again, this starts here with this, you know, coupled backup. We don’t we wanna decouple Active Directory, for example. This is the single most important idea, making sure that that Active Directory is not part of a backup. You wanna back it up separately. So it’s independent of the operating system. So you can just restore AD onto a brand new clean OS that’s hash verified, etcetera, leaving your your compromised or maybe ugly DCs sitting in production, which we’ll get into in a second as well. Anything hiding in those DCs won’t be won’t be brought over into the recovered domain controller. That means even if every domain controller was hit by ransomware, you can still reach a non secure state. You don’t have to rebuild AD by hand from scratch. Never go to greenfield. There’s there’s can’t think of a solution or a reason why anybody should ever go to greenfield unless it was absolutely necessary. Ninety nine point nine percent of the time, you can recover from an attack using your existing Active Directory. Okay? The idea here is we’re gonna recover AD, not the infection of the Melissa’s code, and that’s what I need to restore from. I need to get a customer back up and running. Give me give me Active Directory. Right? All I need is the most recent backup, for what we do here at Semperis, of Active Directory to initiate recovery for a customer. Trust starts with the backup itself. Okay? Recovery is only gonna be as good as the backup you trust, so the principle is about making the backup itself trustworthy. How do we do that? Well, we need the most recently and and readily available backups. So whatever solution you guys have in place for that, you wanna make it sure it’s automated, there’s frequent snapshots, you’re not storing, you know, week old or potentially two week old, data where you’ve lost all that that change that occurred between the backups. You also wanna have immutable and tamper proof once written backups. Right? So they can’t be altered, deleted, not by ransomware, not by threat actor, or even a malicious insider, etcetera. Once it’s there, it’s it’s done. And then you want isolation. You want a virtual air gap, backup storage per sectors that kept away from prod. For example, like an immutable cloud storage. Right? That same attack that took down production can’t touch the backups. That’s the goal. You want that in your design. I’ve seen attackers target backups where they target the infrastructure hosting the backups, just because the identities were synchronized and they could do that. So try to be mindful of those. So as an attacker, if I can get to you just one of those backups, if I get a backup of AD or backup of the main controller, I own you. Right? You wanna remove that circular dependency. There’s a lot of different you know, what are the processes that secretly depend on the very thing that’s down. This is that dependency map I was talking about earlier. So, like, critical business apps, what they depend on, etcetera. If you’re backup, I would consider backups, obviously, critical business app. If that thing needs, like, say, for example, Windows authentication, DNS, Active Directory integration, credentials, If they if they need that to function, then when Active Directory is down, guess what? You can’t even start recovery. It’s the whole chicken and egg thing. Right? Especially when the dumpster is on fire. You don’t want that. The design principle, right, you want the backup and restore engine, whatever is running it, independent of Active Directory. This is where the decoupling pieces are. Right? You want it to be physical, it could be virtual, flexible IPs, OS provisioning, whatever. Because post incidents, you may not have your own usual clean hardware available, so you wanna be able to move it to wherever you can. I’ve seen a lot of cases where we take the backup the existing backup server, we spin up a VM, throw it over here, bring it all up. Within about an hour, we can do the restore. Let’s get going. So keep it back up segmented and keep it non AD integrated wherever possible. K? And now you wanna prove it. K? So you’ve designed all this stuff. Now we gotta prove it, especially before the bad day happens. You wanna try to eliminate as many of those unforeseen circumstances or nasty surprises during an incident as much as possible, and you gotta have confidence in your backup solution or in your restoration plans before the world is coming down because that’s what happens. It’s mat like, it’s it’s can get crazy. It can get crazy. Each backup set that you’ve got is gonna contain everything required to recover the whole forest not just, like, a subset of it. So that means you wanna write backups to more than one location. You wanna run automated verification with notifications if something’s missing, if there’s, a gap, etcetera, and really, really focus on this one. That last bullet point that you see where it says test regularly, that’s probably the most important one on this slide. Teams that drill, and they train for incidents like this, especially in their lab environments with the they practice and practice and practice. They’re not discovering new stuff or new potential pitfalls at three AM during a ransomware event. Recovery is a muscle. You wanna train it. Right? Especially where there’s environments where there’s staff turnover, people may have been brought up on what to do, and then they’re no longer with the organization. So you gotta keep going. Keep rinsing and repeating. Because an attack will happen. It’s not a matter of if, but when. So be ready when you can. And don’t stop at Active Directory. There’s lots of different things that control AD. So the design principles that I just talked about for Active Directory have to extend the whole hybrid fabric. You recover if you don’t, you could recover one platform and then it you know, you miss another one and then bad things start to happen. Cloud Identity, for example, needs first class backup and recovery too. You can use multiple different solutions out there, but you have one. Again, not a sales pitch. But fast secure protection of Entra ID resources and tenant configuration where you have conditional access policies, app registrations, Entra ID, group memberships, etcetera. Right? All this stuff, continuous flexible backup and recovery for things like Okta, where you have policies for ping, app assignments, whatever. There’s gotta be some sort of integration where you’re you’re passing the buck from, Active Directory and you’re encompassing all of these other IDPs. And you wanna start with your most vulnerable and most catastrophic environment. And, unfortunately, for which most companies with all the identity debt people get over the years, that starts with Active Directory. Then you wanna extend that same backup, verify, recover pattern outward towards Entra, maybe Okta, PingFederate, etcetera. But the goal is to have one consistent rehearsed recovery plan across all of them. K? Say the actor hits the cloud first, they pivot to on prem. Again, very common. So in the beginning, maybe we’ll not start with Active Directory. Maybe we’ll start with Entra and conditional access policies. We’ll take control over active or Entra ID and make sure that the threat actor is somewhat contained there before we move to on prem. Depends on skills, depends on experience, depends on people involved in the in the incident. So every like I said, every incident’s unique, but the idea is practice different things so you can be ready. K? So now we’ve kinda got out of the design phase. We’re moving towards the execution pieces of it. So moving from architecture to, like, say, for example, the heat of an incident. These these were to talk about some some pitfalls that I’ve seen from different engagements. We’ve seen it’s in Paris different from in different engagements. Right? Goes from things like reintroducing malware, restoring too fast, having no isolation, etcetera. Right? So, these will somewhat hit home, and I can I can tell stories for for days on this stuff? But this is probably the most damaging execution mistake, and ties really back into the design section, and that is reintroducing malware and persistence. Unfortunately, under pressure, teams are gonna, you know, go back and forth and grab whatever backup they’ve got to go as fast as they can. And if it’s an OS inclusive image, they just reinstated rootkits, ransomware, backdoors, all kinds of malicious fun along with Active Directory. And I I’ve seen I’ve seen threat actors inject malicious code into ISOs that they have found from, like, Golden Image files, for example. I’ve seen threat actors inject malicious code into existing backups, which is why we talked about designing that solution to keep it, you know, as as off to the side as possible. Right? Persistence inside of Active Directory, backdoors inside of AD, different admin accounts, modified ACLs, GPOs, DLLs, pretty much you name it, I’ve seen it. So even a clean operating system can restore a, unfortunately, compromised directory. Right? Does that malware aware pattern? Right? You wanna recover AD objects to a known secure state on a very clean server, hopefully ISO verified, and then if you wanna eradicate or clean up as much of that persistence that you know about, before you bring anything online. So this is working everywhere, inside of that clean environment, so to speak. If you if you skip this process, you can close the reinfection loop, and that attacker simply just comes right. Fast is good, but simply relying on speed means nothing if you’ve actually restored the adversary as well, which is never a good thing. Right? This is where we’re talking about speed without trust. This is the other half of that dilemma. Now it’s a warning for execution. Right? So the instinct that I see a lot of people go through is to reconnect the moment the main restored domain controllers are up. It’s extremely dangerous. I’ve seen attackers come back when we’re playing that game of whack a mole, and they’re watching for the rebuild so they can simply redeploy and re encrypt. This isn’t that one I referred to before where we had to rebuild the forest, multiple times, but I have seen this numerous times, in different engagements and, where even the unfortunately, the CISO was just like, I’m done. Walked out. It was sad, but it actually happened. We did get them back up and running, but once once they understood this pitfall. So there’s two things here. Right? Do remediation quietly. Try to do it in isolation so attackers don’t know what’s going on. Right? If you more than likely, if you started the recovery process, they have an idea that something’s happened. Maybe they’re that they know that they’re that they’re that you guys are onto them, so to speak. But if you’re doing all the work inside of an isolated environment, they can’t see what you’re doing. And so, obviously, it becomes critical to have that isolation environment. The other piece of this is to treat domain controllers, that are online as the start of validation and not the end of recovery. What I mean by that is you wanna balance urgency against verification. So speed only counts, really if the environment you bring back is actually clean. Right? Just because you’ve restored a DC doesn’t mean it’s trustworthy, as as bad as that sounds, as it as unfortunate as that sounds. But it’s just it’s just a fact of the matter. Right? When you have something for isolation, you’ve gotta have and this is where you can start today, right, where you start to prestage that isolated environment. When when things hit the fan, many teams skip this piece because they don’t have they don’t ever have one staged. And the consequence of this is that they end up validating and remediating inside the same network aka production, which is doable. It’s not the preferred method, obviously. I have recovered customers in a prod, but you’ve gotta do it in a certain way. You can’t just restore and go. Without isolation, though, I mean, you you can really get up validating and remediating, what the product the the attackers can see. And so the idea here is to to rapidly build a clean environment that that is trustworthy inside, that’s not that’s completely severed off from any sort of production at all. It can’t go out to the Internet, can’t talk to anything on production. So however you guys wanna design it via firewall, NSGs, for the perimeter inside VMware, I don’t care as long as it exists. Never have your recovery surgery in the same room as the infection. Right? You gotta build the clean room first and have it ready to go wherever you can. This one, I like this slide. I don’t know if, those of you watching are familiar with the manual Active Directory Forest Recovery steps that Microsoft publishes, But this is the documented high level steps of, what happens and what what Microsoft recommends. Right? And this is insane to try to do all of this manually. Okay? There’s literally twenty eight to thirty steps. And under a lot of pressure, that’s a recipe for mistakes. You got mistyped commands, miss you got missed steps completely, configuration updates don’t line up, patching, clean OSs, gotta rebuild the GCs, global catalogs, restoring metadata, doing metadata cleanup, making sure DNS works. It’s never DNS. Right? Each error and thing that you run into during this, it always requires go back to the restart and start over. It sucks. And all this time, the business is bleeding potentially thousands, if not millions of dollars an hour. Again, I’ve seen it. It’s you gotta have something that is that can orchestrate all of this. This this ADFR’s recovery from end to end, so the sequence is consistent, repeatable, very fast. Right? And if you can automate this, I’ve seen a lot of organizations cut their recovery downtime by ninety percent. It’s it’s like a no brainer. This is not just a convenience, but this is risk reduction if you can automate and and orchestrate this. K? Take out the human elements. Right? Making sure that we don’t have to start. And the real kicker for this one, though, is after you’ve gotten through all twenty eight, thirty steps, whatever it is, the forest is still not trustworthy, which yay. But this is what we need to get going. This is another this is the fifth one I wanted to talk about, and that is when lights out or out, dependencies on Active Directory, when things are down, you can’t even talk to each other. This one usually gives people that, oh, I didn’t think about this one reaction. So, like, when identity is down, everything that depends on identity also is down. Right? People are gonna lose just hours trying to assemble and coordinate meetings to have meetings about other meetings because comms are down, decisions can’t be made. Right? So the fix is to make sure you pre establish some out of bound communications, some sort of command and control, port center that does not depend on Active Directory, for example, or other identity systems that may be under attack. This could include your execs, your incident response team, identity team, your I’m your infrastructure, legal teams, communications. They all work from secure scannals with with clear roles. Right? So you will need to decide how you’re gonna communicate during the idea the actual outage before one actually occurs. K? Crisis coordination is part of recovery. It’s not a side task. Alright? So this is a big one. K. So we’ve executed. Now it’s time to hit the home run, so to speak. This helps when we validate and reestablish that trust. This helps protect you guys from a recompromise. You’ve designed clean backups. You’ve executed a careful recovery, but now we have to prove that the environment is actually good, and we can sleep well at night. Sixty percent, roughly, like I said before, have that Active Directory recovery plan. But of that percentage, there’s a smaller percentage that actually maintain dedicated AD specific backups. And if you aren’t one of those, you could be looking at a long time. Remember that manual AD force recovery thing we walked through that’s got those thirty something steps in it? That could take, twenty four, thirty six hours, requires a lot of expertise. That may not be available at, like, three AM. Remember, we’re trying to move fast. Right? So we have to find and evict or contain what’s left behind, and getting that Active Directory back online is not the same as getting a trustworthy directory. But post breach forensics, that’s how you can earn that trust. Right? Different products that are available for out there that not only do we do that, but other companies, you know, you wanna do something, to where you figured out what made the threat actor have may have touched and what identities were used. Because, again, vast majority identity is used. You wanna know what’s there, and then you wanna start doing those, what we call account disposition actions against those accounts. Maybe even go through permissions, look at GPOs, look at stage, SysVol files, different delegations that may be in there. You wanna work the timeline. Identify that try to identify that original attack vector. You wanna try to remove those compromised objects, get rid of bad code, maybe malware, ransomware, close those doors as much possible. You’re never ever ever going to be able to fully shut the door, but you wanna close it as much as you possibly can to keep them from coming back in. And we do that, a lot of times within, starting with a with Active Directory. Right? We wanna do this before reconnecting those critical critical systems. So before we do that cutover to prod, we wanna make sure we’ve got that trustworthy source. Right? This is also something that we specialize in here. It’s pretty unique to the industry. We wanna recover in the right order. This is also a critical one. We don’t wanna try to bring everything back at once. Remember that example where people were talking about, oh, let’s bring up these basic critical apps at the same time Active Directory is trying to come up. No. We gotta sequence it. The concept is this thing called MVC or minimum viable company. This is the smallest set of identity or services needed to get the organization operating again. So you wanna record, recover the core identity first. And then because everything else depends on it, then the most critical business applications, then expand outward and keep going. But the the kicker is getting making sure that threat actors can’t do what they did before you reconnect those users. Otherwise, you’re inviting people back into an environment which you haven’t cleared yet. So establish trust at the center, prove it, and then start to widen the circle. When you do things in stages, it also limits the blast radius. So let’s say you’re in domains, you moved into the applications, you’re starting to notice something’s wrong, you can catch it before it spreads to the rest of the reconnected state. Right? And now we’re getting close to the end here. And to just kinda sum all this up, we’ve got this playbook, concepts of a playbook. Right? High level steps, design, execute, and validate. Right? So these are for the design phase. Right? Just kinda rehash. You want the, active directory as decoupled. You wanna store it on immutable backups and verify whether it be across Active Directory, Entra, Okta, whatever. For the execution, we wanna make sure we’ve got that isolation picked up and going. We’ve got end to end automation for the ADFR’s recovery steps, out of band coordinations, and then trying to avoid those pitfalls that I mentioned previously. And then from the validation, this is the most important piece of it. Having that staged minimum viable company plus forensics and making sure that we’ve done all the necessary hardening steps to Active Directory to confirm that trust before reconnecting critical systems. What I mean by that is before you do that cutover back from isolation in the prod, you wanna make sure that you can trust it. Because, again, you don’t want the threat actor walking right back in. Right? And it’s painful. I’m not gonna lie. It’s painful. A lot of the accounts that get hit or or that are used by threat actors in these engagements, a lot of service accounts tied to massive applications and a lot of different customers. A lot of those are always impacted. You got admin accounts that you gotta do work on. You’ve got, maybe, this is why I like we will always love to refer to group maintenance service accounts. If you guys can move to those or DMSAs, some form of MSA, and not your traditional service accounts, do those. Those are much easier to deal with. And then you’re potentially looking at maybe even a mass password reset. Consider that. If the NTDS DIP file is either if there’s evidence the during the investigation, if there’s evidence that the DIP file was staged, exfiltrated, whatever, If the d, you know, assume breach. Right? If it’s if any of that, then mass password reset is almost a requirement. Right? That’s that can be painful in and of itself, especially with different design limitations that may be in configurations for for, say, Entra. Maybe you guys aren’t using SSO for now, but maybe you’ll want to in the future. It helps along with SSPR, self-service password reset. So, yeah, it can get pretty hairy pretty fast. And so my goal here, hopefully, is to convey how important it is to establish a trustworthy version of AD that that will not allow the threat actor to come back in and do what they did before. Again, not about just getting domain controllers back up and running. It’s knowing you can trust it. Alright? With that, I’m gonna turn it over to some questions. Wonderful. Thank you so much, Tim. Oops. Sorry. No. No. Go ahead. Oh, I was just gonna say thank you so much for that great presentation. We do have some questions in from the audience at this time. And I would also like to remind the audience that if you have any questions, feel free to type them in now, and we’ll get them answered for you. Our first question, yeah, is what’s the biggest mistake you see organizations make when it comes to the ability to quickly recover their minimum viable company? That is a tough one. Not being ready, honestly. Thinking that they can, you know having that mentality of, oh, it’s never gonna happen to me, and then it does. And then they’re freaking out. So not only is it psychological, but it’s, you know, technical by nature. Right? There’s a lot of change that haps that has to happen during recovery scenarios, to move move organizations forward after they’ve been hit. If they’re mentally not prepared for it, it makes, I guess, acceptance of those changes a lot more difficult for them to acknowledge and adhere to, which opens the floodgates right back open again. For example, there it requires some things require an administrative shift. I’m not gonna get into all the technical details right now, but as a as a as a friendly hint to those watching tier model. Most, by default, the domain, when you when you bring a domain up, domain admins can log in anywhere and everywhere. They manage the domain. Right? They’re gods over everything, so to speak. The problem is threat actors notice, and it becomes very, very easy for those keys to the kingdom to be exposed on lower tiered assets, for threat actors to get their hands on, and it’s game over. So people, when when things happen, they have to require an administrative shift to adhere to the new tier model, and that is to only limit what domain admins can log in to, and that’s part of it. And it requires some sort of an administrative shift to where they’re not gonna be RDP ing into, like, say, their SQL servers with domain admin credentials. That can’t happen. Not anymore. So if you start implementing those restrictions, you’re gonna be much better off. But practice. Right? Not having response plans, not having, again, with that mentality of thinking that it’s it’s never gonna happen to you, it that’s again, it’s not a matter of if. It’s a matter of when. So be ready. That’s a good question, though. Our next it was. Our next question up is from John. John wants to know, why would we need a specific backup and recovery solution for identity if our backup solution includes an identity module? Ah, yeah. So this is that, I’m not gonna pick on any vendor names. But there’s there’s lots of backup solutions that are out there that, you know, does the thing that we talked about earlier, then they back up the server. They back up the domain controller. They do system state. They grab Active Directory. They’ve got it all in one big package. Right? But it’s also backing up every other workload that they own inside the organization. And to manage and log in to the backup solution, it’s tied into Active Directory. It’s kind of like all your eggs in one basket. It’s very easy for a threat actor to target those backups. I’ve seen I’ve seen threat actors target backup storage appliances. So we’re not not just the the backup solution itself, but the actual SAN appliance that it lived where the where the where the the backed up data lived and resided and encrypted the blocks directly. You know, that was fun. You wanna you wanna keep it isolated, especially when it comes to your IDPs or your identity platforms in Active Directory because that is very easily compromised during these kinds of attacks. So the more you can protect it so for example, one of the one of the recommendations that we’ll do during incident response is, you know okay. So for example, threat actor hits VMware and encrypted all the VMs. Very common, unfortunately. And, they got into VMware because it was AD integrated. They use AD integrated creds. Once I have access to VMware, boom, it’s game over. I own you. And so it’s like, okay. We’re gonna spin. Now we gotta build a whole new VMware environment, during recovery. And, oh, by the way, you cannot you’re now you’re now you’re gonna use locally stored credentials. You’re not gonna use AD integrated creds anymore, and they learned their lesson the hard way. So you gotta decouple that pieces of it. Right? Even though your backup solution has that identity piece, so to speak, you gotta have it segmented off. It’s too critical. Alrighty. We have another question in. This one comes from Jay. Jay says, you mentioned recovering in isolation. How or when do you recover back up to production? I understand oops. Sorry. I went too far. I understand after the attack is mediated, but don’t you still need, the identity during the time? This is a good question. This is really fun. This is, so this is what we specialize in, and that is keeping production alive, as much as possible, not only for making sure or trying to keep the threat actor at bay, right, not tip them off that what what it is we’re doing, but also to allow the end users to continue working. The most the vast majority of times, organizations that get popped, things the end a lot of times, the end users don’t even know. And so part of the part of the skill set and experience is that, you know, being able to maintain that user base to where they can continue doing functional stuff and maintaining production, is is part of the process. So we do all of our take back inside of isolation, for example. We’ll get our restored Active Directory. We’ll do all of our take back actions, inject tiering, and a whole bunch of other different things that we do, that’s unique to our process, and then we plan a cutover time. During that cutover time is where we’ll tell everybody where the word goes out usually, hey, end users. We’re having a maintenance window, so expect things to be down for a couple hours. And at that point, once everybody’s notified, it’s all planned in a in a perfect scenario, all the production DCs that were running while we were doing all of our work in isolation, all those get shut down. They’re never to come back online again. Then we do the cutover. We start spinning up new DCs based off of the new version of Active Directory that’s hardened and trusted. Reset passwords on things that need password reset, and it just goes. So I can’t give you specific time frames. Obviously, every scenario is different, different applications, different systems, etcetera. But that’s kind of like the three thousand foot view. How we do it. It’s not, it’s never one hundred percent clean. And what I mean by that is, something’s gonna break. So the more you practice, the more you can minimize that that negative impact, if that makes sense. No. That does. Thank you so much for answering that question for us. It looks like we’re about out of time for today, but I did just wanna say thank you so much for your presentation today, Tim. Yeah. It was fun. If you still have questions or if you want to type one in after we finish this webcast, please feel free to do so, and I will send it over to Tim to answer after this webcast. I do just wanna say thank you to Tim for being on with us today, and thank you to Semperis for making this event happen. Thank you again to the audience for today’s webcast and, of course, attending. Have a wonderful rest of your day, everyone.
