Reading view

There are new articles available, click to refresh the page.

The Agentic SOC: Transforming Data into Defensive Velocity

Security Operations Centers (SOCs) are currently confronting scalability challenges on two fronts: structural and cognitive. The day-to-day reality of modern defensive operations is stark: an analyst frequently begins a shift facing a queue deeply saturated with unvetted alerts. To process a single event, the analyst must open the alert, pivot to a secondary console to complete an investigation, manually enrich an IP address, copy a file hash into a third interface, and cross-reference an asset inventory that may not have been updated in months. Following this, they must author and refine queries, waiting for overloaded databases to return historical context.

The actual work of assessing the investigation’s results and moving to decision-making and action has not even begun. This is the administrative burden of the modern SOC. The true threats are not just those that attempt to bypass defenses, but the critical operational hours lost before an active mitigation attempt is even initiated. While analysts are highly trained professionals, the relentless requirement to perform manual data aggregation inevitably leads to exhaustion.

Misdiagnosing the Bottleneck: The Upstream Data Problem

Threat actors operate at machine speed, utilizing automation to pivot laterally across networks in a matter of seconds, frequently disappearing before defensive teams can even log into their terminals. Expecting human defenders to counter automated threat vectors by manually aggregating bad data is an architectural failure.

Every SOC inherits a highly fragmented data ecosystem. Telemetry is continuously generated by diverse sources, including firewalls, cloud workloads, identity providers, endpoint sensors, and legacy systems. This telemetry arrives in disparate dialects, varying formats, and highly inconsistent levels of fidelity. Before AI tools can accurately reason about a potential threat, or an analyst can initiate a logical investigation and run a playbook response, this raw telemetry must be synthesized.

Historically, organizations analysts take on these complex synthesis processes, manually normalizing data points across different vendor schemas. This represents a key misallocation of human intelligence. The asymmetry in modern security operations is not merely a discrepancy in speed; it is an imbalance in how security teams are forced to allocate their finite time. When operators spend the majority of their shifts wrangling data instead of actively investigating threats, the foundation of the SOC itself is inadequate. To achieve defensive velocity, organizations must recognize that fixing the data foundation is the mandatory prerequisite for improving all downstream security functions.

Architecting the Data Foundation with Singularity™ AI Data Pipelines

Addressing the upstream data problem requires the implementation of advanced data pipelines capable of resolving enterprise data chaos before it impacts the detection engine. Frameworks such as SentinelOne’s® Singularity AI Data Pipelines serve as this foundational layer, engineered to ingest telemetry from every source and in every format without requiring months-long integration projects or heavy manual engineering.

Modern pipelines utilize AI to normalize raw telemetry into standardized formats, specifically aligning with the Open Cybersecurity Schema Framework (OCSF). This structural alignment transforms fragmented logs into structured data that is immediately actionable. It eliminates the need for analysts to construct complex regular expressions during critical incidents simply to reconcile how two different software vendors format data, such as usernames or a timestamp.

Efficient data ingestion also requires dynamic, in-flight optimization. Not all telemetry possesses the same analytical value, and storing all generated logs in highly indexed, expensive storage tiers is financially and operationally untenable. Data pipelines optimize data streams by filtering out extraneous noise, trimming excess volume, and routing specific logs based on dynamic criteria. High-value security events are routed and indexed for rapid search retrieval, while lower-priority compliance or operational logs are routed to more cost-effective tiered storage. The result is a substantial reduction in infrastructure costs, a higher signal-to-noise ratio, and a structured data foundation that is completely prepared the moment an investigation is required.

When underlying data pipelines automatically enrich that log with identity and asset information, revealing (for example) that a specific financial director’s laptop in a remote office is communicating with a known botnet, the output transitions from a raw data point into a definitive starting point. Crucially, this enrichment occurs systematically before the human operator ever interacts with the alert. Solving this data problem end-to-end is a primary reason SentinelOne was recognized in the IDC MarketScape for AI SIEM.

Accelerating Detection via Singularity AI SIEM

When a clean, structured data foundation is properly established, the performance of downstream security tools accelerates. Modern detection engines, such as the Singularity AI SIEM, leverage indexless architectures to manage enterprise-scale telemetry. Because the data is normalized and optimized prior to ingestion, these platforms can execute petabyte-scale queries with minimal latency, ensuring investigative results are delivered before the analyst’s attention wanes.

Within this architecture, detection logic is executed continuously against a stream of clean, correlated telemetry. This transforms an ocean of disparate event logs into readable, centralized dashboards that provide immediate situational awareness. The quantitative benefits of this approach are substantial. With AI SIEM, organizations are already executing their queries 70% faster. Adding AI Data Pipelines further augments this workstream, providing cleaner data for AI to run at optimal efficiency. These improvements represent the direct result of ensuring that the data arriving at the SIEM is inherently fit for purpose.

AI SIEM remains a single, comprehensive SKU with customers automatically receiving integrated pipeline functionality for everyday data optimization rather than treating it as a premium add-on. For every unit of paid Data Ingest capacity, customers can process twice that volume through Data Pipelines. A customer with 500 GB/day SIEM entitlement can push 1 TB/day through the pipeline at no additional cost.

Transitioning to Agentic Reasoning Layers with Purple AI

The establishment of a structured data pipeline unlocks the capability for true agentic reasoning within the SOC. Unlike traditional rule-based automation, which executes static responses to predefined triggers, technologies like SentinelOne’s Purple AI operate as a dynamic investigative layer.

When an initial alert is generated, an agentic reasoning system does not simply pause and wait for human triage. It autonomously launches an investigation, comprehensively maps the potential blast radius of the incident, and synthesizes a clear, logical recommendation for containment. Then, the analyst logs into the console and is presented with a fully formed situational briefing rather than a blank investigation screen.

More importantly, an agentic AI layer possesses the capacity to evaluate broader adversarial campaigns rather than isolated security events. In isolation, a minor registry key modification, a singular file write, or a brief outbound network connection may not meet the threshold for a critical alert. Legacy security tools often fail to connect these disparate, low-signal events. However, Purple AI can assemble these seemingly unrelated activities into a cohesive narrative, exposing the overarching strategy of the attacker before a major breach occurs.

This level of autonomous intelligence is strictly dependent on the underlying architecture. Advanced AI algorithms cannot derive accurate conclusions from unparsed, low-quality telemetry. The analytical integrity of the agentic layer is entirely contingent on the principle of data quality; systems like Purple AI require clean, structured data to function effectively, avoiding the fundamental issue of “garbage in, garbage out”.

Governed Hyperautomation and the Human-in-the-Loop

The final component of a modernized, agentic SOC is the deployment of Hyperautomation to execute defensive responses. To counter threats effectively, organizations must deploy automated workflows capable of executing decisions at machine speed. These no-code workflows can be configured to trigger autonomously based on AI triage verdicts, the disclosure of new high-severity vulnerabilities, or specific incoming alerts. By automating the mitigation phase, the SOC evolves from an environment strictly dedicated to passive observation into a dynamic system that actively neutralizes threats.

However, the implementation of automated response mechanisms must be rigorously governed. Executing changes to enterprise infrastructure carries inherent risk. To mitigate this, automated workflows must integrate critical approval steps, ensuring that highly consequential actions are paused until human authorization is provided. The analyst retains the ultimate authority, defining the precise parameters of what processes may run automatically and what workflows require manual judgment.

Redefining the Analyst Mandate via Autonomous Security Intelligence

The strategic objective of integrating data pipelines, agentic reasoning, and Hyperautomation is not the removal of the human operator. Instead, the overarching goal is the restoration of the analyst’s primary function: exercising expert judgment.

By offloading repetitive tasks to technological systems, organizations systematically remove operational friction. The data layer filters out irrelevant noise, allowing the analyst to clearly see the threat. The AI investigation layer removes the administrative grind of data collection, allowing the analyst to focus purely on analytical thinking. Finally, the automated response layer eliminates procedural delays, ensuring the analyst’s decisions are executed rapidly enough to matter. This creates an intelligence fabric, known as Autonomous Security Intelligence (ASI), where data, investigation, and response function concurrently as a single, unified system.

Under this model, the operational output of a single analyst is exponentially multiplied, allowing one unburdened professional to accomplish the work of ten while still owning every critical decision. While the alert queue will perpetually require attention, the fundamental nature of the work fundamentally changes. The timeline of a manual initial triage to active investigation compresses from a multi-hour ordeal into a matter of minutes. The data arrives clean, the investigation runs automatically, and the response mechanisms are prepared. The hours previously consumed by administrative waiting are directly reallocated to strategic decision-making.

Conclusion

When defensive systems are finally architected to operate at the speed of the modern threat landscape, the role of the human operator transforms. Analysts are no longer forced to act as passive passengers, grateful to be carried by fragmented tools. They are elevated to the role of pilots, operating with full situational awareness, retaining their judgment, and actively directing the defensive posture of the organization. This is the paradigm of the agentic SOC, and it is entirely predicated on the foundation of clean, structured data.

Contact us today to learn more about how SentinelOne is leading the way forward with Agentic SOC.

 

Missed incidents, persistent threats, and response gaps: Insights from compromise assessment projects

The following analysis presents the key findings from Kaspersky Compromise Assessment engagements performed in 2025. A compromise assessment is an independent, expert-driven service that examines whether a target network has been compromised. The service combines threat intelligence analysis (including darknet sources), tool-aided endpoint scanning, a systematic review of security event logs and network traffic, and, when necessary, an initial incident response and digital forensic investigation.

This report focuses on missed incidents – threats that remained undetected for weeks, months, or even years.

Key trends observed during compromise assessment engagements

  • Proactive compromise assessment decreases the number of missed high-severity incidents. The highest proportions of high-severity incidents were revealed in organizations that requested our compromise assessment service after containing a known incident. The lowest proportions of high-severity incidents were observed in organizations that conducted regular audits. Of all the incidents discovered, 20% were found manually, while enterprises missed 60% because of the absence of high-confidence alerts from the tools in place.
  • Nearly a third of discovered incidents took over three months to detect. The longer a threat persisted in the target environment, the greater the likelihood that an incident would be severe. 30.8% of all discovered incidents and 52% of high-severity compromises had historical activity spanning over three months. The oldest incident discovered in 2025 had gone undetected for four years.
  • Malicious files often remain in backups and are restored after incident response activities. 40% of all discovered web shells resided in backups and went unnoticed until a proper compromise assessment was conducted.
  • Threat actors rely on remote management tools and LoLBins. These types of tools were found in all compromise assessment engagements that resulted in an incident detection.
  • Monitoring tools and controls are not self-sufficient; operational maturity makes the difference. Monitoring tools must be configured and adapted to the changing threat landscape. Furthermore, human analysts need to review low-confidence alerts. A lack of continuous monitoring and threat hunting activities increased the likelihood of high- and medium-severity incidents to 84–86%. At the same time, high‑severity incidents were rare among organizations with in-house capabilities to reverse-engineer malware.
  • Communication issues lead to missed incidents. Nearly a third of the compromise assessments revealed communication issues that impacted incident response activities.
  • The incident response playbook is not set in stone. For incident response to be efficient and effective, playbooks must be updated as new artifacts are discovered. Treating the incident response plan as a living document reduces the risk of missing threats.

About the Kaspersky Compromise Assessment service

Our global compromise assessment portfolio spans several regions. In 2025, around 71% of the incidents we identified affected our customers in the META region, while the APAC and CIS regions accounted for the remaining 29%.

Geographic distribution of incidents identified during Kaspersky Compromise Assessment projects in 2025 (download)

Our service was requested by organizations from a diverse set of sectors. The government sector accounted for around 29% of incidents, followed by the education (19%) and financial (17%) sectors.

Distribution of economy sector incidents identified during Kaspersky Compromise Assessment projects in 2025 (download)

Detection logic families

Our compromise assessments operate on a continuously updated catalogue of indicators of attack (IoAs). Because the raw set of IoAs is too granular for high-level reporting, we map them to a concise set of detection logic families. The statistics indicate that three detection families dominate the incident mix:

  • Credentials from dumps: 12.4% of all incidents;
  • Specific living-off-the-land (LOTL) tools: 11.2 %;
  • Specific malware families: 11.2 %.

These three detection logic families represent high-fidelity indicators of attack that reliably signal infrastructure compromises ranging from dormant, disk-based malware to persistent and multi-stage attacks.

Distribution of detection logic families (download)

Reasons for requesting Kaspersky Compromise Assessment services

Analysis of our compromise assessment engagements that took place in 2025 reveals a clear correlation between the stated purpose of the engagement and the risk profile of the findings. General audits dominate the portfolio with 56% of requests, followed by authority reporting engagements (19%), post-incident checkups (17%), and acquisitions (9%).

Statistics on the reasons behind CA project requests (download)

When the findings are classified by severity, the post-incident checkup category exhibits the highest proportion of high-severity incidents (40.7%). The full breakdown is shown below.

Incident severity breakdown by service engagement reason
Incident severity (%)
High Medium Low
Reason for service Acquiring new company 28.6 42.8 28.6
General audit 27.7 36.7 35.6
Report to an authority 30 46.7 23.3
Checkup after a cybersecurity incident 40.7 25.9 33.4

Post-incident checkups are frequently initiated after an initial incident response (IR) effort. The elevated share of high-severity findings suggests that IR activities, which are typically limited to containing a known incident, do not provide a complete view of the broader environment. Consequently, other threats may remain undetected until a full compromise assessment is performed.

Merger and acquisition-related assessments are proactive assessments performed when a company acquires another entity. This involves the target’s network being scanned for hidden threats before the two environments are merged. These assessments demonstrate a balanced distribution of severity: 28.6% low-severity, 42.8% medium-severity, and 28.6 % high-severity. This reflects the mixed risk posture of target environments of acquisitions, which are often evaluated for both known vulnerabilities and hidden malicious activity. Similarly, other proactive approaches like general audit assessments or assessments driven by the need to regularly submit a compliance report to a regulatory authority, share almost the same ratio. This indicates that regular, proactive and compliance-oriented assessments tend to reveal substantive issues earlier in the attack lifecycle, reducing the likelihood that they will evolve into high-severity incidents.

Organizations that conduct regular audits have the highest rate of low-severity findings (36%) and the lowest rate of high-severity issues (28%). We can assume with medium confidence that continuous, proactive compromise assessments are more effective at limiting the emergence of high-severity compromises than reactive, incident-driven evaluations. The data collected in 2025 are consistent with this hypothesis. Integrating regular, third-party compromise assessments into governance processes can therefore reduce the probability of unexpected high-severity findings and improve overall risk posture.

The following case study illustrates the impact of relying on a reactive rather than proactive approach. It describes a persistent threat that remained dormant on a client’s network and was only discovered after a comprehensive compromise assessment was performed following initial IR activity.

Case study: Dormant threat uncovered only by a compromise assessment

A midsize enterprise suffered a high-severity intrusion that was contained and remediated by the IR team within the defined scope of the initial alert. Following containment, the organization requested a check to determine if any additional footholds existed elsewhere in the network. To address this need, the organization engaged Kaspersky’s Compromise Assessment (CA) service, which performed a full forensic review of the environment beyond the scope of the initial incident.

Compromise assessment experts collected forensic metadata, historical security event logs, and Active Directory configuration data from the entire infrastructure. Threat hunting queries were executed against the aggregated telemetry, focusing on persistence mechanisms, lateral movement artifacts, and anomalous process activity. As a result, a number of severe threats were detected and reported; for example, malicious persistence:

  1. A cron job that recreates a web shell
    A critical Linux system (web server) had a cron job that automated fetched a copy of a PHP web shell from a public GitHub repository and placed it in an online directory. Even if the file was removed by security personnel, the cron job would simply download it again, giving the attacker a persistent remote code execution point on the web server.
  2. A live reverse shell
    On a server hosting a published web application, the process list showed a bash reverse shell.It was run by a user with the username “apache,” which was the account used to run the web application. This may indicate that the attacker exploited a vulnerability in the web application to gain remote code execution, allowing them to establish a reliable command and control channel that bypassed the firewall because it was initiated from inside the network.
  3. ClipBanker data stealer persisting via Windows registry
    A ClipBanker variant was detected on a user’s workstation machine maintaining persistence by adding itself to the registry key HKU\S-1-5-21-[REDACTED]-500\Software\Microsoft\Windows\CurrentVersion\Run\9Er6IIp.

    This was done after adding the malware’s folder to Windows Defender exclusions and applying hidden and system attributes to the file to hide it from regular users.
  4. Malicious WMI event consumer with deceptive alias
    A malicious WMI event consumer was detected that downloads and executes a PowerShell script. It created the alias “Kaspersky” for “Invoke-Expression” in an attempt to blend in as legitimate activity in the hope that a quick glance at the script would not raise suspicion. Kaspersky’s Cyber Threat Intelligence confirmed that the downloaded script (no longer reachable) was a weaponized payload used to spread the infection further.

The IR containment was rapid, focused and effective in addressing the specific incident that triggered the alert. However, the broad-scope compromise assessment revealed multiple backdoors across the environment, each using a different persistence technique: cron jobs, scheduled registry runs, and WMI subscriptions. The infected hosts were outside the original IR scope, so they remained unseen until a comprehensive hunt was conducted.

Incident response excels at stopping the bleeding and ensuring business continuity after a known incident. A compromise assessment provides a health check that determines whether any other wounds exist. By pairing timely IR with regular, full network compromise assessments, the organization had both the reactive agility to contain incidents and the proactive visibility to eradicate malicious persistence wherever it was hiding. The investigation uncovered additional undetected footholds, providing a clearer view of the environment and reducing the likelihood of a repeat incident.

Missed long-term incidents

The statistics on the mean time to detect (MTTD) incidents identified during compromise assessment projects are concerning. Many incidents go unnoticed for extended periods. For example, in 2025 we identified an incident that was approximately four years old!

Such prolonged detection times can lead to severe consequences, as 30.8% of incidents have historical activity spanning over three months. These incidents can range from dormant malware to persistent threats, highlighting the need for robust detection and response mechanisms.

Severity distribution of incidents by MTTD (download)

The relationship between detection latency and incident severity was analyzed by grouping findings according to their MTTD:

  • For incidents detected within the first month, severity is more or less evenly distributed among the low, medium and high categories.
  • However, as the MTTD increases, the severity of incidents shifts towards higher severity. Notably, a high proportion of incidents that took between 30–60 days to be detected are medium-severity incidents (78.57%), while those detected between 60–90 days are predominantly high-severity (71.43%).
  • Among incidents detected after 90 days, a significant proportion are also high-severity incidents (52%).

Overall, 52% of high-severity incidents are only identified after 90 days of going undetected. This represents a concrete risk: the longer an incident goes undetected, the higher the probability of severe compromise. Organizations that integrate continuous detection, threat hunting activities, and regular compromise assessments can reduce MTTD, limit threat escalation, and lower their overall risk profile.

The following case study highlights the importance of timely detection and response to prevent incidents from escalating into high-severity events.

Case study: Four-year-old crypto mining activity on domain controllers

In May 2025, our compromise assessment experts identified three domain controllers on a customer network that were infected with malicious files. The files had remained hidden for almost four years. They were created in the C:\Windows\Fonts\Mysql directory, abusing its unique characteristic whereby only font files in this directory are visible to regular users. Files with the names nei.bat, dl1host.exe, bat.bat, cmd.bat, and a spoofed svchost.exe were found there. These files were created in June and July of 2021.

Kaspersky Threat Intelligence confirmed that these files are part of a crypto-mining campaign called NSABuffMiner, which spreads via the SMB protocol by exploiting the EternalBlue (MS17-010) vulnerability. A patch was released for this vulnerability in March 2017, four years before the initial compromise. This was more than enough time to patch the systems. This underscores the importance of implementing effective patch management operations and staying informed through threat intelligence news feeds.

Based on the organization’s request, the malicious files were collected along with a forensic image for analysis and revealed the following:

  • bat.bat and cmd.bat generate random IPs and scan them with a lightweight port scanner renamed taskhost.exe to locate live hosts with SMB port 445 and NetBIOS port 139 open and looking for vulnerable machines.
  • Discovered vulnerable IPs are handed to helper scripts named bat, poab.bat, load.bat, and loab.bat that execute the malware mance.exe, Eter.exe, and puls.exe to inject the malicious DLLs Eternalblue2.dll and Doublepulsar2.dll into lsass.exe and explorer.exe, enabling lateral movement.
  • Persistence is then established by creating scheduled tasks to execute the propagation and infection scripts, and services are created to execute the crypto miner, with the names MicrosoftMysql, MicrosoftFonts, and MicrosoftMSSql. Other scheduled tasks were also observed with the names At1 and At2 and created for the same purpose.
  • After successfully compromising the machine and installing the persistence mechanisms, a cleanup task is performed to delete temporary files and dropped malware.

Because of the lack of proper monitoring and threat hunting procedures, the organization was unaware that a mining operation had been hijacking their resources for four years, running on their domain controllers.

Unintentional malware preservation

An issue that is frequently discovered during compromise assessment activities is that of web shells remaining or being restored on target systems. Based on data collected during 2025 compromise assessment engagements, 64% of web shell incidents were classified as high-severity findings, 7% as low-severity (possibly legitimate files, but potentially compromised), and 29% as medium-severity findings requiring eradication.

Web shell incident distribution by severity (download)

One way web shells persist is through infected backups. The distribution of discovered incidents in our projects shows that 60% of the web shells were located on active systems, while 40% were stored in backups. Restoring such backups can reintroduce the threat long after the initial infection.

Web shell location (download)

Another common issue is asset inventory gaps, which were observed in 25% of engagements. This resulted in untracked devices, particularly cloud-only Linux web servers that are not joined to Active Directory, evading routine scans.

Asset inventory issues (download)

An attacker can plant a web shell on such a cloud server, and that server never appears in the inventory, though is still regularly backed up. As a result, the web shell may persist on the cloud server for a long time. If it is occasionally deleted, the backup server later restores the infected files, exposing the web shell to third parties again. This demonstrates that without a complete and up-to-date asset inventory, detection capabilities are significantly impaired.

One case was observed in which the web shell was located on an internal file server (not a web server) within a .rar archive at the following path: D:\backup\[redacted_for_privacy].rar/wwwroot/<…>/[redacted_for_privacy].aspx

During the investigation, the server administrators indicated that the folder had been copied from a different server that was offline at the time of the assessment. Because of poor asset inventory, the company’s security team did not detect the infection of this server. As a result of the backup procedure, the web shell was copied to the internal file server. Forensic analysis of the offline server revealed that the adversary had introduced a backdoor to the majority of the Windows servers in the environment, configuring the local administrator account with an identical password.

The technique involved using PsExec to execute a .cmd script across all the servers listed in a .txt file; the script altered the local administrator password to a common value:

Legitimate, yet suspicious: LoLBins and remote management tools

In 2025, nonstandard remote management (RM) utilities were observed in all compromise assessment engagements. Living-off-the-land binaries (LoLBins) were also present in every engagement. These findings highlight the ongoing challenge for security operations centers (SOCs) that must distinguish between legitimate administrative use and malicious abuse.

The observed remote management utilities span both proprietary platforms, such as TeamViewer and AnyDesk, and freely available tools, including PsExec, VNC servers, and open-source RM frameworks. These binaries are used daily in many environments for troubleshooting, software deployment, or remote support. However, the same capabilities – creating a new local admin account, copying files to a remote share, or launching a network port scan for diagnostics – are also typical of attacker post-exploitation activity. Our analysts frequently encounter cases where a legitimate sysadmin action resembles a lateral movement step. This makes the mere fact that “a remote management tool was executed” insufficient to classify it as an incident. Instead, the incident must be judged against an organization-specific baseline of expected usage. Establishing that baseline requires a deep, contextual understanding of who is authorized to run the tool, from which endpoints, and under which circumstances – a resource-intensive process on a case-by-case basis.

LoLBins, binaries that are part of the operating system or commonly installed utilities (such as certutil, bitsadmin, regsvr32, and wmic), were also present in every assessment. While these files are trusted system components, threat intelligence confirms they are often repurposed for lateral movement, data exfiltration, and persistence. The graph below shows the severity distribution for incidents involving riskware or a LoLBin binary. The relatively high share of medium- (40%) and high-severity (31%) findings underscores that misuse of legitimate utilities is often the vector that enables a compromise to progress beyond the initial foothold.

Severity distribution of incidents involving riskware or LoLBin involvement (2025) (download)

To address the potential use of LoLBins and remote management tools by attackers, we recommend a multi-layered approach that goes beyond static deny lists:

  1. Formalize a policy that enumerates the remote management tools authorized for use. The policy must be coupled with a requirement to forward software operational logs to a central log management platform (SIEM or dedicated log collector). Continuous monitoring of these logs enables a SOC to detect deviations from authorized usage patterns.
  2. Periodically perform a software inventory audit to identify unauthorized remote management tools. Consider collecting data from the following registry keys on all hosts:
    • HKLM\Software\Microsoft\Windows\CurrentVersion\Uninstall
    • HKLM\Software\WOW6432Node\Microsoft\Windows\CurrentVersion\Uninstall
    • HKEY_USERS\*\Software\Microsoft\Windows\CurrentVersion\Uninstall
    • HKEY_USERS\*\Software\Wow6432Node\Microsoft\Windows\CurrentVersion\Uninstall
  3. Enrich the hashes (MD5/SHA-256) of every executed binary with a functional category, such as “Remote Access”, “Golden Image”, or “Security Software.” Correlating the category with the execution path makes it possible to hunt for instances where a “Remote Access” binary runs from a non-standard location, such as %TEMP% or a user’s Downloads folder.
  4. Deploy detection rules that capture known LoLBin abuse patterns, such as certutil -decode, bitsadmin -transfer, regsvr32 -i <dll>, wmic process call create. These rules should be continuously baselined against the organization’s normal activity. The baseline is derived from a period of verified legitimate use and refreshed whenever new legitimate use cases emerge. Alerts are generated only when observed behavior diverges from the established norm, thereby reducing noise while preserving sensitivity to genuine abuse.

Impact of not having continuous monitoring and proactive threat hunting

Analyses of recent compromise assessment projects reveal a systematic blind spot in organizations that follow the security-by-purchase model to defend their networks. Without continuous human monitoring or a dedicated threat hunting program, the severity profile of detected incidents becomes heavily skewed toward a higher impact:

Incident severity breakdown, where 24/7 monitoring or threat hunting is absent
Control type Low-severity Medium/high-severity
No continuous monitoring 14% 86%
No threat hunting 16% 84%

Often, the problem is not a lack of tools, but rather a lack of operational use of those tools. Many enterprises deploy next-generation security solutions and then let them run in “set-and-forget” mode, or they rely exclusively on an alert-driven workflow. The following issues are common in such organizations:

  • Alert fatigue: high false positive rates drown analysts in noise, forcing them to triage superficial indicators rather than conduct deep, contextual investigations.
  • Fragmented analyst assignment: without a dedicated hunting team, the same analyst may be tasked with dozens of unrelated alerts, limiting the time available for the hypothesis-driven exploration required to uncover stealthy footholds.

The practical consequence is that adversaries retain an extended dwell time, enabling continued lateral movement and data exfiltration before the organization becomes aware of the breach. This pattern represents a measurable risk exposure that translates directly into business impact. As the following example illustrates, merely purchasing security controls does not guarantee detection; continuous monitoring, regular alert validation, and structured threat hunting are essential to reduce dwell time and limit business impact.

Case study: Secure by design without continuous monitoring

The enterprise invested in security controls and assumed that the environment was secure by design. However, security controls require proper configuration, continuous tuning, and active monitoring to be effective. The tools had been installed, but no one was ensuring that the security controls were configured effectively, there was no analyst reviewing the alerts they produced, and no schedule existed to review the collected logs.

The organization opted for Kaspersky’s Compromise Assessment service. Historical security logs were collected and investigated as part of the assessment procedures. The goal was simple: to determine what had really been going on in the network over the previous few months.

Log analysis revealed clear evidence of malicious activity. Activities related to Impacket behavior were discovered that led to the deployment of Cobalt Strike and Mimikatz on several critical servers, including the domain controllers. These activities were three months old at the time of detection, and the enterprise was unaware of them because there was no effective 24/7 monitoring in place.

Impacket is a collection of Python scripts for network protocols and low-level network packet manipulation. Attackers can abuse it to move laterally into the network. The following are examples of its artifacts detected in the network:

The attacker used Impacket to execute a PowerShell command that downloaded an executable from a command-and-control server. This server was found to be associated with Cobalt Strike. Cobalt Strike is a post-exploitation tool that provides capabilities for remote command execution and lateral movement within a compromised network. The execution was set up via a scheduled task that attempted to masquerade as a legitimate Google Chrome update task.

The timeline assessment confirmed the presence of a Mimikatz binary and a memory dump associated with the same incident on the compromised system, confirming that a credential theft operation had indeed taken place.

The organization was completely unaware of the breach. The activity had gone undetected for three months because the deployed controls were never monitored. Upon learning of the findings, a full-scale incident response was initiated to eradicate the footholds, rotate credentials, and harden the security of the environment.

Security controls are not self-sufficient. Deploying a firewall or an EDR solution does not automatically protect you. Without proper configuration, baseline tuning, and, most critically, continuous log monitoring and threat hunting, those controls become merely decorative. Always-on monitoring, either performed internally or delegated to an external managed security service, can turn weeks-old compromises into minutes-old alerts by correlating events, hunting for anomalous use of penetration testing or hacking tools, and escalating suspicious activity.

Incident response action statistics

An analysis of historical compromise assessment projects reveals a persistent discrepancy between the best practices described in incident response playbooks and the operational realities of executing them in unprepared, often legacy-affected environments. The figure below shows how frequently each response action was required during the initial response phase of a compromise assessment.

Incident response actions required after compromise assessment (download)

The distribution highlights three frequently observed patterns:

  • Forensic analysis accounts for the majority of cases, with around 59% requiring at least one forensic package collection and analysis.
  • Remote eradication, i.e., file or registry key removal, was reported in 39% of cases.
  • Plans evolve as the investigation proceeds; 39% of engagements required a mid-engagement plan update, reflecting the iterative nature of incident response.

Why forensic collection is the default entry point

Forensic package collection and analysis was the most frequent response action, occurring in 59% of cases. The prevalence of forensic package collection can be explained by two observable factors in CA engagements: (1) the targeted organization’s limited historical visibility and (2) the fact that a substantial proportion of incidents were older than 90 days at the start of the assessment. In many cases, native logs had already been rotated or purged, forcing investigators to rely on residual artifacts (e.g., MFT entries, registry hives, filesystem timestamps) to reconstruct timelines.

Our observations suggest that remote forensic package collection is effectively a prerequisite rather than an optional convenience. The graph below summarizes the reported ability to collect forensic packages, categorized by incident severity level. It highlights that, in a significant proportion of high-severity cases, the affected organization lacked this capability.

The organization’s ability to collect forensic data by incident severity (download)

Containment: The remove files/registry keys paradox

Response execution and eradication actions, such as file or registry key removal (reported in 39% of cases), were also common. However, they highlighted a notable gap in execution practices. While many organizations reported having EDR capabilities for remote removal, execution was often delegated to IT teams or MSPs via ticketing systems. This can introduce delays and reduce the precision of the removal process. Malware removal is a surgical process, particularly in multi-stage, fileless, or persistence-heavy scenarios. Capability alone is insufficient without expertise, sequencing, and planning, especially when artifacts may exist in shadow copies, backups, hidden paths, or downloader chains.

Communication failures: An additional operational overhead

A notable organizational finding emerged regarding communication. In 32% of projects, internal communication issues at the assessed organization materially impacted response execution. Below are the typical blockers:

  • Unclear action confirmation – system administrators could not quickly confirm whether a suspicious file was legitimate.
  • Delayed owner validation – ticket escalations stalled while waiting for system owners to respond.
  • Compromised communication channels – email accounts or ticketing portals may already be under the attacker’s control in the event of a suspected domain compromise.
  • Staff turnover – loss of knowledge about historical configuration baselines.

These findings suggest that regular tabletop exercises are required to test not only technical playbooks, but also human and communication workflows, as well as operational level agreements that govern and facilitate communication between different teams, and standard operating procedures for proper documentation.

The iterative nature of response plan updates

The need to update response plans based on new analytical input arose in 39% of cases, emphasizing the inherently iterative nature of incident response. Early-stage plans cannot realistically account for all variables. Examples of the most commonly observed causes for updating the response plan are listed below:

  • Reverse engineering results that reveal previously unknown command-and-control (C2) servers or behaviors.
  • Forensic discoveries, such as hidden scheduled tasks, shadow-copy artifacts, or dormant DLLs.
  • Traffic analysis outcomes that expose additional lateral movement paths.
  • Human constraints – unavailable system owners, changes in management processes, or supervisor approval.

Based on our experience, teams that treat the IR plan as a living document – incorporating each new artifact, reprioritizing actions, and reissuing the playbook before the next containment step – reduce the risk of missed eradication steps. Conversely, strict adherence to an initial, evidence-limited plan can increase the risk of overlooking persistent footholds.

Distinguishing real attacker artifacts from penetration testing leftovers

Finally, distinguishing attacker activity from penetration testing artifacts remained a recurring challenge (12% of cases). Compromise assessments frequently uncover remnants of legitimate testing tools, which can create uncertainty about whether a detected artifact originated from a malicious intrusion or a legitimate penetration test. Contributing factors:

  • Poorly documented penetration test report and artifact cleanup.
  • Overlapping toolsets (e.g., SharpHound) used by both red team operators and adversaries.
  • Running compromise assessments and active penetration testing projects simultaneously, which degrades analyst focus and increases false positive rates. Although correlating findings with penetration testing reports is essential, compromise assessments are human-driven investigative processes, and confusing analysts with overlapping “legitimate” attack signals leads to misinterpretation and weaker outcomes.

Incident response maturity and its effect on severity

Our data show a correlation between the presence of internal digital forensics or malware reverse engineering capabilities and the distribution of incident severity categories. Across the 2025 compromise assessment engagements, the distribution of low-, medium- and high-severity findings differed markedly between organizations that possessed these capabilities and those that did not. The data below illustrate this correlation and provide a basis for assessing the business value of expanding internal response skill sets.

Incident severity split for cases requiring digital forensics, based on an organization’s capabilities (download)

Organizations capable of analyzing digital forensic artifacts independently experienced half as many high-severity incidents and a higher proportion of low- and medium-severity cases.

Incident severity split for cases requiring malware analysis, based on an organization’s capabilities (download)

The presence of a dedicated reverse engineering resource correlates with a total absence of high-severity cases in our sample set; the majority of incidents were rated as medium severity, with a significant proportion of low-severity outcomes.

The analysis of this correlation indicates, with medium confidence, that the observed shifts are unlikely to be caused solely by sample size effects. Rather, they are more likely to reflect a genuine operational phenomenon: internal digital forensics and malware analysis capabilities contribute not only to SOC processes, but also to cyber-resilience in general.

Case study: In-memory LionTail infection on critical Windows servers

During a compromise assessment, a persistent in-memory threat was identified on several critical servers. The activity was attributed to the LionTail framework, a sophisticated set of custom loaders and memory-resident shellcode implants. LionTail takes advantage of undocumented Windows HTTP.sys driver behaviors to covertly deliver and retrieve payloads via inbound HTTP traffic, effectively blending malicious activity into legitimate network flows.

Several observed variants are attributed to the Scarred Manticore actor, which generates a unique implant per compromised host and performs data exfiltration while carefully masking command-and-control communications within normal-looking traffic.

Detection was achieved through static memory signatures discovered within the scrcons.exe process. Although scrcons.exe is a legitimate WMI host binary located under C:\Windows\System32\wbem, it is frequently abused to host injected payloads, making it an attractive target for stealthy in-memory operations.

The response plan comprised a number of actions, the most critical of which are highlighted below:

  • Collection of volatile memory dumps for in-depth analysis.
  • Acquisition of full forensic disk images from affected systems.
  • Detailed analysis of the collected artifacts and subsequent updates to the incident response plan.

Executing these actions proved challenging for the organization because of its limited digital forensics and reverse engineering capabilities. In incidents dominated by fileless memory-resident threats, these capabilities are not optional – they are essential. Without them, organizations risk losing critical evidence, misjudging the scope of the compromise, or failing to fully eradicate advanced implants that leave minimal traces on disk.

While our specialists were able to complete the investigation and contain the breach, the case revealed a readiness gap. It demonstrated the operational risk of depending on external assistance during high‑impact incidents and reinforced the necessity of in‑house forensic and reverse‑engineering maturity to achieve timely, confident and comprehensive incident handling.

Solving the root cause problems

Upon completion of a compromise assessment engagement, the focus shifts from incident response to a consulting phase. The final workshop focuses on preventing recurrence of incidents by identifying underlying deficiencies that allowed them to go unnoticed. The recommendations are actionable and tailored to the environment. For the purpose of this report, they have been grouped into a limited set of high-level categories.

Root-cause category Share of incidents Typical findings
Insufficient detection fidelity 60.7% • No high-confidence alerts were generated by the EPP/EDR or related log sources.
• In 9.4% of cases, the product was mis-configured or out of date or malfunctioning.
Missing alert-driven monitoring 35.9% • Alerts that could have indicated compromise were generated, but an incident was not declared.
• Signals with high uncertainty (e.g., heuristic web shell detections) required analyst validation.
Deficient vulnerability and configuration management 28.2% • Evident misconfigurations (e.g., disabled audit logging, over-permissive service accounts).
• Known vulnerabilities left unpatched or unmitigated.
Lack of structured threat hunting processes 27.4% • Low-fidelity alerts were never reexamined after initial dismissal.
• High-volume telemetry remained unchecked due to staffing constraints.
Inadequate security awareness programs 25.6% • Credential leaks from personal devices of employees or contractors accounted for 27.2% of incidents where inadequate security awareness was identified.
• Social engineering attempts were successful because of insufficient user training.
Absence of documented policies/processes 23.9% • No formal incident response playbooks, change management procedures or data handling guidelines were available.

Common observations on root causes

The detection health check was the most frequent corrective action. In more than half of the cases where alerts were missing, a simple verification of sensor health and rule relevance was recommended to fill the gap. Without such validation, immediate attribution of the failure to the product capability could not be made.
Human analysis is still essential for low-confidence alerts. Automated pipelines alone cannot compensate for rules prone to false positives (e.g., generic web shell heuristics). Embedding a manual triage step was recommended to reduce the dwell time for incidents.

Process hygiene (vulnerability management, threat hunting, security policies) accounts for a substantial proportion of the root causes. Even mature organizations exhibited gaps in routine activities that could be mitigated with disciplined workflows. The absence of documented policies/processes was the root cause of 23.9% of cases.

A modern example of a policy gap is the use of generative AI development tools that operate without clear data handling rules. During one project, we identified a macOS workstation that executed the Claude Code (Anthropic) command-line assistant as a VS Code extension. The tool automatically captured filesystem snapshots to enrich its language model prompts. These snapshots included full directory listings and absolute paths to several Excel workbooks containing internal confidential data:

Parent command line Command line
/bin/zsh -c -l source /Users/[REDACTED]/.claude/shell-snapshots/snapshot-zsh-[REDACTED].sh && eval ‘ls -lh “/Users/[REDACTED]/Documents/[REDACTED]/”*.xlsx‘ \\< /dev/null && pwd -P >| /var/folders/[REDACTED]/claude-[REDACTED] ls -lh /Users/[REDACTED]/Documents/[REDACTED].xlsx /Users/[REDACTED]/Documents/[REDACTED].xlsx /Users/[REDACTED]/Documents/[REDACTED].xlsx .. [REDACTED]

The organization was advised to conduct awareness sessions for employees on the risk of exposing confidential internal data to generative AI tools, and to develop a policy governing the use of such tools with confidential information.

Lack of detections: Causes and impacts

Compromise assessment engagements repeatedly show that insufficient detection fidelity is a significant contributing factor to high-severity incidents. In cases where the target organization’s detection coverage was rated low, 52% of incidents were classified as high severity and 15% as low severity. This suggests a correlation: limited visibility appears to increase the proportion of incidents that evolve into high-severity compromises.

Incident severity distribution when detection coverage was insufficient (download)

A common assumption is that engaging a managed security service provider (MSSP) improves detection maturity. The data, however, show a more nuanced picture. Even when an MSSP is engaged, 26.5% of incidents related to low detection coverage remain unidentified, and roughly 50% of MSSP-supported projects have basic Windows audit gaps (e.g., missing event log collection or disabled audit policies).
These findings suggest that outsourcing alone does not guarantee effective detection; active governance and continuous validation are required. Detection should be treated as an evolving capability that requires continuous testing, measurement, and refinement, irrespective of whether it is managed internally or by a third party.

Statistics of missed incidents due to lack of detection capability with or without MSSP (download)

The analysis of root causes of missed detections reveals several recurring themes. In many environments, the technology is present but poorly operationalized. The main issues are:

  • Absence of endpoint protection platform (EPP) health check – nearly 50% of incidents escalated to high severity in engagements where the EPP health check was weak or absent. This reflects the classic “installed-but-not-enforced” risk, where agents are present but not tuned, updated, or validated.
  • Threat intelligence gaps – when there was no functional threat intelligence feed or platform, about half of the incidents reached high severity. Without curated indicators of compromise and contextual enrichment, analysts rely on generic alerts and may overlook known malicious behaviors.

The underlying issue is an alert-driven, set-and-forget mindset: organizations assume that deployed tools will automatically protect them, even though the tools are not continuously tuned, validated, or enriched with threat intelligence.

Incident severity breakdown where there was no EPP health check or threat intelligence
Missing control High-severity Medium-severity Low-severity
EPP health check 48.3% 36.7% 15%
Threat intelligence feed 50% 40% 10%

Detection failures are rarely caused by a single missing control; they emerge from weak configuration, insufficient telemetry, and an absence of regular checks of controls and processes to ensure they are functional, especially in outsourced models. A hybrid monitoring approach that combines internal ownership with external MDR or MSSP support consistently proves to be the most resilient model when roles, expectations, and performance metrics are clearly defined. Detection must be treated as a living function, not a procurement outcome.

The following example illustrates the real-world consequences of control gaps by walking through a severe incident that persisted undetected for months simply because the organization lacked the necessary detection capabilities and security tools.

Case study: In-memory PurpleFox infection evades conventional endpoint protection

During a compromise assessment engagement, memory was scanned on the target hosts using the threat hunting rule set. Two hidden objects were identified:

PurpleFox drops specially crafted DLLs and forces svchost.exe to load them. From there, it installs a kernel-mode driver that gives the attacker persistent and stealthy execution capabilities, as well as the ability to pull additional payloads. This results in the loading of the XMRig miner.

The deployed EPP solution monitored file creation, registry modifications and network connections. However, its memory inspection module was disabled. Additionally, the signature set applied at the time of the assessment was not up to date. As a result, no alerts were generated for the injected DLLs or the miner’s shellcode. The compromise assessment team identified this detection gap during the memory analysis phase and documented the missing in-memory inspection capability in the final report.

The organization’s security operations were outsourced to an MSSP, which collected the logs and forwarded them to the SIEM solution. Because the logs never contained alerts for in-memory activity, PurpleFox activity was not identified.

Insufficient vulnerability management: A catalyst for high-severity compromises

In the 2025 compromise assessment engagements, more than half of the threats identified and linked to insufficient vulnerability management practices or missing patches were classified as high severity. The most frequently observed consequences were the deployment of web shells that enabled persistent remote code execution and the exploitation of misconfigured Active Directory instances.

Severity distribution of incidents due to improper vulnerability management (download)

The root causes of missing patches are multifaceted. They include inadequate asset inventory management (25% of projects) and the absence of formal vulnerability management processes (41% of projects). Moreover, 86% of organizations that claimed to have a vulnerability management program still exhibited exploited misconfigurations during compromise assessment engagements. These findings suggest that robust patch management, comprehensive asset inventory practices, and structured vulnerability management processes are critical for preventing high-severity incidents.

Case study: How overly permissive GPO-based software distribution goes wrong

During multiple compromise assessment engagements, a high-impact misconfiguration was consistently observed: a Group Policy Object (GPO) was used to point to an executable in a shared folder and run it on every workstation via a scheduled task. The access control list (ACL) on the share was set to “Everyone – Full Control”.

Given that any authenticated domain user can write to the share, an attacker who compromises a single low-privilege account can replace the legitimate binary with a malicious payload. The next scheduled task run propagates the payload automatically to all endpoints that receive the GPO. This provides:

  • Elevated execution context: the scheduled task typically runs under the SYSTEM or local administrator account.
  • Automatic lateral movement: the malicious binary propagates without requiring additional network exploitation.
  • Privilege escalation: a compromised low-privilege account can lead to domain administrator code execution.

Vulnerability management procedures that include systematic GPO and share permission audits would have flagged the writeable ACL as a high-severity finding, enabling remediation before exploitation. Remediation typically involves restricting the share permissions to “Authenticated Users” with read-only access and limiting modifications to certain privileged accounts. Incorporating these checks into the baseline security controls reduces the attack surface, demonstrating the tangible risk reduction achievable through disciplined vulnerability assessment and penetration testing (VAPT) practices.

Conclusion

In 2025, Kaspersky Compromise Assessment helped organizations reveal a persistent detection gap: 30.8% of all incidents and 52% of high-severity compromises had historical activity spanning over three months. Of all the incidents discovered, 20% were found manually, while 60% were missed by enterprises because of the absence of high-confidence alerts from existing tools. The oldest missed incident identified by the Kaspersky Compromise Assessment team in 2025 was four years old.

Post-incident checkups produced the highest percentage of high-severity findings, while regular proactive audits, compliance-driven audits, and audits performed before merging two networks tended to reveal issues earlier. This indicates that purely reactive investigations often miss hidden persistence. The top high-level recommendations for immediate improvement in 2025 for all projects were:

  • Run a comprehensive detection engine health check within 30 days of project closure, prioritizing telemetry integrity and rule relevance.
  • Introduce a Tier 1 alert validation team that reviews all low-confidence events on a defined schedule.
  • Ensure robust 24/7 monitoring augmented with threat hunting capabilities focused on baselining, low-fidelity alerts, and emerging adversary techniques.
  • Reevaluate the vulnerability management pipeline to ensure continuous patching and audit log activation across all critical assets.
  • Update security awareness curricula to address credential leakage from personal devices and reinforce secure BYOD practices.
  • Ensure periodic tabletop exercises are run to test technical playbooks and sharpen the team’s skills and communication workflows.
  • Establish operational-level agreements to govern and facilitate communication between different teams and standard operating procedures used for proper documentation.

Addressing the root cause categories systematically will reduce the likelihood of future blind spots and improve the overall security posture of the engaged organizations.

The Autonomous SOC, Revisited: What 18 Months on the Road Has Taught Us

This post revisits SentinelOne’s Autonomous SOC maturity model, first introduced in “Autonomous SOC Is a Journey, Not a Destination” (December 2024).

When SentinelOne® introduced the Autonomous SOC maturity model, we made a deliberate choice: describe a journey, not promise a destination.

The industry had no shortage of vendors declaring that AI would transform security operations. We thought the more useful contribution was a framework for understanding what that transformation looked like, at what pace it was realistic, and what conditions each stage of progress required.

Security teams found the model useful. Not as a marketing claim, but as a map. CISOs and SOC leaders started placing their organizations on it, asking what it would take to move forward.

What happened next was telling. By RSAC 2026, ‘autonomous SOC’ appeared in vendor keynotes and product launches from companies that hadn’t used the term twelve months earlier. Add in pseudonyms like Agentic SOC and AI SOC, and the list explodes. Fast adoption brings loose definitions. For us, it’s worth being precise about what the concept means and what it doesn’t.

Here is what SentinelOne has learned from 18 months of real-world Autonomous SOC deployments.

What Held Up

Today, the progression still maps accurately to where organizations are and what separates each stage from the next. That accuracy holds even for a framework built before most organizations had meaningful AI deployment experience. The inflection points reflect real operational transitions at each maturity step.

The “journey not destination” framing has proven more important than we anticipated when we wrote it. In early 2026, Gartner published guidance to help buyers evaluate AI SOC claims more critically, noting that vendor credibility in this space depends on honest representation of where the technology is:

“Some vendors exaggerate capabilities (like being able to deliver a fully autonomous SOC), risking buyer trust and harming the reputation of legitimate solutions.”1

A maturity model is structurally honest. It reflects where you are, not where a vendor wishes you were. Gartner’s research found that while 40% of organizations are actively evaluating AI SOC capabilities, only 18% have actually deployed2. The gap between evaluating and deploying is rarely about technology. Most organizations cannot advance because they lack a clear view of where they stand or what the next stage requires.

When security leaders use the model as a reference point, the evaluation conversation changes. The question shifts from “does your product make my SOC autonomous?” to “what would it realistically take to advance, given where we are today?” A feature list cannot answer that question. An honest vendor can.

Watch our webinar on why most AI SOC deployments stall here.

What We Underestimated

The levels were always sound. What we underestimated was how much organizations needed to build before they could operationalize them. Customers understood where they wanted to go. But achieving Partial Autonomy (Level 3) requires a data foundation, a workflow architecture, and AI readiness that most teams were still building when we first published this model. That’s a fact about where most security organizations were in 2024.

The transition from AI-Assisted Operations (Level 2) to Partial Autonomy (Level 3) is primarily a governance problem, not a tooling one. The tools are capable. What most organizations are missing is an understanding of the foundation of data and trust that Partial Autonomy (Level 3) requires, including the role humans play in building it.

When analysts work with AI assistance, they leave traces. Which queries they accept. Which results they act on. Which steps they modify or override. Over time, the system learns which investigation patterns the team trusts, which AI recommendations get acted on, and where analyst expertise is required – the kind of institutional knowledge that only comes from doing the work. Partial Autonomy is built on that record, not installed on top of an existing stack.

The path from AI-Assisted Operations to Partial Autonomy starts earlier than most organizations realize. It begins before they’re thinking about autonomy at all. Every assisted workflow is building toward what comes next.

What Holds Organizations Back

The primary barrier between AI-Assisted Operations (Level 2) and Partial Autonomy (Level 3) is accountability.

Consider how the automotive industry defined its equivalent of Partial Autonomy – SAE Level 3.

Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles3

The designation applies only within specific, defined operational conditions. Outside those conditions, the human must take control. What qualifies a system for L3 is defined before autonomous operation begins: explicit parameters, a defined scope, and clear conditions for human override. Governance precedes autonomy.

Consider Waymo. It is the most capable autonomous system deployed at scale today — L3+ — operating without a safety driver under defined conditions. The vehicle is remarkable. But Waymo’s primary innovation is the organization built around it: the cloud infrastructure that keeps cars in autonomous condition, the human operations that handle exceptions the system cannot cover. The more autonomous the system, the more organizational maturity it required to build. High autonomy is an organizational capability.

The same logic applies in security operations. Accurate AI is the foundation. What makes Partial Autonomy legitimate is what gets built on top of it: defined rules of engagement, pre-approved policies, audit trails, and a clear organizational answer to who is responsible when an AI verdict is acted on. That accountability sits with the security team. When automation fires, it fires because someone made a deliberate governance decision to allow it. That is what makes it auditable, defensible, and durable.

Gartner’s readiness criteria for AI SOC deployments require that operational workflows be established in playbooks before AI is introduced4. In practice, the organizations that advanced most consistently treated that requirement as a sequencing discipline, not a box to check. They defined their rules of engagement before turning on automated response.

The second learning was the attacker asymmetry. Defenders who stall between AI-Assisted Operations and Partial Autonomy have often done the validation work. The AI logic checks out. What remains is the decision to extend that trust to autonomous action — and that decision takes time. Attackers move differently. They deploy, observe what works, and iterate. Governance is an externality. Trial and error with no consequences for failure is a significant operational advantage. The gap between a defender’s trust-building timeline and an attacker’s operational tempo is structural. It compounds.

Why High Autonomy Stays on The Horizon. And Why That Matters Less Than We Thought

Eighteen months of deployment have also changed how we think about the upper end of the model.

When the original post was written, High Autonomy (Level 4) was described as dependent on a level of AI reasoning we hadn’t yet seen in production security environments. That remains the right framing. What’s changed is how close that horizon has become. Two years ago, asking a model to reason through a multi-stage attack, correlate signals across data sources, and produce an auditable verdict required significant scaffolding and produced inconsistent results. That’s no longer true. The gap between where AI was and where High Autonomy requires it to be has narrowed substantially.

High Autonomy still requires more than capable models. Institutional trust takes time to build. Accountability structures have to go beyond controlled tests to survive real incidents. Human oversight has to be redefined from reviewing individual actions to governing a system’s behavior within a defined scope. Those are organizational problems, technology doesn’t solve them. That work is already underway at Partial Autonomy (Level 3). What the road to High Autonomy requires is only visible from Partial Autonomy. Organizations that haven’t operated there yet are planning for a destination they haven’t seen. The knowledge of what it takes is path-dependent, and it emerges from operation, not from design.

As organizations move deeper into Partial Autonomy, the distinction between levels matters less in practice. What security leaders actually want is relevant control: governance over the decisions that matter, without being burdened by the ones that don’t. You cannot be responsible or accountable for a system that asks you to review everything.

Control over the right decisions is what matters. An analyst reviewing every alert has maximum control and minimum leverage. A system that acts autonomously on well-understood threat patterns, surfaces only the ambiguous and novel cases for human judgment, and maintains a complete audit trail, gives the analyst control over exactly what deserves their attention. That is a better and more focused version of human oversight.

High Autonomy, seen through this lens, is AI that has earned sufficient trust within a defined scope. The remaining human decisions are the ones that require human judgment, because the governance architecture evolved to allocate human attention correctly.

In the same way, a pilot does not manually adjust every control surface for the duration of a flight. They set the destination, define the parameters, and monitor the instruments. The system handles thousands of micro-corrections that would be impossible to manage directly. The pilot’s job is to govern the conditions under which the aircraft flies itself, not to manage every control input directly. Nobody describes this as a lack of pilot control. It is a better allocation of pilot judgment. And it works because of the environment surrounding the autopilot: pilot training standards, airline operational doctrine, air traffic control, and regulatory frameworks. The technology is one layer of a much larger system.

The governance work done at Partial Autonomy is the same work that produces High Autonomy. Organizations investing in it now are not waiting for a future capability release. They are building the foundation on which High Autonomy operates.

What This Means for the Road Ahead

The first step toward Partial Autonomy is a policy decision. Define the conditions under which your organization will allow a system to act: which response actions, against which threat types, within what scope, under whose authority. Write it down, however rough. That document is the actual starting point. Without it, the tooling is irrelevant.

The work at Partial Autonomy is real, meaningful, and available now. Security teams that define accountability structures before deploying autonomous systems, build a record of AI efficacy in their specific environment, and treat governance as a prerequisite rather than an afterthought, are the ones that reach and sustain Partial Autonomy. They are also the ones best positioned for what comes next. That work produces a more integrated SOC — data, AI, and response operating as a unified system.

High Autonomy remains the north star. This clearly articulated ideal state stops organizations from settling too early. It is the same function that “zero trust” serves as an architectural principle: no organization fully achieves it. Every organization is better for pursuing it.

The tools are capable. The frontier models have advanced significantly since our maturity model was first introduced. The capability gap that once made waiting feel reasonable has narrowed. What remains is the institutional work. That work is always harder than buying a product, which is why vendors who are honest about it are worth paying attention to.

SentinelOne customers operating the Autonomous SOC are seeing it in their numbers: 75% faster investigations, 4x more threats handled, 42% fewer false positives5. Read the IDC Business Value Snapshot.

References

1 Gartner, “AI SOC Agents: Harnessing Innovation, Managing Expectations,” Kevin Schmidt, Alex Tytarenko, Steve Santos, 25 February 2026. G00841784.

2 Gartner, “AI SOC Agents: Harnessing Innovation, Managing Expectations,” Kevin Schmidt, Alex Tytarenko, Steve Santos, 25 February 2026. G00841784.

3 SAE International, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” SAE Standard J3016_202104, April 2021. https://www.sae.org/standards/content/j3016_202104/

4 Gartner, “AI SOC Agents: Harnessing Innovation, Managing Expectations,” Kevin Schmidt, Alex Tytarenko, Steve Santos, 25 February 2026. G00841784.

5 IDC Business Value Snapshot, “The Business Value of SentinelOne Singularity AI SIEM,” Michelle Abraham and Matthew Marden, May 2026, sponsored by SentinelOne. #US54435826-BVS.

Third-Party Trademark Disclaimer 

All third-party product names, logos, and brands mentioned in this publication are the property of their respective owners and are for identification purposes only. Use of these names, logos, and brands does not imply affiliation, endorsement, sponsorship, or association with the third party.

❌