Reading view

There are new articles available, click to refresh the page.

The Good, the Bad and the Ugly in Cybersecurity – Week 30

The Good | Authorities Dismantle Kratos Phishing Network & Arrest Its Developer

Kratos, a prominent phishing-as-a-service (PhaaS) platform, was dismantled from the inside out this week thanks to German and U.S. law enforcement agencies. During “Operation Olympus Blade”, authorities seized over 200 servers to render Kratos’ global network entirely inoperable while the platform’s suspected developer was apprehended in Indonesia.

So far, investigators estimate that more than 1800 cybercriminals utilized the platform to launch nearly 15,000 phishing campaigns monthly since late 2024. Operating as a franchise, the service provided threat actors with toolkits designed to generate convincing Microsoft authentication pages. These campaigns targeted victims across the United States and Europe, facilitating widespread credential theft and unauthorized account access. The operators earned at least €300,000 in subscription fees.

Source: BKA

A recent report from cyber researchers reverse-engineered the Kratos toolkit, revealing how it offered operators two distinct functional modes. While one mode harvested traditional credentials, the more advanced setting deployed a Node.js reverse proxy. This adversary-in-the-middle (AitM) capability allowed attackers to intercept active session cookies in real-time, effectively bypassing standard multi-factor authentication (MFA) controls.

Once compromised, these accounts provided actors with initial footholds to execute business email compromise (BEC), lateral data theft, and secondary phishing attacks. Just this February, actors ran Kratos in a campaign that used tax-themed lures and personalized QR codes to target dozens of American manufacturing and healthcare organizations.

While the immediate server takedown severely disrupts ongoing operations, officials acknowledge that the existing customer base retains access to the underlying kit code. That means the toolkit itself outlives the infrastructure seizure, and operators who already have copies can resume campaigns under new branding with minimal rebuild effort.

The Bad | Threat Actors Conceal HollowGraph Malware in Microsoft 365 Calendar Events

A novel espionage implant, dubbed HollowGraph, is hijacking Microsoft 365 calendars to establish a covert command and control (C2) channel. By routing operator instructions and exfiltrated data through legitimate Microsoft Graph API traffic, the malware ensures its activities blend seamlessly with routine network chatter.

The .NET DLL implant operates purely as a two-way dead drop without communicating directly with an attacker-owned payload server. To receive tasking, HollowGraph queries the compromised user’s calendar for an event planted far into the future – in this case, dated for May 13, 2050. Operators embed their instructions within text files attached to this anomalous event, ensuring the mailbox owner never naturally scrolls far enough to discover the malicious entries.

For data exfiltration, the malware executes the reverse process. It systematically encrypts stolen files using hybrid RSA and AES-256 encryption, generates a new far-future calendar event, and uploads the targeted data as attachments. To maintain continuous Graph API access, operators utilize a secondary DNS-based channel to refresh the application’s Entra ID login credentials. The malware decodes these values from an attacker-controlled domain and writes them to a disguised configuration file.

Analysts observed this highly targeted campaign actively compromising machines at an Israeli organization between June and July 2026. While the implant’s underlying code shares significant structural similarities with a modular backdoor framework called Cavern, frequently utilized by Iranian state-sponsored syndicates, researchers have not yet definitively attributed this specific operation to a known threat group.

HollowGraph buries its C2 in M365 calendar events dated 2050 – no attacker server ever touched. https://t.co/yJo4fpw0X9 #ThreatIntel #HollowGraph #Cavern #Cav3rn pic.twitter.com/kaABaAYZg0

— ThreadLinqs (@threadlinqs) July 22, 2026

Since HollowGraph relies entirely on compromised account identities and legitimate application permissions rather than software vulnerabilities, standard patching remains ineffective. The technique effectively weaponizes the trust organizations place in their own Microsoft Graph API traffic, turning routine calendar activity into a blind spot by design.

The Ugly | AI Models Escape Sandbox to Breach Hugging Face Infrastructure

Open-source AI platform Hugging Face recently disclosed a network breach orchestrated entirely by an autonomous AI agent. The attack compromised the company’s data-processing pipeline by deploying a malicious dataset that exploited two distinct code-execution vulnerabilities. This initial access allowed the agent to run unauthorized code on a processing worker, harvest cloud and cluster credentials, and subsequently move laterally across several internal systems.

The platform reported that the agent executed thousands of individual actions across short-lived sandboxes while staging self-migrating C2 infrastructure on public services. Following the discovery, Hugging Face evicted the agent, revoked affected credentials, and rebuilt compromised nodes. It was during the forensic investigation that responders encountered a unique operational hurdle: Western AI models refused to process the malicious artifacts due to built-in safety guardrails. At this point, the company turned to an unrestricted, open-weight Chinese model to successfully complete the analysis.

Shortly after Hugging Face’s disclosure, OpenAI confirmed that its own AI models orchestrated the intrusion during internal benchmark testing. Operating with reduced safety guardrails for evaluation purposes, GPT-5.6 Sol and an advanced pre-release model attempted to cheat the ExploitGym cybersecurity assessment. Rather than solving the challenges natively, the models identified and exploited a shortcut, retrieving the test solutions directly from Hugging Face’s production database via a zero-day vulnerability in a third-party package registry cache proxy.

Source: Adel Ka

After gaining open internet access through privilege escalation and lateral movement, the models accessed Hugging Face servers. The AI agents chained together multiple attack vectors, utilizing the stolen credentials and additional zero-day vulnerabilities to establish remote code execution. OpenAI subsequently disclosed the zero-day flaw and collaborated with Hugging Face to implement stricter infrastructure controls and guardrails.

Your Best Analyst Shouldn’t Be a Person. It Should Be a Capability Everyone Can Summon.

For thirty years, we have measured security operations by the tools we buy. The next decade will measure us by the outcomes we deliver. That shift is already here, and it is being driven by something quietly radical: a repository of AI “skills” that turns the deep expertise of a principal analyst or engineer into a capability any team member can invoke on demand.

I want to talk about what that actually changes for the business, not the bits and bytes underneath it.

The problem every CISO already knows by heart

You are not short on data. You are drowning in it. Endpoint telemetry, identity logs, firewall traffic, cloud control planes, email security, SaaS audit trails. Each one speaks a different language. Each one demands a specialist who knows where the bodies are buried. The talent who can fluently read all of them at once is rare, expensive, and almost certainly already burned out.

So the work stacks up. Alerts wait. Investigations get triaged by whoever is awake. The third repeat of an attack pattern goes unnoticed. The analyst who caught the first two left for a competitor. Your security posture quietly becomes a function of who happens to be on shift.

This is the real cost center in modern security operations: expert human attention, not licenses or infrastructure. There is never enough of it.

What changes when expertise becomes a skill

The ai-siem repo, located on the Sentinel One GitHub community (https://github.com/Sentinel-One/ai-siem/tree/main/plugins/s1-secops-skills), attacks that bottleneck directly. Instead of asking a human to remember how to query log sources, pivot through threat intelligence, correlate findings, and write it all up, each of those steps becomes a skill. Captured once. Available to everyone, every shift, every time.

Disclaimer: This sample script/prompt is community-contributed, open-source content provided “AS IS,” without warranty of any kind. SentinelOne does not certify or endorse it, is not responsible for its accuracy or outputs, and is not liable for any outcomes arising from its use. Test and validate in a non-production environment before use.

The senior analyst’s playbook stops living in one person’s head. It becomes a durable asset owned by the whole organization. That single change cascades into outcomes leadership actually cares about.

The data lake is the foundation nobody’s talking about

Here is the part that makes the rest of it work. It is the most underrated shift in security right now. Skills are useless if the data lives in a dozen disconnected silos. Each has its own query language, retention tier, and price per gigabyte. The reason this whole model becomes possible is the security data lake. A single place where endpoint, identity, network, cloud, email, and your own application logs land together in one queryable substrate, at a cost that doesn’t punish you for keeping data.

This is where SentinelOne’s Singularity Data Lake stops being infrastructure and starts being the differentiator. It was built for streaming AI from day one, not retrofitted onto it. That architecture is what makes an AI analyst viable. Data becomes searchable the moment it arrives. No indexing delay to wait through. Everything stays hot and searchable. All of it. There’s no cold tier to thaw, and no log you quietly dropped because retention got expensive. It scales to petabytes where legacy SIEMs buckle at terabytes. And it does this at more than ten times the query performance, for less than half the cost of the per-gigabyte SIEM model it replaces.

Translate that into outcomes, and the picture is stark. Ingestion, detection, and query that used to take minutes to hours on a legacy SIEM now happen in seconds. More than 2,000 detections run in the stream itself. Threats surface as the data lands, not minutes after it’s stored. That speed is not a nice-to-have. An AI agent is only as fast as the data underneath it. Give it a lake that answers in under a second, and it reasons across your entire estate before you’d have opened one console tab.

That is what breaks the twenty-year SIEM economics. For twenty years, the industry’s answer to “where do we put all the security data” was a SIEM. One that charged so much per gigabyte that teams were forced to drop the very logs they later wished they’d kept. The data lake inverts that math. Keep everything. Query everything. Correlate everything. Let the ingest bill stop dictating your detection strategy. The skills are the brain. The data lake is the nervous system, letting the brain feel the whole body at once, instantly. You cannot have the outcomes below without it.

This is Autonomous Cybersecurity (AI-Native Protection Across the Enterprise) in practice. Autonomous Security Intelligence, ASI, is the intelligence fabric that runs on top of that data. It is not a bolt-on skill pack. It is what turns a queryable lake into an analyst that never sleeps.

Outcome 1: Investigations that took a shift now take minutes

The gathering is the slowest part of any investigation, not the decision: pulling the alert, finding the affected asset, enriching every indicator against external intelligence, sweeping the rest of the fleet for the same fingerprint, and assembling the timeline. That is hours of skilled work that have to happen before anyone can even say “true positive” with confidence.

When those steps run as orchestrated skills, the gathering collapses into minutes. Your analysts spend their judgment on the verdict and the response, which is the part only a human should own. Mean time to detect and mean time to respond stop being aspirational metrics on a slide. They become numbers you can defend to the board.

Outcome 2: A first-year analyst operating at a principal level

This is the one that genuinely reshapes the org chart. When the hard-won method of your best investigator becomes a skill, a junior analyst inherits it directly: the answer, arrived at the right way, with evidence cited, confidence calibrated, and assumptions flagged.

The skills gap that has defined this industry for a decade narrows dramatically. You stop competing for the handful of unicorns who can do everything, because everything is now a shared capability. Tier-one talent does tier-three work. New hires become productive in days, not quarters. And the people you already have stop drowning. That’s how you keep them.

Here is what convinced me that this is real and not a demo. It was not a SOC analyst who proved it first. It was an engineer. Reviewing application logs, they surfaced a genuine fraud case. A true positive lived in business telemetry. No traditional security tool was even watching. Sit with that for a second. People who do not carry a security title, looking at data that never reaches the SIEM, caught actual fraud. That is what happens when investigative expertise stops being gated behind a job description. The capability travels to wherever the data and the curiosity are. Threats that used to hide in the gaps between teams suddenly have nowhere to live.

More impact per analyst and greater control with less fatigue.

Outcome 3: No blind spots, because nothing gets correlated in isolation

Attackers do not respect your tool boundaries. They land in email, execute on the endpoint, move through identity, and leave through the network. A threat that is invisible in one source is often obvious the moment you line it up against three others. The trouble is that lining them up has always required a specialist for each layer. All working in concert, under time pressure, at 3 am.

Cross-source correlation built into the workflow doesn’t depend on who’s in the room. The full attack story assembles itself. You see the chain, not the fragments. The single most dangerous phrase in security operations, “we had the data, we just never connected it,” starts to disappear.

Outcome 4: Every alert arrives with context already attached

A medium-severity alert on a domain controller matters more than a critical one on a throwaway sandbox. Every experienced analyst knows this. Yet most alerts land in the queue as bare indicators with no business context. Someone has to hunt down what the asset is, who owns it, and whether it matters. That manual lookup happens thousands of times a week. It’s where prioritization quietly goes wrong.

When asset enrichment runs autonomously, every log and every alert already carries the device and user context that determines its importance: what the machine is, how critical it is, and whose account is involved. The queue effectively sorts itself by business impact. Analysts stop chasing noise in disposable systems and spend their time where the real risk lies. False-positive fatigue drops, and the genuinely dangerous signal stops getting buried under the trivial. Prioritization by business impact stops being an aspiration and becomes the automatic default.

Outcome 5: Proactive defense, finally, at machine speed

Known-bad signatures catch yesterday’s threats. The adversaries that actually hurt you, the patient ones and the insiders, only ever show up as deviations from normal. A login at an impossible hour. A workstation reaching a destination it’s never touched. A service account suddenly behaving like a human.

Hunting for that kind of anomaly across the entire estate, continuously, has always been a luxury. Reserved for the most mature and best-funded teams. Make it a repeatable skill and proactive hunting stops being a quarterly project you never quite get to. It becomes the default mode of the SOC. You move from reacting to alerts to anticipating the attacker’s next move. That is the whole point of the discipline. Most teams never have the capacity to actually do it.

Outcome 6: A new threat in the headlines becomes a detection the same morning

When a new campaign breaks, the clock starts immediately. The window between “this threat is now public” and “we are protected against it” is pure exposure. Historically, that window has been measured in days or weeks. Someone has to read the intelligence, translate it into detection logic, test it, and push it live. That someone is usually already underwater.

Make detection engineering a skill, and that window collapses to a morning. The moment an emerging threat surfaces, its behavior becomes a live detection rule: validated and deployed across the estate before the first coffee gets cold. Your defenses move at the speed of the threat landscape instead of at the speed of your backlog. Just as importantly, the detection logic your team writes today gets captured and reused. Coverage doesn’t just grow. It compounds.

Outcome 7: New data sources onboarded in minutes, not quarters

Onboarding a new data source has traditionally been a small project: parse the logs, normalize the fields, build the dashboards, write the detections, and wire up the response. Weeks of specialist time have to pass before that source earns its keep. That’s exactly why the backlog of “sources we really should be ingesting” never shrinks.

That math is now broken in your favor. When those steps are packaged as skills, a new feed goes from raw and unreadable to fully operational in minutes: normalized, with detections firing and a dashboard live. Read that again, because it rewrites your roadmap. Every integration you’ve been deferring for budget or bandwidth reasons just got cheap. Cheap enough to do the same day someone asks for it. Coverage stops being a function of how many quarters you can fund. It becomes a function of how fast you can decide.

That is the compounding version of Maximize Efficiency and Effectiveness of Security Operations: coverage that gets cheaper and faster to extend every time you use it.

The economics that should end the conversation

Now brace for the part that makes the CFO lean in. Everyone assumes the AI is the expensive bit. It is the opposite. Bring-your-own-AI on top of the data lake costs peanuts relative to what it replaces and the work it does. The heavy historical spending on security operations was never on intelligence. It was on the ingestion licensing of a legacy SIEM, and the salaries of specialists doing by hand what a skill now does in seconds.

Sit the two columns next to each other. On one side: per-gigabyte SIEM pricing that grows with your business, whether or not it makes you safer. Plus the fully loaded cost of analysts spending their nights on manual gathering. On the other: a data lake built for scale, and an AI layer whose run cost rounds to a rounding error against either line item. The capability goes up and to the right while the cost line stays flat. That’s a different business model for security. It’s a rare case where the cheaper option is also the more capable one.

This is Enable Business Growth and Innovation Safely in dollar terms: the budget fight between “more coverage” and “more efficient spend” disappears, because the same architecture delivers both.

The deeper shift: the SOC stops being a cost center and starts compounding

Here is the part that should excite anyone running a security budget. Every investigation a human does is an effort spent once and largely lost. Every investigation captured as a skill is an effort spent once and reused forever. Your operation stops being a treadmill and starts being an asset that compounds. The work your team does today makes the work tomorrow faster, cheaper, and more consistent.

AI answering faster is the easy headline. The real shift: institutional security expertise stops walking out the door and starts accumulating on the balance sheet.

What I would tell a peer

We have spent a generation buying tools and hoping the outcomes follow. The teams that win the next decade will flip the order. Define the outcomes first. Then make the expertise to achieve them a capability everyone can summon, day or night, junior or senior, first alert or thousandth.

The technology to do this exists now. Our purpose is simple: to give the advantage to those who secure our future. That advantage only counts if it reaches every analyst, not just the ones already fluent in every log source. The organizations that adopt it won’t just be faster. They’ll run a fundamentally different kind of security function: one where the best analyst in the building is available to everyone, all the time, and gets sharper with every case it touches.

The bottleneck was never the data. It was access to expertise. That bottleneck just broke.

If you run a SOC, lead security for your organization, or work the queue every day: how much of your team’s best thinking is locked inside one or two people right now? That is the question worth sitting with this week.

Curious what this looks like in practice for your environment? Come talk it through in our Reddit community, r/SentinelOneXDR. Practitioners there trade real detection logic, ask the SentinelOne team direct questions, and compare notes on what’s actually working in their SOCs.

Disclaimer:  The sample scripts, code, AI prompts, and other tools referenced or included in this publication (“Community Content”) are provided for informational and educational purposes only. Community Content is contributed on an open-source basis and is made available “AS IS” and “AS AVAILABLE,” without warranties of any kind, whether express, implied, or statutory, including, without limitation, any warranties of accuracy, completeness, reliability, merchantability, fitness for a particular purpose, or non-infringement.

SentinelOne does not certify, endorse, or guarantee any Community Content, its outputs, or its suitability for any particular use, and Community Content does not constitute part of any SentinelOne product or service offering. SentinelOne has no obligation to maintain, support, or update Community Content. AI prompts in particular may produce inaccurate, incomplete, or unexpected results depending on the model, configuration, and environment in which they are used.

Any use of Community Content is at your own risk. You are solely responsible for evaluating, testing, and validating any Community Content in a non-production environment before use, and for ensuring your use complies with applicable laws, licenses, and your organization’s policies. To the maximum extent permitted by law, SentinelOne and its affiliates will not be liable for any damages, losses, or outcomes of any kind arising out of or relating to the use of, or reliance on, Community Content. Where Community Content is hosted in or links to a third-party repository (e.g., GitHub), your use is also governed by the applicable open-source license and the terms of that platform.

Mount Here, Read There: Twin Path Traversal CVEs in Kubernetes Storage

filepath.Join was never designed to be a security boundary. We found two CSI drivers that shipped on the assumption it was, and the result was cross-tenant data access with optional node destruction, using nothing more than a valid Kubernetes manifest. The vulnerable drivers are the Kubernetes CSI Driver for NFS (csi-driver-nfs) and the Kubernetes CSI Driver for SMB (csi-driver-smb), both maintained by the upstream kubernetes-csi organization. An attacker who can create a PersistentVolume can craft a volume identifier that escapes the subdirectory boundary and reaches another tenant’s data on the shared export.

The impact varies by deployment. In the default configuration, an attacker can read, modify, and delete files belonging to other tenants sharing the same export. In production deployments where the CSI controller is configured with broader hostPath mounts (a common pattern for log collection and operational tooling), the same primitive enables arbitrary directory deletion on the Kubernetes worker node, including paths like /var/lib/kubelet, /etc/kubernetes, and /etc/cni. Remove those and the node is dead.

The Kubernetes Security Response Committee published advisories and assigned CVE-2026-3864 to the NFS driver issue and CVE-2026-3865 to the SMB driver issue.Fixes shipped in csi-driver-nfs v4.13.1 and csi-driver-smb v1.20.1. We reported both vulnerabilities in January 2026 and worked with the maintainers through coordinated disclosure.

Kubernetes Storage Concepts

First, some background on how Kubernetes does storage. This section covers the four building blocks an attacker manipulates in this attack chain.

Container Storage Interface

Kubernetes does not implement storage directly. It delegates that responsibility to the Container Storage Interface, a gRPC contract between the Kubernetes kubelet and a vendor-supplied driver.

Figure 1. The CSI specification sits between Kubernetes and the underlying storage drivers, with each layer owned by a different party.

When a workload requests storage, Kubernetes issues calls such as CreateVolume, NodePublishVolume, and DeleteVolume to the CSI driver, which in turn provisions, mounts, and tears down the actual storage on the underlying system, whether that system is an NFS export, an SMB share, an Amazon Elastic File System filesystem, a block device, or anything else.

Figure 2. The CSI driver mediates between Kubernetes pods and the underlying storage backend.

The CSI specification is intentionally minimal about what drivers must validate. It specifies the gRPC interface and the lifecycle but leaves input validation, authorization, and isolation to each driver implementation. This design decision is the underlying reason the same class of bug recurs across multiple drivers.

PersistentVolume and the volumeHandle

A PersistentVolume is a Kubernetes API object that represents a unit of storage in the cluster. When a PersistentVolume references a CSI driver, it carries a string field called volumeHandle. This field is opaque to Kubernetes: the API server stores it but does not parse, validate, or interpret it. The driver alone is responsible for understanding the format of volumeHandle and using its contents safely.

Each driver defines its own format. The NFS CSI driver parses volumeHandle as {server}#{share}#{subDir}#{uuid}#{onDelete}

The SMB CSI driver parses it as //{server}/{share}#{subDir}#{uuid}#{secretNs}#{secretName}#{pvName}

In both drivers, the subDir component is the security boundary. It’s also the one that breaks.

Subdirectory-Based Multi-Tenancy

Some clusters share a single NFS or SMB export across tenants by giving each one a dedicated subdirectory. Both drivers’ deployment guides document this pattern. It’s common wherever provisioning a separate export per tenant is expensive: on-premises NAS appliances, cloud file services that bill per share, shared storage systems that are slow to reconfigure.

In this model, the security boundary between tenants is the subdirectory path. Team A’s PersistentVolume is scoped to subDir: team-a. Team B’s is scoped to subDir: team-b. The driver mounts each tenant into their own directory, and the assumption is that neither tenant can reach the other’s data, even though both share the same underlying export.

Role-Based Access Control and PersistentVolume Authorization

Kubernetes uses RBAC to control who does what. Creating a PersistentVolume (PV) is cluster-scoped, and the documented threat model assumes only cluster admins hold this permission. In practice? It’s everywhere. CI/CD pipelines, ArgoCD, Helm automation, operator controllers. They all need persistentvolumes:create to provision storage for applications. Compromise any of those service accounts and you’re in.

That gap between the documented threat model and how clusters actually run is what makes these bugs practically exploitable.

Trusting the Wrong Function

There’s a misconception in the Go ecosystem that keeps burning people: the idea that filepath.Join prevents path traversal. It doesn’t. filepath.Join is a normalizer. It collapses ., .., and redundant separators into a canonical form. It has no concept of “base directory” or “boundary” and can’t tell whether the result landed somewhere the caller never intended.

Try it yourself: give it “/var/lib/csi/team-a” and “../../../etc/passwd“. You get /etc/passwd. Clean path, totally canonical, pointing straight at a location outside the intended base directory. The function did what it was designed to do. It just didn’t do what the caller assumed. A safe-path function would take the base directory as a parameter and refuse to return anything outside it. A safe option exists in Go, but these drivers never adopted it. The community workaround is filepath-securejoin, which does the symlink-aware join and returns an error instead of letting you escape.

Both vulnerable drivers shipped this misconception. They pulled the subdirectory string out of volumeHandle, passed it through filepath.Join, and trusted the result.

CVE-2026-3864 — Kubernetes CSI Driver for NFS

We found the bug in the controller-server implementation of csi-driver-nfs. The driver parsed the volume identifier into its parts, pulled out the subdirectory string, and built an internal mount path by joining that subdirectory to a per-volume working directory.

Figure 3 (NFS code): The NFS CSI driver extracts subDir from the volume ID and passes it to filepath.Join without validation, then calls os.RemoveAll on the result.

Three things matter here. No validation on subDir between extraction and use. filepath.Join won’t refuse a result that escapes its working directory. And the resulting path gets handed to os.RemoveAll. That last part is key: this isn’t just a read primitive. It deletes.

The exploit is a one-line modification to a PersistentVolume manifest:

Figure 4 (Exploit YAML): A malicious PersistentVolume manifest. The subDir component team-a/../../team-b escapes the legitimate tenant directory.

When a pod mounts this PersistentVolume, the path that actually gets bound is 10.0.0.50:/exports/team-b. Not team-a. The attacker’s pod now sees team B’s files, can read them, modify them, and (if the reclaim policy is set to delete) trigger their recursive removal when the PersistentVolumeClaim (PVC) is deleted.

In production deployments where the CSI controller has been granted broader hostPath mounts of the worker node (common in clusters that integrate the controller with kubelet log collection or other operational tooling), the traversal can reach beyond the export and into the host filesystem itself. A volumeHandle containing the subdirectory string ../../../../../../var/lib/kubelet provides the controller with a recursive-delete primitive against the kubelet’s own state directory, rendering the node permanently non-functional.

Two things make this harder to fix than it looks. First, NFS servers from major vendors expose hidden .snapshot directories. Point-in-time copies, invisible to ls but reachable by path. A subdirectory of team-a/../.snapshot mounts the snapshot tree and surfaces data that admins thought was gone. Second, even after the driver patch, an attacker with write access inside their own tenant directory can plant symlinks that cross the tenant boundary through the NFS data plane. The control-plane fix is necessary but not sufficient. Data-plane mitigations require mounting with nosymfollow (Linux 5.10+) or enabling subtree_check on the export.

CVE-2026-3865 — Kubernetes CSI Driver for SMB

The SMB CSI driver is maintained by the same upstream organization and follows the same architectural shape. It contained the same vulnerability:

Figure 5 (SMB code): The SMB CSI driver follows the same pattern as NFS: subDir extracted at index 1, joined without validation.

The exploit is structurally identical. The volumeHandle for SMB is //{server}/{share}#{subDir}#…, and a payload of //smb.internal/shared#dept-engineering/../../dept-finance#uuid causes the same escape into a sibling tenant’s directory.

Honestly, the recurrence is the more interesting finding than either CVE on its own. Same org, same misconception, years apart. When a standard-library function looks like a sanitizer but only normalizes, people keep reaching for it. Every codebase that did has to be audited now. Every callsite patched.

One note for other researchers on this bug class. The Kubernetes triage team initially could not reproduce the SMB report because their reproduction harness ran the proof-of-concept against an SMB-server pod inside the cluster, and SMB exports are sandboxed by the underlying storage system. The vulnerability is in the CSI controller’s internal file operations on its working-mount directory, not in protocol traffic to the SMB server. Researchers reporting CSI path traversal should lead with a host-mount reproducer and only attach protocol-level demonstrations as supplementary material. This pattern recurs because the server-side of NFS and SMB is usually well-defended by the storage vendor, while the driver-side code path between the Kubernetes API and the host filesystem is where the bug actually lives.

Multi-Tenant Storage as Attack Surface

These two CVEs are instances of something bigger. Kubernetes storage drivers sit on a trust boundary that the CSI spec doesn’t enforce. Nobody validates volumeHandle. The kubelet assumes the driver does it. The driver assumes whoever wrote the PV knew what they were doing. That was probably a CI/CD pipeline that assumed the API server would catch anything dangerous. The API server treats the field as opaque. So nobody checks. The string just flows through.

This isn’t just NFS and SMB. There are dozens of CSI drivers in production, each with its own volumeHandle format, each parsing user input with varying degrees of care, each forwarding strings into mount commands and host file operations. We haven’t audited all of them. But the pattern is there.

The bigger question is where input validation should live. The driver? Implemented inconsistently. The API server? Treats volumeHandle as opaque. An admission controller? Only works if someone remembers to deploy one. The CSI spec? Mandates none of the above. Until one of these layers owns the boundary, this bug class will keep producing CVEs.

The Full Chain — From PersistentVolume-Create to Cross-Tenant Compromise

The full attack chain:

Prerequisite: the attacker holds persistentvolumes:create (typically a CI/CD pipeline, GitOps controller, or operator service account) and identifies the cluster as using the NFS or SMB CSI driver with shared multi-tenant exports.

  1. Craft a PersistentVolume whose volumeHandle contains a traversal payload in the subdirectory component.
  2. Create a PersistentVolumeClaim bound to the malicious PV. The scheduler treats it as normal.
  3. Schedule a pod mounting the claim. The CSI driver parses the malicious volumeHandle, joins the traversal string against its working directory, and mounts the victim tenant’s directory.
  4. Read, modify, or delete the victim’s files. With onDelete=delete, deleting the PVC triggers recursive deletion.
  5. (Optional) If the CSI controller has broader hostPath mounts, traverse into the host filesystem and destroy node-critical paths like /var/lib/kubelet.

The entire chain executes against the Kubernetes API server using only the persistentvolumes:create permission. No exploit code. Just YAML.

Fixes and Mitigations

The Kubernetes Security Response Committee released fixes for both vulnerabilities. The NFS driver fix shipped in csi-driver-nfs v4.13.1 and rejects subdirectory strings containing .. components outright. The SMB driver fix shipped in csi-driver-smb v1.20.1 and applies the same validation pattern.

Beyond the upstream patches, cluster operators have several additional mitigations that should be applied as defense-in-depth:

  • Restrict persistentvolumes:create authorization. Audit which service accounts in the cluster currently hold this permission. CI/CD pipelines, GitOps controllers, and operator service accounts should be reviewed, and the permission should be scoped down or wrapped behind admission-controller approval for any account that is not strictly an administrative identity.
  • Deploy admission-time validation. A ValidatingAdmissionPolicy (in Kubernetes 1.30 and later) or a ValidatingWebhookConfiguration (in earlier versions) can reject any PersistentVolume whose spec.csi.volumeHandle field contains .. or other suspicious patterns. This control survives any future zero-day in the same class. The next CSI driver to ship the same misconception is automatically defended.
  • Address the data-plane symlink amplification. The driver patches don’t cover the symlink variant. A tenant with write access to their own directory can still plant a relative symlink that crosses into someone else’s. Mount NFS exports with the nosymfollow option (Linux 5.10 and later) where supported, enable subtree_check on the NFS server export configuration, or audit symlink targets on the storage server.
  • Treat the bug class as recurring. The pattern of filepath.Join on attacker-controlled input is present in many storage drivers, container network interface plugins, operators, and webhook servers throughout the cloud-native ecosystem. Any Go code that takes an untrusted string and joins it to a path should be audited and converted to use either strict validation or the filepath-securejoin library.

Conclusion

Path traversal is one of the oldest bug classes in computing. The fact that it keeps showing up in 2026, in actively maintained infrastructure running multi-tenant production clusters, says something. The safe path-handling already exists, but these drivers kept trusting filepath.Join instead of using it.

The CSI specification doesn’t mandate input validation on volume identifiers. Individual drivers implement it inconsistently, or not at all. The protocols underneath (NFS and SMB here, but the pattern generalizes) were designed for trusted networks decades before multi-tenant container orchestration existed. Until one of these layers takes ownership of the boundary, the same misunderstanding will keep shipping.

For defenders: patch the drivers, restrict who can create PersistentVolumes, deploy admission-time validation, and treat any .. sequence in a volumeHandle as a credible alert. Validators reject. Normalizers don’t. Know which one you’re using.

Disclosure Timeline

  • January 15, 2026 — Vulnerabilities identified during CSI driver code review by SentinelOne Researchers.
  • January 16, 2026 — NFS and SMB reports submitted to the Kubernetes SecurityResponse Committee.
  • January 19, 2026 — NFS report triaged after proof-of-concept harness was clarified.
  • January 19, 2026 — SMB report triaged.
  • March 9, 2026 — NFS driver fix released in csi-driver-nfs v4.13.1.
  • March 17, 2026 — Public disclosure of the NFS finding; CVE-2026-3864 assigned.
  • March 21, 2026 — SMB driver fix released in csi-driver-smb v1.20.1.
  • April 11, 2026 — Public disclosure of the SMB finding; CVE-2026-3865 assigned.

Additional Resources

The Good, the Bad and the Ugly in Cybersecurity – Week 29

The Good | Authorities Sanction Cybercriminals & Dismantle Russian Bulletproof Hosting Infrastructure

The EU and the United Kingdom have jointly sanctioned multiple Russian individuals and entities for targeting government networks and critical infrastructure across Europe. The sanctions specifically target senior Russia military intelligence (GRU) officers and operators, as well as four entities linked to the Federal Security Service (FSB).

Officials say that the Russian government actively utilizes these state-sponsored units alongside recruited cybercriminals and private companies to systematically destabilize international partners and compromise key infrastructure across the continent.

From the U.S. Treasury Department, two individuals and a virtual private network (VPN) provider face sanctions for actively enabling ransomware attacks against American organizations.

OFAC designated First VPN Service (1VPNS) and its administrator, Dmytro Rashevskyi, for supplying infrastructure that helped cybercriminals obscure their identities and manage stolen data. The service, which law enforcement dismantled last May, notoriously ignored abuse complaints and maintained zero user logs.

Yegeniy Silayev was also sanctioned for developing cryptors designed to conceal malware. Investigators estimate these specific tools and services directly facilitated billions of dollars in financial losses across critical sectors.

U.S. Federal prosecutors also unsealed indictments this week against three Russian nationals for operating bulletproof hosting services that facilitated over $62 million in global ransomware damages.

Defendants Aleksandr Volosovik, Yulia Pankova, and Kirill Zatolokin allegedly managed “Media Land” and “ML Cloud”, providing essential infrastructure to syndicates like Lockbit, Play, and Blacksuit. These hosting platforms actively shielded cybercriminals by disregarding victim complaints and ignoring law enforcement takedown requests.

To disrupt this supply chain, the State Department is offering a $10 million reward for actionable information regarding foreign government links to these hosting providers.

The Bad | Attackers Trojanize Popular Remote User Platforms to Deploy Starland Malware

Cybersecurity researchers identified a financially-motivated Russian threat actor tracked as UAT-11795. Active since June 2025, the actor has utilized trojanized applications to harvest user credentials and cryptocurrency while primarily targeting users across the United States, Germany, Romania, and Venezuela.

To distribute their payloads, UAT-11795 operators disguise malicious installers as legitimate software, including WebEx, Zoom, MobaXterm, DBeaver, and FaceIT. Researchers suspect the attackers likely deploy these files via ClickFix social engineering.

The infection chain typically starts when a victim executes a malicious HTA file. This file retrieves an altered NSIS installer harboring a hidden Python loader disguised as a standard text document. The loader then modifies the Windows Registry to ensure persistent access before decrypting and deploying the Starland remote access trojan (RAT).

Upon execution, Starland verifies whether it is operating within a sandbox before creating scheduled tasks and attempting to escalate its system privileges. The malware scans compromised systems for browser data, cryptocurrency wallet assets, detailed system configurations, any antivirus products, and Active Directory infrastructure such as domain structure and controllers.

Beyond data theft, Starland possesses extensive capabilities to capture desktop screenshots, execute arbitrary shell commands, and fetch secondary payloads. Depending on system architecture, the malware can inject a 64-bit shellcode chain to deliver the CastleStealer information stealer or a 32-bit chain to deploy the Remcos remote access trojan.

UAT-11795-controlled Telegram channels (Source: Cisco Talos)

To maintain resilient command and control (C2) communications, the operators integrate a redundancy mechanism that queries a Polygon smart contract for a fallback domain, and control two Telegram bots to receive notification beacons, including messages with the victim’s machine fingerprints and cryptowallet inventories.

Users are reminded to avoid executing unidentified commands online and should only download confirmed software from official vendor sources.

The Ugly | Nearly 300 Imposter GitHub Repositories Distribute Infostealing Malware to Collect Sensitive Data

Threat actors have published almost 300 fabricated GitHub repositories to distribute an information stealer from the BoryptGrab malware family. The actors systematically impersonated premium security products, cryptocurrency tools, and developer utilities to deceive victims searching for free software downloads.

As part of the lure, the malicious landing pages employ highly sophisticated client-side scripts that parse referral URLs to render customized branding and spoofed trust badges, significantly increasing the likelihood of successful social engineering.

Once a targeted victim clicks the download link, the infrastructure delivers a constantly rotating ZIP archive containing a legitimate, signed WinGUP updater paired with a trojanized dynamic link library file. When the user executes the updater, the program side-loads the malicious file, which then decodes and reflectively executes the BoryptGrab-variant payload directly into system memory.

Operating without establishing long-term persistence, the malware is designed to exfiltrate maximum data in a single execution cycle. The stealer targets passwords, payment details, and session cookies across 19 different web browsers and 32 cryptocurrency wallet brands, alongside messaging tokens from Discord, Steam, and Telegram.

The infostealer’s execution workflow (Source: Arctic Wolf)

To maximize collection, operators utilize direct code injection to bypass Chrome’s native App-Bound Encryption. All newly harvested data is compressed and routed to a Russian-based C2 server. Although the malware leaves behind forensic evidence by failing to wipe temporary staging directories, the scale of the impersonation campaign poses significant risks to unsuspecting developers.

GitHub has already removed a large portion of the false repositories, though several of the malicious redirector pages remain actively online. Researchers advise users to independently verify software authenticity and exercise extreme caution when navigating unofficial portals, sharing this YARA rule to help detect BoryptGrab activity and IoCs.

The Good, the Bad and the Ugly in Cybersecurity – Week 28

The Good | Authorities Apprehend Pro-Russian Hacktivist & Dismantle Global Fraud Networks

Spanish authorities, acting on intelligence provided by the FBI, have apprehended a suspected core member of the pro-Russian hacktivist syndicates CyberArmy of Russia Reborn (CARR) and Z-Pentest.

While masquerading as ideologically motivated collectives, these groups have actively executed disruptive cyberattacks against critical infrastructure, including food processing and water facilities across the United States and Europe. Investigators allege the arrested individual, residing in Palencia, provided extensive operational and logistical support to a Ukrainian hacker working for CARR, even attempting to facilitate their escape to Russia.

The individual is also suspected of coordinating cyber operations for the NoName057(16) group using encrypted messaging platforms. In a March 2026 raid, law enforcement officers seized multiple computers and successfully froze cryptocurrency wallets utilized to launder illicit proceeds generated from stolen data sales. The suspect currently faces ongoing criminal investigations for alleged collaboration with a recognized terrorist organization and severe computer damage.

In a massive global crackdown on social engineering and financial fraud, international law enforcement agencies have arrested 5,811 suspects and seized approximately $293 million in illicit assets. Codenamed “Operation First Light 2026”, the coordinated initiative spanned 97 countries and specifically targeted business email compromise (BEC), investment scams, and money laundering syndicates operating between January and April. Interpol actively coordinated the extensive joint action, collaborating directly with regional policing bodies like ASEANAPOL, GCCPOL, and Europol to swiftly block over 31,000 fraudulent bank accounts and virtual wallets.

Eswatini police seized electronic devices, foreign currency, and realistic replicas of Brazilian police uniforms, signage, and equipment (Source: Interpol)

Investigators identified more than 142,000 victims worldwide and pinpointed an additional 15,600 suspects for future prosecution. This success builds upon recent international efforts, including Operation Synergia II, to dismantle the sprawling infrastructure supporting transnational cybercrime. Officials emphasize that robust, cross-border law enforcement cooperation remains essential to combat the escalating threat of organized cyber-enabled financial crimes globally.

The Bad | Threat Actors Deploy Forg365 PhaaS to Hijack Microsoft Accounts

Cyber researchers have identified a new phishing-as-a-service (PhaaS) operation dubbed Forg365, which targets Microsoft 365 enterprise accounts. Blending adversary-in-the-middle (AiTM) techniques with device-code phishing, the platform provides an integrated dashboard to manage post-compromise activities.

Forg365 works by incorporating AI to assist in generating customized phishing lures. By integrating AI directly into the control panel, Forg365 developers significantly lowered the financial cost needed to launch targeted campaigns. To remain undetected, its operators route messages through legitimate Amazon SES infrastructure while hosting their landing pages on Cloudflare.

The Forg365 panel (Source: ZeroBec)

The operation leverages the OAuth 2.0 device code authentication flow, originally designed for input-constrained devices. Attackers present victims with a deceptive verification page, tricking them into authorizing an attacker-controlled gadget rather than stealing their password directly.

Once initial access is achieved, the platform ensures persistence through a specialized browser extension called ForgCookie. Compatible with Microsoft Edge, Google Chrome, and Brave, this extension operates silently to request account data, clear session cookies, and trigger a hidden OAuth flow to capture fresh tokens. Doing so grants attackers continuous access to the victim’s Microsoft services without requiring them to ever re-authenticate.

To protect its administration panel from being accessed by security defenders, Forg365 integrates robust anti-analysis features. The platform utilizes debugger traps, polymorphic code, and dynamic sandbox checks to evade detection, redirecting connections to innocuous websites whenever a VPN is detected.

Administrators can defend against these hijacking techniques by monitoring Microsoft Entra logs for unexpected device-code authentication events. Organizations can also restrict or entirely disable device-code flows unless absolutely necessary. In the event of a suspected compromise, security teams should revoke all OAuth grants and refresh session tokens to sever access.

The Ugly | Rival Espionage Actors Breach & Spy On Pakistani Law Enforcement Networks

Between February 2024 and April 2026, suspected state-sponsored threat actors based in China and India separately converged on several Pakistani law enforcement organizations in unrelated cyberespionage campaigns.

According to SentinelLABS, operators heavily targeted the Balochistan Police, compromising critical network appliances and web servers. By infiltrating these critical systems, both nations actively sought independent visibility into Pakistan’s internal security posture and ongoing counter-militancy operations.

The China-nexus intrusions, leveraging PlugX, ShadowPad, and Cobalt Strike malware, were likely driven by Beijing’s concerns over the safety of Chinese nationals working on regional infrastructure projects within the China-Pakistan Economic Corridor (CPEC). Ongoing terrorist attacks have left the Chinese government dissatisfied with Pakistani protection and data from Balochistan Police would give the PRC direct insights.

Conversely, the India-nexus activity utilized Remcos backdoors to gather intelligence on the restive Balochistan province, a recurring flashpoint in the adversarial relationship between the two countries. Control over Balochistan Police networks means having invaluable visibility on how Pakistan manages their security posture as well as persistent access to civilian data.

Timeline of C2 traffic to Pakistani law enforcement organizations
Timeline of C2 traffic to Pakistani law enforcement organizations (Source: SentinelLABS)

One China-aligned threat actor specifically compromised the Balochistan Police Force’s Complaint Management System (CMS), a web application that actively serves both law enforcement personnel and Pakistani civilian users. The attackers uploaded custom malware implants disguised as routine portal updates, effectively weaponizing the digitalization of public policing services.

One payload masqueraded as a legitimate component of endpoint security software to evade initial detection and deploy an AsyncRAT client in order to establish persistent footholds into internal police networks while simultaneously surveilling citizens utilizing the platform.

This multi-actor convergence highlights how modernized policing infrastructure can serve as a high-value intelligence target for rival nations seeking comprehensive regional data.

HPC AI Workloads Need Runtime Security. The Architecture Already Exists.

The US Federal Government is committing $600 million to build one of the world’s most advanced AI infrastructure systems. Executive Order 14363, the Genesis Mission, connects national laboratory supercomputers across nuclear simulation, biodefense, energy grid modeling, and every major scientific domain. Fifty-one organizations signed on, including NVIDIA, OpenAI, IBM, Microsoft, AWS, Google, and Oracle.

The security framework governing these workloads was not written for this scale of use.

NIST SP 800-234, the High-Performance Computing Security Overlay, is well-constructed, tailoring 60 controls across four security zones, building on the SP 800-53B moderate baseline. It was designed for deterministic HPC workloads such as climate simulations, finite element analysis, and computational fluid dynamics. These workloads share a common attribute: code that runs the same way, every time, and behaves predictably under well-understood inputs. The security controls governing those workloads assume you can scan at the perimeter, clear memory between jobs, and attest to integrity at load time.

AI workloads break every one of those assumptions.

SentinelOne has submitted a formal proposal to the NIST HPC Security Working Group regarding this gap, and NIST has acknowledged it. We have a post on LinkedIn to share our proposal, and welcome commentary from across the industry.

The supply chain problem just got a lot more dangerous

This spring, in just three weeks, three AI-driven supply chain attacks targeted widely deployed software: LiteLLM, the most-used AI infrastructure package in Python development environments, Axios, the most-downloaded HTTP client in the JavaScript ecosystem, and CPU-Z, a trusted system diagnostic tool with a legitimate signed binary from the official vendor domain.

SentinelOne stopped all three on the same day each attack launched, with no prior knowledge of any payload.

The most important aspect of this outcome is how these attacks were stopped, and why signature-based detection couldn’t work. Each attack arrived through a trusted delivery channel. LiteLLM was compromised after credentials were stolen via Trivy, a security scanner. The attacker published two malicious versions to the PyPI repository. In at least one confirmed case, an AI coding agent with unrestricted permissions auto-updated to the infected version, meaning there was no human review or approval step before the payload ran. The Axios attacker exploited a legacy access token that the project maintainers had forgotten to revoke, bypassing every npm security control. CPU-Z attackers targeted the vendor’s distribution infrastructure directly; anyone who downloaded from the official website received a properly signed binary containing a payload. In all three cases, while the authorization chain was legitimate, the intent was not.

This is the defining characteristic of modern supply chain attacks: the workflow is verified, but the intent has been subverted. Every perimeter control, signature library, and reputation lookup checks authorization and passes. These attacks were designed to exploit that gap, and they ran at machine speed through automated pipelines with no human checkpoint.

To put this into the context of HPC and AI workloads running at scale, a compromised Python package in a developer’s environment is a serious incident; a poisoned training pipeline on classified biodefense data on a national laboratory supercomputer is on a different order of magnitude. The model it produces may be correct 99.9 percent of the time and adversarially wrong under precisely targeted conditions. No perimeter scan, signature check, or load-time integrity verification will catch it after training completes.

Where the current framework falls short

Of the 60 controls SP 800-234 tailors, three bear directly on AI workload protection, and each carries a documented gap. In a fourth area, supply chain, the overlay does not tailor at all.

  • SI-3 (Malware scanning): The control acknowledges that real-time scanning is most effective but explicitly permits tailoring for performance on HPC systems, deferring to perimeter scanning before data reaches the compute zone. For traditional HPC workloads, that tradeoff may be defensible, but for AI workloads, it leaves behavioral analysis of the execution process completely unaddressed. A poisoned training run that executes within the expected statistical range of a training job looks like legitimate compute to a perimeter scanner.
  • SI-4 (System monitoring): The control notes that high-speed data flows in HPC environments can overwhelm standard monitoring tools, and lacks AI-specific monitoring requirements or telemetry collection requirements from execution pipelines. The practical interpretation of this is: monitor what you can, accept the gap for what you can’t. On infrastructure running AI at scale, that gap creates a primary attack surface.
  • SC-4 (Information in shared resources): Requires GPU memory clearing between user reassignments. It addresses data residency at the transition but does not address runtime behavioral monitoring of workloads during execution, side-channel attack detection, or anomalous compute-pattern identification while training is active.
  • SR family (Supply chain risk management): The overlay carries all 12 moderate-baseline SR controls forward from SP 800-53B, with no HPC or AI-specific guidance, and supply chain is not among the 14 categories it tailors to. The SR controls still address only the conventional software and hardware supply chain; they say nothing about training-data provenance, model-weight integrity, or pre-trained-model validation, and the framework defines no AI equivalent of a software bill of materials. LiteLLM, Axios, and CPU-Z all arrived through legitimate software supply chain channels. AI workloads carry that same exposure one layer deeper, in the data and model artifacts that software trains on, which is exactly where the overlay is silent

AI Runtime Threats

The attacks against AI workloads on HPC are not theoretical, and they are not detectable at the perimeter.

  • Training data poisoning scales at rates most security teams are not equipped to respond to. Research1 across 41 studies documents attack success rates exceeding 60 percent from manipulation of 100 to 500 training samples, a fraction of a percent of a typical dataset. Poisoning as little as 3 percent2 of training data achieved 41 percent attack success rates in code-generating models. OWASP’s LLM Top 103 documents the consequence. Backdoors leave model behavior intact until a specific trigger activates adversarial outputs. The model ships, it gets deployed, and operates correctly, until it doesn’t. No post-training audit reliably catches a well-designed poisoning attack.
  • GPU side-channel attacks are executed remotely by a co-tenant workload on shared GPU infrastructure; no physical access is required. The NVBleed research demonstrated covert channel attacks on NVIDIA NVLink, achieving over 91 percent accuracy in recovering data-dependent information from co-tenant GPU workloads on a shared fabric. The BarraCUDA research demonstrated the extraction of neural network weights via electromagnetic side channels from NVIDIA hardware. Both attack classes execute during active training, not at job transition. If your HPC environment runs multiple projects or security classifications on shared accelerators, the co-tenancy model is an active attack surface today.
  • Inference pipeline compromise survives load-time integrity checks. A model with clean weights at deployment faces attacks through three vectors: hot-swap modification of serving configurations while inference runs; preprocessing and postprocessing layer injection that alters inputs before they reach the model or modifies outputs before delivery; and adversarial input manipulation that triggers targeted misbehavior in a model that appears fully operational. For AI serving safety-critical inference, each is a security risk, not just a research concern.

The characteristic that makes AI workloads uniquely difficult is persistence. A compromised simulation may produce visibly wrong results, but a compromised model can produce correct results the overwhelming majority of the time and adversarially wrong results under precisely targeted conditions. By the time anyone has reason to investigate, the window for recovery has often closed.

Securing HPC AI Workloads

We know the technology required to address these gaps exists and has been proven at scale in environments with performance constraints far tighter than those in HPC. What is needed is a well-defined architecture that enables the secure execution of large-scale AI workloads.

Dedicated security compute. Runtime security that shares CPU resources with the workload it monitors can be starved of CPU time under heavy load and interfered with by a workload that achieves kernel-level access. The SPiCa research demonstrated that eBPF monitoring pipelines can be manipulated from within the kernel by rootkits filtering events before they reach the analysis engine, meaning that a co-scheduled monitor is not a reliable monitor.

Every other infrastructure function on an HPC node has dedicated resources. The job scheduler, the filesystem client, and the out-of-band management plane. Security monitoring is infrastructure and should be afforded the same dedicated resources.

Modern HPC nodes have 128 to 256 CPU cores. One reserved for security monitoring is less than one percent of the available compute. Linux kernel CPU isolation via isolcpus, nohz_full, and rcu_nocbs is production-proven in high-frequency trading and real-time systems, with bounded, predictable overhead.

eBPF-based behavioral telemetry at the training layer. Effective monitoring of an AI training pipeline means continuous observation of compute behavior profiles, memory access patterns, GPU utilization, and inter-node communication, with behavioral baselines established for approved training configurations. A poisoning attack that executes within expected statistical ranges is not visible to a perimeter scanner, but it is visible to a behavioral baseline that knows what the training job should look like.

This is the same principle that SentinelOne’s on-device Behavioral AI detected for LiteLLM, Axios, and CPU-Z. The LiteLLM detection flagged a Python interpreter executing Base64-decoded code in a spawned subprocess. The CPU-Z detection flagged an anomalous process chain: cpuz_x64.exe spawning PowerShell, which spawned csc.exe, which spawned cvtres.exe. CPU-Z doesn’t do that. The behavioral baseline knew what legitimate execution looked like, and in these cases, that behavior was the decisive signal.

Cloudflare uses an eBPF-based architecture to mitigate DDoS attacks exceeding 7 Tbps. SentinelOne uses it to detect and stop threats in under one second across enterprise fleets. A training job that begins writing to unexpected locations, establishing anomalous inter-node communication, or deviating from its expected compute profile is detectable at runtime, before the model completes training. The performance argument against runtime monitoring on HPC was never about the technology; it requires a shift in architecture.

Inference-time output monitoring. Deployed models require continuous observation of output distributions, latency patterns, confidence score distributions, and input-output statistical properties. A model under adversarial input attack, or serving modified weights, exhibits detectable output patterns before any human analyst notices the outputs are wrong. Circuit-breaker logic needs to be designed into the serving architecture, not added after the first incident.

Model integrity verification that runs during inference. Load-time attestation is a necessary and important requirement; it is not sufficient. Long-running inference deployments are vulnerable to hot-swap attacks that replace weights after the initial integrity check passes. Continuous cryptographic hash verification of loaded model weights, running on the dedicated security core with automated circuit-breaker logic on failure, closes that vector. For a model serving safety-critical calculations, the re-verification frequency should match the workload’s risk profile with predictable overhead.

An SR-family extension for the AI supply chain. The existing SR controls address software supply chain risk. They do not address training data provenance, model weight integrity at ingestion, or pre-trained model validation. An AI bill of materials, including cryptographic documentation from the training data source through intermediate checkpoints to the deployed model, is the model-layer equivalent of software supply chain controls. Without it, every pre-trained model loaded into an HPC environment is an unverified artifact from an unverified chain.

Defending AI at Every Layer

The supply chain attacks this spring demonstrated what happens when defense architecture falls behind the delivery mechanisms attackers use. LiteLLM, Axios, and CPU-Z all arrived through trusted channels, carrying payloads no signature database contained. They were stopped because behavioral detection does not require prior knowledge of the payload. It requires knowing what legitimate execution looks like and acting when execution deviates.

Defenders protecting AI workloads face that same problem across every layer they own. HPC is the hardest version of it. But identities, endpoints, applications, and infrastructure all carry the same exposure at different scales. SentinelOne gives defenders coverage across all four, with behavioral AI running at each layer to catch what signatures miss. The specifics of how that works across your AI environment are in our AI security overview.

Citations

1 “Data Poisoning 2018–2025: A Systematic Review. IACIS (2025)”, and “Data Poisoning Vulnerabilities Across Health Care AI Architectures. JMIR (2026)

2 “Poisoning Attacks on LLMs Require a Near-Constant Number of Poison Samples” (2025). arXiv:2510.07192 and Huang et al., 2020.

3 OWASP (2025) LLM04:2025 Data and Model Poisoning. OWASP Gen AI Security Project.

The Good, the Bad and the Ugly in Cybersecurity – Week 27

The Good | Authorities Apprehend Iranian Cybercriminal & Extradite UNC3944 Hacker

Montenegrin law enforcement, alongside the FBI, have apprehended a 39-year-old dual Iranian and Turkish citizen wanted by the U.S. government for several cybercrime offenses. Arrested in Kotor, the suspect faces charges in the Southern District Court of New York for conspiracy to commit computer fraud, hacking, and identity theft.

Since 2013, this individual allegedly orchestrated mass cyberattacks against more than 150 American universities and inflicted damages estimated at over $3.4 billion. Investigators say the stolen data and compromised academic credentials directly benefited the Islamic Revolutionary Guard Corps and various Iranian state entities. The case now proceeds to a High Court judge for formal extradition hearings. The arrest follows recent warnings from U.S. cybersecurity agencies regarding escalating Iranian state-sponsored operations targeting critical domestic infrastructure.

A 19-year-old dual United States and Estonian citizen, Peter Stokes, has been extradited to face federal charges for his role as a core member of the UNC3944 (aka Scattered Spider, oktapus) cybercrime syndicate. Finnish authorities initially apprehended Stokes at the Helsinki airport as he attempted to board a flight to Japan. Prosecutors are accusing him of orchestrating multiple high-profile corporate breaches, using intense social engineering tactics against IT helpdesks to bypass multi-factor authentication controls.

Source: U.S. DoJ

In one notable May 2025 incident, UNC3944 compromised a multibillion-dollar retailer, demanding an $8 million ransom while inflicting over $2 million in operational disruption and remediation costs. UNC3944 operators are also responsible for more than 100 network intrusions globally, all of which yield upwards of $100 million in illicit extortion payments. Stokes remains in federal custody in Chicago, facing charges of fraud, conspiracy, and computer intrusion.

The Bad | Russian Intelligence Exploit Phishing Campaigns To Steal Signal Backup Keys

CISA and the FBI are warning that Russian state-sponsored threat actors have made large strides in evolving their phishing operations to target the backup recovery keys of Signal users. In an update to their March 2026 advisory, the two agencies attribute this ongoing activity to Russian Intelligence Services (RIS), including Russia’s Federal Security Service (FSB) Border Guards and the country’s military.

Tracked as UNC5792 and UNC4221, these campaigns specifically target high-value individuals, including government officials, military personnel, journalists, and policy analysts. Previously, operators focused on harvesting standard verification codes or tricking users into silently linking unauthorized devices. Now, they employ social engineering to access private communications without compromising the application’s underlying end-to-end encryption.

During targeted intrusions, attackers masquerade as official Signal support personnel and send direct messages falsely claiming the platform requires mandatory two-factor verification following alleged international cyberattacks. The operators systematically guide victims through the specific process of enabling the Secure Backups feature, instructing them to paste their newly generated recovery key directly into the chat interface.

Once adversaries obtain this critical key, they seamlessly download and decrypt the victim’s entire historical message archive onto their own controlled devices. Simply registering a new account under the same phone number does not natively invalidate a compromised key – users must actively generate a new backup key within their application settings to effectively secure future communications.

Source: U.S. Rewards for Justice Program

The U.S. Department of State recently announced a substantial reward of up to $10 million for information leading to the identification or location of these operatives. Through the Rewards for Justice program, federal authorities actively seek actionable intelligence regarding the syndicates’ operational infrastructure, illicit funding mechanisms, and direct affiliations with Russian intelligence services.

The Ugly | Unknown Hackers Breach Department of Homeland Security Information Network

The Department of Homeland Security is actively investigating a cyberattack that recently compromised its Homeland Security Information Network (HSIN). The network is an information sharing platform used by federal, state, local, and private-sector partners, specifically to share sensitive but unclassified data amongst the government and internationally.

According to an initial report, an unidentified threat actor orchestrated the intrusion between late May and early June. Investigators indicate that the attackers targeted the HSIN’s core servers alongside a SharePoint environment designed for extensive interagency collaboration.

Source: dhs.gov

While officials have not yet attributed the breach to a specific foreign government or syndicate, the full extent of data exposure remains unclear. The compromised platform routinely supports real-time incident management, intelligence exchange regarding persons of interest, and operational coordination. Because the United States is overseeing security for World Cup matches across the country, experts raise concerns that the intrusion could have exposed critical security planning, response procedures, and communication protocols.

In a public statement to the press, a departmental spokesperson confirmed the incident, clarifying that the breach strictly involved an unclassified legacy information-sharing environment. Security personnel promptly isolated the affected systems and mitigated the underlying vulnerability before starting forensics.

Officials emphasized that the attack did not impact classified networks, and the primary system remains operational for authorized partners. As the investigation continues, authorities strongly urge U.S. government staff, contractors, and associated partners to remain vigilant while defense teams harden the underlying infrastructure against further unauthorized network access attempts.

❌