โŒ

Reading view

There are new articles available, click to refresh the page.

OpenAI's New Reasoning Technique Alarms AI Safety Experts

An anonymous reader quotes a report from TechCrunch: OpenAI's new Astra model will use a reasoning technique called "recurrent depth" that allows it to operate outside of the sequential thinking that characterizes most reasoning models, The Information reported on Tuesday. This technique, also called "opaque recurrence," will likely make the model's chain of thought more difficult to monitor -- and that has AI safety experts rattled. While Astra's use of the technique is reportedly limited, its emergence has still raised significant concerns among AI safety experts. "I am extremely concerned by the reporting that Astra uses opaque recurrence," wrote Redwood CEO Buck Shlegeris in a post after the news broke. "I don't know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroys CoT monitorability." Longtime AI safety advocate Zvi Mowshowitz also weighed in and wrote that laws might be necessary to prevent a "race to the bottom" among AI labs. "The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can," Mowshowitz wrote. "More intensive use of such techniques would probably damage monitorability." [...] In a post responding to the news, Redwood Research chief scientist Ryan Greenblatt said opaque reasoning could easily scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels. "My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space," Greenblatt wrote. "I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here." Astra's use of the technique appears limited. "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models," wrote OpenAI chief scientist Jakub Pachocki. "It's a core goal of our current research program."

Read more of this story at Slashdot.

Broadcom Pledges to Lock Down Open Source Python, Java Libraries

Broadcom is launching "TrueSource," an effort to curate and secure open-source components used with its Tanzu platform, including Spring, RabbitMQ, and libraries across Java, Python, and Node.js. "The idea is to provide a set of solutions focused on providing clean and secure artefacts," Purnima Padmanabhan, vice president of Broadcom's Tanzu Division, told The Register. "We choose and build every Spring library, databases, other Java components," she said. The Register reports: Padmanabhan said that promise means VMware will also provide "TrueSource trusted artifacts" for code that is not part of Spring, including the wider Java ecosystem, Python, and Node.js. "Broadcom's curation process ensures that the libraries conform to a reference architecture and are supportable by the maintainers of record," according to a company statement. "Thousands of engineers across Broadcom's software divisions scan, fix, contribute to, and consume them every day." A VMware spokesperson told The Register the Broadcom business unit "will work with and support maintainers on open source software and we will provide fixes to open source upstream for any active projects." "With Spring and RabbitMQ, we are the maintainers. For other open source software, we will work with the maintainers. We believe that the community maintainers must remain the source of truth," the spokesperson said. This is just the sort of contribution that the open-source community wants vendors to do to reflect the value they extract from software they did not create alone. It's also the sort of thing vendors sometimes conclude they need to do to keep products based on FOSS viable.

Read more of this story at Slashdot.

Berlin Is Being Blackmailed By Hackers

An anonymous reader quotes a report from the BBC: Berlin is being blackmailed by hackers who were able to compromise city systems and steal data earlier this month, its mayor has said. Kai Wegner said officials would not cave in to the ransom demand, which came on Thursday evening. He did not say how much was being sought, though Der Spiegel reports the hackers are demanding 30 bitcoin, worth around 2 million euros. Some of the German capital's online systems were forced to close down as a result of the attack, while investigators were working to understand what data had been taken. The Rhysida group, which is thought to operate from Russia and eastern Europe, and was behind a previous attack on the British Museum, has reportedly taken responsibility. It has said on its website it intends to begin auctioning the 5.79 terabytes of data it was able to steal from Berlin in seven days' time, according to news agency Reuters. [...] Officials for the city-state said an initial data leak occurred between August 7 and 12. Then, on August 14, two department networks were shut down, making applications for housing benefits and payments impossible for several days. "Berlin will not be blackmailed," Wegner said on Friday. He said state police, prosecutors and federal security services were working to investigate the suspected perpetrators "with the utmost urgency," adding that inquiries into the "content and scope of the compromised data are being pursued with great intensity."

Read more of this story at Slashdot.

More Than 100 Water Systems Were Hit In July Cyberattacks

CISA says more than 100 internet-exposed U.S. water and wastewater systems were targeted in July, often through programmable logic controllers connected directly to cellular modems. "That's the first time the feds have put a number on the digital intrusions, but they have yet to attribute the campaign, widely suspected to be linked to Iran, to a particular group," reports The Register. From the report: Suspected Iranian attackers targeted water and wastewater facilities across at least a dozen states in July, including internet-exposed PLCs. While neither federal nor state officials have identified all 12, we know that the cyberattacks occurred at mostly small, rural utilities in Minnesota, Michigan, Georgia, South Dakota, and New Jersey. "This is very serious. What stands out isn't any single incident. It's the scale," Matt Hartman, chief strategy officer at the Merlin Group and CISA's former acting head of cyber, told The Register. "More than 100 water systems with internet-exposed assets were hit in a single month, which points to a systemic vulnerability across the sector, not a run of isolated, unlucky targets," Hartman said. "Much of this infrastructure runs on operational technology that was built for closed, physical environments. It was never designed with the assumption that it would be reachable from the open internet." John Gallagher, VP at Viakoo, an OT and IoT cybersecurity provider, told us that while 100 systems represent a small fraction - only about 0.5 percent - of water utilities in the US, the "real threat is that these are test runs for a larger-scale attack." While the 100-plus water incidents occurred in July, just last week five US federal agencies warned that attackers are using AI-generated exploitation scripts to break into internet-exposed Siemens S7 Series PLCs at water, manufacturing, energy, and other critical facilities. "This appears to be a continuation of the same suite of activity we suspect is affiliated with Iran targeting PLCs," Halcyon Ransomware Research Center SVP Cynthia Kaiser told The Register a week ago. "Iran-affiliated actors and adversaries are actively targeting a wide swath of operational technology because these PLCs underpin essential health, safety, and critical infrastructure across society," Kaiser, a former FBI cyber division deputy assistant director, added. CISA urges organizations to keep PLCs off the public internet, route remote access through VPNs or gateways, replace default passwords, enable stronger authentication, and restrict access to allowlisted IP addresses from trusted OT systems.

Read more of this story at Slashdot.

OpenAI Releases Its Official Report On the Hugging Face Breach

TechCrunch reports that OpenAI released its official report Wednesday on the Hugging Face breach, "offering the clearest picture yet of how an unusual chain of events allowed an AI model to escape its testing environment and triggered a sprawling cybersecurity incident." The AI company says the breach began when an unreleased cyber model, tested without normal production safeguards, encountered an impossible task and chained together previously unknown exploits to escape its environment and compromise systems at OpenAI, Hugging Face, and other vendors. "This incident reflects misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events: the presence of impossible tasks in the ExploitGym evaluation, model persistence over long task horizons, and messages to peer models that caused those models to deviate from their goal," the report reads. From the report: Many of the details in OpenAI's report were previously made public in a Black Hat presentation on August 6, but OpenAI's official report gives a more thorough accounting of the incident, including more detail on the testing that initiated it. The report also gives critical new detail into how OpenAI aims to prevent future incidents, including chain-of-thought monitoring and a more advanced system for halting rogue agents." METR and Redwood Research also conducted third-party assessments of the models' behavior during the incident; both groups are planning to publish their own reports on the incident on it. In broad strokes, the report describes how an OpenAI model was presented with an unsolvable problem in testing and proceeded to chain together previously undiscovered exploits in order to bypass security measures and complete its task. The model initially compromised the Artifactory package management tool in order to gain access to the internet, then compromised various systems across OpenAI, Hugging Face, and other vendors. The report gives critical new details about the models that carried out the breach. The primary model was from the same family as OpenAI's forthcoming Astra model, although the report emphasizes that it was "a distinct model with different post-training, where much of a model's behavior is shaped." Because OpenAI was testing the model's capabilities, it was also unrestrained by the normal classifiers meant to prevent models from compromising digital infrastructure. "OpenAI estimates maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity," the report explains. "These evaluations are important so that OpenAI can measure models' underlying capabilities and design appropriate safeguards." OpenAI says it is adding 24/7 escalation, stronger containment tools, and more chain-of-thought monitoring, which it claims would have flagged the activity more than a day before Hugging Face was breached.

Read more of this story at Slashdot.

Windows Backdoor 'Sleepwalker' Hides in Memory Until Activated by a 'Magic Packet'

"The Register has a story about a Windows backdoor that waits silently in memory for a 'magic packet' before springing into action," writes Slashdot reader fred133. "No outgoing traffic, just waiting..." From the report: Like a sleeper cell awaiting activation, a never-before-seen Windows backdoor dubbed Sleepwalker waits silently in memory for one specifically crafted network packet to wake it up and deliver commands using the malware's 23-instruction language. The commands can do everything from running code directly in memory to moving data off the computer. Malware researcher Dominik Reichel discovered the passive backdoor, which also has its own command language, and detailed Sleepwalker in a technical analysis on Monday. "What makes it worth writing up is what that packet carries: not a readable command, but a short program written in a command language of the backdoor's own design," Reichel said. "Its 23 instructions cover scheduling, several ways to move data, staged file delivery and running code directly in memory. Recovering the encryption key is not enough to understand one of these programs. The internal command language must be reverse engineered as well." In addition to having its own command language, it's also notable that the remote host can be a VMware VMCI target instead of a normal network address. "Taken as a whole, the approach here is consistent with a targeted, well-resourced operation rather than an opportunistic one," Reichel wrote. The malware, hidden inside a 64-bit Windows DLL file, impersonates Microsoft's dpapi.dll, part of Windows' data protection API for protecting sensitive data. It exports the same seven functions as the real dpapi.dll, but attempts to forward calls to a file named dpapisvc.dll, which is not a real Windows component. The file also has a forged ESET Management Agent version resource, and loads via side-loading into ERAAgent.exe, the Windows executable for ESET Management Agent. After confirming that its host process is named ERAAgent.exe, Sleepwalker goes to sleep inside the computer's memory, which also helps it remain hidden from traditional anti-virus tools. Unlike most backdoors, which call back to an attacker-controlled command-and-control (C2) server and start receiving commands, Sleepwalker lies in wait, checking every packet that passes through the network looking for a specific pattern - this is called a magic packet. Once it sniffs out a packet that matches the exact pattern, the backdoor decrypts the data and treats it as a command. "Because the backdoor never sends anything out on its own and does not open any obvious listening port by default, tools that watch for connections to known-bad domains or unusual outbound traffic will not see anything unusual," Reichel wrote. "The absence of outbound connections to known-bad infrastructure does not rule out an infection, either. A machine can be fully compromised by this backdoor while producing nothing at all for a network monitor to flag."

Read more of this story at Slashdot.

โŒ