Reading view

There are new articles available, click to refresh the page.

Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost

The performance gap between frontier AI models from US tech companies and the best open-weights models from Chinese companies has closed to just 4.4 months, according to a Mozilla report. That explains why many companies are shifting to the significantly cheaper open models for routine work—and helps reveal a narrow band of workloads where frontier models are worth the cost.

Most organizations should ideally be using open models as the default for the majority of their work, according to the latest State of Open Source AI report from Mozilla, published on September 15 and shared with Ars prior to publication. The report highlights how a leading open model, Moonshot AI’s Kimi K3, achieves a composite AI performance score on the Artificial Analysis Intelligence Index that is just three points behind Anthropic’s Fable 5 closed frontier model, all while costing just 30 percent of the latter.

“[A Closed model] earns its premium in a few places: expert professional work, high-intensity retrieval, and long context,” Raffi Krikorian, chief technology officer at Mozilla, said in an email to Ars. “We see the decision to pay for closed [models] as workload-specific rather than organization-specific.”

Read full article

Comments

© Imen Ben Youssef / Hans Lucas / AFP via Getty Images

Claude users found ways around safeguards for bioweapons research

Anthropic said it stopped multiple attempts by scientists this year to use its technology for research that could help develop biological weapons, as experts increasingly fear the threat that AI poses to public safety.

The startup gave five examples of times actors “circumvented controls” and made other efforts to “obfuscate” the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran.

“We hope that by sharing these examples, we spark a conversation within the AI industry and with governments about emerging biological risks and how best to counter them,” Anthropic said in a report about efforts to use its models for malicious activity.

Read full article

Comments

© Getty Images | picture alliance

Four major AI models suffer rare overlapping downtime

Cloud-based AI models operated by OpenAI, Anthropic, xAI, and Google suffered a rare and overlapping set of significant service interruptions over a period of hours Thursday morning.

Anthropic first reported a "partial outage" related to "elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5" at 9:23 am (all times Eastern). The company reported that it had "identified the cause" of the error roughly 15 minutes later, before reporting that "a fix has been deployed" and the issue was resolved by 12:16 pm. A separate incident report indicated "elevated errors on requests to Claude Sonnet 5" for a brief period just after noon.

OpenAI, meanwhile, reported "elevated errors across ChatGPT and Codex" were resulting in "degraded performance" as of 10:43 am Thursday morning. A mitigation put in place a little more than half an hour later led to the issue being marked as "resolved" by 12:55 pm.

Read full article

Comments

© Getty Images

“Zlibrary my beloved”: Anthropic staff chats extolling piracy cited in Sony suit

Some of the world’s leading music publishers think that Anthropic got off too light in a historic settlement where the Claude maker paid authors $1.5 billion after admitting to pirating more than 7 million books to train AI.

“$1.5 billion is obviously not a large enough settlement to deter infringing conduct by a company that has parlayed such mass infringement into a staggering $2-trillion-dollar valuation,” music publishers said in a lawsuit filed Friday.

Music publishers suing Anthropic include Sony, EMI, and Warner Chappell. They alleged that Anthropic’s illegal torrenting also included “thousands upon thousands” of their copyrighted musical compositions.

Read full article

Comments

© Steven Errico | Photographer's Choice RF

Bill Gates in his own words: How he’s using AI, and why he’s worried about the future

Bill Gates, shown here in April 2025, released a memo this week warning that the world isn’t ready for AI. (GeekWire Photo / Kevin Lisota)

This week on the GeekWire Podcast: Bill Gates published a new essay warning that the AI industry is crossing the safety lines it set for itself, and that nobody is preparing for what’s coming. At age 70, he also uses AI more than most people half his age, and he finds it enthralling, as you’ll hear on this week’s show, with highlights from our interview with him.

Along the way, we dig into his three proposals: new institutions for managing the transition, a category of jobs reserved for humans, and a tax on the use and purchase of AI and robots.

The change in his own tech usage: “I joke with people that I used to have Claude-like people that I would send email to, but they were so slow, and there were some topics they didn’t actually know. … It’s three a.m. I want to understand sodium batteries, and now there’s no reason to go to sleep. Here we go. Yeah, it’s crazy.”

How he uses AI specifically: “If you’re a curious person, this is a mind-blowing time. When I’m working on malaria, nutrition, my poor humans that I work with always get these long conversations from me, where I paste in — me, Claude, me, ChatGPT. Sometimes I do it if there’s three of us: Claude, ChatGPT and me, debating these things.”

On where personal agents are headed: “We will get to a point where you won’t buy things yourself. You just won’t. … You won’t go to those applications. You’ll just go to your personal agent. … From a productivity point of view, we are in heaven.”

What has surprised him: “I was shocked by ChatGPT, and I was shocked by Claude Code. Those are both things where I went, oh my God. … I did not expect that a statistical machine would essentially learn to read, and the idea that the code is better than human code. Those are two stunning thresholds.”

On writing this essay: “It’s very unnatural for me to think that innovation may be a net negative if it’s not managed properly. The more I wrote the memo, the more I was like, Jesus, we really need to get our act together here. Even though this may come across as negative, that’s the truth. If we don’t step up, the negatives will substantially outweigh the positives.”

What AI leaders say privately: “You’re in this perverse period right now where people in the AI industry who are willing to say that AI might have some negative effects are told, ‘Hey, you’re hurting our PR while we’re trying to raise trillions of dollars.’ … I know they’re all worried. Or all of them that I know, which is basically everybody but Elon.”

On losing control of AI: “The wake-up for the memo is that the bad stuff thresholds are all being crossed. Even lack of control that I thought would be many years from now, we’re seeing lack of control. … These are people who are super expert on the thing, going, well, maybe we won’t be able to control these things. What kind of risk have we chosen to run here?”

On how fast robots are coming: “What’s weird about AI is it’s better at doing jobs across the entire economy, including physical jobs when the robots come — which you can guess when that is, but my view is it’s only a couple of years.”

Is he still an optimist? “I don’t think being pessimistic is helpful. I do think, wow, this is sure an interesting time. I’m the guy who in my 30s thought people in their 50s or 60s didn’t understand anything. So it’s kind of bizarre if a guy who’s 70 comes and writes a memo that’s actually helpful. … But I am very concerned. And honestly, when you get people one-on-one, so are they.”

Related headlines and links

Subscribe to GeekWire in Apple Podcasts, Spotify, or wherever you listen.

Edited and produced by Curt Milton. Music by Daniel L.K. Caldwell.

Trump blacklisting of "woke" Anthropic deemed illegal by federal judge

The Trump administration's blacklisting of Anthropic was illegal, a federal judge ruled in an order vacating government directives against the use of the firm's AI technology.

The government illegally retaliated against Anthropic by designating it a supply-chain risk to national security, said yesterday's ruling by Judge Rita Lin in the US District Court for the Northern District of California. The maker of Claude AI technology was barred by the US after it refused to drop restrictions on the use of its products for lethal autonomous warfare and mass surveillance of Americans, the ruling said.

"The undisputed record shows that the challenged actions constituted unlawful retaliation in violation of the First Amendment," Lin wrote in an order that granted key portions of Anthropic's motion for summary judgment.

Read full article

Comments

© Getty Images | picture alliance

Claude Plays DOOM

Large language models (LLMs) are generally thought of as machines that accept textual prompts and spit out textual content. However, if you’re creative in the way you interface with them, you can get them to do a wider range of tasks. For example, [Andrea Ricci] figured out how to get one to play DOOM.

For this project, [Andrea] began by porting the game to the SCINTIX P4. It’s a rather interesting device, being a single board designed in the Raspberry Pi CM4/CM5 form factor, but carrying an ESP32-P4 and an ESP32-C6 instead. The game runs on the P4 and is displayed on a 1024×600 MIPI DSI panel, but it’s only stepped through a few frames at a time. These frames are then passed to Claude Sonnet via a WebSockets setup. With only the same information as a human player would get, the LLM has to figure out what it’s looking at, and then respond with movement and fire commands to play the game.

It’s quite interesting to watch the system play—the LLM mostly accurately describes the game world, navigates down corridors, opens doors, and shoots at enemies. There is a bit of work behind the scenes to enable it to see and understand the game world—namely, using a depth fan across the field of view so it can figure out where walls are and how not to bang into them. There’s also an ASCII automap used to allow the system to keep track of where it has already been. But fundamentally, the LLM is playing the game without any other sort of additional assistance.

We’ve seen some other great ways in which AIs have been whipped up to play various games, like Trackmania.

❌