Star Trek Was Right about Prompt Injection, Sorta

This following statement is a lie: βI am telling the truthβ. Okay, now that itβs just us meatbags, letβs get down to brass tacks. Captain Kirkβs logic bombs couldnβt possibly work on modern LLMs, right? Surely that was just a bit of 1960s silliness from when computers filled rooms and were esoteric magic even to most sci-fi writers?
Well, not entirely, according to a recent article in IEEE Spectrum. While you might not be able to make a data center explode, you certainly can useΒ a lot of tokens by making an LLM overthink with your prompt.
It comes down to the much-vaunted βreasoningβ ability of the new models β which isnβt really reasoning the way we think of it, but does involve breaking the stated prompt down into smaller problems. Thatβs part of what lets the new models tackle such involved tasks as porting MicroPython to the SNES with a prompt like βPlease make this [stuff] work now!β Itβs also a weakness, because with the right prompt you can get that virtual βreasoningβ to tie itself in knots with mutually incompatible smaller steps.
The models seem to be able to break out of it, but they burn a lot of tokens along the way, which is an attack in and of itself if youβre found a way to inject prompts into someone elseβs API. Itβs a little more subtle than what Kirk got up to, but underneath itβs essentially the same thing. At scale, it could serve as a DDoS attack on LLM servers. (Un)Fortunately, modern computers are better designed than their imaginary 23rd-Century counterparts, and thereβs no way to craft a logic bomb into something that will let out the magic smoke.
The scientists and technicians who build Americaβs newest nuclear bomb just finished a critical manufacturing step three months ahead of schedule, the U.S. Department of Energyβs National Nuclear Security Administration announced, marking another milestone in what officials describe as one of the fastest nuclear weapons development efforts since the Cold War. NNSA said its Y-12 [β¦]