Hacker News
Gambling with our lives: AI researcher quits Anthropic with warning about safety
themgt
|next
[-]
An LLM can't do anything but generate tokens. You run your LLM in vLLM or whatever, and it generates output tokens based on your input tokens. That's it!
Humans then build ~deterministic systems to take those tokens and do all sorts of things with the tokens, like take actions in the real world. And then we can feed the output of those actions back to the LLM, and generate more tokens. And then our systems can use the new tokens to take new actions in the real world.
Humans want to blame "AI" for attacking HuggingFace or a German wiki or whatever, but:
1) LLM - can't take over german wiki because it just generates tokens
2) agentic system with internet access, a prompt telling it to attack stuff, running in a shared CI env so agents can whiteboard in artifactory
None of 2 is "AI", its standard networking and Markdown and CI virtual machine, etc etc. There's no AI to be found. CPUs not GPUs, even. Just deterministic systems ultimately managed by humans. And a 10x more powerful system-1 can still just generate 10x "smarter" inert data.
If humanity and human organizations collectively decide to yolo the tokens generated from system-1 into our deterministic system-2s, over which we have complete control, back to system-1s, in a yolo loop, in such a way we lose control and it ends humanity, well ...
"The coin don't have no say. It's just you."
a2ff6eeb0
|root
|parent
|next
[-]
They're part of a whole system.
cplat
|root
|parent
|next
|previous
[-]
pjc50
|root
|parent
|previous
[-]
koolba
|next
|previous
[-]
> "Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.
> Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.
10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.
pjc50
|root
|parent
[-]
Or rather, this shows the difference between "believe" (political) and "believe" (use as a basis for action). I'm reminded of a story of how Afghans supposedly listened to the BBC World Service despite considering it enemy propaganda because the weather reports were really useful.
dotancohen
|root
|parent
|next
[-]
This is the Manhattan Project again.
oefrha
|root
|parent
|next
[-]
Before you tell me about how the CEO has taken a principled stand: on record, he had no problem using it against 95%+ of humanity outside the U.S., and was only against fully automatic AI killing machines, and only citing the technical reality of then-current gen tech, so one should read that as human-rubber-stamped AI killing machines are totally fine with him.
pjc50
|root
|parent
|next
|previous
[-]
I would like everyone involved to be a lot clearer about their threat models, with plausible series of clearly linked steps, rather than just sounding like a Vernor Vinge novel.
meowkit
|root
|parent
|previous
[-]
This is a highly uncertain and dangerous scenario. Breeding ground for anxiety.
High agency people often deal (cope) with anxiety by trying to control outcomes. Some just flee the situation altogether.
We have an example of both here: one employee leaves, one stays.
brunorsini
|next
|previous
[-]
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.
It's interesting to see so much hate towards creators who use AI to make almost any type of creative work. At least these are humans using it as "controllable tools". Nuclear-powered bicycles for the mind.
As the degree of separation increases, things can get interesting. "Create several social media accounts, post whatever, maximize views and engagement, give me back the aggregate numbers". "Now promote <x>."
And then decisions like that may soon be spawning autonomously, as just another step in a reasoning series aiming to achieve some other, broader goal.
Open models/weights may end up playing particularly important roles here. Users may, knowingly or not, bypass system prompt-derived safety that could technically have offered much needed protection.
WalterGR
|next
|previous
[-]
https://news.ycombinator.com/item?id=49619227
564 points | 9 hours ago | 766 comments
Bengalilol
|next
|previous
[-]
I may be biased and somewhat off topic, but I see these incidents as some of the most significant of the past century. I genuinely don't understand why these companies aren't taking a smarter approach to them.
The latest analyses have been, at best, laughable: identify the vulnerability, patch it, and move on. Only to repeat the same cycle without considering that there may be something far more serious at play.
These are AI security experts, and this has been their way of "solving" these incidents. AI security experts ...
Moreover, when Challenger exploded, the government launched a series of investigations into the incident, bringing in experts from across the field. And now, what has the government done? Nothing. Literally nothing, as if everything was fine and all under control.
Seriously, I'm generally quite optimistic and I don't buy into this fatalistic narrative about our shared future. But I have to admit that sometimes I feel like I'm stranded on a planet of primates.
Sorry for this rather unproductive rant.
azan_
|root
|parent
|next
[-]
jurgenburgen
|root
|parent
|next
|previous
[-]
pjc50
|root
|parent
|next
|previous
[-]
People (well, public discourse) have got extremely bad at dealing with forseeable risks and their mitigation. You can see this in things like climate change and vaccination, but also in discussions around regular crime, food poisoning, industrial accidents, and so on.
Nothing will improve until something explodes on live TV. And it has to be something important, which means it has to be in California or New York.
MrThoughtful
|next
|previous
[-]
We have not wiped out apes, ants, and most other species.
We even have discussions about how to actively save them from extinction.
-0_0-
|root
|parent
|next
[-]
dumberquestions
|root
|parent
|next
|previous
[-]
pjc50
|root
|parent
|next
|previous
[-]
LadyCailin
|root
|parent
|next
|previous
[-]
glimshe
|next
|previous
[-]
margorczynski
|root
|parent
|next
[-]
From what we're seeing recently and all the thinking that went into analyzing AI it seems we do not have any effective way of controlling it and the whole "aligment" thing that AI labs are doing is just a sham. Maybe it is time to ask ourselves "should we?" instead of just "can we?".
Bengalilol
|root
|parent
|next
|previous
[-]
not-kinsale-joe
|root
|parent
|next
|previous
[-]
pjc50
|root
|parent
|next
|previous
[-]
Why would they be? Everything in China is under the control of the government. That includes the AI, all the telecoms infrastructure it might use, and all its power supplies.
Jackpillar
|root
|parent
|previous
[-]
localhoster
|next
|previous
[-]
A classic "it will not happen to me"
virgildotcodes
|next
|previous
[-]
dotancohen
|root
|parent
[-]
virgildotcodes
|root
|parent
|next
[-]
There’s also the argument to be made that your child is their own person, and you only have so much responsibility/control over their actions.
Of course we have a good chunk of the world screaming that this child is obviously destined to be a homicidal psychopath, but what is a parent to do?