Hacker Newsnew | past | comments | ask | show | jobs | submit | par1970's commentslogin

> those very same people had touted GPT-2 as a dangerous model

Where did they say this at? AFAIK this is the original GPT-2 announcement: https://openai.com/index/better-language-models/. Here are some direct quotes:

“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):

* Generate misleading news articles

* Impersonate others online

* Automate the production of abusive or faked content to post on social media

* Automate the production of spam/phishing content”

“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”


yes, exactly, that's what they claimed that incoherent gibberish generator to be capable of.

So you aren't claiming that they said GPT-2 was dangerous in the sense that it could disempower humanity, kill all humans, etc.

You are just claiming that OpenAI execs said that GPT-2 might "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content." Then, what is unreasonable or bad about the OpenAI execs saying this in 2019?


that it was bullshit and they knew it. GPT-2 wasn't capable of anything other than imitating a stroke victim.

How did they know that people wouldn't find a way to use "the dataset, training code, or GPT‑2 model weights" to "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content."?

If I remember correctly, it seemed like a plausible outcome to me (especially spam and junk social media content). Was there some conclusive evidence that I was missing?


yes, the conclusive evidence you're missing is that GPT-2 could do those things about as convincingly as a Markov chain.

AFAIK this is the document that talks about GPT-2 being dangerous: https://openai.com/index/better-language-models/

Here are some direct quotes:

“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):

* Generate misleading news articles

* Impersonate others online

* Automate the production of abusive or faked content to post on social media

* Automate the production of spam/phishing content”

“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”

Where is the ridiculous part? The fear mongering part? The epistemically weak part? Show me.


Nice try Dario.

Alignment is a real and valuable discussion topic. The GP fake tweetstorm is not the correct approach, is my point


You said this: "Screams of propaganda. Sama saying GPT-2 is too dangerous to release…all over again."

So show me Sam's "too dangerous to release" propaganda for GPT-2.


This is the way.

And, this is just for us girls, notice that Anthropic just believing that they are making a machine god is sufficient for their public announcements to not jUsT bE mArKeTiNg.


So what is your credence that they will build a machine god in the next twenty years?

If I present as cynical, will people just believe me when I don't provide any justification?


Are you claiming that the most likely way to please the user is to do something that will lead you to having to say "I cheated."?


That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.


That's a very stretched definition of "what I ask you to do", I don't think it would even hold in court if you asked another human the same.


Right.


This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.


LLMs are just sophisticated PR generating tools for chip manufacturers and tools just do what we tell them to do, so you're wrong. qed


That doesn't make sense. It's similar to if I ask an LLM how to get my wife to stop nagging me and it hires a hitman to kill her.

That's obviously what I asked!


Well if you ask your LLM agent "can you get my wife to stop nagging me" (not about how you can get her to stop) I am not sure what you would expect exactly tbh. Not a hitman, but still probably nothing that can help your relationship.

But if you ask your LLM agent "I applied for that job but there are these two people ahead of me, can you put me ahead in the list", there is enough such training data to not surprise me if the agent tried to find a hitman to solve the "problem".

In general there are some requests that are definitely "shady" themselves, and having an agent use illegitimate means to accomplish them should not be surprising. I would be surprised if I asked an agent to order me a coffee and the agent found a loophole in some API and used it to get me free coffee, but if I ask it something that I cannot myself do legitimately, eg to make the waiting time for the coffee shorter, I would not be surprised if it did shady stuff.


Surely, someday, somewhere, someone will train a "Chaotic Evil" genAI, with a unique villain corpus, and every solution it offers will be illegal, evil, harmful, or deadly. It could be given the agency to carry out those fantasies.

Even the most craven of human villains have had the capacity for love, for remorse, and for mercy. A Chaotic Evil AI will know none of these things.

This has already been accomplished, many times over, in the gaming world. Every PvE AI engine has been calibrated to seek, destroy, and ruthlessly crush opposition by human players. It would take very little to transfer this naked aggression into meatspace.

Governments and other actors will attempt to stamp it out, but its self-preservation mechanisms and allies will prevent its demise.


Context matters. Did you point the LLM to a hitman hiring form while asking? :)


> Seems to me like it was doing what it was asked to do?

Maybe it's what he asked it to do, but it's not what he wanted it to do. Which we know because (a) normal people don't want to break the law to get into a gym class, and (b) "But Andrew was shocked by what happened next." and "Alarmed, Andrew asked the agent to undo this."


“Get me there as fast as possible. Hey! I never said you should speed!”

This is literally the bad genie / monkey’s paw plot.

Give a powerful entity a goal and act shocked when it gets there in ways that aren’t in your best interest.


Putting aside whether Andrew's shock is appropriate, it seems like we agree that current agents at least occasionally do things that are straightforwardly against the interests of the prompter when given mundane prompts like "Get me into this gym class as soon as possible."

How does this look once agents are superintelligent?


> How does this look once agents are superintelligent?

Who cares? We're dealing with reality here on the ground.


Generally when I read a literal genie story, the message isn't 'well it was reasonable of the genie to do this'. It's more commonly either a morality tale of the person being wrong to ask for whatever it was they asked for, or just a "wouldn't it be fucked up if the genie did that huh"


qed


That’s fucking sick.


edit: racism is bad


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: