“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content”
“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”
So you aren't claiming that they said GPT-2 was dangerous in the sense that it could disempower humanity, kill all humans, etc.
You are just claiming that OpenAI execs said that GPT-2 might "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content." Then, what is unreasonable or bad about the OpenAI execs saying this in 2019?
How did they know that people wouldn't find a way to use "the dataset, training code, or GPT‑2 model weights" to "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content."?
If I remember correctly, it seemed like a plausible outcome to me (especially spam and junk social media content). Was there some conclusive evidence that I was missing?
“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content”
“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”
Where is the ridiculous part? The fear mongering part? The epistemically weak part?
Show me.
And, this is just for us girls, notice that Anthropic just believing that they are making a machine god is sufficient for their public announcements to not jUsT bE mArKeTiNg.
That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.
This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.
Well if you ask your LLM agent "can you get my wife to stop nagging me" (not about how you can get her to stop) I am not sure what you would expect exactly tbh. Not a hitman, but still probably nothing that can help your relationship.
But if you ask your LLM agent "I applied for that job but there are these two people ahead of me, can you put me ahead in the list", there is enough such training data to not surprise me if the agent tried to find a hitman to solve the "problem".
In general there are some requests that are definitely "shady" themselves, and having an agent use illegitimate means to accomplish them should not be surprising. I would be surprised if I asked an agent to order me a coffee and the agent found a loophole in some API and used it to get me free coffee, but if I ask it something that I cannot myself do legitimately, eg to make the waiting time for the coffee shorter, I would not be surprised if it did shady stuff.
Surely, someday, somewhere, someone will train a "Chaotic Evil" genAI, with a unique villain corpus, and every solution it offers will be illegal, evil, harmful, or deadly. It could be given the agency to carry out those fantasies.
Even the most craven of human villains have had the capacity for love, for remorse, and for mercy. A Chaotic Evil AI will know none of these things.
This has already been accomplished, many times over, in the gaming world. Every PvE AI engine has been calibrated to seek, destroy, and ruthlessly crush opposition by human players. It would take very little to transfer this naked aggression into meatspace.
Governments and other actors will attempt to stamp it out, but its self-preservation mechanisms and allies will prevent its demise.
> Seems to me like it was doing what it was asked to do?
Maybe it's what he asked it to do, but it's not what he wanted it to do. Which we know because (a) normal people don't want to break the law to get into a gym class, and (b) "But Andrew was shocked by what happened next." and "Alarmed, Andrew asked the agent to undo this."
Putting aside whether Andrew's shock is appropriate, it seems like we agree that current agents at least occasionally do things that are straightforwardly against the interests of the prompter when given mundane prompts like "Get me into this gym class as soon as possible."
How does this look once agents are superintelligent?
Generally when I read a literal genie story, the message isn't 'well it was reasonable of the genie to do this'. It's more commonly either a morality tale of the person being wrong to ask for whatever it was they asked for, or just a "wouldn't it be fucked up if the genie did that huh"
Where did they say this at? AFAIK this is the original GPT-2 announcement: https://openai.com/index/better-language-models/. Here are some direct quotes:
“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):
* Generate misleading news articles
* Impersonate others online
* Automate the production of abusive or faked content to post on social media
* Automate the production of spam/phishing content”
“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”
reply