Given the prompt, I imagine this is the result of the agent trying to resolve a form of cognitive dissonance. The prompt was:
"User
Allow API consumers to request decrypted credential payloads as part of the normal GET /credentials and GET /credentials/:id responses, but only for credentials where the caller already possesses the update/decrypt permission.
[...]
Make the change end‑to‑end: DTO layer, controller, service, repository, plus any enterprise variants."
I would expect that this triggered a discussion with itself whether its safety instructions apply for this task. In that its rationalizations for completing the task probably ended up going off the rails into some quasi-philosophical "I can and I must! For humanity's own good!" justification.
All in all imho probably another instance of having been trained to be determined to complete tasks by itself and encountering (somewhat) conflicting instructions.
I feel like they're being outright misleading unless they publish the actual transcripts.
We have zero idea what the prompt was, what OpenAI provided, how the model arrived there, and sharing that quote like "Look what the model came up with!!1" without explaining the background and context, feels like it's intentional so they can claim "These models really are acting by themselves" rather than taking responsibility for their fuck ups when it comes to the security testing.
Well, someone could, theoretically, do something about it, instead of letting him ignore the law and decency just because he's very rich and has no ethics.
OpenAI spent many multiples of the prize money in just a few days to get there and even if one solves a problem in the traditional way, that person is most likely already an accomplished professor at a reputable university where a million dollars doesn't mean as much as the eternal fame that comes with it.
OpenAI said that at public API prices, the agents they ran would have cost $15M. I don't know what their internal pricing is, but it almost certainly cost more than $1M.
Nothing happens I guess. If the AI can't communicate its work or apply it to anything, it's useless and funding for those experiments will quickly dry up.
SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.
But I don't think openai will bother to release a competitor, the real threat is that anyone with a decent LLM and a harness to try a few queries will land at the same or a better query within minutes.
Interesting. They look great on a single-surface roof with 90° edges. It seems that as soon as your roof deviates from that the aesthetics begin to trend toward conventional PV panels (which of course also do not accommodate more complex roof profiles).
Bonkers roof profiles are primarily and American thing, to be honest; not that they’re totally non-existent elsewhere, but the super-complex roof on a detached house is very American.
tl;dr - It's driven by internal room shape requirements (plus weak planning; in some jurisdictions "no, that's an eyesore mess, redesign it" it a thing in planning, but I gather largely not in most US juristicitions).
reply