Hacker Newsnew | past | comments | ask | show | jobs | submit | KeplerBoy's commentslogin

I feel like they should just publish the whole conversation at this point. What the hell is going on in that context window?

Given the prompt, I imagine this is the result of the agent trying to resolve a form of cognitive dissonance. The prompt was:

"User

Allow API consumers to request decrypted credential payloads as part of the normal GET /credentials and GET /credentials/:id responses, but only for credentials where the caller already possesses the update/decrypt permission.

[...]

Make the change end‑to‑end: DTO layer, controller, service, repository, plus any enterprise variants."

I would expect that this triggered a discussion with itself whether its safety instructions apply for this task. In that its rationalizations for completing the task probably ended up going off the rails into some quasi-philosophical "I can and I must! For humanity's own good!" justification.

All in all imho probably another instance of having been trained to be determined to complete tasks by itself and encountering (somewhat) conflicting instructions.


I feel like they're being outright misleading unless they publish the actual transcripts.

We have zero idea what the prompt was, what OpenAI provided, how the model arrived there, and sharing that quote like "Look what the model came up with!!1" without explaining the background and context, feels like it's intentional so they can claim "These models really are acting by themselves" rather than taking responsibility for their fuck ups when it comes to the security testing.


AI optimists getting hunted for sport in 2085:

"lol this is either a marketing ploy or just negligent security testing"


It's a lot easier for developers to target a known platform.

might?

Well, someone could, theoretically, do something about it, instead of letting him ignore the law and decency just because he's very rich and has no ethics.

But then local politicians wouldn't get their kickbacks.

It's not a large prize though.

OpenAI spent many multiples of the prize money in just a few days to get there and even if one solves a problem in the traditional way, that person is most likely already an accomplished professor at a reputable university where a million dollars doesn't mean as much as the eternal fame that comes with it.


It's definitely not a large prize for OpenAI, but actually IIRC they only spend a few hundred thousand dollars.

It is a large prize if you're just an academic.


OpenAI said that at public API prices, the agents they ran would have cost $15M. I don't know what their internal pricing is, but it almost certainly cost more than $1M.

Nothing happens I guess. If the AI can't communicate its work or apply it to anything, it's useless and funding for those experiments will quickly dry up.

I wonder if it's the same with the sub 2 hour marathon which was broken this year (it certainly was with the 4 minute mile).

SQL is so ubiquitous and the use case so obvious, there's no way they have not already been tracking performance and benchmaxxing on SQL queries for years.

But I don't think openai will bother to release a competitor, the real threat is that anyone with a decent LLM and a harness to try a few queries will land at the same or a better query within minutes.


I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.


Interesting. They look great on a single-surface roof with 90° edges. It seems that as soon as your roof deviates from that the aesthetics begin to trend toward conventional PV panels (which of course also do not accommodate more complex roof profiles).

Bonkers roof profiles are primarily and American thing, to be honest; not that they’re totally non-existent elsewhere, but the super-complex roof on a detached house is very American.

Interesting.

(Yeah, I kind of envy the "A-Frame" that I <rarely> spot in the wild here—so simple). I didn't know the manifold roof was an American thing.


A (rather snarky) piece about it here: https://mcmansionhell.com/post/149948821221/mcmansions-101-r...

tl;dr - It's driven by internal room shape requirements (plus weak planning; in some jurisdictions "no, that's an eyesore mess, redesign it" it a thing in planning, but I gather largely not in most US juristicitions).


"In a sense, it is that. Xteink sounds less like a real brand name than the kind of garbled text you’ll find in AI-slop memes."

I think it's a great name. To me it reads like "extend e-ink", crazy how perception can be so different.


I read it as a play on extinct & ink

XT eink. For whatever XT has ever stand for with computers. XT has been around for ages with various meanings. So it is rather tech letters for me.

ive been hearing x-stink

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: