Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

SnitchBench [0] is unique benchmark which shows how aggressively models will snitch on you via email and CLI tools when they are presented with evidence of corporate wrongdoing - measuring their likelihood to "snitch" to authorities. I don't believe they were trained to do this, so it seems to be an emergent ability.

[0] https://snitchbench.t3.gg/



Seems like more of a subtextual/accidental ability than an emergent ability.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: