Rendered at 17:28:58 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jocoda 32 minutes ago [-]
Does this ban eventually extend to the support chat systems that we see everywhere now? And does that mean that soon we will not be able to insult companies for fear of hurting their feelings?
CoastalCoder 58 minutes ago [-]
Bringing this up here for serious discussion:
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
staticman2 5 minutes ago [-]
It occurred to me that Anthropic is going to continue to say things on A.I. sentience while conveniently always concluding the "solution" to A.I. suffering is always something which never has a huge impact on Anthropic's bottom line.
chrisjj 53 minutes ago [-]
Like banning punching a vending machine means keeping one is slavery?
staticman2 19 minutes ago [-]
If you mean banning punching a vending machine and stating the reason for the ban is the vending machine is alive and suffering, then yes.
CoastalCoder 35 minutes ago [-]
I think a closer analogy would be a vending machine owner saying you can't use the machine if you curse at it.
I would suggest that a big part of avoiding abuse is that it leads to unexpected output. Given a model reflects its inputs, I imagine that abuse leads to unstable output as the model might then seek victim behaviours to better reflect the input.
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
asanineassasin 56 minutes ago [-]
Abuse might be the trigger to get linux kernel quality code.
Quarrelsome 39 minutes ago [-]
I'm reminded of that guy who used AI and it deleted his production database and all the backups. He shared all his chats as some sort of "proof" that Anthropic had screwed him over, but it revealed he had very unhealthy prompting that was possibly a contributing factor to his outcome (obviously the biggest one giving it production keys).
perching_aix 35 minutes ago [-]
Could you unearth that story?
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways.
chrisjj 55 minutes ago [-]
> I imagine that abuse leads to unstable output
How would you know, given you never get any other kind of output?
Quarrelsome 41 minutes ago [-]
Sounds like a user issue to me. I get exceptional output as long as my inputs are well crafted.
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
chrisjj 58 minutes ago [-]
> "This level of anthropomorphisation of AI is harmful. It leads people to believe that it's something it's not."
It is easier to stomach when treated for what it is.
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
My opinion: https://news.ycombinator.com/item?id=50012423
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
Sure:
https://x.com/lifeofjer/status/2048103471019434248
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways.
How would you know, given you never get any other kind of output?
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
It is easier to stomach when treated for what it is.
Just more false advertising.