As the agi nurse fails to wipe his ass the intelligent HN commenter shouted: "You're moving the goalposts. You never said it needs to wipe my ass SUCCESFULLY"
Your coding agent is also a very real-world version of the paperclip optimization thought experiment, yes. Have you never seen it reward hacking? Editing tests to pass instead of fixing the code?
It knows what you want, it can even tell you, and it absolutely doesn't give a shit.
> I think Congresspeople hearing that EU AI Act is forcing secret codes into the infrastructure of American technology across all industries is sufficient.
So, you think it's good to disconnect words from their actual meanings (lie) to low-information people! I doubt this will do much to congress, but it certainly teaches us something about the sort of mind who would suggest it.
You're assuming they're training the model to maximize the watermark signal, on top of already adding the watermark. I suspect that would hurt model performance quite a lot, and simply be unnecessary... the watermark tech works well enough as it is.
As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations.
> As far as I know, anthropic aren't intrinsically motivated by watermarking (if anything it hurts sales, and seems indifferent to safety(?)) they're simply doing it to fulfill the EU obligations.
They are. They want to reduce the amount of LLM generated text they feed into their next model training.
Also, how would you watermark a sentence with just 3 words for an example? This exactly why it became so verbose.
That would be a terrible tradeoff. The ship has already sailed and a lot of public AI content will not be their own. Deliberately making their product worse to reduce identifiability of AI inputs by 25% just doesn't sound worth it to me. Is that what you would pick if you were in charge of anthropic and wanted to maximise the company's product?
And what wisdom do you think they would be missing if unable to distinguish three word written pieces? Keep in mind that most sources are not inherently trustworthy just because they rate as human written, too. You need some other way to rate text in all cases.
Perhaps, but there are certainly now catchphrases and words that can indicate it was written with AI i.e. load-bearing, idempotent, etc. Style and structure are in and of themselves, a fingerprint.
https://panda.substack.com/p/salvage-fuzing-and-co-orbital-n...
reply