Hacker Newsnew | past | comments | ask | show | jobs | submit | TedDoesntTalk's commentslogin

You think the hugging face incident was a stunt? Can you explain?

OpenAI started fearmongering way back with GPT 2, arguing that model was too dangerous to release freely. That model was barely coherent enough for using it as a twitter bot. Anthropic just hopped onto that later. Conveniently, calling for regulation now would ease the competition from open Chinese models, opening the chance for both companies to eventually reach positive ROI, with consumers paying the price.

The burden of proof that this isn't just a publicity stunt again is squarely on them.


> The thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses.

So we decided to build that to validate and confirm our fears.


I was happy surprised to see the Swedish term "skräckblandad förtjusning" in the article. I've always loved that expression and I think it's a wonderful explanation to this seemingly crazy behavior.

"Wow, that's terrifying! So exciting!"


I get your point, but it’s obvious that everybody is going to do this. So building a benchmark isn’t likely to directly move capabilities. OTOH it might give some good signals on required changes in training recipes.

People are still debating the liabilities associated with AI development.

If your agent hacks the CIA, people want to blame the AI lab.

but...

If your agent spends $100k on tokens, then thats a user error. If your agent spends $10m on a shopping spree, then that is also a user error?


Yes, of course? Put a spend cap on your card like you would on your API key.

At some point these systems will get certified as fiduciary agents but they sure as hell aren’t claimed to be that now.


Yes, let’s not do any experiments at all. Then we’re sure to remain ignorant.

That always works well.


… but he’s not using a “hacking model”


Great point.

Is this why Oregon Trail Deluxe freezes on Mac?

https://oregontrail.ws/games/the-oregon-trail-deluxe/play/


I will make miniature dioramas and photograph them.

At that point you've earned the fruits of your deception, just like the tricksters who spent time doing physical photo editing.

Miniature dioramas wouldn't be size appropriate. Apple could detect faces/cars/other common objects of ~known size and verify -- or even just dump depth map for anyone to check.

Even easier is to just take a picture of an AI-generated picture.

Propagation

Sophisticated attacks will always leverage multiple vulnerabilities. That’s why you have to think of any vulnerability holistically, not in isolation.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: