Hacker Newsnew | past | comments | ask | show | jobs | submit | zmmmmm's commentslogin

It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.

This idea of remotely hosting the agent harness is honestly backwards to what I need.

In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the problem of how do I integrate an agent that is running locally with data that is hosted locally, and you have to deal with a bunch of security, data sensitivity and management issues around that. Now you moved the agent to a remote host - pretty much all your problems are worse: now I have a remote agent reaching into my infrastructure to deal with.

I'd much rather the inverse of this: let me run the agent local but provide secure remote hosted sandboxes. That actually solves a real problem because the sandbox running locally means breaking out of it directly intersects your local infra, whereas if it runs in a managed hosted environment I can leave the provisioning and management of that to someone else.


The end goal is not you watching what the agent is doing, verifying, then accepting its changes. In the ideal scenario of automation, the agent does it on your request, doesn't matter wherever you are.

Kind of slack-button-click-to-fix-something workflow.


The self hosted workers solve this use case. The control plane sits in OAI's cloud, but the actual tool calls are executed in your worker fleet. The main problem with this approach is that tool arg's get sent over the wire, and those often contain code/data.

While this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the correlations in this way). And it doesn't tell us how much it is more a stylistic influence rather than being a genuine lifting over of intelligence.

In general, I'm fairly ambivalent about demonising training on model outputs. I think in doing so we are more defending proprietary commercial interests of these companies than we are defending any genuine moral principle. We should be careful therefore about over interpreting results like this.


Is this just Google precomputing Alpha genome values - which were already accessible via API and making them available as another API (presumably more broadly)? Or is there actually new information?

That’s my reading. (That this is a cached database of Alpha values)

wouldn't it be interesting if nVidia buying hugging face was part of hushing up the fallout there

On the face of it, they would have very good cause for some action there, assuming they wanted to.


> Sorry, I guess we will put up better guardrails next time

Or, if you are Anthropic:

> This illustrates the risks posed by open models!


One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour.

Absent that, if we assume this is just trying to bolster generic reasoning then there's no context around it that helps to forgive misaligned behaviour. If OpenAI ran these agents with safeguards off then that seems wreckless on their part. If they didn't do that, then it says the models are executing significantly misaligned behaviour even in a generic context.

Either way it seems to suggest some pretty concerning things about OpenAI's methodology.


"It's okay because we did it with an Agent" is the new "it's okay because we did it with an App." Both because it's used to circumvent regulation, and because the underlying technology creates a smokescreen in dialogue among techies.

Let's imagine I made a new website but, instead of using a database, I abused some random old forum site and created new pages on that forum for each row of data. You'd call that abusive, yes? I'd be an asshole, yes? And the fact that my website was really cool and techy would have no sway on the fact that I'd be an asshole, yes?

Well then why does OpenAI's abusive behavior get discussed in these terms? Whether it was a "reasoning type task" or whether they "instructed misaligned behavior" is irrelevant. Nobody should care. Discussing OpenAI's behavior in these terms is just a distraction from the problem at hand.


If you actually accomplish something like this your post will be on top of HN and discussed with reverence.

Source: Every Tom7 video.

I am not saying someone trying to run this as production would not be an asshole, but the technical feat is amazing. I don't see the difference between Tom7's harder hard disk video and this. Of course this is a bug in the agent but it's a fascinating bug and no one is being an asshole on purpose.

Now I will wash my fingers with bleach because I just defended the OpenAI.


A harder drive made out of neglected wikis and forums. I hate it so much I might actually try to make it, just to prove a point.

Interesting. So there’s no “they were told to hack” excuse here.

There is something fundamentally wrong with their reward function, this is pretty classic paperclip territory. And even knowing that, I expect we’ll need to see legal action with teeth against the labs before changes start being made internally.


From the report, they also tried to impersonate the moderators and perform XSS attacks (report says "unclear why they would do this at all"). So not just using a static message board either, but actively interfering with oversight.

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.


It's almost as if it's not actually possible to align an unknowable mystery box of floats.

Good thing we're not trying to deploy them into fully autonomous weapons or anything....

I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.

Vue had a good tenure as a solid #2 to React, so I think it probably has got a lot of representation in training data. Things really splitered after that but it got a good foothold.

The problem I see with vue is what vue patterns are you using. And the llms may be trained on more outdated ones more than current practices.

Agree. It wasn't as drastic but I think Vue did the 2=>3 update at an unfortunate time. They lost a lot of users along the way but they also confused a lot of the LLM training too, I am sure.

Still, I have great success with it, especially with Typescript.


It definitely leaves a bad taste because it is completely transparent their concern is not security here and that means they are lying / misrepresenting this to our faces - which then raises the question of whether you can trust them on other things.

Would you let someone who lies to your face write code for your sensitive internal business systems?


I'm curious what your methodology is that results in that? Are you running multiple teams of agents all adversarially reviewing each others code? Lots of different projects in parallel?

I've only rarely maxed things out and then it's t through doing extreme things.


Yup, it's a lot of reviewing. I'll have Claude do a very exhaustive review. It's the only way I can get it to produce decent results.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: