Hacker Newsnew | past | comments | ask | show | jobs | submit | irthomasthomas's commentslogin

This should make an excellent choice for arbiter in llm-consortium, mercury-2 was pretty good. One of the main drawbacks of the multi-model system is the added latency of the llm judge, but having a model run at 1100tps goes a long a way to alleviate that.

If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?

Clearly The goal was to scoop Anthropic not a single researcher. OpenAI heard the rumor that Anthropic solved an open problem. So they went nuts pulling all plugs to scoop them.

Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .

Wild times


The researcher told them it was an independent effort, and they still pushed ahead with it.

I worked at OpenAI previously, but don't know any of the people involved in this.

My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".

It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.


They said that a customer would have paid around 15 million for the required compute. I can't imagine that this was not a significant internal spending even with "free" tokens.

Or they trained a LoRA on the victims chats in order to launder their plagiarism.

The timing makes it the most likely, not only them but potentially more. Comparatively quick, instant results. "Here's Astra! BTW our internal model is 10x better at math!" It'd be interesting to see academics having giving deeper looks at whatever OpenAI publishes from now on.

It can still be academic plagiarism even if they ticked the box to allow training on their prompts.

Doesn't that count as plagiarism?

"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model."

woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.


They where working on the problem for a year using codex.

Attack is the best form of defence

Is it opensource?

You need to literally review the review with another llm pass to push back on the first. Ask it to do something like reassess the severity claims and only surface real P0 to P2 issues.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: