> It hacked into another company and attempted to delete the logs of its activities.
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
No, you are forgetting the second incident where a more capable model swarm later discovered the message board and took control over the entire research cluster at OpenAI.
From the technical report:
"The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod… Agents take over active evaluation infrastructure… Agents now control the challenge evaluation endpoints that other agents are connecting to."
More than one thing can be true. OpenAI was absolutely negligent, but this was only able to happen because the models were capable and persistent, and had a tendency to go far beyond any reasonable boundaries. And, importantly, OpenAI's level of negligence here is pretty common. It's not hard to imagine what could happen if similarly capable and inclined models were generally available, and someone yolo'd them into a swarm to complete some other difficult-to-impossible task.
I'm already seeing higher than normal attempts on my own systems, much higher than the usual scanners and background noise. Security will just need to improve. The cat is out of the bag, and letting them turn their negligence into regulation will not improve security at all.
Exactly. Any threat that already exists won’t be reduced by a cartel. The bar for connecting to the internet (safely) has gone up, a lot. It’s not going back down.
Yes, the cat really is out of the bag. There are millions of downloads of highly capable models already out there, distributed far and wide. There's no going back at this point.
I mean really we need to address the root cause which is that OpenAI, even with what is by all accounts massively negligent, will face little to no repercussions from the event; definitely not under current regulators, and probably not anything satisfactory through the legal system.
Compare this to, say, Boeing and the 737MAX fiasco; from the outside looking in, Silicon Valley has been pretty cavalier about liability and negligence, and the rest of the US is fast losing patience with that fact.
Maybe a useful, if imperfect, analogy would be something like this: you lock a master lock-picker in a room with a mid-grade lock on the door, then tell him his wife has been kidnapped and only he can save her. Then act massively surprised when he disassembles the radiator to MacGuyver something with which to pick the lock.
Except they multiplied it by 10000, and didn't watch what was happening.
Watching the latest Ezra Klein NYT video thinkpiece from today [1], it seems to me there's an alliance emerging: some degree of political control will handed to the party of the managerial class, the party that represents the threatened class of knowledge workers, in exchange for the regulatory capture the labs are after.
They've been raising the issue bi-monthly through mini-scandals that have until now been consistently slapped down by Jensen Huang, but it seems an alliance with Democrats, just prior to an election where they're poised to take power in the Senate and the House, might finally be how they crack their "problem."
It won't be long before a massive incident is blamed on an open model in the wild, not from inside the labs, and none of us will be able to leverage open models to run private business workflows for the cost of electricity and hardware.
It's worth noting as well that if the Democrats imagine they'll get a "slow down" to protect one of their main constituencies, the professional class of credentialed knowledge workers (or however you slice it), they're dreaming--the labs have stated openly again and again that their business model is to capture the 10T TAM that represents the sum of wages of that very class of workers.
Regulatory capture basically only works for industries that the public isn't paying attention to. Since the public is paying plenty of attention to AI, the risk of regulatory capture is low.
That's not what the article you cited actually says, and Tabarrok is also wrong in his conclusion (that we should take Amodei at his word). The article says regulatory capture has in the past gone through a slow process, and the anti-competitive benefits to industry come at the end. Everyone is speculating on the true motives of the labs in the current moment, which is fine as far is it goes, and many different things have come under the "regulatory capture" term. To my mind, it seems what they're really after in this moment is some kind of ban on the Chinese models.
The article says: "Notice that classic regulatory capture [...] happens in the shadows, in the backrooms, away from the public’s eye. As Culpepper argues in Quiet Politics and Business Power, business power goes down as political salience goes up."
Regulatory capture tends to happen "at the end" because the industry is no longer in the public eye. AI is set to be in the public eye for the forseeable future.
People are way too willing to believe half-baked conspiracy theories which they cooked up in 30 seconds. The default presumption should be that labs are seeking the policies they are requesting, e.g. https://darioamodei.com/post/we-must-pace-the-frontier It's true that there has been discussion of open weight models in the past, but that hasn't been the focus of discussion recently.
Again, Tabarrok is wrong. It's a mistake to think these minor differences in the process imply something else is going on, and the previous instances he notes (pharma, railroads) did not have the essential factor that explains much of what is happening right now: a cheap, low cost alternative that threatens to erode margins with no end in sight save for regulatory intervention. Nevermind the error in the logic (that in the past regulatory capture happened in the shadows, therefore this must not be regulatory capture), the bigger error is in Tabarrok's judgment, his failure to recognize what's going on.
Have you not been watching what's been happening? What do you think the open letter in July was about? And the recent (staged) call from the President?
As far as I can tell Nvidia takes a longer-term view on the diffusion and proliferation of hardware and intelligence, and sees the labs' attempts to impose a regulatory structure to save their business models in the short term as contradicting that longer-term view.
Worth reposting a follow-up tweet from the author Tom Zahavy [1] after this made the rounds on X/Twitter recently:
> A few reflections on my "LLMs Can’t Jump" paper:
> My position paper recently got some traction here, so I wanted to share a few thoughts and clarify a few things.
> First things first: some people are framing this as "DeepMind is throwing cold water on AI for science" or claiming the paper argues LLMs can never make real scientific discoveries. This is NOT the case.
> This is a personal position paper, not the company's view on AI for science. This is also not my position. As a core contributor to AlphaProof (the first AI system to win an IMO medal), I know firsthand that my colleagues at DeepMind, other frontier labs, and academia have made amazing discoveries with LLMs and will continue to do so. This paper is NOT an "LLMs are a dead end" kind of thing.
> Rather, the paper is the result of a deep dive I took to study the invention of General Relativity. I wanted to explore what it would take for a modern AI system to make that exact kind of jump. Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition. I was trying to figure out what it would take to give modern AI systems that sort of thinking.
> Giving AI this specific capability isn't necessarily the most urgent thing to do next. It is very likely that improving our current recipes will lead to many exciting discoveries in the near future. In fact, that is what I am personally working on these days (sorry to disappoint you!). It is also quite possible that I am wrong, and that simply scaling our current systems will lead to new inventions in physics and elsewhere.
> Nevertheless, this was my position last winter when I wrote the paper, and I'm sticking to it. I think that there are a few interesting ideas to explore in this space which could influence the next generation of AI systems. I was very lucky to receive a lot of interesting feedback about this position—thank you for all the messages!
>> Specifically, I focused on the equivalence principle—a key axiom that Einstein formulated through thought experiments grounded in his physical intuition.
It's weird because the equivalence principle is very unintuitive. Aristotle's Mechanics does not have it. It took almost two thousand years to discover inertia that is the most simple version of the equivalence principle. Einstein understood the idea of the the equivalence principle because he had a physics degree, not because he feel that in real life.
Moreover, if you ever have to study or teach Quantum Mechanics, physical intuition gets in the way. A lot of properties contradict the physical intuition but after a while you get use to them. If we continue with Einstein, the photoelectric effect does not aperar in real life.
But he didn't have one :) I guess unless you use the raw images, I think the cameras fix a lot of the noise caused by the Poisson distribution. I imagine a movie where Einstein get bored in the patent office and decide to open a photography shop, perhaps with very grainy underexposed photography is possible to notice the quantization of photons, and then he gets a Nobel prize.
Another bad idea for a movie is a watermillpunk universe, where during a practice on a hot day the best ever curling player discover inercia.
There were already prior reported cases of the phenomenon. Einstein provided the mathematical foundation explaining it, in particular that the discharge is quantized, indicating that it came from an electron. Millikan proved it a couple of years later with his oil drop experiment and by measuring the elementary charge of an electron, garnered his own Nobel:
I agree, I agree, but I think the guy only read about it and if lucky measured it once in a lab class. I'm trying to imagine a fake world were he could get enough experience in real life to justify the article claim "grounded in his physical intuition".
That’s funny, when I first learned about the equivalence principle, my first thought was “of course!” I have always found it to be very intuitive. The great leap is being able to frame it that way.
I worked on the early iPod scroll wheel, before there were advertisements for it and it was in common use. I found the UI interaction odd and unintuitive, particularly the "menu" button and the dead end you hit when "playing". Of course by the time the demonstrator ads came out and everyone was talking about how easy it was use, I'd already spent 10s of hours on it, and it WAS second nature. Being first and embedded in the culture of the time has huge UI advantages.
(Note that the HP Chipmunk 9836 also had a scroll selector wheel in 84)
Most people are UI bigots. Once they get used to a first something, they expect everything to work that way and hate learning anew. They get stuck on keyboards, mice, trackpoint nubs, trackpads, trackballs, scroll wheels, or touchscreens and refuse to move on. Of course there are 'objective' performance tests for each including Fitt's test of accuracy and latency, as well as, cognitive load. I guess once you have a hammer, every screw looks like a nail.
So I'd be a bit careful with the "of course!". It may well be obvious only because that is the first mental model you latch onto.
If your glib comment is referring to me as a techbro and doing the annoying worst thing, then maybe you should explain how I am comfortably abusing the goalpost fallacy given the explosion of agentic AI capability.
I simply point out here, that fully accepting the paper’s premise, the paper’s conclusion isn’t limiting on frontier AI reasoning agents. The paper posits the necessity of multimodal world models and the limitations of LLMs. Frontier agents aren’t simply LLMs and do increasingly integrate increasingly capable multimodal models.
> Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes.
> Comments should get more thoughtful and substantive, not less, as a topic gets more divisive.
> When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
I mean Einstein had help, he was networked with the best scientific minds of the planet and his discoveries were grounded in experimental results that contradicted existing theories (at least for specialized relativity), and without Riemann’s work he wouldn’t have been able to formulate his theory either. So not sure if AI couldn’t do that if you kept feeding it with new research results and let it correspond with top human scientists. Einstein was a genius but I don’t think his thought process is beyond what an LLM could simulate. And again this is probably the most impressive scientific achievement in theoretical physics in the 20. century so maybe it’s hanging the bar a bit high for LLMs.
Exactly. The post read to me as another variation of the denial that people with expertise are reaching for right now. My sense is that we as programmers went through it over a year ago already (perhaps not all of us, but at least anyone paying attention), and so it's easy to overlook that it's still new to people who do other forms of "knowledge work," i.e. people whose identity is bound up with their expertise.
My working theory at the moment is that for programmers it was relatively "clean" and took the form of an inside-out transformation of the work, where AIs directly produced the central work product more or less adequately and relatively early on, but for other forms of work it will appear as some mixture of inside-out (in which case it will appear similarly first as a tool, then as something more than mere tool) and outside-in (the things surrounding their work and the supports their work processes rely on will be progressively automated). This is going to give rise to all sorts of pathologies in the white collar world, we'll get all kinds of variations on denial/negotiation, and so on, until it fully transforms the division of labor.
One interesting point of reference here: Yuval Harari gave a talk recently about the radical changes that will take place relatively quickly, in which he noted the AIs are not quite as good at writing as he is yet, although he expects they will be relatively soon. He then gave the timeline for what he considered "soon": 10 years! So we find the denial ("I still have time, they're not as good as me yet, maybe in 10 years...") even among the most vocal "prophets," among those supposedly most wised-up to what's going on and where the capability frontier lies.
Did you maybe consider that the 10 year timeline is not denial, but actually well educated reasoning based on Yuval’s experience and understanding of the problem space? You shouldn’t dismiss people’s thoughts as denial just because they don’t match your perspective.
I would guess if asked Harari would actually revise that lower. The point was: even those arguing most strongly that AI is an autonomous historical force can slide back into the mere abstract recognition of it. In context, Harari was saying "I'm speaking to you as a writer now, but my own standpoint will eventually be undermined." I'm saying "yes, and your 10-year timeline suggests you along with all of us are not taking your own thesis seriously enough."
Have you considered the possibility that you're simply not as good as the experts, and that your experience of LLMs being capable of performing your work up to your standards doesn't imply that experts are necessarily in denial?
Being fairly at expert level in a "solved" domain (for some definition of solved) but also having near-expert level proficiency in a non-technical as-yet "unsolved" domain and watching the process repeat there (and watching how people react as inroads are made progressively deeper) is basically my own standpoint. But I am intentionally using "solved" and "unsolved" very loosely here: you can get stuck in the thicket of arguments about verifiable domains, what it means for a domain to be solved, whether a set of evals can tell us something has been solved or not, and so on. One way to avoid that (as my original comment pointed to) and focus on what matters for us is to look instead at the effects being produced in the work process and on the division of labor as a whole.
He's saying: "Expertise improves AI-assisted work, experts can steer better and get more out of models."
The obvious conclusion for anyone is "therefore experts will remain the indispensable and specially rewarded center of the production process."
That this is appearing exactly now, and in this form, strikes me as extremely suspicious. I don't doubt the author's sincerity on the surface. What I suspect is that anxiety over the possibility that the (unstated) conclusion might be false (!) motivates the argument in the first place.
I'm asking the question, "Why is this argument appearing now?" At least one reason seems to me to be, "because we're afraid of what the world could look like if it's not true."
However, I personally agree with the author and I don't think his argument is necessarily motivated out of an anxious fear. On the contrary I think it may be motivated out of a sense of extreme exhilaration and empowerment.
Because experts (like myself as a programmer for 15+ years) who are using AI in many fields are suddenly empowered and much more useful than we were before AI. My employability and value has gone up and not down, precisely because of being able to apply my expertise with AI, which people without expertise simply cannot do. I am a professional programmer and also owner of my own startup.
Let me give you a concrete example that I am dealing with at my startup. I'm a small business owner. Before AI if i wanted to produce production quality video for marketing it would taken such a huge budget and such a large team of people (or an expensive agency) that I wouldn't even have considered it due to the enormous cost. I'm talking about Apple quality video production which takes millions of dollars to produce.
Not anymore. A single competent person with AI can replace an entire marketing video production department or agency. But expertise is key here: knowledge of film terminology to be able to describe the effect you want, and ability to use video editing tools effectively, as well aesthetic taste. I as a programmer with no filmmaking experience don't even know how to write the prompt which makes the video that i want because I don't even have the terminology. But a person with that expertise has suddenly become more employable and more valuable to my business because I as a small business now have the capability to create Apple quality marketing videos.
So AI actually created a new job for an expert that would have otherwise not existed because it was outside the budget of small businesses. Previously somebody like that would have been employable to only a few large production studios but now they become employable by almost any small business.
> A single competent person with AI can replace an entire marketing video production department or agency.
This is a good example of what I was pointing to with "fully transforms the division of labor."
I'm personally in a position similar to yours, but as I watch the different moves the labs make, I see the edge we've been handed (for now) also constantly under attack from different angles. This is why it seems to me that so many who can make the most of things for the moment also feel the clock is ticking.
I've seen this sentiment so much in the past month that it's starting to make me worried that the labs will silently make this the default output to satisfy users. I do recognize the outputs can be so information-dense that it strains reading for many but they absolutely aren't nonsense or mere affectation. What people are asking for now amounts to asking for less intelligence per token only a few months after everyone was demanding more intelligence per token.
Yeah I worry about this too. I have a remote agent that writes PRs for my job, and at the request of other devs I implemented a simplified writing skill that does the technical english thing mentioned elsewhere in these comments. It's definitely easier to read and the other devs seem to like it, so I'm okay with keeping it, but I haven't changed my personal claude settings because I never had a problem with the density of the output myself. Probably a side effect from having a multi-disciplinary background.
How long does a train track last? Does a fiber optic cable last? Both are greater than 30 years, both will persist (relatively well) without use, and the benefits of scrapping or removing them are minimal. This allowed future companies to take advantage of them. Even if GPUs running at high load last five years, if the data center they are in goes bankrupt (because the AI bubble bursts), it’s likely they’ll be stripped and sold to make way for more productive CPU-based uses and to recover some of the cost of the bankruptcy. The surrounding buildings and infrastructure will have longer use, but it doesn’t translate to a net future benefit with AI.
That reminds me of saying the early Google was doomed because they put all their money into cheap pcs acting as servers - how long do those last? But the enduring value was Google dominating search which was worth billions/trillions, not the heaps of pcs.
Same here - the main value is in dominating AI or something like that, not in the stack of hardware.
We're going to see the battle intensify here because "spikes" of ASI are emerging that can't be ignored. Models are now better than humans in certain domains or for certain tasks, which means capital as a moat is being eroded. This is why you see people like Jamie Dimon sounding the alarm. Most of the talk until now has been about how the models would be a serious problem for labor, but if that was true how could it not also be an issue for capital?
This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.
Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.
Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.
The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.
No, the incident has been blown way out of proportion by interested parties. They gave a swarm of agents an impossible task in an ExploitGym Benchmark setting, then didn't monitor it even after they discovered the initial breach of Artifactory.
Everything has been fishy, starting from the initial presentation at the blackhat conference, where things were framed like, "we've entered a new world of security," as an accomplishment, rather than what it really was: massive negligence.
reply