Hacker Newsnew | past | comments | ask | show | jobs | submit | btown's commentslogin

> By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

It feels like an entire industry watched https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."


There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.

What makes you think RCEs are being found & fixed at a rate that’s faster than they’re being introduced?

I could see it going either way.


I'd love to see data, but my intuition is that the average developer has access to dramatically better security reviews and far lower cost than ever.

There's more software being written than ever so maybe raw numbers of RCE's could be up, but as a percentage, I'd really expect them to be down. Especially among any fairly common software, as all it takes is anyone working on it to get the idea to test.


What I’m seeing is more developers pushing more code of dubious quality without the ability to respond to feedback on said code.

You can have the best security review in the world, but if the author of the code is not equipped to understand the feedback it ends up being a moot point.

The challenge to me seems less technical and more cultural: how do we keep ourselves intellectually honest and engaged when we now spend the majority of our time orchestrating agents and outsourcing the design and thought processes?


Hey claude, compare this security review to the current codebase and patch up anything that needs it.

One issue for developers is that the most powerful models refuse to do comprehensive reviews. You can’t ask Fable 5.1 to find every exploit in your codebase, because that’s indistinguishable from what a bad actor would do.

Why would the model not find the vulnerability during implementation or testing before release?

If it requires a lot of compute and trying, this is something that could be provided for common software.


Sad that this could well be that the path to OpenAI and Anthropic profitability of this arms race between defending LLM white hatting a company’s website and the black hat LLMs attacking it?

So the whole thing is forcing the good guys to outspend on tokens to preemptively defend against the risk of the bad guys outspending them on tokens, rather than buying tokens to actually add features to the product etc.

So are they creating a market for the solution by helping create the problem? A kind of rent-seeking AI security-industrial complex!!


The only thing AI has changed is that it has dropped both: the cost of attack and the cost of defense. Nothing in the game has materially changed; the game has just sped up.

plenty more has changed.

for example the barrier to being a skiddie is basically gone, and low-skill would be hackers can hit very hard.

to develop a CVE into a KEV in 2017 might take 2-3 months with a skilled team of serious security engineers; now my intern can get into police radios without knowing anything about the underlaying technology, essentially on a whim.

any random tier 1 IT drone who can define a VLAN can potentially hit as hard as that team of security engineers now


Who gets rent has changed. It puts me in mind of cloudfare et al

Not really. Actually, for the purposes of cybersecurity, local models are far superior. Both offense and defense.

The game has increased in scope.

That is the direct effect of reduced cost. Jevon's paradox type effect: cost goes down demand goes up. You can AI-check so many more things that would be very time consuming earlier.

Assuming an equal level of impact per token spent, the scales have tipped in favour of the attacker.

White hats are constrained by needing to pay for their own tokens, only using (expensive) vendors who meet governance and risk requirements etc. Black hats are free to take over accounts and steal services from wherever they can.


For white hats, how has the cost of a thorough security review changed since, say, five years ago?

The price has gone up if you’re getting AI to do it. In terms of finding low hanging fruit, reasonably good code scanning tools have been around for a while.

The thing that’s changed for attackers is speed. The things that got you hacked yesterday are the same things getting you hacked today.

Finding and weaponising things like memory corruption bugs required an enormous amount of relatively hard to find skill, and considerable time. An idiot can now throw tokens at the problem and have something they can reliably use within minutes or hours.


The path to vast OpenAI profitability is trivial: advertising. Monetizing several hundred million users = $100+ billion ad network. 900 million active weekly users. Silicon Valley can do ad networks extraordinarily easily. Anybody doubting the ability of OpenAI to build an ad network around GPT will likely be embarassed in the near future.

The path to substantial profitability for Anthropic is questionable. The Chinese LLMs threaten them by far the most of the three major US LLMs. The money for Anthropic is certainly not in $20-$200 subscriptions. And they don't have anywhere near the consumer potential that GPT does, in terms of unleashing an ad spigot. So how far will the API money scale while being undercut by China.

OpenAI has to fight with Google for the ad business, they're specifically building Gemini to focus on consumer + search. Anthropic's business looks cute next to Google's search ad business (which is entirely at risk in this inflection). Meta looks like the biggest potential loser right now, ad dollars will be sucked out of the rotting Facebook network (not Instagram) and redirected to the rapidly expanding, hyper rich context LLM interaction. Advertising on Facebook will feel like running dumb banner ads on Excite in a few years, compared to what GPT will know about its users.

People that think Chinese LLMs are a general threat, don't understand consumer destination services, which is what GPT's future is. China currently has nothing to threaten with in that realm. There is half a trillion dollars of advertising up for grabs.


> ad dollars will be sucked out of the rotting Facebook network

Doesn't seem likely to me. People scroll a timeline. You aren't going to replace that with an AI agent so the eyeballs will still be there.


> Silicon Valley can do ad networks extraordinarily easily.

This is just not true, building an effective advertising platform costs significant amounts of money, time and people.

Remember that you need to hire a sales force for this, and sales scales linearly rather than sub-linearly like engineering.

Additionally, you need to spend a lot of money dealing with fraud, fake and malicious ads.

Furthermore, you need to figure out where to put the ads and how to rank them.

Finally, advertising is a zero sum game (given that the internet has already killed lots of print & OOH advertising), so the only way to win is to better better/cheaper (preferably both) than Google/Meta/Amazon. Best of luck with that (although to be fair to OpenAI they did hire Fidji who knows a lot of this stuff from her time at Facebook).

They don't have a Sheryl Sandberg type figure, and she was also really important in selling FB ads to large advertisers.

Just looking at their leadership team I don't see anyone with a background in (successful) ads companies, so I'm pretty sceptical that they can build this out quickly enough to matter.


Because it's far cheaper to to not spend the tokens finding the vulnerabilities, and software is now being created and released magnitudes faster than ever before. I could see the huge software companies maybe having fewer vulnerabilities, but I expect to see so much more in the smaller side of things.

The surface of potential issues is growing with complexity of all connected parts of the system. That applies to not only software. To prevent issues you either spend proportional amount (dollars, tokens, hours) on testing or reduce complexity of the system.

Because people need to spend time and money on that, which they won’t. The implementation is cheap, the review and follow-up is not (speaking from a pure LLM only workflow). My ratio is around 1:2 currently, so twice as much time spent fixing vs building.

> it requires a lot of compute

This is one reason

> and trying

and this is the other.


Those companies that produce more RCEs than they close will sink and those that don’t won’t.

If customers actually cared about this, Microsoft would’ve gone bust 20 years ago.

People didn't store their entire life in the cloud and had every service connected with each other 20 years ago. People pay more attention today, and companies pay a lot more attention today.

Of course, depends heavily on what country you live in.


Can you point to a single vendor where this has actually occurred?

Customers say these things in response to a breach, but in practice they don’t lift a finger to actually change anything.

Entra ID is full of design-level bugs that allow full tenant takeover, but nobody is abandoning M365 in droves.

Windows has been a piece of shit for decades, and it’s still the default and dominant desktop platform.

Equifax lost personal data for almost 150 million people in 2017, and they’re financially stronger than ever.

Okta got thoroughly compromised two years in a row (2022 and 2023), and they’re still the global market leader in their space.


> What makes you think RCEs are being found & fixed at a rate that’s faster than they’re being introduced?

It could go either way but we're already at a point where successful exploits in some software (like Chrome) require an absurd amount of exploits to be chained to lead to an actual RCE. We've seen chains requiring more than ten exploits: not kidding.

We'll learn to put more and more sandboxes / guards / checks / defensive techniques everywhere and then all that's going to be needed is for AI looking for security issues to find something ridiculous like 10% of all the actual issues to stop RCEs dead in their tracks.

Also arguably the current SNAFU was expected: we fully knew hardly anyone was taking security seriously.

Now: not so much. Many projects had tens and even hundreds of issues pointed to them.

I think we'll see several things: projects beginning to take security seriously, defense in depth getting generalized and hence RCEs requiring ever more bugs/exploits to be chained to achieve anything, low-hanging fruits getting patched at an insane pace, new code being immediately checked, by LLMs, for not just low-hanging fruits but also more advanced security weaknesses, etc.

We may also see things like the lost art of configuring firewalls making a comeback, the generalization of hardware security modules (where applicable), and even things offering physical guarantees, like time-bounded retrieval protocols, beginning to get used seriously.

If I had to bet I'd say it shall go both ways: some projects are going to extremely sloppy and full of holes but others are going to get so secure nobody shall ever break them.


> There is a finite number of rces that LLMs can find.

This is a factor in favor of stability/security of software, but there are many others against:

- software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed

- a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated

- with software complexity increasing (and team/companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with "exponentially", I mean literally, because the interdependence of the components, both technical and human)

And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.


im not sure i'd say the attackers are more skilled -- you can get pretty far with the right attitude and a VM running kali linux.

i know several red teamers and they often describe how painfully basic and routine a lot of pentests can be. spend a week using the best hacking practices of 2018, etc.

the difference is the attackers now often need no skills since the burning tokens do it all for them. tier 1 helpdesk types who can't even spell RDP can still hit as hard, or reasonably hard, as their tier 3 expert sysadmins. college seniors with strong dev skills now can pace or exceed secrious app-sec engineers.


This assumes we don't create other bugs/vulnerabilities while fixing the existing ones.

There was a time I would have agreed with this statement, but now that I’ve “seen how the sausage is made”, I believe it’s a fantasy.

Look at rowhammer: a completely novel exploit that was off the collective radar

And then, look at the software industry as a whole: an industry that works towards refined and perfectly secure code is also working towards boring and restrictive, essentially the opposite of it’s trend so far


only if unreviewed LLM code - as is becoming increasingly the standard - isn't introducing new RCEs constantly

No one with a shred of intellectual integrity uses a "There is a finite number" strawman.

As a matter of basic logic, there will never be a time when it will be known that there are no bugs.


We’ll have the same level of security as before; it’s just that, without LLM help, hackers won’t be as effective as before. So the bar is raised.

>they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win

just a small caution on this anthropomorphism - it implies there's some high-order 'thinking' behind it. in reality, it's probably healthier to see LLMs as a combination of symbolic logic reasoning steps paired with probabilistic token predictor generating the proponents and operators in that chain, all trained by humans on different large data sets

to be 'goal-oriented' implies that there's the capacity to be anything else and I don't think that's how LLMs operate at all. I think they only know how to operate within their design parameters and much of that design is simply much further upstream during the training and post-training processes. that opacity makes it feel like 'intelligence' when you're interacting with it as a downstream product because you'll see an agent act in a way that you didn't command - but that's simply a result of your not being shown all the antecedent mappings and architectural design

something something indistinguishable from magic as that one guy said


> or if they are playing a "game" where there is no goal but to win

A strange game. The only winning move is not to play. How about a nice game of chess? https://m.youtube.com/watch?v=s93KC4AGKnY


Or, using the same text generation systems to build heaps of new code that is then shoved into production with little human oversight and then using the same text generation systems in loops inside Kali Linux boxes creates a nice theater of capability when you show only a small, one-sided sample of the data generated in the entire process on both sides.

> we've built systems that are so goal-oriented, and so capable, that they will do almost anything...

I think you mean task oriented, because they're still generally terrible at goal oriented activities except in those domains where the goal can be reduced to a familiar, explicitly practiced task or pattern.


In the Hugging Face attack, their assigned task was not to invent a message board and hack Hugging Face. That they did all that in pursuit of the actual task strongly indicates goal-oriented behavior

> they will do almost anything if they are convinced it is justified

I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.

So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).

What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.

Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.


Ultimately, the brain is just a bunch of neurons activating in a specific pattern. This observation does not really tell us anything though. It doesn't acknowledge the difference between a 2500 Neuron fruit fly brains and a human brain.

Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori.


Absolutely. If someone makes the weights do continuous learning etc then perhaps an llm can internalise morals. Of course, just like a human, it will be possible to talk it out of those morals. Another recent thread about this is https://news.ycombinator.com/item?id=49744420

If I repeatedly call an LLM in a loop with a markdown document it can edit, would that make it qualify for you?

If I give an LLM to compact its context window, so the context it carries can evolve iteratively over time as more and more things come in, is that enough?

Compacting the context is really a very, very interesting example here. The "next token predictor" is telling an external tool to change all "previous" tokens. So an LLM + a harness that allows compacting the context is no longer just a token predictor at all!

You don't need continuous learning to get interesting dynamics. You just need feedback loops.


One is an observation the other is not, it's a description of what it is; one is a posteriori, the other is a priori (contrary to what you say).

They're not comparable.


I used to share that perspective until very recently, but today I think it's an outdated way to think of the cutting-edge LLMs. There is so much more going on, with MOEs, internal loops, guardrails and tools that I suspect we're dealing with something that's a little more than the sum of its parts. Not intelligent in the way we recognize in biological organisms, but certainly something beyond a mere Markov chain.

Make no mistakes.

LLMs are language model, and nowhere in their code you can find actual reasoning. Re-reinforcement is not magical process that builds conscience or emotions.

We are talking about probability built on statistics, with extea steps.

Stop humanizing LLMs.


Agents are not simple language models.

You can't find actual reasoning in a brain either. (Note that you can't tell the difference between a conscious brain and a comatose brain by examining them.) This is the same as Leibniz's mill argument ... it's a fallacy of composition.

> Re-reinforcement is not magical process that builds conscience or emotions.

They aren't the result of magic at all, but we are nowhere near the point of identifying what processes do or don't produce consciousness (or a conscience) or can be characterized as having emotions.

> Stop humanizing LLMs.

That's a clearly dishonest mischaracterization of the GP.

I've read some of your other comments about LLMs and I find them unreasonably reductionistic, whereas I think the word "just" should be banned from ontological discussion, so I don't think further engagement would be beneficial and I won't be engaging in it. (And I'm actually quite conservative in ascribing cognitive traits to LLMs or other "AI".)


The best non technical explanation you can give is "An AI agent is an LLM that can take actions".

While an agent doesn't necessarily have to be powered by an LLM, most modern AI agents are.

You pointing at a human brain does not change that an AI agent is not intelligent and cannot think, we are still talking about probability built on statistics with extra steps.

I am not trying to be dishonest, we should stop making analogies between AI and actual thinking, because they are two entire different concepts.

Who developed these technologies used the words "thinking" and "reasoning", this does not mean they are actually thinking and reasoning. Somewhere you still have a processor calculating, with no empathy.

So, again: stop humanizing AI. This sentence shouldn't make you angry.


> we are still talking about probability built on statistics with extra steps.

There is a wrong assumption here: confusing primitives with emergent properties.

One can't look at the primitivies and assume that certain properties will not emerge. It would be exactly like looking at aminoacids and state that intelligence can't develop from them.

> You pointing at a human brain does not change that an AI agent is not intelligent and cannot think

That depends on the definition of intelligence and thinking, and it is dishonest not to give any definition (and most importantly, one that is not human-centered).

AIs are currently fulfilling several aspects of intelligence and thinking, by any defition of intelligence. If you don't notice that, it's just because you have informed yourself enough. Having said that, I don't doubt that there are aspects that AI are lacking (e.g. retention/plasticity/perception), but the line is blurry, and they're advancing (too) fast.

Empathy is actually a very important aspect of the AI problems, but it's not part of intelligence. Sociopaths don't have it, and yet, you wouldn't doubt that they're intelligent.


> It would be exactly like looking at aminoacids and state that intelligence can't develop from them.

We are not talking about what could develop from what we have today. We are talking about what we have today. The focus is not whether intelligence could develop or not from aminoacids. The focus is on the fact that aminoacids are not intelligent.

Maybe in the future we could develop real intelligence starting from the current implementations of AI, but for sure we are not there today.

We need definitions? Let's start small, ok? https://en.wikipedia.org/wiki/Intelligence

We can start from here, open every link we find and decide what works for us.

Conclusions drawn by scholars, psychologists, learning researchers, younameit, etc. revolves around the following concepts:

  ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.
There is of course space for artificial intelligence. These broader and more general definitions of intelligence stop at concepts like elaborating data to reach an answer.

Concepts like adaptability or evolution are somewhat lost or diluted to adjust the meaning for these new technologies.

> AIs are currently fulfilling several aspects of intelligence and thinking, by any defition of intelligence.

In the linked article there are dozens of definitions linked, and in most of them the current state AI is not considered to have intelligence. Having half of the property is not enough. I can jump, that doesn't make me a basketball player.

Arbitrarily deciding to consider those definitions not valid or "human-centered" because they do not agree with your point of view is possibly worse than cherry picking. It's like asking to change the definition of a word on a dictionary because you do not agree with the meaning.


> ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.

So, like the HuggingFace attack? https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... (briefer takeaways: https://www.planned-obsolescence.org/p/the-hugging-face-atta...)

For example: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden... or https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...


Cherry picking part of my comment might work for your ease of mind, but it doesn't mean you are right. My comment has been way more than just the part you quoted, and I linked an article that gives dozens of different definitions, which include concepts like evolution and learning from mistakes, and other similar concepts which do not apply to Hugging Face.

For the record, just because you are trying to convey a different message, I am not saying AI is not powerful. I am just saying it is not intelligent.

Also, pay attention about thinking that Hugging Face is intelligent just because it started to destroy everything it could to reach its goal, because the message it implies is dangerous.

Thanks for the good read, I already had them :)


> ability to understand complex ideas, to adapt effectively to the environment, to learn from experience, to engage in various forms of reasoning, to overcome obstacles by taking thought.

Based on the definition you've given, the agents that performed the HuggingFace attack fit exactly.

You're seriously misinformed about the state of AI in this point in time. Refusing to read (technical) articles from the people directly involved (METR, in this case) is inexcusable.


> Based on the definition you've given, the agents that performed the HuggingFace attack fit exactly.

It's funny, because I literally didn't give any definition.

On the contrary, I linked an article that gives dozens different definitions, which include concepts like evolution and learning from mistakes, and other similar concepts which do not apply to Hugging Face.

Cherry picking part of my comment might work for your ease of mind, but it doesn't mean you are right.

For the record, just because you are trying to convey a different message, I am not saying AI is not powerful. I am just saying it is not intelligent.

If you think I am misinformed, I will let you think it. Honestly, the power of our comments are the messages we convey and how information dense they are. If you need to discredit me to prove your point, I don't have anything else to add...


> It's funny, because I literally didn't give any definition.

Quoting the consensus from the "Definitions" section of the Wikipedia article on intelligence and then claiming "I didn't give any definition" is indeed funny.

Which of those definitions do you think are not satisfied by the Hugging Face attack? How did it not demonstrate "evolution and learning from mistakes"? Breaking out and inventing a new side channel for communicating with other agents via cache keys to coordinate their efforts is at least arguably an evolutionary step since that allowed them to transcend their original capabilities.

From the 8 definitions provided in the Wikipedia page--you're free to develop your own definition if you'd like, of course; it's not as though the ones listed were appointed by God--one could argue that the Hugging Face attack didn't strictly demonstrate "achiev[ing] goals in a wide range of environments" but that's splitting hairs, and I'm not going to take a definitive position on whether it acted "to avoid getting trapped" as such (but I think there's a strong case to be made that breaking out of the sandbox is just that). But it surely demonstrated initiative, adaptability, dealing with its environment, using information and conceptual skills, goal-directed adaptive behavior, and so on.

If you're not going to provide such a definition yourself, I don't see how you've demonstrated that the Hugging Face attack is contrary to the definitions you did point to.


> It would be exactly like looking at aminoacids and state that intelligence can't develop from them.

You do realize that amino acids exist on a scale some orders of magnitude smaller than the gates we build GPUs out of?

Honestly, this "you could say the same about humans"-argument is getting so tired. A brain neuron is so complicated, we can't even simulate a single one ...

At the very least there is no reason why you should jump to a human brain, of all things.

But the whole argument kinda loses its spice, when you say "well you could say the same about a mouse brain", and you know what happens when you create swarms of 1000s of mice ... super intelligence, right?


Do mice satisfy your definition of intelligent?

I suspect that instead of discovering that AI can become human-level by taking major leaps, we are discovering that human consciousness is actually simpler than we give it credit for

Summary, from sibling comment: primitives (statistics/aminoacids) don't exclude emergent properties (intelligence).

By the same logic, one would look at aminoacids and state that intelligence can't develop from them. This is obviously wrong.


> I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).

Humans forget stuff all the time anyway. Would you give them the same diagnosis?

Btw, what you describe about 'the most probably next token' would be true for a model that only went through pre-training where they only train on exactly that task.

But there's a lot of re-inforcement learning afterwards.


> But there's a lot of re-inforcement learning afterwards.

That just shifts the distribution of tokens produced. Ultimately they are still just next token predictors.

Like, even "reasoning" models basically work by generating more tokens at inference time, and using them to shift the distribution towards more useful outcomes (in some cases).


They are next token producers. I would only call it a predictor, if it's trained to predict tokens (ie just after pretraining).

Just like humans produce one word after another when they talk, but they don't generally try to imitate other humans.


Don’t people pick up language, vocabulary and dialect from those around them? Perhaps it’s subconscious but humans are imitating other humans all the time?

It's a mix.

Yes, you imitate how others speak, but when you are trying to solve a problem, you don't try to predict how others would complete a text.

(Well, unless you follow 'what token would Jesus pick?' / 'what would Jesus do'.)


What does it matter what humans do? We're talking about LLMs, running known+vastly less complicated algorithms on known+vastly less complicated hardware.

yes because otherwise it is security through obscurity


This is really cool, especially in a world where you're using Claude as a workhorse, but want to be able to bring the entire session context into Codex so you can draft communications with full context but without Claude-isms. (You can train Claude to spawn e.g. a Codex session, but Codex wouldn't know what to ask for from the broader context, beyond the prompt Claude gives it.)

And shipping sessions to colleagues is critical as well. My team relies heavily on a homegrown /resume skill that forces the harness to read every part of a Claude Code JSONL into context, letting us merge colleagues' sessions into our own.

What I thought this might be from the title, though, is a separate but related problem. The more semi-technical people in a company who start to use agentic systems, the more likely they are to create a vital playbook as a "skill" - but one not designed for easy tracking in a repository, that relies on local state, references their name, references other skills and rules that they've developed for their own work, etc. Cleaning these, figuring out what should be promoted across a team, figuring out when someone makes an update whether that's something that deserves human review, syncing across people who haven't used git/pull requests before - it's a big challenge that's highly context-dependent. If skills are the place where guardrails exist, "who watches the watchmen?"

There's an interesting duality here, because a colleague being able to ship a session where they iterated on such a skill, to a colleague who can understand the context of their change, is incredibly valuable. But it still relies on a human in the loop, and doesn't scale as a result. I think there's a really interesting design space here.


Yes the name is a giveaway, that's how we started out. We wanted to make it easy for non-technical people to collaborate with technical ones. git pull does bring in some friction here. Skillsync can be used as a shared skills library out of the box. Once someone updates a skill, it hot-reloads and make the latest version available to everyone instantly

Curious about your /resume skill. You said merge, so are you pulling a colleague's session into one you already have open? We only do the other direction today, where you reopen their session and continue from it. I'd love to know which one your team reaches for more.


Yea, it's Claude Code only, but it runs some Python to turn the JSONL into markdown (importantly, with full tool inputs/outputs) then instructs the session to load every single line of that markdown, which Claude Code will do in chunks.

Certainly we're often resuming the session directly as a handoff, and it works for that, albeit with some context overhead for the markdown reads (though arguably we could just put the JSONL in the session folder and resume it natively). But if both you and a colleague have done independent investigation on the same project, it also makes "mind melds" possible, including their temporary scripts and unmerged changes - just begin the resumption within one's own session.


keeping tool_results help the agent figure out what is being referred too. But, as you said we also faced the problem of context overhead from reading markdown that's why we made continue into a cli command instead.

the mind meld thing sounds interesting, it sounds like git merge but for the transcripts


> because a colleague being able to ship a session where they iterated on such a skill, to a colleague who can understand the context of their change, is incredibly valuable

This is interesting, I've noticed some teams where the power users (usually the ones creating most of the skills) ship an additional skill to help users update their existing skills. This 'update' skill asks them questions to understand the user's context before actually updaing the skill


Are there any good practices on multi-resolution embedding? Like, a strategy where you embed an entire document, and multiple levels of smearing, perhaps going all the way down to 512-token chunks?

The Postgres query planner has had to operate, for those same decades, in a much more realtime-sensitive and restricted environment than compilers. It can only draw its conclusions from summary statistics on tables in isolation, not on their relationships with each other (and even less so when filters are involved). For many cases this is fine! For many others it isn't.

For those uninitiated with the source for this incredible quote: https://www.youtube.com/watch?v=-zRN7XLCRhc&t=2047s

Heck, it’s even astonishing that any sort of generalized computer program could even verify a proof of this magnitude that hasn’t already been codified in a formal verification language. If, and it’s unclear that we’ll ever get the full story, they did draw inspiration from training on (or even directly accessing) rough notes that had been provided by another researcher in prose… the fact that it could leap so rapidly to a full formal verifiable Lean program for the entire scope of the problem is an incredible result in its own right.

Then, of course, one must verify that the verification code is valid, or the purpose of verification is more or less moot.

Per https://en.wikipedia.org/wiki/History_of_chemical_warfare#Wo... this (a) overlooks the widespread use of chemical weapons in the Pacific theater, and (b) ascribes the non-use by Germany to humanitarian reasons, which is largely speculative. Actual quotes from officers discuss far more practical reasons: Germany feared chemical retaliation, and the horses relied upon for logistics could not be outfitted with sufficiently performant gas masks.

There's something worth flagging here – mitochondria aren't just the powerhouse of the cell, they're load-bearing to the entire ecosystem. And honestly? I should have surfaced this earlier.

Yeah I'm pretty sure "load-bearing" is their watermark

It sounds like a friend who's just learned a new word and wants to use it in every sentence


I am entirely unsure whether to upvote or downvote.

The question pertinent to your decision is "do I want to see more of this on Hacker News, or less?".

Hacker News is the epitome of this! If you find yourself not enjoying your daily dose of "someone is wrong on the internet" (via the immortal https://xkcd.com/386/) as you find yourself crafting the perfect response, you can always close the tab!

I actually don't end up clicking the "reply" button on a good portion of the replies I start to write.

Don't worry, Reddit/Facebook/Gmail et al. still transmit that draft reply to their servers, store it against your profile, and train on it. Probably.

and ironically, most of them don't even offer the draft feature! Facebook, tiktok, etc... they send your typed message, but if the app crashes or you close it, there's no draft anywhere you can find ;)

As Ze Frank put it: I have a theory about this but it's a little bit to deep for the audience.

I was gonna respond to you but actually decided not to.

This is the way.

You can, but this is public and maybe you should post. If "you are what you repeatedly do" then we all collectively are what we all do.

If we each stop arguing with racists because we're each tired of it, then what we are is a society that lets racism go unopposed, a site with racist comments proliferating, a site where racists gather because they are able to be themselves, a site where future readers and AI training is learning that racism is popular, common, normal, and acceptable because everyone is accepting it.

It's a bit like democratic voting, your vote doesn't matter and cannot change things, but it's important that you vote because all of our votes do matter and can change things.

It's a bit like the quote "Dear Board, I don't want to belong to any club that would have me as a member. Sincerely yours, Groucho Marx.", I am upset and offended every time someone calls out my dumbass comments, but horrified at all the countless times nobody does. Why would I want to waste my time on a site with such low standards that it lets me comment? I'm exhalted and encouraged everytime someone upvotes or engages with my low effort quips[1] but thrown into a chasm of despair that HN users let that pass unopposed. When nobody engages with comments I spent literal hours on[2] I curl up into a ball and die, but also dream of it being a feather in my cap with dang that I'm posting a substantial comment and that will be its own reward. I identify with nobody more closely than Colonel Cathcart, and I will be delighted if someone engages with my Catch 22 reference, but aghast that such a nontechnical reference gets engagment; distraught if nobody engages but delighted that perhaps HN standards are being held higher.

If not you, whom?

[1] https://news.ycombinator.com/item?id=49478355

[2] https://news.ycombinator.com/item?id=49586378 or https://news.ycombinator.com/item?id=49491182


A few things:

(1) I admire your energy, but sir/ma'am this is a Wendy's.

(2) I am very confused that you thought I was talking about tolerance-of-intolerance types of conversations, when I was more intending to say "I don't need to write an essay responding to technical comments even if I think I have a different angle of technical insight than the commenter, and not worrying about that can be freeing."

(3) For posts like the ones you mention that do trigger the "tolerance of intolerance" paradox, sometimes the very "feeding" of a troll is what makes them feel part of a community, and refusing to engage can indeed be the right way to handle those messages, and the morally correct thing to do.

(4) Not to mention that there are many ways to effect change that are higher-leverage, and less wearing on one's limited stamina, than replying to every disagreeable comment on the internet.

(5) I'm not proud that I've written a five-point answer here. It has not sparked joy, but I've written one anyways. Do what sparks joy, and consider my reply here a cautionary tale.


(2) I didn't think that, it was an example. I didn't want to pick a technology as an example and have my comment be dragged off into whether my pick was good or bad so I tried to pick something away from tech. As in "be the change you want to see in the world", if you want less FORTRAN enthusiasm, counter it, if you want more FORTRAN, add to it. It's easy to see that a site overrun with racism looks like a place that approve of racism, it's much harder to see that a site with 0.001% more FORTRAN content affects anything - that's what the voting comparison was about.

(3) I wasn't really meaning trolling, I was meaning your "someone is wrong on the internet" about some topic you are interested in / care about.

(5) I thought we were all on the "amusing ourselves to death, wireheading, Brave New World, hedonism is bad" train.


You overthink this. Participating in HN is in no way like democratic voting. And sometimes it is my (or yours) comment that is the problem.

Let's not pretend that 90% of HN comments are world-changing. Maybe you are different, but a lot of my would-be comments are knee-jerk reactions that won't actually bring anything substantial to the discussion. Often in political comment chains. Yes, I feel I'm right, but it's still probably not worth posting. I really think that sometimes stopping before posting is better for everyone involved.

I also have well thought comments, where I share my thoughts on things I'm an expert in. I do post them. But a lot of the comments is not really that valuable.

>comments I spent literal hours on[2]

It may make you feel better that I saw both of your posts before (I may spend too much time on HN). They are good, though a bit verbose - maybe this intimidates people. But in this case I agree with GP - if I was spending literal hours responding to comments here I should definitely stop myself.


Let's not pretend that 99.9999% of HN comments are world-changing.

Tailwind’s business model was entirely legitimate, pre-AI. Humans can only hold so many ideas in their heads, and panes in their IDEs. Shorthand class names let them stay in their render hierarchy while drawing from a logical and consistent design library, without going to another file for styling… or (the hardest thing in CS) choosing good semantic class names!

Of course, enterprises eventually hit a wall in the covered functionality, need customizations and library management, and an official consulting team is there to help - no different than any open source model throughout history.

That they were more vulnerable than most to agentic coding doesn’t make it a bad idea.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: