Hacker Newsnew | past | comments | ask | show | jobs | submit | benob's commentslogin


More info please. Can I add this into birdnet go for more species detection?

Lots of people in non-tech jobs are dropping their laptop/desktop for their phone for all work-related tasks. They love having more screen real estate, plus their company pays for the premium.

A natural next step is to use the reasoning traces to jailbreak the stronger models (https://arxiv.org/pdf/2603.12277)


+1


Maybe the correct UX could be to list the locations of settings data, their size and ask the user whether they want to leave them, put them aside in a dedicated folder ("the attic", "the basement" or whatever), or remove them


Windows uninstallers sometimes just have a checkbox


People will keep hiding reasoning because it allows prompt injection https://arxiv.org/pdf/2603.12277, in addition to facilitating distillation (you don't pay the full cost of RL)


Or they could store the reading traces and validate the user hasn’t edited them server-side? They could sign reasoning traces so they can’t be counterfeited?


both openai and anthropic actually do this and let you as a harness store only the encrypted contents of thinking traces to be passed again on further turns but they only have the key to unencrypt it so its jibberish for final users

https://developers.openai.com/cookbook/examples/responses_ap...


yeah the article says that. but parent talked about just signing it, idk if it's still really moveable then either though...


Why would the model bother to check signatures if it doesn’t check tags?


The latter ("facilitating distillation") is a real concern from the labs I'm sure, but the first part I don't understand what you mean, or you misunderstand that paper, it's about "prompt injection" via manipulating the reasoning, not about the model reasoning by itself and somehow that leading to more prompt injections. They're quite literally maliciously rewriting the reasoning as the model reasons, not just showing it to the end-user, two very different things.


Yeah, signing would prevent the paper's attack, but seeing the "raw reasoning" is helpful in figuring out what triggers refusals or (especially) what in the system prompt is guiding a specific refusal.

Gemini (the web UI) used to show raw reasoning, or at least a more detailed summary of its reasoning, less than a year ago. Complete with markdown and weird spelling idiosyncracies, so I'd lead towards "real reasoning", but who knows. The reasoning block leaked the system prompt way more often than the response block did, and you could figure out why it would refuse a request through the reasoning, even if the response itself refused to elaborate. This is, presumably, why they stopped showing it. No loss for them, just prevents "pesky users" from low hanging fruit snooping.

(Gemma 4's reasoning and output remind me strongly of what I remember Gemini 2.5/3's reasoning to be, as an aside. I guess that's obvious, but Gemma 3 felt like a totally different model, while 4 feels very Gemini-ish.)


What matters for this injection strategy to work is to follow quite closely the style of the reasoning. It's particularly effective if you copy reasoning from the same context. If you cannot see the reasoning, you cannot duplicate it's style.

That said, including instances of the attack in training is already a good countermeasure.


Token-based reasoning also seems like it would be inefficient, there’s no reason it has to be English or even human understandable.


There have been models trained for latent space reasoning. It has a lot of advantages but the big drawback is the reasoning is completely opaque.


What use is the reasoning if its unintelligible to a human?


Even if it is intelligible, reasoning styles (and hence reasoning effectiveness) differ between models.

For example, gpt-oss loves reasoning in the style of "I be caveman, hungry, need food, need coconut, will search coconut now, eat when find." Giving that to a model unaccustomed to that style could cause it to respond like that.

Reasoning interpretability helps debug why models fail (and some labs will give it to you if they trust you not to distill or be hacked), but there are also conflicting goals like token-efficiency, so interpretable reasoning doesn't always mean "pretty sentences".


Reasoning improves the output of the model. Originally it was a prompting method called chain of thought but later models were trained to do the reasoning steps without the human prompting for it. Being able to see the reasoning steps is just an accidental benefit.

https://arxiv.org/abs/2201.11903


The next step is to build software without bugs


Longest definition and semi-columns are strong biases for right answer. Also, my run contained a lot of adjectives for which it is pretty obvious that noun definitions do not match.


It may be a clever move. By using the same models as android (contractually?), they can compete on the user experience which they typically handle better than android phone providers.


And papers on bias amplification in ML predate LLMs. I remember this specific one which was a spotlight paper at EMNLP:

Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints, Zhao et al.

https://arxiv.org/abs/1707.09457


The bias concerns in Gebru's paper cover pre-LLM systems. For all we know, modern frontier models might mitigate many of the concerns the paper brings up. It's hard to know. The logic used in summaries like the one we're commenting on is conclusory: centuries of prejudice are encoded in the total corpus of human language, language models are trained on that corpus, ergo language models must be biased.


Does changing the date fix it?


No. There are standards in EU and US that enforce these things. Changing time would be too trivial bypass and manufactures would have coensequences.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: