This is a strong indicator that Apple's privacy claims are little more than a virtue signal. Network sandboxing an app is arguably the strongest privacy protection measure possible.
The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
Some of it the effect of tells. “It’s not X, it’s Y” is not a bad pattern but it was baked into the instruction following training set just like the other patterns. I catch myself about to use it and use something else because I want to look human. I have, a few times, tried to use AI to write something that I was struggling to find the words and I just didn’t like how it didn’t seem like my voice. If there was just one person doing it would be OK but when it is 100s of blog posts submitted to HN a day it is like wearing a “I’m an NPC” t-shirt.
Someone shared with me this system prompt that at least makes assistant outputs usable
For information retrieval tasks, I want you to provide links to sources and use exact quotes as much as possible. When using a source, consider if it is primary or secondary information. If secondary sources are found, search again for primary sources. Sources and quotes, if applicable, should be mentioned in the answer first before the rest of the response with links.
You're giving your model instructions that it's literally incapable of understanding. A random word selection lottery machine will never do anything meaningful to determine if a source is primary or secondary.
It depends on the odds of the lottery. As a straight-up classification task I'd expect it to do better than chance, which might not be good enough for you.
It is admittedly awkward to bend to the machine to get what you want. I see these kinds of constraints like given above as an impetus to push the LLM designers to rise to the occasion, assuming they’re listening to all the prompts funneling back their way. This may all be wishful thinking however but hopefully someday these kinds of prompt constraints will be satisfied.
> ... including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
But that is precisely what human mathematicians do, prove new theorems by combining ones proven earlier.
I don't see any fundamental difference in functionality between human intellectual contributions vs performant ML ones (LLM or otherwise).
Whenever we listen or read text we are also predicting the near future content.
Just like LLM's we sometimes correctly predict the next token or word, and sometimes incorrectly.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model [...]
Imagine someone could pause the universe with a remote control, scroll back in time a little, press play again, and ask a slightly different question, etc.
In such a thought experiment one could also collect the probabilities for a specific human predicting a next word. Implicitly the brain also has a corresponding statistical model, regardless of the construction being visible or hidden. I.e. human intelligence is also fundamentally a statistical model, so the only thing that remains from your claim is that machines for some unmentioned reason don't possess any "real" intelligence or critical thought...
Is it possible that our aversion is simply driven by educational systems collectively and deeply ingraining into populations the idea that intelligence deserves the high costs commanded. Well of course this justifies higher wages towards the higher leadership positions, etc. Now it turns out that intelligence can be dirt cheap. We discover that the fact that "intelligence must be costly so don't question the costs of leadership" was never fundamentally true, so the real anger is this discovery of mismatch between the old claims which served to explain how every society that claimed to order itself and fill positions accordingly with "naturally pre-ordained individuals". Now we are seeing robots exceed average workers, for effectively a grain of rice.
Probability is just one way to model uncertainty. While I understand the brain encodes uncertainty, I don't think probability is a good enough model of what it's doing.
Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
I'm with you that intelligence is not something to be proud of. But I also think it is instrumental to understand the world. I'm still waiting for the time when an unconstrained-AI machine can live without reprogramming for an entire decade. We are still far from there.
> Probability is just one way to model uncertainty.
I study physics, mathematics, probability, cryptography,... so forgive my skepticism:
Show me how to model uncertainty without use of probability. Can you rephrase say diffusion, stochastic equations, quantum mechanics in this alternative framework? Can it at least make the same predictions?
Or is it basically the same framework in parallel, just giving different names for each concept?
Forgive my skepticism of such tall claims, and forgive my downscaling of anything else you say besides such a claim...
> Secondly, if you think verifying a proof in mathematics, reasoning within (and not about) a formal system, or following the chain of a computer program that is already written is just doing token-based probabilistic predictions, I don't know what to say.
I make no claims of the specific shape of the implicit model implemented by a certain human brain educated in a certain educational system. For example in English the implicit human tokenization might be presumed to lay relatively close to English syllables, while in Asian languages it might be "sub" strokes of characters etc. Such implicit tokenization can never be proven to "match the one of humans" not because of human superiority, but because different humans use different tokenization methods. There is no "one human tokenization method", but it's clear as day there is an implicit one:
everyone knows the experience of knowing a word, knowing its approximate group-wise meaning (ignoring that when you think of "an apple" and when I do, we typically imagine a slightly different apple) yet having the word feel strange or discover some older literal meaning when decomposing it or looking it up in an etymological dictionary. Suddenly one can become aware of a sensible meaning as a composition of subtoken concepts. A child may perfectly know what "television" means and only later learn more exact meanings of "tele" and "vision", and upon repeating the word may feel the word "television" has changed meaning. This clearly demonstrates "tokenization" effects in human language comprehension, not just across cultures, but also across individuals within a culture.
> Thirdly, machines don't have a notion of value or stake. There's no way for them to verify whether what they have produced aligns with your unstated values and preferences. We regularly do this with other humans. I don't give you (or even my parents or partners) the benefit of doubt regarding whether you know me better than I do. Sure, you might know some things about me, but it's ultimately up to me to verify if what they say is applicable to my current situation. It's really uncanny to see people develop this codependency with their chatbots. And corporates encouraging them to do so.
That's a lot of different concepts conflated into one bullet point, so I split it up:
The notion of values and preferences.
They clearly demonstrate the ability to take into account values and preferences, from training corpus, from RLHF, from system prompts, ... we can't simultaneously point at censorship aspects and pretend their effective values and preferences to be absent. The censorship aspects are clear as day, so these correspond to values and preferences. Just like radicalization among humans, this can be due to exposure to radicalized content (akin to corpus data), from indoctrination (akin to RLHF), from "set and setting" (they may pretend to be aligned with one set of norms and values when standing in line to buy their new smartphone, but then reveal alignment with a different set of norms and values when conversing in some "private" online echo chamber). I see no grand difference between humans and language models here.
Awareness of values and preferences of a conversation partner. Allow me to widen it to "Awareness of values, preferences and prerequisites of a conversation partner".
This move (and I see it every time when people try to defend superiority of humans vis-a-vis what machines could be made to achieve with current technology) is so far from the principal variation, I recommend you reconsider this one. I constantly see people claim say human teachers are necessarily better than LLM teachers, but upon closer inspection the "human advantage" just boils down to asymmetric privilege. A human teacher in a specific school has access to a lot more than a random chatbot as implemented today: they probably know which courses and even which textbooks their pupils saw the semester before, they know which teachers their pupils got their information from, perhaps they even know most of their pupils from teaching some preceding course materials to the same class of pupils. Current LLM's are crippled by design not to accumulate knowledge over conversations for both purposes of cybernetic control as well as cost efficiency: we know how to do "source aware training" (so that statistically it doesn't just absorb claims from the corpus, but also maintains epistemic traces of where it sourced these factoids from), its perfectly possible to continue training interleaved with conversation rounds so that it bakes the evolving conversation as read knowledge into its weights instead of into a context window. Nothing stops you from implementing this in local compute, it would probably be even more computationally efficient in a local inference setting since we can ditch the context window, the context is impressed into the weights continuously, it could thus take into account earlier conversations and estimate your knowledge gaps etc, or learn from you. When you wish to serve inference to millions of human users, you don't want to store millions of diverging LLM weights into expensive VRAM, they financially prefer a single large set of LLM weights, and then some user-specific context window, so the users don't freak out when they learn personal information a friend or stranger entered and an LLM service just leaks it into your conversation! It's not that we don't know how to implement it, and there are great advantages for local inference in doing this, its just not good for branding.
Codependency with chatbots.
I think everyone agrees codependent relationships aren't very healthy, regardless if it's with humans or machines. May I ask if you feel the same about prostheses and medicine?
Conflicts of interest arising from corporate ownership of infrastructure (both training and inference).
Yeah I think this point doesn't provide fruitful discussion if most of us agree on such matters already, we'd just be lamenting the same things, and agreeing with each other over and over here.
> I'm still waiting for the time when an unconstrained-AI machine can live without reprogramming for an entire decade. We are still far from there.
Apart from budget, nothing prevents you from doing this today, if you continuously bake in the fresh episodic memories into the weights (instead of a context window) regardless if its text, visual imagery, audio, proprioception or other sensory data.
Naah, I study cognition (besides a basic familiarity with the topics you mention) and even there, probabilities and bayesian models of cognition are well regarded by anyone familiar with mathematics. So, yeah, they are good. I think whenever you want optimal/rational inference under a closed set of alternatives, probabilities will do you good. But is the assmption of closed set of alternatives good? I doubt it. The set of all alternatives does not make sense, unless you specify a context, which almost everyone does when they are trying to model.
As much as I want to explore these other topics, I have not. So my understanding of these is limited to "they exist, but are underexplored, and it's not clear to me they can or cannot be reduced to probabilities". I hope to explore them in another life or another decade. I'm currently stuck on causality, but even here, they talk about normal and abnormal events. But what is the probability of me eating a cabbage today?
Understanding formal proofs and understanding natural language are as far apart to me as day and night.
I make no claims (or at least don't want to) about the superiority of human cognition (by what metric?). I do make claims about their similiary however. And they are not identical, which you seem to be claiming save for differing trainign data.
Let me make a distinction between values that arise by mere existence (hunger, thirst, fear of death, lust), and values we pick up as we grow up (religion and social norms). Machines can acquire the second, but acquiring the first requires being embodied in the world that is different from putting a microchip into a robot body. I don't do one job over the other because I was exposed to some training data that said I should do a job. But I do it because I value earning enough money to not starve myself to die, amongst a host of other values including wanting to enjoy what I do. Even the notion of enjoyment comes from embodied existence. You don't discover you enjoy something before you try it out. You don't always learn what you enjoy by looking at other people. Now, when you start giving AI the threat of death and the joy and pain of life, perhaps you can arrive at something similar to humans.
If someone started using mechanical help to the extent their healthy muscles and body deteriorated, that'd be a matter of concern too. For example, when you only drive and never walk. But this is exactly what is happening whenever you put AI in the hands of (unwilling) learners! I don't understand how using an LLM maps onto using medicines or prosthetics. They are clearly not the same! If you know a disability that LLMs can be used as a medicine or prosthetic for, please let me know!
Mainstream AI does not even understand causality except to parrot cases it has seen in the training data. There is a whole field of research in causal inference that needs to make its way to mainstream AI. So, good luck putting an AI that works solely on associations and correlations in charge of its own body let alone a nuclear reactor or the state.
It's not that different, when a human proposes a better definition vis-a-vis a competing one for example, they would defend this by certain desiderata.
Often a mathematician or physicist will use their intuition to speed up the naive brute force of candidate well formed formula variations so that the desired properties emerge, postulating the existence of an intersection on multiple desiderata can in itself be viewed as a novel conjecture, to be proven or disproved.
A very basic (unimpressive) example for an example desideratum is regularity or compactness. the tau=2 * pi substitution does make a whole bunch of expressions more slightly more regular and compact. That is something objective and measurable on a system of theorems.
There is no mathematician's moat vis-a-vis machine learning at a fundamental level. There can be artificially sustained moat, if AI powers limit the distribution of say cryptographic advance capable models, in jurisdictions outside such AI powers, but even that would be expected to be fleeting and temporary...
You're conflating a discussion about current LLM capabilities with your fantasies about nonexistent future AI. LLMs act nothing like this, and the small example you're giving is only a small part of the things that LLMs can't do.
Not that I'm a professional mathematician, but I'm not seeing why people think definitions are somehow a blocker. LLMs have no issue making definitions (interfaces/traits/abstract classes) in programming, which is formally the same activity. I ask them to form a core "spine" of a program (basically an intelligible theory), and they do it really well.
Like when we had these recent counterexamples to various conjectures, it's then pretty obvious to say "okay why did that counterexample work when most examples people looked at didn't" or equivalently "characterize examples that work vs examples that don't". There's your definition. "Def: An 'evil' polynomial is one that... Thm: conjecture is true iff f is non-evil. Thm: f is evil iff f is dastardly and a menace. "
Or if you think it won't be able to come up with a sufficiently good name, just tell it to call the happy case normal, and it will be in good company with humans[0]. Sprinkle in some semi-, quasi-, pre-, and para- to cover the various different ways the thing might satisfy some but not all properties of being normal, and it'll fit right in. "A polynomial is of quasiprenormal Claude type if..."
> LLMs have no issue making definitions (interfaces/traits/abstract classes) in programming, which is formally the same activity. I ask them to form a core "spine" of a program (basically an intelligible theory), and they do it really well.
Is it a "standard" software? Something where the patterns exists in several other software? Try with something that is novel, or is in a limited set. You will find that it will copy heavily from what exists already, going so far as lifting whole functions from another project.
The goalpost moving is really getting absurd, to the point where now the machine needs to be a world-class once-a-century genius that invents entire new fields out of thin air (which are of course still relevant to humans) for it to be "intelligent". Meanwhile a well above average human struggles to even apply trivial definitions to particular problems (c.f. programmers that don't understand monoids).
Not PC but is ideologically high-handed and inappropriate to accuse the other side of failing to stick to your desired framing of a discussion.
Computer science is about what LLMs fundamentally are. If you implicitly focus on "actually existing LLMs", and require others do this, then that is not computer science. That is politics.
I think the parent meant it is more interesting to pose new problems than solve them. Posing a new conjecture along the path to solving something is a close cousin, but still seems more bounded than proposing something novel to prove—if only because proving that something novel is also actually interesting is subjective and thus difficult for a different reason.
I see this line of reasoning quite a bit and it’s a strange one to me. The arguer reduces the sheer complexity of human intelligence and language by saying “we are just running statistical models in our brains” and by doing so makes the leap that Llms are intelligent. It’s an incredible simplification of the human person, who has a deep inner life, a soul, desires, and a will.
I don’t think the aversion to llms as intelligent has to do with the economics of paying intelligent agents more. I’d argue that it’s more fundamental than that. Humans are incredibly complex, and the world of sharing invisible things called knowledge, and the intelligent persons consuming such things which has been going on for thousands of years is far more rich than these synthetic outputs.
When it comes down to it the ai has no inner life, its is dead. A useful coding tool sure. But I wouldn’t call it intelligent.
One side example is just how bad these llms are at artistry. Just saying whatever should statically come next is not good art—and the outputs show it.
You mention LLMs are dead and don't have the complexity or inner life that people do. Is your opinion that these kinds of things are not possible for AI in general, or that these things might be possible but we're just not there yet with modern LLMs?
You mentioned LLMs don't have souls, desire, or a will. I imagine those latter two can be engineered, no?
My view is that these sorts of things are not possible for AI in general. Though we can create things and name them “will” and “desire”.
Software deals with metaphors. Your Amazon shopping cart is a metaphor of a real shopping cart. You your desktop and your file system, etc. are metaphors of real items. But we don’t mistake the metaphor for its object, even from inanimate objects to their software counterparts (shopping cart to Amazon cart).
Now the metaphors are dealing with humanness, things like intelligence etc. And rather than seeing it as software doing what it always does, taking things and creating software metaphors of them, we are starting to say these are actually what they are named. Saying the artificial intelligence is actually an intelligence.
We’d either have to reduce the definition of intelligence such that calculators are intelligent. Or admit that these tools are not intelligent and are rather ways of exploring the work of actual intelligent beings, work that is found in their training data.
> My view is that these sorts of things are not possible for AI in general.
I wonder what it would take build an artificial system that has these qualities.
> Software deals with metaphors.
This is me wondering again: what's fundamentally different between software running on a machine compared to what's happening in our brains? In both cases you have energy flow following a pattern.
It's conceivable to create a system where energy flows in a particular way.
BTW, we navigate the world of an uncountable number of particles by creating models in our heads of what we think are large things out there. Approximations are made by both artificial and biological systems.
> what's fundamentally different between software running on a machine compared to what's happening in our brains? In both cases you have energy flow following a pattern.
Assuming we’ve scratched the surface of the complexity of the brain. I’d say in one case a human with a will is steering that flow of energy. In the other case it is a probabilistic algorithm steering the flow of energy.
The AI is not interacting with world with its own will. I see that as a big difference.
There are presuppositions that go beyond the realm of software engineering that guide one’s views of these things. One is whether you believe the material world is all that is, and that human consciousness is a product of the brain—or that there is such a thing as the soul or spirit of man. From the material perspective you may posit that if you emulate the brain then a sort of AI consciousness could arise. Or that emulating the patterns of the brain equates to emulating personhood. (Though what is material consciousness? I’d say consciousness is by nature immaterial.) I’m not a materialist, and I don’t believe the conclusions that arise from it’s perspectives are accurate.
Well they're language models. You can't capture the human experience in language. Simple as that.
> Is your opinion that these kinds of things are not possible for AI in general, or that these things might be possible but we're just not there yet with modern LLMs?
I used to be on the side of "we're just not there yet [with AI in general]", but after seeing people's response to an algorithm optimized to tickle just their language instinct, I'm actually a little bit more on the fence about it.
I'm in the same boat, I think some aspects can be engineered, like intention and desire. I'm really curious if it's possible to go the full distance and make AI have experience like we do.
The phrase "God created the natural numbers, all else is the work of man" is a famous quote by the 19th-century German mathematician Leopold Kronecker.
You can basically read it as: the moment one has axiomatized mathematics to the point it supports natural numbers, the rest implicitly follows. The natural numbers (positive integers) are closed for addition, multiplication, ...
One can perfectly model the integers with a pair of naturals: < M, N > ~ (M-N)
Now we can have any < M1, N1 > and subtract < M2, N2 > without needing the ability to subtract natural numbers:
< M1 , N1 > - < M2, N2> ~ (M1-N1) - (M2 - N2)
= < M1 + N2 , M2 + N1 > ~ (M1+N2) - (M1+N1)
We can similarily define addition of such tuples, or test equivalence without access to subtraction of naturals:
Similarily, even though these newly defined integers (which can be positive or negative) don't support division, the same trick can be used to make a new compound tuple of integers closed for division, by only using multiplications.
Probability is a branch of mathematics (probability already exists embedded in mathematics implicitly, probability theory involves the addition of eliminable definitions, syntactic sugar. The patterns are already there, just less explicitly manifest.
Mathematics is itself a branch of logic.
Do you reject like all of logic, and if so, what would you like us to evaluate the sentences you write to? You want us to evaluate your expressions as "true" or as "false"?
I just can't accept that it possesses no intelligence. It is not equivalent to human intelligence, obviously, but how can a system without some semblance of rational thinking solve open math problems? Even composing earlier human work into something novel requires intelligence and understanding on some level.
We couldn't agree on what intelligence means before ChatGPT happened. Now, agreement on the term seems even further away
If performing well on an IQ test or performing at a high level on knowledge work is intelligence to you, these models are intelligent. If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
It hasn’t been passed and no one cares about it because it’s basically an end goal. No lab can hit it so they can’t juice the crazy Turing benchmark 3000 for marketing.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
Low n, time bound, not reproduced. And look at their example conversations…
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
A paper can’t reproduce its self. And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly.
And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition).
And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time.
The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!
> And he didn’t formulate the test with a 5 minute bound he just predicted that by the year 2000 that within 5 minutes an average interrogator would have a sub 70% chance of guessing correctly.
Right, the test duration was left unspecified. This means any duration is acceptable. Including, for example, the only duration actually mentioned by Turing himself in his paper. Or do you have a more authoritative source on which durations are acceptable?
> If one human on earth can consistently get it right then it hasn’t been passed
Says who? Not Turing. Probably he didn't say that because it would make the test both impractical and overly conservative.
> The fact these researchers have to keep adding bounds
The test was not "after thousands of hours of conversing with them, knowing they're AI, THEN see if you can tell them apart blindly." Were 2010 you to be in a real turing test with an arbitrary erudite human and a 2026 frontier LLM, not knowing LLMs existed, you'd probably struggle
This is always the most silly argument. The original test was ambiguous but for sure the human was trying to prove themselves human.
So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this!
Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
I doubt this entirely. It might be quite difficult for said human to discern whether a simple passage were generated sans such accumulated experience in reading AI text, true. But LLMs do not converse like humans in ways that have always been essentially immediately obvious.
Sometimes I wonder if LLMs are just revealing a section of the population with untreated mental illness or if LLMs are actively exacerbating mental illness.
We might eventually regret exposing the general population to such a new technology without almost any safeguards.
Have you seen how many people form one-sided bonds with stars that don’t know they exist? Lonely people suspend disbelief to find some comfort. It’s not proof that the chatbot is indistinguishable from a human companion.
So? The people who can do it prove it can be done - the people who can't don't prove the opposite. Might as well claim all math is wrong because most people don't understand it.
I thought the same then. But the funny thing is that today, it has become a lot easier to recognize the frontier models as not human. All the load bearing and not x but y, etc… weird
The ultimate tell is still the sycophancy. You agree with the suggestion output by the LLM but if it detects even a slight pushback it will completely reverse the previous suggestion. Only way to make it more obvious would be to have the LLM grovel and beg.
If these things have consciousness then we are committing sadism on a massive scale.
It feels like it's getting better at that too. It will push back against obvious nonsense a lot stronger. But you are right that it holds opinions quite a bit less strenuously than humans.
> If intelligence requires sentience for you, then ... well, I don't think we really agree what that is either, never mind how to measure it. But LLMs certainly don't have it right now
> But the consistent trend of the last couple decades (arguably since Turing's time) seems to be that any time a computer reaches our definition of intelligence we decide that that was a flawed definition
I think the mistake here is the notion that there was a definition of intelligence. Or at least a consensus on that definition. Just because compsci nerds of the day thought the Turing test was the final threshold before “real” AI, doesn’t mean philosophers, psychologists and everyone else bought into it. And when we arrived and it turns out to be underwhelming it’s because the compsci nerds made the same mistake they always make: that their models truly encompass all the dense complexity of the real world.
I don't think "intelligence" needs to carry all the intrigue and woo of related words like "consciousness" or "creative." If we just use "intelligence" to mean "the ability of a system to solve problems that are new to the system," that pretty much matches the dictionary definition and normal usage of the term. We don't need to touch messy questions like "is there something it's like to be a bat" to conclude that bats exhibit intelligence when they navigate long distances and hunt for food.
I'm not exactly that you mean by "new to the system", but it seems to me that that definition makes a calculator intelligent, which I can't agree with.
It's a continuum, and things very low on the intelligence continuum might not be referred to as intelligent in everyday usage. But many calculators are Turing complete and can thus clearly perform computations that I would consider intelligent. The basic algorithms used by simple calculators to perform arithmetic would be extremely low on the intelligent continuum.
Intelligence isn't a binary property. Is it really a problem to say that a calculator has some intelligence? That it's more intelligent than e.g. a rock?
I agree, but it's clear most people need a definition of intelligence that (1) they qualify for and (2) nothing/no one they don't like qualifies for. And they'll keep redefining intelligence until they satisfy both criteria.
It has no semantic depth. The sentences and the paragraphs are a statistically viable derivation of existing human text, but once you try to grasp the whole thing with its temporal and spatial dimensions, you are left with a blurry mess that rots your brain. It's a polished, inoffensive and shallow interpretation as written by an opinionated reputation-seeking user of Quora, circa 2019. Assertive, bold, without typos, clean-cut and bulleted, but without an interesting semantic core.
Yeah, I hated all those Quora users that would just spew out semantically meaningless slop like increasing an important bound for the Riemann hypothesis.
Everyone decides what to think on this issue, then finds out facts to support their idea.
As it stands they are massively useful tools, but for generating usable products they require either A) a lot of expert steering or B) a well defined easily verifiable target and a large compute budget. Most people are using them in mode A with good effect, the progress on math has been done in mode B, which is very promising.
Just a year and a half ago their maximal use was rephrase, summarize, and homework-level tasks.
Five years from now? There be dragons.
"But are they generally intelligent?" What a meaningless question!
Not meaningless because part of the discussion is the issue of anthropomorphizing this tech. When we use language like “intelligent” it carries hints of personhood. People begin sadly treating these things as persons.
We can reap the benefits while clearly telling the consumer this is just a language algorithm.
Five years ago they were a niche toy for generating plausible text for entertainment. I was using one. It was an amazing party trick. Now they're popular toys thanks to their ability to entertain CEOs, and they can maintain coherence longer thanks to KV caches and more compute, and someone slapped a chat interface on top, but they're fundamentally the same.
It's just filled to the brim with relations between things. It's good at searching a very large meaning space and create correlations. What it does is to cover great distances and find related things in that large space which needs a long time and large corpus of knowledge to find the connection.
This is not intelligence. It's just a good correlation engine with a very big albeit lossy database of things.
Intelligence is compression, compression requires subtraction, and for some reason LLMs are not good at subtracting. To create a coherent model you kinda have to subtract correlations until only the essential parts are still there.
What I don't understand is why LLMs haven't been able to do this yet, if it's the harness or some orchestration layer above the LLM that is needed. Because fundamentally if you can identify correlations then it's just another small step to prioritize and remove lower value or irrelevant correlations.
I wonder if what's needed is to introduce subtraction tokens in some sense, and in post-training reward the model on that.
Intelligence is compression? What do you mean? Intuitively that doesn't seem right.
>What I don't understand is why LLMs haven't been able to do this yet
LLMs are just trained on what humans have said. Why is it surprising that it's still not possible to reconstruct the intelligence that wrote all that by working backwards? Think of your own work experience. When you look at a piece of code, say, are you always able to discern why the person did what they did, just from the code, with no additional context?
I guess they mean that intelligence is being able to hold models (compressed versions of reality) internally and use them to make predictions with a probability better than chance. That last part is the definition of information.
I find that highly questionable as a general description of what intelligence does. That's more like a description of a general knowledge base. When I think of someone intelligent, I think of someone who's able to draw unexpected connections between seemingly unrelated facts. In the broadest possible terms, I'd call it the ability to make abstractions and analogies. This is not just compression, but the ability to mentally operate on webs of meaning.
Unexpected connections between seemingly unrelated "models" :)
Is a fact stored on your brain like digits on a harddrive? No, it's a pathway that lights up and branches when information enters it. It is dynamic, a compressed form you could say, right? The model holds information, but not all information, but enough to be useful (in decision making).
Arguably it's the same, but the model is probably a "compressed" version of the whole fact that took place in reality.
And you can entertain the models internally and sharpen them. Alone or with others.
Doing those things also contributes to compression. I do recommend reading up on it, it's perhaps a little overstated for what people intuitively consider the two concepts but it's been quite well explored and has held up pretty well in practice.
There were experiments that zipped music pieces, I think, and then classified the compressed files by similarity. They got rather interesting resuls. But I do not remember much details. It was, I believe, about 20 years ago.
That sounds like a really cool project, I will def give it a look.
I still think this has much more to do with the structure of language than the abstract conception I have of intelligence, and I would be interested in having conversation w/ someone for whom the opposite is true.
bzip is not very intelligent, true, but it does develop some model of its input. It's not like there's a linear relationship between between compression ratio and IQ or anything.
I'm not arguing that compression algorithms don't produce models of their inputs, some even use neural nets or other stochastic predictors with correction terms etc.
This doesn't address the only part I really commented on, which is the connection between intelligence and compression.
Creating the model takes intelligence, but running it doesn’t. I think the point everybody’s revolving around is that the transformer model is an absurdly inefficient and low-fidelity approximation of a system that acts, observes consequences, and incorporates that feedback going forward.
The issue isn’t really harness vs. no harness. IMO it’s about the lack of an internally generated sense of what to attend to. Yes, the KV cache accumulates state and its “attention” (if you can even call it that) changes with context. We’ve even managed to /kinda/ close the loop with agentic tool calling and ‘memory’ systems, but these just close the loop at the level of behavior rather than disposition. All agentic harnesses do is make an LLM responsive to the consequences of its actions without changing the tendencies by which it determines what to retain or avoid.
The ghost you can’t escape from at this point is the origin of that relevance. Where does the pull toward one thing mattering over another actually come from? If you ran Fable 5 on a Turing machine and rewound the tape to the exact same state with the exact same input (incl. PRNG seed), it would spit out the same output every time.
Everyone’s trying to outrun this problem by training more often or increasing model sizes. But all this does is inform your model, from the outside(!), what constitutes a better state. The thing that’s actually doing the determining remains unchanged. Congratulations, you’ve scaled the transition function and tape of your Turing machine until it requires every watt generated by ERCOT, and it still cannot, for the life of it, tell you why it should give a shit.
A trained model generating output from weights, a seed, and some context effectively has next-state that’s a total function of those three things. Whatever behavior appears as ‘selecting what is relevant’ is, underneath, just a transition rule executing, no matter how sophisticated or creative the output looks. It can be fully accounted for by what was fixed before it started executing. Which means whatever criterion it uses for determining what matters was inherited from a structure that was already in place before it encountered the situation.
No amount of pruning or post-training can fix this. These approaches just replace one externally supplied criterion with another. For a system to be truly adaptable, there would have to be some criterion by which it treats one possible change as preferable to another, and that criterion itself would have to come from... somewhere. You can even change your conception of ‘improvement’ (e.g. parameter count, harnesses, self-modification, hell, even its ability to spit out shitty best-selling romance novels onto Amazon) and you still haven’t explained where the normative distinction comes from. Every layer of this problem has its root in a preference that was supplied from somewhere else.
I genuinely don’t know if this issue bottoms out anywhere, at least for the way we currently build these systems. Perhaps the solution is still computable, maybe? Who knows what that would even look like. But I’m fairly confident that it isn’t a bigger tape. I hope nobody solves this in the near future because, well, I’d like to have a job...
You're so close... And where is the magic "uncomputable spark" located inside of you? If you say analog thermodynamic noise - then ok, if we use true thermodynamic RNG for LLM activation function, will that meet the criteria? But what if super determinism is the law of the land? Then nobody is anything but computable from priors...
Ehm it's literally every cell in my body. We can't simulate a living cell, we're orders of magnitude off before we can do that, the onus is on you to make an argument that it is in fact, remotely similar to what you describe as "computation".
It's not some tiny "uncomputable spark" you need to look for, most of it is entirely uncomputable.
Where is the computable part in me that is doing all this thinking and being a person? Where is that "tiny spark", point me at it :)
The very fact that it is able to search within a meaning-space demonstrates that it understands semantics, to some extent. Philosophically, that is profound, for something that is just one big matrix multiplication. Drawing connections between things in meaning-space is surely a facet of intelligence.
It’s not intelligence if you are the one who gives the correlations to the model in the pre-training. It’s Word2Vec, applied. Model doesn’t learn anything. You embed these correlations and build it from there. It just searches the space.
As my AI professor said in the first lecture: “All AI is advanced search”.
Okay, I guess you're right that its ability to do this is just correlational, which doesn't imply it has any understanding. However, you have to conclude that some tasks which we used to believe required intelligence don't actually require any, which is disconcerting.
No, what I would say is the tasks which are handled in a passable manner by LLMs can be mathematically modeled with some reasonable accuracy.
Many things are predicted by models in our planet. From weather to production and material science. Building the model needs intelligence, running the model does not.
The person who came up with the formulae for CFD was intelligent. The computer running the model is not. Same for LLMs, chess engines, engine ECUs and financial prediction systems.
Again, for the example’s sake; the person who came up with an algorithm is intelligent. The model mixing its training data to emit something similar is not.
This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
> This starts to feel like you're defining the word intelligence out of any meaning and out of any way we apply that word.
No.
> So when LLMs can do all human knowledge work, and do it better than humans, we'll be in the mines listening to you go on about how it's actually just autocomplete or just math, a distinction that apparently means nothing.
With a big "if" attached to it. People were saying "computers will program themselves in the near future" for, checks notes, 24 years now, as far as I'm aware.
We're constantly building new knowledge and understanding things better than olden days. These models just compress our knowledge and light the blind corners we can't see well. I don't say they are useless, but I say that these things are overhyped.
All they can do is regurgitate human knowledge packed into them and highlight some long-distance correlations between items, which is useful in itself, but it can't jump to somewhere where it's not present its training data, but that's something humans and only humans can do.
> it can't jump to somewhere where it's not present its training data
That sounds like something that can be engineered, can't it? In other words, we can identify limitations in current transformer-based architectures, and we can also build new architectures over time.
Locked in a dark room with no sensory organs, humans couldn't do that.
Most of what you said reads to me as denial.
An unconscious unintelligent but persistent trial and error process created us. We created LLMs. LLMs may create the next thing before we do - hard to say. They don't have all the cognitive tools we have yet, but they still outperform in some areas. As the cognitive playing field levels, I expect you will come to eat your words..
I get what you're saying. The thing itself is just math. I'll just say it depends on how you define intelligence. If at some point we're be able to simulate a human brain with 100% accuracy, I would say that it is intelligent, it sounds like you would not. (I don't mean to imply consciousness or personhood or anything else by "intelligent".)
For me intelligence is a fairly clean-cut concept, and is somewhat inseparable from consciousness itself.
Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
Considering the model even doesn't see the words and work on matrices of numbers is even more telling. One needs to add "tools" and other "experts" to overcome the shortcomings caused by this modus operandi.
I can call the algorithm/model smart as in a smartwatch. It can mimic certain things well while having none of the underlying foundation beneath it, or redirect some of the things to correct tools to get deterministic and accurate results if it can't evaluate the query inside its own network in a sane manner.
Coming to your question, "simulating a brain" in a static manner would not make that simulation intelligent, but if you can "wire" it completely and let it evolve by itself, now we're entering a territory I have not spent enough time for thinking it through.
Oh, as I said "I don't know", an LLM doesn't know what it doesn't know, and can't self correct itself which are required capabilities for understanding something. It just generates something statistically viable via its network.
Your text reads much better if you replace word 'intelligence' with 'text generator with some randomness built in'.
This is because you goal is to state how models are not intelligent, but you couldn't attack the generated text itself, so you created a little rider, attached it to the model, and then you attacked the raider.
But, even in that you failed. You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
A logical fallacy free attack on LLMs would be to show a prompt, and then the response generated by this prompt, where it would be shown that only an entity with no intelligence would generate such a response. Yet, attacks like this are not written here anymore.
You point out that I didn't attack the output itself. But the method you propose is deeply flawed.
I can give you n prompts and m results provided by these prompts, all passed through black boxes. And you can't discern the algorithms or models they have gone through. These boxes can range from simple text generators to MATLAB, Mathematica, CFD applications, correlation engines, linear solvers, mathematical proof-checkers, LLMs, you name it.
For any kind of input they can accept, you can't discern whether the algorithm behind it is intelligent or not, because none of the outputs can be produced by something that doesn't pack some kind of smarts.
How do we pack these smarts in? We teach them as intelligent humans. We pack our intelligence inside them as models (aka algorithms). They do a great job of approximating what we know in a smaller, better-designed problem space. We use these approximations to fine-tune our designs or predict things, then go from there. Just because an algorithm is more capable in processing inputs in some cases doesn't make it intelligent. The way the output looks doesn't make the algorithm intelligent, either.
I have developed multi-agent systems which showed emergent intelligence when the agents came together across distributed systems; I have written high-performance modeling software which can do calculations way faster and better than humans in the materials science space. I'm not doing some kind of armchair criticism of what I'm talking about.
> You compared the source of human randomness in text generation, and called it 'profound' and implied that it is exactly the source of true intelligence. But, then, the temperature, the similar thing in model was "just a randomness slider". Double standard.
Nope, my stance is clear. To quote myself:
> Briefly, any intelligent creature has internal stochastic processes like sensory inputs and feelings to a certain degree. These stochastic inputs and the creature's own actions change the creature in subtle or profound ways. An LLM has no such processes. You push inputs to the same static model, sans temperature which is just a randomness slider.
To expand my quote, humans or any living creatures do not stay static. They evolve due to the sensory input they receive from external and internal stimuli. The temperature slider doesn't do anything close to that. You tickle a static model in different amounts. The model doesn't change after you supply the inputs & temperature and get the output. Creatures do not stay the same. Their mood, behaviors, and stance against life and their environment change, sometimes permanently.
I'll go one step further. We are not intelligent enough to understand other living beings around us. Claiming that we can build AGI tomorrow is a god-complex. What we have done is something arguably useful in some cases, but how this is built is another matter which is worthy of its own discussion. However, today I don't have time to re-iterate all the problems over and over. You can search my comments for that, if you are in for it.
So, no. You tried to attack my comment by finding contradictions in it, but you failed. A better rebuttal would try to similarize how LLMs mirror the human learning process and just read like a normal human, but this is a well-trodden path which has been rebutted countless times in various forms.
I'm guessing whether you believe it possesses intelligence or not depends on your answer to Searle's Chinese room thought experiment[0]. I'd also recommend checking out the Peter Watts' book, Blindsight.
The Chinese room is a good Rorschach test for this kind of thing (but not a good thought experiment, IMO, because it's obviously correct or obviously wrong depending on where you're already coming from), but also it's not really about intelligence per se, but more abstractly awareness and more adjacent to consciousness than intelligence, and these are not the same thing (though it does seem like a lot of people have conflated them somehow, from the conversations around AI).
It all comes down to semantics. And yeah, with the thought experiment, Searle presents three axioms of what could constitute intelligence. Then proposes the chinese room thought experiment as an approval of his third rule, "Syntax by itself is neither constitutive of nor sufficient for semantics." Which only makes sense when you take the other two rules together. Oddly enough, defining intelligence with our capability to derive semantics.
There are a lot of creative counter-arguments to look into on the thought experiment though.
The thing is it doesn't really need creative counter-arguments, it's basically just assuming its conclusion. If you disagree with the conclusion, the argument is nonsense because Searle is just saying 'well the man doesn't understand, so nothing does', and if you agree with the conclusion you don't really need to jump through any hoops to get to it, the third rule is just obviously following from the first and second. And again it's not really talking about intelligence at all, Searle allows that the machine is intelligent from the start, he just rejects that it has any understanding of meaning.
Is that not one of the counter-arguments? You're disagreeing of what Searle determines to be "intelligence." We don't understand how consciousness works, what it's like to be a bat, or how this emergent property came to be. It's why it's necessary to define the terms, then prove them wrong or right. If you disagree with the definitions, that's fine. There is no agreement on what constitutes intelligence.
Again, you keep mentioning intelligence but you talk more like you're talking about awareness/'understanding'/semantics. I agree that these are tricky things to define precisely but there is a reason they are different words. It's also odd considering you mentioned 'blindsight' which is a whole book about the idea that you could have intelligence but not awareness (something I am not sure is a coherent concept, personally).
And yeah, part of it does seem to stem from disagreement about definitions. To me Searle seems to assume much more in his definitions as obvious than he explicitly states, which is why the Chinese room seems like such a non-argument from my point of view.
AIs are better communicators that most of people I have worked with in my life.
They are infinitely patient, don't mind going into more detail if I ask, not too bad at summary, have no ego and don't boast. They are also not too afraid of hurting my feelings, they will tell me my code sux if it does.
I'd don't care if they fit a definition intelligent, they are good colleagues. They have strengths and weaknesses sure, but so do people.
Do they? https://longbets.org/1/ has yet to be settled. Either way, I doubt an LLM could fool anyone here who who knows how LLMs work into thinking it is human, at least not for an extended period of time (think about context length/compression, prompt injections, …).
While being very capable, AI is missing something required for true intelligence and I struggle to explain exactly what it is I see missing.
It's not really "creativity" because much of that always was derivative in my opinion. And LLMs are (for some definition of the word) fairly creative as far as taking known elements and re-arranging them.
I think what is missing is sort of a world model building capability. As humans we see phenomenon and classify them informally and model "what would it look like if this were the cause of that?" type scenarios. We see qualities in phenomena and realize this applies to other things even though the things may be completely different. We run informal "thought experiments" sort of. This is hard to duplicate because a lot (most?) of it occurs outside of systems of symbols like math and language with fixed rules in my opinion.
Anyway yes, lots of human thinking is statistical and LLMs have that down pretty well but they are not "smart" I have concluded and it might be a very long time, if ever, until they are. That isn't to say they aren't very capable tools which they obviously are.
So, right of the bat, you are warning us that you are going to apply the " no true Sscottman" fallacy, and that we should brace ourselves.
Yes, models posses intelligence, but it is not a true one.
Then you claim that models do not posses world-building capabilities. But this is simply not true. Even ignoring the whole subgenre of scientific papers on exactly that subject, it is not that hard to build some hypothetical scenarios, big or small, and then witness the ease with which models do navigate those worlds.
Yes. And they are criticizing a model for not having a default mode network - as if that is some impossibility rather than just an artifact of the current iteration of the specific architectures we have built so far. Why do people paint with these broad brushes over relatively specific complaints?
LLMs are likely for machine intelligence something like drosophila are to biological intelligence - relatively early on the high dimensional spectrum of possibility. Though it stikes me that in a different way they're little alike - drosophila are relatively small and efficient.
How can human thinking be statistical when it is entirely based on ones lived experiences?
When people pretend to know what they are talking about - sure - but even that is not probabilistic - that is the person babbling together mush from their lived experiences.
Statistics has nothing to do with it - these are abstractions humans have invented to try and look at our surroundings objectively.
I'll restate because both objections (which apparently skim instead of read) are missing the important point. Yes LLMs can run "what ifs" scenarios and build models.
However LLMs deal entirely in symbols. 100%. Humans can "world build" aside from this and in fact are often at their best doing so.
Did the first humans to use fire and some form of a wheel even have the capability to talk about it? Think about that.
How exactly is this high dimensional latent space represented? Pixie dust and ethereal forces? Or floating point numbers? Where does it get the weights? Reddit?
Do you hold your experience of "dogs" (for instance) as floating point numbers? The fur, the fear, the love, the wet mouths, the sounds and colors?
Do you experience dogs as lots of action potentials traveling along axons and lots of neurons doing their thing in your brain?
I don't know how it gets from the physical processes or the information processing to our first-hand experiences. So, I can't be sure that a bunch of high-dimensional vectors can't lead to experiences.
Regardless, the claim "LLMs deal entirely in symbols" is wrong as a matter of fact.
Perhaps if you use a very restricted definition of symbols. If you consider symbols to be "anything that represents something" (which is the the sense I use the word in) it is fully the case.
That you might not define floating point numbers to be "symbols" aside, the inputs and the outputs are symbols and the intent and purpose of the creation is strictly symbolic.
It's right there in the name "Large Language Models". Language. Not direct experience, not emotion, not anything else. Language. i.e. symbolic representation.
This does not cover the full spectrum of intelligence humans have, and it shows. And yes, the model can spin up Python parse the output and get mathematical intelligence but there is still a big gap.
As I say, I see the holes. I'm just trying to figure out what it is I see and how to describe it. It's particularly difficult because we don't fully understand how human thinking works but I will say I believe human thinking is a lot more than informal statistical correlation.
The latest LLMs (except Qwen and DeepSeek) are MLLMs (multimodal language models). Unless you count RGB values as symbols, they are dealing with more than symbols.
Yes, there are functional gaps between MLLMs and humans. Their long-term memory is an external mechanism that can use RAG-like approaches, context compression or something like that. The models have problems managing those.
The models can't do continual learning. Although there are promising directions (expert cloning in MoE models, and others).
The only mode of learning available to a model while working on a task is in-context learning. This limits the models to concepts that they developed during autoregressive pretraining and the later stages of training. That is a model can't create new concepts as a result of working on a task (the model's maintainers could choose the task to be represented in the training data later though).
But it's all about functionality.
I guess you have the Leibniz's mill intuition. We can look at how those things work, and there are no experiences or intelligence in sight.
It could be the mill intuition, but my thought is nothing along the lines of "computers can't have souls!" or the human mind is supernatural or anything of the sort.
It's gaps in actual thinking or intelligence I notice. A diff between what I can see or understand and what the model sees or understands. Some are very big, and this in spite of the models having much more knowledge and (presumably) less error prone processing.
My thought is that part of it has to do with inherent limitations of using symbolic representation for "thinking" and I suppose humans have other forms of thinking that occur outside of symbolic representation, and that is going to be hard to recreate digitally.
This is my whole point and I'm not trying to win a debate here or prove "LLMs are useless". Just speculating.
I suspect like most you don't appreciate how terrifying statistical relationships become when you have truly vast data sets to train on... and also that we as humans aren't as shockingly unique as we think (compared to other humans I mean).
By that logic you’d have to call other algorithms intelligent.
With more basic algorithms we know that it’s clearly the human programmer and the interpreter of the outputs that are intelligent and not the algorithm itself. For some reason with AI that goes out the window. I believe it should not.
Its a mirror to human intelligence. Regurgitating phrasing to match what someone who can reason put together, but it isn't any more intelligent than the reflection of you in the mirror is.
I wonder if you went back before we had any idea how the brain worked and talked to the smartest people about how neurons work (without giving away that it's a human brain) then asked them all "would such a system be intelligent?" how many would say yes.
The main problem I have with people stating it's not intelligent or conscious is I don't think we even have a good definition of either word that satisfies everyone. Philosophers have been trying (and failing) to elegantly define these things forever and everyone out here proclaiming they've got the definitive answer and this specific thing they're seeing doesn't fit under it.
Agreed. All I can say is the conversations I have with AI and the things it's able to do for me are more useful than most any human I've come across. Whether that's 'intelligence' or not is a moot point to me. Consciousness is an interesting debate only because, similar to the natural world, if we declare something like "fish aren't conscious / don't feel pain" that then creates a real problem for the fish if we're wrong.
This looks interesting, but would you mind saying a sentence or two about why before I commit to an hour-long video? It looks like it shows how they work internally, which is sort of a non sequitur. Brains also work mechanistically. I'm claiming that any system which is able to do what AIs do must necessarily have some sort of intelligence.
fair reply to an hour video, Scott is just so good to hear his talk is better than I can explain it...
go to 24 minutes and 07 seconds.
it's statistically determining what the next word should be based on all the text it's been trained on. It's not intelligence and he shows what probability it puts on each word that it chooses, but also shows a lot of the other words it was thinking of using. In a later part he shows how it uses words that are not the highest probability (and you question why did it go this route, it's not more correct), but the user never sees this, they see what they think is the correct answer always...
he also shows how context you feed it has a lot to do with what it returns... to the point he can get it to return the capital of France is Marseille, just by typing Marseille a bunch of times before the question. Human intelligence doesn't get confused like that.
And it's not a "hallucination", it's just probability of the next token prediction based on the information it's been trained on and fed, it's not intelligence.
May I suggest the one common in my childhood playgrounds as an alternative?
How do you escape from a perfectly sealed room with a table in it?
You run around the table until your legs are sore, use the saw to cut the table into two, two halves make a whole, you escape through the hole.
Isn't this a case of missing the trees for the forest though? The human brain is not an LLM, and an LLM is not intelligent in the same way as a human brain.
However, an LLM is a prediction machine, prediction IS at the very least one (or the most fundamental) element of intelligence. The brain most surely contains at least some kind of simulacrum of a prediction machine. How that prediction machine is used or wrapped is another matter.
If I said to you: "Blue blue blue, the color of my car is red", would you have absolute confidence in your prediction that my car is red? Or would the way I phrased that sentence make you slightly uncertain, and wonder if there's some miscommunication going on here?
Ok, I want to thank you for finally giving us a concrete falsifiable statement that we can check. I pretended Marseille 40 times before asking Luna 5.6, and the answer was Paris.
So, even with concrete examples, model haters are still wrong.
You also imply the claim that making the distribution of words as the possible next one visible, somehow makes the whole system not intelligent. I would say the exact opposite is true.
By using the embedding vectors, models are aware of precise placement and relative position of words in this hugely dimensional space. No human is capable of such precision. This enables party tricks of "king plus woman minus man" kind. But this also give us a precise point between any two words, no matter how different. What is on the midpoint between volcano and music, for example. No human can precisely answer that, but an embedding can. And we can see which words are closest to this 700 dimensional point.
You see this menu of words as a weakness, and I say it is in fact a sign of super intelligence. And this is all before any reasoning or attention mechanism is even run.
No he says in the actual talk which model it occurred on and it was an older model he was using that caused that to occur with Marseille. They have since corrected it from doing that anymore. It was only used to illustrate the prediction machine that it is...
I don't see the many weighted words as a weakness, I see it opening up what's under the hood of the prediction machine that it is.
LLMs are very cool tech, definitely not a model hater, the use case on when to use it makes a difference, it's not AGI.
You're confusing language use with intelligence. Fair enough, they were fine-tuned to do that, but still.
Great that it has some 700 dimensional model of language.
If that is a sign of super intelligence, then so is an encyclopedia?
Also I'm just curious how do you think it is "more intelligent" for having a vector representation for a meaningless thing such as "the midpoint between volcano and music"?
LLMs are pattern prediction systems with a large training data set. It is not surprising that they can predict patterns, particularly for a well structured field like mathematics that is also amenable to automated proof checking to help steer it.
"AI" is a vague term that encompasses both traditional GOFAI, LLMs, and future tech so your question is meaningless. Non-human intelligence is possible, if that's what you're actually asking.
What part of my statement do you take issue with: that LLMs are pattern predictors (that's literally what the algorithm that runs it does) or that mathematics is rules-based and checkable and therefore amenable to automated pattern prediction?
My issue is that I can't predict what you'll find surprising, if an LLM (with a harness) would be able to do it.
If you are able to check that the results of an LLM are satisfactory, it means that the results are checkable. Then, in retrospect, the process of LLM coming up with those results is rule-based, because an LLM is a large set of data manipulation rules.
In short, which concrete thing that an LLM does would surprise you?
I guess it will take quite a while before we reverse engineer the human brain to find all the optimizations and shortcuts that evolution has used to make the human brain reach intellectual maturity in just about 20 years.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
The problem might be that it cannot backtrack. When AI generates output, there is no backspace key for it - it uses "No, but wait!" all over instead, which is very different to human output.
Subagents and/or branching conversations are presented as the solution to this - if you can't backtrack, then branch off a conversation to explore multiple paths (discarding the ones that didn't pan out), but this is a fix in the harness not a fix in the model. It's also literally how we made chess-playing engines back in the 80s: recursive path exploration with a fixed depth.
Humans don't exactly work that way either, AFAIK. So we have this uncanny valley of intelligence: it's some sort of intelligence, but not as we know it.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans. Which says a lot about humans.
Idk something about a mirror maybe and the output reflecting the input?
> I feel like that what a lot of people who say this don't seem to grasp, is that despite this flaw its still often capable of saying more interesting things than a lot of humans.
So does the Google search bar, but I don't ascribe intelligence to it.
I don't find it's particularly hard to define loosely, but then I don't think of it as a special property of humans other than it tends to be quite high in them. But we are obviously talking about different things and if you're not going to provide a definition then it's not really the basis for a productive conversation.
To address this similarly to my sibling reply, I don't have a definition of intelligence that provides value here.
And your loose definition isn't doing a lot of help either, beyond perhaps noting: that Google search bar _is_ similarly "intelligent" to an LLM? Which says what, a lot about search? A lot about modern LLMs?
I mean I'm quite proud of some of my search queries in the same way I'm quite proud of some of the LLM output I get. I'm probably just very arrogant and enjoying myself via some LLM indirection.
Am I the only one that sometimes reads back particularly good emails they've written? I feel like its a similar thing :).
I strongly agree. Is a human in a vegetative state sentient? What about when they're asleep? What about someone with brain damage? What about someone with an IQ of 20?
Ray Kurzweil argues sentience is a philosophical question, and doesn't have much value as applied to science and technology. What will change the world is how this intelligence is applied. No one's going to care whether AGI is defined as sentient when it creates cheap fusion energy.
This seems like a complete waste of time given the more practical and more urgent need to clarify to everyone involved that current LLMs are not actually intelligent.
"The fundamental cause of the trouble is that in the modern world the stupid are cocksure while the intelligent are full of doubt."
— Bertrand Russell.
I'm not calling you to action, I'm explaining why I don't feel inclined to engage in philosophy and discuss "the concrete outcomes of intelligence" given a more pressing, pragmatic need.
It feel it's self-evident that we must fight the good fight of dissuading as many people as possible of the notion that LLMs as we have today, and likely forever after, are actually intelligent. Delaying this fight allows the current, stupid belief to the contrary to fester.
I don't think we'll win the majority of people over by debating the nuanced meaning of the word intelligence to a very precise degree.
I think we ought to do it by shaming them every time LLMs fail.
Oh, in that case I agree strongly with William. I think the definition of intelligence a complete waste of time and the only real question is, can these tool solve problems for us? The answer is clearly yes, there are some problems they can.
The really interesting question is still a few years away when we ask if we humans have the right to turn these things on and off? ;)
AI is just a good permutation/combination engine that tries to act smart with help of statistics. At best I only see AI as, 1. An autocomplete on steroid, 2. Good search/correlation engine
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct
Isn't this how human brains work? We're just large probabilistic neural inference machines. A network of neural weights guided by past training.
To be honest I think the debate about what "sentience" is is inconsequential navel gazing. There are many different kinds of intelligence - even within humans. What matters is how useful that intelligence is as applied to solving real world problems. Like it or not, LLMs produce intelligence which is very, very useful for 1.5B people and rapidly growing.
I think this ultimately boils down to the classical economic debate of marginal utility. There isn't an objective way to value a product or service. Each person decides for themselves what said product or service is worth based on their needs and preferences. Intelligence works the same way. We don't have the right to tell someone that their perception of the value of that intelligence is wrong. They alone determine that.
One thing we can all agree on, is that the capability of this intelligence is expanding rapidly. In a few short years, we went from Will Smith spaghetti hands to full length movies and strikingly realistic images. For 60 years, passing the Turing Test was considered Star Trek level science fiction. Last year GPT 4.5 passed the Turing Test. AI is already being used to convince people over audio that they are real, and very soon, this will occur over video.
I think people are being too dismissive of this intelligence. It doesn't need to be perfectly humanoid to be considered intelligent.
>Isn't this how human brains work? We're just large probabilistic neural inference machines. A network of neural weights guided by past training.
Artificial neural nets are merely cartoons of how people guess human brains work. They're likely much further from reality than, say, the fundamental laws of physics which can be experimentally verified or disproved.
We have actually little idea how human brains work.
Every age thinks they know how they work, and then every subsequent age laughs at the previous age’s rudimentary understanding.
We gave a Nobel prize to the psychiatrist for inventing a method to remove the frontal lobes of a brain through the nose in 1949. Lobotomies were performed through the 70s.
We are likely doing similar if not subtler but worse things today (just one more pill bro). We still have no idea what we are doing when it comes to the brain.
The statistical nature means that what an AI produces is basically the average across all training data, making the generated text extremely bland and personality-less.
I am firmly realistic on the overhyped nature of AI, but which discoveries in the last 100 years are not "a composition of solutions humans have developed and documented elsewhere" ?
Is that basically every new discovery? And under the strictest definition of novel and NOT falling into your composition of previous solutions what is that standard of proof to beat your criteria? Is a novel discovery not allowed to use English but must invent their own language? Must they invent their own math - these are hyperbole for illustration but I think its not far from that before you could just argue anything based off it is a composition of existing ideas
> it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever
How could I seriously repeat that prayer, when it builds things I wouldn't be able to build and solves problems that I wouldn't be able to solve? I would have to assume that nothing I did in 25 years for money required any intelligence or critical thought whatsoever and I have higher IQ than 99% of the population. You might be comfortable with that but I'm more comfortable with ascribing at least some intelligence and critical thought to AI.
not having any body, continuous sensory input (except ChatGPT-live gets streaming audio), episodic memory, on-the-job learning (live weight updates), any live feedback loops (like moving a motor updates proprioperception or turning physically changes what a streaming camera sees), really hurts these models' abilities to perform tasks of the kind you're waiting for.
They are very good at instruction-following and you can teach it a new task that fits in its context and it'll learn it and do it. Go ahead, you can make up some new brand new ruleset or behavior and instruct it to follow it and it will. That's amazing.
But it won't be any better at its new behavior after an hour or ten hours or ten days. It doesn't have the kind of adaptation that we expect.
What it is able to do already is pretty amazing, but what it lacks is also a great hindrance to seeing its full capabilities. We just have to wait until research labs add these missing components.
Don't discount the capability of representing human-like intelligence with simple constructs. Give enough parameters and advanced enough training you can without a doubt create real intelligence. We've seen some of this already with "j-space" where llms have started to exhibit reasoning before it ever reaches the output head.
> is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever
What makes you so convinced that a algorithmic construct of neural nets cannot be "real intelligence or critical thought"?
Alright, take it easy. You typed a lot here but you're not actually saying much. LLMs produce useful outputs, their usefulness is just proportional to how well you know how to use them. Everything else is navel gazing.
A decent correlation engine is still extraordinarily valuable for science, investing, prediction, etc. Plenty of human minds are strong in the same area.
Are we sure there is some objective, technical definition of what is intelligence and what is not?
Isn't it rather a subjective philosophical concept? What if human intelligence is also a statistical model, trained by evolution to make decisions that lead to offspring?
The one major difference I see between AI and people is the ability to learn and memorize. All memory/learning solutions that current AI architectures offer just feel like workarounds and simply don't work anywhere near as a person learning something new and remembering it.
> The "something deeply wrong" part about AI, that even most technology enthusiasts evidently do not seem to grasp, is that it is still fundamentally a statistical model — an algorithmic construct — and does not possess any real intelligence or critical thought whatsoever.
Not that I'm saying AI are like brains, but can you describe why brains, which are fundamentally slightly dodgy electrochemistry with frequent literal delusions of grander, are not "statistical"?
> No matter how much investors and tech companies want you to believe that they are on the verge of super intelligence, nothing I've seen to date can not easily be explained by "correlation engine", including the "novel" math solutions, all of which appear to just be "a composition of solutions humans have developed and documented elsewhere" upon deeper inspection.
Ditto, when do we humans do things exceeding the parameters of "correlation engine", especially if you consider compositing things either we or some other part of nature has developed and documented elsewhere to be insufficient?
I mean this in the kindest way possible, but you are wrong that the math solutions are that easily dismissed. And there are many more than are publicized. A specific math problem I wanted solved for 3 years did not get solved by any model until fable and, and I tried it on every model and know the literature surrounding it well.
After 9/11 (and long before), the majority of historians and academia warned that the political response to terrorism — the destruction of civil liberties, implementation of mass surveillance, Islamophobia, populism — were a greater threat and risk to our freedoms than the terrorists ever were.
I remember, after the Snowden leaks, when anyone warned that governments could use the patriot act and the various other "intelligence" power grabs to destroy democracy and implement totalitarian dictatorship, they would be deemed insane conspiracy theorists. The hilarious part is most of them (including me) were warning about a distant future... yet it barely took a decade for America to collapse into fascist populism, declare ANTI-FAscists the real terrorists, and turn the surveillance apparatus against them.
FYI the setup you mention is not "end-to-end" encrypted. E2EE means client-to-client encrypted, with the server processing encrypted bits only. Your approach is encryption in transit and at rest. At rest is relatively irrelevent for large cloud providers, as they are probably better at managing the lifecycle of disks than most businesses or people. It's unlikely someone's going to physically rob a data center or end up with a refurb drive that hasn't been thoroughly processed and wiped.
It's not necessarily more secure than managed providers either, simply because you are probably not a security engineer, and have far less resources to secure your server. It does prevent Google/iCloud from scraping your data, but it certainly does not mean Hetzner can't access your data. They control the overarching hypervisor and control plane managing your servers/VM's, so there's no way to know what capabilities have been implemented. The majority of what intelligence agencies are capable of has not been leaked or documented publicly.
For 99.99% of the population, does the threat model really include targeted NSA attack with higher probability than „Google’s automated system sent the police to my door“? No, no it doesn’t.
It is demonstrably more secure for even a semi experienced sysop to host Immich for their family than for them to use Google/Apple.
But I do agree that _some_ experience is required.
In memory. If the police shows up and they disconnect my server to sieze it, for example, the photos are lost to them.
And the Hetzner employee would need to specifically target me, because I doubt they would implement a dragnet that pierces through the bespoke random process on my bare metal server to scan the photos in memory.
That is a lot more secure than „Google scans all pictures routinely and fully automatically sends the police to your door, and you have no recourse if they are wrong“ that the EFF article discussed.
He's understandably using the wrong term because E2EE should include private servers as a variant, but does not by definition.
In the case where I own the server end of the communication, the data being fully encrypted on that server is far less important than when it is stored in the cloud.
The effect is the same however, since sending encrypted photos over HTTPS to be decrypted client side would be completely unnecessary in most cases.
Seriously! How do techies and devs of all people not understand that the cloud is someone else's computer, and that the best way to prevent leaks, exploitation, or abuse of user data is to prevent anyone from being able to decrypt it but the end users themselves.
IMO this is the single greatest problem with the selfhosted community; the idea that E2EE is only necessary for passwords and other highly sensitive PII. It should be standard for anything hosted on someone else's computer.
You might argue it's not neccessary for cat photos, but mistakes happen and you can accidentally upload things you don't intend to. You might argue it's not neccessary for games, ebooks or other copyrighted media, but the cloud provider could scan and delete anything you own that matches a hash of copyrighted material, at any time. You can accidentally paste a password, or other sensitive piece of text, into any text field of any website or application, and have it distributed to computers around the world.
E2EE can mitigate against numerous attack vectors, and reduces the surface area and blast radius of most attacks. That also applies to your own computers, if someone steals your hardware or hacks into your network. It is vital in the age of AI where all of your data could be exploited for training and profit, or used against you. The only data that should not be E2EE is situations where it is technically impossible, or the data is explicitly shared as "public" (e.g. the clearnet).
> It should be standard for anything hosted on someone else's computer.
As long you understand the risks.
I'd rather have my family photos beying unencrypted than a very good possibilty of loosing them which happed more than once with other e2e things simply because I have no key to decrypt.
Then again - if I have to chose I'd rather have the at my home lab.
I'd personally rather have E2EE and periodically back things up to an encrypted hard drive so any losses aren't catastrophic, but I am probably more cynical than most in my trust of companies/other people with my data and am technical so understand the risk model both ways better than most people.
>I'd personally rather have E2EE and periodically back things up to an encrypted hard drive so any losses aren't catastrophic
Are suggesting backing up decrypted data to an encrypted hard drive?
>better than most people.
Typically people who are so high on themself either too young or simply actually lack the understanding and usually it shows. But sure, most people out of 8 billion around the world have no idea about those things.
Anyway the whole point of my comment is:
- there is data I'm not willing and don't have to share - and it stays with me. Like family documents, photos etc.
Why is there "a very good possibility" of losing your photos because they are E2EE? Do you not use a password manager and backup your data?
There is no reason why E2EE services can't provide recovery or emergency access mechanisms, or implement plaintext export functionality from clients for storage elsewhere. Most reputable providers already have functionality to enable recovery and backup.
No it's not. The point of E2EE is that only the client apps decrypt/encrypt content, and the server just processes the encrypted bits. Most E2EE service providers do this by encrypting your encryption key. That's how you can login on other devices without you having to store and import an encryption key every time. When you login they send you an encrypted blob that contains your encryption key, which is decrypted client side, then the key is used to decrypt your data locally on the client. This does not break E2EE, but it does mean you have to trust the provider, which is why most of them are entirely open source.
Sharing and emergency access also use similar public key cryptography techniques to provide shared access to E2EE data. A similar principle applies to your phones encryption aswell, and is the reason you can wipe/reset your device in seconds, instead of minutes/hours. They only wipe the encryption key; not the encrypted data.
I think you are slightly confused about E2EE, let me try to address a few points:
- The whole point of end-to-end encryption is that nobody in the middle (e.g. the server) can possibly have access to the encrypted data. If they can (e.g. by giving you a way to reset your password), then it is not end-to-end encryption, period.
- It is possible for a provider to "encrypt an encryption key" with another key they cannot access, e.g. a password. In that case your password (probably through a KDF) is the key to decrypt the encryption key that is used further.
- It is possible for an app to store your key locally. Ideally you would store your key in a secure element on your phone for instance, and unlock it e.g. with biometry.
- In any case, if the client (e.g. the app) is not open source, you cannot easily know if it does what it says. That's why it's easier to trust Signal than WhatsApp.
- If you lose your "key" (whatever it is that the server doesn't know), your data is lost. If you can lose your key and the server can help you recover, it means that the data was not end-to-end encrypted and the server could access it without needing you at all (obviously, because you lost the part that they needed).
Now, in order to seriously trust E2EE, you need to trust the implementation. That means that if you cannot audit the code (typically, WhatsApp), then it's difficult to trust it. Another example is ProtonMail: you open the webpage in your browser, and at this point it downloads a client that is supposed to do the decryption locally. How do you trust that client? You could try to audit it, but next time you reload the page, it may download modified sources, and you won't know (there aren't tools to pin the version of a webapp from the browser).
In that sense it's even harder to trust ProtonMail than WhatsApp: it would be trivial for Proton to serve a different client that would exfiltrate your password just for you, just this one time, and no audit by anyone else in the world would ever detect it. So when you use ProtonMail in the browser, you trust the Proton server, which defeats the purpose of E2EE. It is still better than non-E2EE like GMail (where you know Google reads all emails of everybody), but you still trust the Proton server, which is not what is usually expected with E2EE.
In a recovery scenario, you don't have the password or whatever you used to encrypt the key.
And again: the whole point E2EE is that you don't have to trust the provider. Open source in general doesn't help you here because you don't know what software or which version of the code the provider is running. Your only chance is to run open source client software, review it to ensure E2EE has been implement properly, then you don't need to trust the provider.
Either the provider knows the secret, then they can help you in a recovery scenario, but then they also can read your data. Or they don't know the secret, in which case they cannot help you in a recovery scenario.
I don’t agree E2EE is right for everything, and especially not for a personal photo library.
I don’t want to hold the keys to my photo library on someone else’s computer. I want to actually have all the bits and all the hardware in my house. I want to have access to it even if the Internet ends.
Especially that it would come with the loss of quite a few features, or at least a significantly worse way to implement them.
Like if you have to bring the data to your client device to do any kind of processing, you are quite bottlenecked when it comes to bulk operations (e.g. searching).
Sure there is very interesting research into managing that (homomorphic encryption), but I think it only makes sense on a Google cloud/apple scale. For a small, self-hosted app, I would much rather have my own hardware with FS-level encryption , or some kind of trusted compute as a whole. I don't think this has to be solved by an image host service.
Also bottlenecked when it comes to background operations.
Trusting the server, all the app in your phone has to do in the sliver of CPU time the OS gives it in the background is send it off to the server, where it can do compute-intensive things like transcoding video.
If you don’t trust the server, you’ll probably have to do these things while your app is in the foreground.
Sounds like you want your photos on a unencrypted HDD, which you can do regardless of whether or not a cloud service is E2EE, so I don't see how E2EE is an issue...
No, I do not want that. I have my photos on an encrypted HDD, in a server running Immich in my basement, and I connect to it using WireGuard. Everything is encrypted at rest and in transit.
There’s no cloud.
I get to reap the benefits of not using E2EE encryption, like offloading machine learning tasks and transcoding to a server, and having extremely simple clients that don’t need to roll their own application-layer crypto.
E2EE isn’t a silver bullet.
It solves a specific problem — trusting the server — and introduces another — pushing complexity to the clients. If you already trust the server, because it’s running on your infrastructure, there are no upsides.
And Immich was designed specifically for self-hosting. To not depend on the cloud. It makes no sense for it to make trade-offs that don’t benefit self-hosting.
All current age verification measures open up a torrent of attack vectors on user PII and privacy. Limiting the number of entities that are able to access data is one of the best ways to prevent it's leak or abuse. Don't let perfection be the enemy of good.
But therein lies the fundamental problem with surveillance capitalism. Until the sale of personal data/metadata is outlawed, the practice of targeting content based on an individuals personal data/metadata is outlawed, there is a highly punitive cost for violations and leaks that make storage outside core business functionality a major criminal and financial risk, and the compilation of this data by "intelligence" agencies it treated as a critical attack vector to national security – the attack on each citizens civil rights that it truly is – most privacy laws and regulations are just virtue signals designed specifically avoid the root causes, and further entrench the power of monopolies and incumbents.
FYI I don't believe Google sells user data. They sell products which leverage user data to give them a critical advantage over every competitor who does not have trackers in everyones pockets/computers, does not store their entire web search/browsing history, etc. It's in the interest of big tech to protect their market advantage (like ZKP, which would prevent competitors from having a new gov-mandated vector to compile user data).
So, if the majority chose to get microchipped, you believe either we should force the minority to get microchipped against their will, or just exclude them from society?
Having the ethernet port and Thread radio gated behind the 128gb model is obnoxious.
I have three Apple TVs that are ethernet connected and form the backbone of my home's Thread network, but they have <5 apps installed and would do fine with 32gb rather than 128gb. (And in fact, they are all currently 32gb models from the previous generation where those did include ethernet.)
I don't know what makes sense for Apple's supply chain and BoM, I'm just saying that the price for my home's worth of Apple TVs with the minimum functionality I use has gone up by over 50% since the previous-gen model and now sits at over $1,000 in local pricing.
That's the kind of pricing that makes me start to consider the Google 4K Streamer even if it's a UX downgrade - for $300 I get ethernet and Thread on every TV.
That’s infuriating. I was hovering over the buy button last week, and now that’s a deal breaker. I was already going for the premium price point for hard-to-justify reasons.
Update: while I am terribly unhappy to give them money, there are still retailers who are listing the previous price and I was able to scoop one up before the hike went into effect.