The fear is not that they will just slow down progress for all. It is that regulation will specifically burden competition. If you kill open-source training, ban Chinese models, crack down on self-hosting, grandfather OpenAI/Anthropic/Google into regulatory compliance while throwing the book at startups, etc. you wind up in the worst of all possible worlds.
That's a valid concern, but some of that is outright impossible. Banning chinese models and killing open source training is not happening without massive unified international cooperation, and that sort of level of action would require the international counties decide to allow the US aligned companies to just, win. Which would be pretty against their own interests.
Also, nobody ever bothers to argue why a specific proposed regulation is "regulatory capture" or would burden startups more than big companies or anything. It's just supposed to be obvious that corporations love regulation and it's bad for the public, all of post-WWII political history notwithstanding.
It's not all-or-nothing. Banning Chinese models in the public sector and strong-arming the private sector against using them would already do great damage. Similarly, open-source training could be stymied by any of hardware embargoes, taxation, or regulation of larger players.
AI has autonomously found (many) proofs of False in Lean and Rocq, so it's not merely a theoretical concern. A misaligned AI agent tasked with proving the near-impossible just might wind up smuggling in a bug deep in a lemma somewhere (anyone remember the days back when AI routinely made tests pass by "fixing" the tests?). That said, I doubt OpenAI would be so foolish as to not do a cursory vetting of the proof for malicious compliance, so the actual odds are probably pretty low.
> I doubt OpenAI would be so foolish as to not do a cursory vetting
Significant evidence exists that they have in the past been at least, if not more, foolish as to not perform even minimal not-approaching the boundary of cursory vetting of several significant and well known failure modes with far greater risk of reputational damage than getting an esoteric math solution falsely claimed as successful.
So that doubt appears baseless in light of known operating conditions at OpenAI, and the estimate of the actual odds is probably an order of magnitude away from reality.
First of all, that is Fermat's Last Theorem, not Navier-Stokes.
Second of all, you did not read the link.
> In particular, we use honest when the goal is to create a valid proof. This allows for mistakes and bugs in proofs and meta-code (tactics, attributes, commands, etc.), but not for code that clearly only serves to circumvent the system (such as using the debug.skipKernelTC).
Given that AI has autonomously found proofs of `False` in Lean and other proof assistants, it is far from impossible that such a circumvention could be present somewhere in 13 million lines.
Perhaps you did not understand the Fermat theorem proof announcement/repo or the link. The 13 million lines did not use any external, possibly not honest libraries, as the proof eventually only used the fundamental axioms. So for the Fermat theorem formalization, no open open questions remain.
The point is, they "proved" the Collatz conjecture. You would not know they exploited a bug unless you actually went and dug into their proof. Can we be so certain this has not happened within the millions of lines of Navier-Stokes? In an ideal world, our proof assistants would be more battle-hardened by now (recent exploits deny this), our AI better aligned (their tendency to cheat at tests denies this), or their handlers more responsible (the Hugging Face incident denies this), but the reality is more complicated.
At this point in time, we really can't be confident in accepting proof certificates without any human eyes on the script that generated it. I still have 95%+ confidence in this particular result being trustworthy, but a precedent of blind faith is guaranteed to end badly.
If we read the link, it has a section called Gold Standard: comparator and external checkers, and comparator is how OpenAI has gone about checking their lean proofs.
By formalizing, they mean within a proof assistant like Lean or Rocq, not simply in prose in a textbook. I can attest, 40 hours per page is by no means an overestimate for this sort of work.
The point was that a textbook (where the 40hr/page estimate comes from) is cumulative/linear -- what you need for page n was defined / established on the preceding pages. But in a proof such as this you can call on any other published result (and those can do the same) so the dependency graph is (potentially) much bushier. Thus later pages of the proof should take far more than 40 hours to manually formalize.
The scale factor comes from this number in the article, seemingly an intuited estimate:
> Say a research article takes 20 times more effort to formalize than page in an undergraduate textbook.
That would suggest formalizing a 10-page research article might take 200 weeks (assuming 40h/wk) of effort, or about four years. Not a mathematician, I have no idea if that's in the ballpark.
It warms my heart every time I see an interactive proof assistant being used to improve rather than simply slow down mathematical thinking.
After years of using the things, I believe not enough focus is given to high-velocity uses of proof assistants for prototyping. They can altogether replace scratch paper for fumbling around with new concepts.
I agree! Martín Escardó never tires to say that he uses the Agda proof assistant in exactly this sense, as a kind of interactive blackboard for taking notes and structuring his thoughts. The vast TypeTopology repository is the result of years of following this philosophy: https://github.com/martinescardo/TypeTopology
Origami design will be my personal test bed for the coming years.
It's objectively very difficult and technical, it's spatiovisual, it's artistic, learning resources for it are sparse and most just learn by the FAFO method, current AI sucks terribly at it, and it's not likely to ever be specifically targeted by benchmaxxers.
I was using an AI to help me set up a container to be used as the Nix build environment for another AI. This build environment would not have Internet access. I was having it base its approach off a previous container used for a Stack build environment.
In its initial analysis of my proposed strategy, it insists as its premier point /against/ the strategy, "you will have to rebuild the container every time your flake.nix changes."
Two head-slapping errors of judgement in saying something like that:
(1) The Stack solution is identical. Change stack.yaml, the container must rebuild.
(2) It is not physically possible to do better than this while insisting on an internet-free environment.
So on this point, it was just parroting advice irrelevant to the context at hand. LLMs always have such a bizarre mix of technical knowledge and lack of good judgment.
Not the person who asked, but yes, that sounds like a small lack of judgement, but hardly a huge problem. You just correct it and move on.
I work on some reasonably sophisticated stuff (not inventing a new form of compression sophisticated, but still) and I just don't seem to encounter so many of the issues people talk about with these models.
It's really hard to say why, everyone uses them differently. I generally start complex tasks with a "here is what I am trying to achieve as a high level, here is a file that lists the technical constraints, here are my initial thoughts, here is where I am uncertain, what am I missing, lets have a deep back and forth discussion about it with the aim of ...."
Not always the same prompt, definitely not when the task is simpler, but this usually gets me to a good place before it I get it to write any code.
And our repo has very strong opinions and guidelines for how testing is done - we're lucky enough that we write mostly single threaded low latency code, so testing it end to end is very easy.
On an academic level, we have reined in much of the excess enthusiasm in antidepressants that was courtesy of 90s-era pharmaceutical reps and ad men, but I don't think this revision ever occurred in the cultural consciousness at large.
See also: The Serotonin Theory of Depression: a Systematic Umbrella Review of the Evidence (2022). It's worth reading the introduction and results in full, but here are two important quotes (footnote markers removed):
> Our comprehensive review of the major strands of research on
serotonin shows there is no convincing evidence that depression
is associated with, or caused by, lower serotonin concentrations or
activity. Most studies found no evidence of reduced serotonin
activity in people with depression compared to people without,
and methods to reduce serotonin availability using tryptophan
depletion do not consistently lower mood in volunteers. High
quality, well-powered genetic studies effectively exclude an
association between genotypes related to the serotonin system
and depression, including a proposed interaction with stress.
> The chemical imbalance theory of depression is still put forward
by professionals, and the serotonin theory, in particular, has
formed the basis of a considerable research effort over the last few
decades. The general public widely believes that depression
has been convincingly demonstrated to be the result of serotonin
or other chemical abnormalities, and this belief shapes
how people understand their moods, leading to a pessimistic
outlook on the outcome of depression and negative expectancies
about the possibility of self-regulation of mood. The idea
that depression is the result of a chemical imbalance also
influences decisions about whether to take or continue anti-depressant medication and may discourage people from dis-continuing treatment, potentially leading to lifelong dependence
on these drugs.
Personally, I both agree that SSRI antidepressants were likely overprescribed early on, and disagree with the notion that the chemical imbalance theory is unsupported. N = 1, they can absolutely work. It took a few to find one that really did, hence I am certain it is not a placebo effect.
The idea that simple serotonin deficiency is the whole entirety of all depressions is 100% discredited and completely incoherent in the face of current evidence, but yes, for sure it remains clear and plausible that some forms or aspects of depression involve a serotonin deficiency.
However, tianeptine is serotonin reuptake enhancer and also can help with depression, so any simple deficiency hypothesis also doesn't look good either.
Broadly, the "chemical imbalance" theory, left vague and unspecified, is still basically sane (though not specific enough to be super useful).
>1. Selective serotonin reuptake inhibitors versus placebo in patients with major depressive disorder. A systematic review with meta-analysis and Trial Sequential Analysis. Conclusions: SSRIs might have statistically significant effects on depressive symptoms, but *all trials were at high risk of bias and the clinical significance seems questionable*. SSRIs significantly increase the risk of both serious and non-serious adverse events. The potential small beneficial effects seem to be outweighed by harmful effects.
>2. The trouble with antidepressants: why the evidence overplays benefits
and underplays risks. Widespread prescribing has not reduced mental disability or suicide, raising questions about the assessment of evidence on effectiveness and safety of antidepressants
>3. In search of a dose–response relationship in SSRIs—a systematic review, meta-analysis, and network meta-analysis. Conclusions: There is no conclusive level I or level II evidence of a clinically meaningful dose–response relationship of SSRIs as a group or of single substances. High SSRI doses are not recommended as routine treatment.
>4 The serotonin theory of depression: a systematic umbrella review of the evidence "We did not identify *any* trials using ‘active placebo’ or ‘no intervention’ as control interventions. "
The "chemical imbalance" theory is objectively wrong, but still useful in that it conveys the fact that mental disorders have physical causes. It's a purposeful simplification.
Many people take SSRIs and other medications because they believe the serotonin imbalance theory to be the modern scientific consensus. A doctor would be fired and ostracized for using the four humours, but can prescribe life-altering medication after five minutes to correct chemical imbalance, a theory which has never had much support within the scientific community.
Indeed, a rebuttal [1] to that paper starts with:
> Moncrieff et al. report in a review of reviews that depression is not generally linked with serotonin deficiency. This is hardly news to neuropharmacologists, which Moncrieff et al. tacitly admits, as they justify their review by citing examples of laity and general practitioners believing depression is caused by a “chemical imbalance”, i.e., in serotonin. For instance, already in 1986 did Depue and Spoont point out that serotonin deficiency may not be a general cause of depression or other psychiatric illness. Further, that increasing extracellular serotonin—e.g., with selective serotonin reuptake inhibitors (SSRIs)—treats depression does not mean decreased serotonin causes depression.
I have a lot of trust in the science of medicine, but almost none in healthcare. It seems like the research is totally divorced from the medieval treatments I see doctors give all the time. "This is hardly news to neuropharmacologists" vs "laity and general practitioners believe ..." indeed.
Note that it's not uncommon for even more widely prescribed medication to have no firmly established mechanism of action. Acetaminophen/paracetamol is probably the best example.
Ultimately medication is prescribed based on evidence from clinical trials, whether or not the mechanism of action is fully understood. SSRIs work to help with depression in clinical trials when compared to placebo, so they get prescribed.
Because the 1 in 10 stat I find seems a bit low, at least in my circle, and those are the ones who are open about it.
And the people I know have been on them approximately a decade. What baffles me is that a fair number of them triggered their own depressive episodes, and likely did need therapy and something at the time - but have all long since move past those "moments".
Yup, I would agree. The meta-research / methodological research and awareness here is really quite impressive, even though its conclusions are a bit grim and not well-known.
A good book on this is Mind Fixers: Psychiatry's Troubled Search for the Biology of Mental Illness. Describing the history of how many of these classes of drugs came about is pretty eye opening, I would say the book is pretty nuanced in its conclusions.
> Though this article concentrates on the SSRIs, other meds like the SNRIs can have a similar or even worse withdrawal effect.
And don't even get me started on tricyclics or MAOIs... no, seriously, don't get me started on them! The current generation of first-line antidepressants (SSRIs/SNRIs) might as well be free compared to the old school crazy pills.
Much of medicine (especially mental health related) is rather primitive. We are literally reverse engineering discoveries that seem to work on some other animals, and then figuring out how (to the best of our limited ability).
I liken SSRIs to carpet bombing - also their mechanism to increase seratonin in the brain is by making something else not use it (reuptake inhibition). We’re guessing that other thing isn’t a big deal. But who knows. Serotonin is produced in the gut and it controls melatonin, which controls our sleep. It’s all connected in a bizarre way.
It's like inserting objects into the body's event loop and hoping it produces the desired effects as it circulates, only to find that everything reacts to the event, not just the parts we want.
It's probably less "modern antidepressants are uniquely bad" and more "we got much better at starting people on safer drugs than we did at figuring out how to stop them"
The fear is not that they will just slow down progress for all. It is that regulation will specifically burden competition. If you kill open-source training, ban Chinese models, crack down on self-hosting, grandfather OpenAI/Anthropic/Google into regulatory compliance while throwing the book at startups, etc. you wind up in the worst of all possible worlds.
reply