We are a student-led initiative. Our sponsors don't pay us and don't have a say in our decisions. All of our funding goes toward our judges and participants.
Hey! Thanks for answering, i appreciate it. Dont get me wrong, this is what you should be doing, understanding what these models are good (and most importantly bad) for. My observation is that this is very very valuable for the labs, and in an ideal world they should be paying you to do this, not just the tokens and "prize" for the winners.
I am also a bit frustrated seeing maths go in the direction of prompt enginnering. I am afraid of a world were a math phd student cannot go one week thinking about a problem without prompting an LLM to give him/her an invented answer. Something is lost along the way.
For me maths is not Lean, or formal systems, or an agent reasoning about formal systems to join literature from different fields. I see the value of it, but i think it will make it way more difficult for students (and profesional mathematitians) to see beyond that. And i see us heading into a reality were those who think like me will in practice remain a minority for quite a few years/decades because the low hanging fruit of LLMs will be to vast to ignore.
Agreed. I think we share the same concerns about LLM use. We just drew different conclusions.
Mathathon's goal is to reshape rather than stop LLM use. Can we set high standards for LLM use? Can we highlight the roles of a mathematician beyond proof generation? Can we redesign our incentives to promote these standards and roles?
I'd love to hear your thoughts on how to improve this event. We're very open to criticisms.
Thanks again for answering! As I told you I think the event is a good idea from your side. I would try to go beyond proof prompting and verification, and discussing the questions that you wrote above is the way to go.
I am mathematitian that is working as a software engineer. I have see first hand what these models are doing to SE. Its not that writting well thought, and compact code is not a good idea anymore, or that it doesnt beat LLM code, but that the people that see the value are a minority. If you write a piece of old school code in an LLM repo, it doesnt really make a difference, because old school code requires a team effort.
In maths the situation is not exactly the same. Probably reading a good piece of well thought math inside a book of LLM assisted proofs will stand out so much for the carefull reader that there will be no question about the value.
However, if the only way to get a position at a university is to print as much papers as possible, then who would risk printing 1 paper instead of 5 for doing old school maths?
So the solution to this is building a culture and community effort around these topics. And for that the topics need to be openly discussed.
Good luck with the organisation! And thanks again for the conversation!
1. IMO the hard part isn't prompting. It's selecting the problem and understanding
the solution. It'd be especially exciting if a participant formulates their own conjecture, proves it with AI, then generalizes it to a new theory.
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
Why not specify 2-3 problems? Won't you have people show up having already spent a bunch of time on their self chosen problem? Which sort of defeats the point of seeing what you can do in a short period of time?
Solving a problem is the easy part with AI. Selecting a problem worth solving is half of the challenge.
We allow prior work as long as it's labelled. We record chat logs, so it's easy to verify what's prior work. When we evaluate the significance of a result, we focus on the part produced at the event.
it seems like a person could cheat by finding an open problem that's AI-solvable ahead of time (by trying many problems), and then pretend to do it for the first time in the competition. You can't verify past chat logs so make sure they didn't already realize it would work.
In your final report of this experiment, can you also report the following: How many problems were NOT solved by AI in spite of attempts. What is missing in the debate of AI in math is that only the successful cases are reported. To get a balanced picture of the effectiveness of AI, what is also needed are cases where it failed.
Is this hackathon only for those with formal math backgrounds?
I've seen a few instances of AI assisted advances math and cs this year that were _not_ published by authors with formal backgrounds in those fields (or even institutional affiliation). Which makes me wonder if they would have a place at the event.
Yes. To us, solving a problem with AI doesn't matter as much as selecting the right problem and understanding the proof. A participant must have a good mathematical intuition.
Serious question: Are you sure solving abstract mathematical problems with AI is responsible use? You will likely put mathematicians out of jobs, and I doubt solving the Collatz conjecture is urgent or will save lives. It also robs a future Fields medalist of the pride of doing all by themselves.
All things AI seems to assume that more and faster is better, but there is no justification of that assumption. As a biological counterexample, a tree grown quickly will likely not be as healthy or strong as one grown slowly.
I dunno, I like math and I still think automatically solving problems is worth doing. If problem-solving is a hobby then it will still be a hobby after all this. If it's about doing things for humanity--which it often is since many thousands of people are paid to do it--then doing it more efficiently benefits humanity. If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There's just no version of this where "not solving the problems" is morally justifiable.
Now, I think it's the case that professional mathematics spends way too much money on open problems and way less than it should on pedagogy, exposition, mastery, etc. But that has always been a problem, even decades ago (I've been complaining about it my whole life). AI just finally puts pressure on the world to do something about it. I find it relieving, honestly. And I'm an AI skeptic in many other ways; it's not an AI-maximalism thing. I genuinely think the state of the field of mathematics has been something of a disaster for a long time (thanks, largely, due to the academic incentive structure which heavily favors novel results, no matter how esoteric).
> If the value on the other hand comes from people being extremely good at math, rather than from them solving novel open problems, then we should pay them to do that, which is independent of whether the open problems are solved or not. There's just no version of this where "not solving the problems" is morally justifiable.
I think the value comes from having people who have built very good intuition in a way that allows them to give explanations that make their ideas (new and old) accessible. Of course the most cutting edge math has become completely inaccessible for even many mathematicians, but at the same time, we've made massive progress in this regard. A hundred years ago college students might barely see calculus, and only serious researchers would see something like group theory. Today, many college students that aren't even math majors learn group theory and (we hope that) this gives them cognitive tools that they can apply in other situations (ability to axiomatize a concept, abstract reasoning etc).
The concern with how these tools are being used is that the current push by AI companies to solve math problems by chucking LLMs at them and producing a proof in Lean, and then using that as currency in the media to increase their stock value, undermines this process because it produces "proofs" without producing the understanding that actually allows humans to think better. All of this is then marketed as being the same as doing mathematics which it manifestly is not. If mathematicians lose the media war though, we'll have a generation of people who believe "math has been automated" and are unlikely to put in the effort to learn how to think for themselves.
To the extent that this event is intending to help the mathematical community find ways to use LLMs in pursuit of improving human understanding and intelligence, as the organizers seem to say it is, I think it's a very laudable goal. But I don't really see how this event is supposed to do that. It sounds a lot more like another fundraiser for team "isn't it cool that AI can produce useless chunks of computer code that compile to prove statements that the vast majority of the people commenting on these results don't even understand." For example, if it's really about finding ways to use LLMs to produce mathematics that improves human understanding, why is there even a requirement to solve a new problem? Why not make it explicitly about using LLMs to produce pedagogical content? Or if you really want it to be a new problem, why not add a requirement that the final product has to be accessible to a broad audience (say relying only on material in the undergraduate curriculum)?
P.S. There's another scenario where "not solving the problems" is morally justifiable: the scenario where the "solution" provides very little value (say because of what I said above - the solution just being a Lean artifact that adds very little to anyone's understanding), and the cost of solving the problem is extremely large. I know a lot of AI people are effective altruists, but before they could smell the IPO money I didn't see any of them talking about how if they had $20 million the most effective thing they could do with it is spend it in an extremely environmentally costly way in order to prove Navier-Stokes. Back before AI I seem to remember these folks talking about like... mosquito nets and malaria treatments?
If mathematicians aren't solving problems people are having (which your comment seems to imply), then putting them out of their job with AI is not a bad thing. Of course mathematicians are solving problems, just in a very different way than other professions.
Steam engines actually replaced what a horse does with a machine that could do the same job. So far LLMs aren't doing this - they're producing Lean proofs but very little to actually aid in understanding (again so far pretty much all of these proofs have been extremely difficult optimizations of known techniques). Mathematicians do solve a problem that people have: they build theories that explain the world and give us mental models to navigate questions in science, technology, etc., this just isn't a problem that LLMs solve.
So the issue isn't so much that LLMs will replace mathematicians, but that AI companies bragging constantly about how their machines "solve math" will convince people who don't understand the value of math research to no longer fund it, or students who don't yet understand why learning math is useful for developing their brains that it's a waste of time. That could put mathematicians out of a job without providing a useful replacement.
Motto: a mathematician's job isn't to solve the Hodge conjecture, it's to understand why the Hodge conjecture is or isn't true, and turn that understanding into something that makes it easier for the next person to grasp/use/enjoy.
LLMs absolutely have the potential to make this job easier, but the way in which these companies are using them right now risks being antithetical to that goal.
Don't respond to a strawman argument with another strawman. The post you are responding to ignored the reasons given in the second paragraph. They're just trying to score points by preaching to the choir, not engage with the concern.
> All things AI seems to assume that more and faster is better, but there is no justification of that assumption.
is good argument?
of course faster discovery without human in the loop is better. is that not what humans have been optimizing for the past few thousand years ? faster mobility, faster communication, faster medical recovery etc.
everything modern civilization has to offer is because of a rush to get better and faster. for example, discovering penicillin 2 years early would've saved ~15 million people more.
FWIW even Astra hasn't been able to solve the problems I care about, which are less about proving theorems and more about understanding the right way to think about already existing stories (and thus permitting extensions to new contexts). However it's been a more than a capable interlocutor to test my ideas with and see if they actually have any content. It's also great for parsing possible mistakes in long technical arguments that at least my brain isn't wired to verify completely satisfactorily. Personally, I think it's good to know what is made trivial (meaning depending only on token expenditure) vs what remains a real hard kernel.
I don't think AI will solve more math problems in a world with Mathathon than the counterfactual by EOY. People will use AI in math anyways. What matters is: can we encourage them to do so transparently and with full understanding of their results? Can we change the incentives in academia to reward problem selection and verification over proof generation?
It’s a noble goal to change the incentives, but how will you prevent the headlines from this event being “students prove Collatz conjecture with Claude” and instead be “students give great explanation of Collatz conjecture proof”?
You're right that we can't. We'll be responsible in our press releases and award prizes based on explanation, but we don't control the headlines.
However, this is already an improvement over the current state, where results are announced by headlines alone.
As I've commented elsewhere in this thread - if your goal is really to have people use LLMs to further mathematical understanding for humanity, instead of to brag about "solving" open problems, then why is there an explicit push for the Marathon to involve open problems at all? The whole thing could explicitly be about generating pedagogical content, or writing mathematical theories that simplify known results (example project idea: Kevin Buzzard wrote a lovely article on the issues that he had formalizing Grothendieck's definition of a scheme. They have since been wildly successful using AI to do formalization all the way to FLT, but no one has gone back and written a new Hartshorne that takes the insights from the formalism into account, and writes a clearer introduction to schemes that is simultaneously formally rigorous in ZFC. Using an LLM to attempt this would certainly contribute far more to human understanding of math than writing a new arxiv paper would).
I suspect you will find there is less appetite at the funding level for this kind of thing though, because what your funders really care about is generating headlines in front of their IPOs, and this kind of thing wouldn't generate the same headlines. I would be pleasantly surprised to be proved wrong of course.
EDIT: A more cynical point that I should add - I also suspect your funders would have less appetite for this kind of marathon because LLMs don't seem to be very good at this yet, which kind of points to the whole problem: so far, LLMs seem good at producing Lean proofs but not very good at the rest, but that fact is being lost in the media narrative, and "the rest" is actually the part that matters.
I get your point and agree to some extent, but you can't understand the proof without significant background in Maths so it will just allow mathematicians to solve issues faster than not have the opportunity at all.
Why do people need to understand proofs? If Amazon improves package routing with new advances in graph theory, my cat doesn't need to understand it to benefit from better shipments of cat food.
Similarly, humans don't need to be involved in scientific advances to benefit. We just need an aligned AI to take over the scientific thought for us. AI is already better than all but the top tier of humans at doing mathematics, it's writing most of the posts on the front page of this website, and it's doing the bulk of programming at many startups.
> humans don't need to be involved in scientific advances to benefit.
I agree with you on this point in isolation, but I think it's missing an enormous amount of context. Humans can absolutely benefit from science they weren't involved in and don't understand - I have no idea what a "histimine" is but I benefit from my allergy medication in the springtime.
That said, we're already living through a time where, on the whole, measures of intelligence, literacy, critical thinking, etc. are falling (at least in the US). That is a problem, which risks being exacerbated by AI, and the broader point is that we should be figuring out how to use these tools to produce knowledge that benefits humanity while also maintaining incentives for people to use their brains. Going back to my allergies: while I don't understand how my allergy meds work, my life is better, and I'm a better spouse/parent/friend/citizen etc., because I've taken the time to understand how other parts of the scientific and mathematical world that do interest me work. The current AI push to just throw out LLM-generated Lean proofs of everything under the sun to get headlines and pump up their IPO valuations (which this Marathon seems, intentionally or not, to be participating in), doesn't appear to be considering this alignment between what we get from AIs and how we can maintain our incentives to do human science. It seems more like measuring you-know-whats while risking that the message the broader public takes away is that math "has been automated" so what's the point in using your brain anymore?
Well, for one, most math proofs don't have any practical applications, so a proof that no one reads is basically a digital paperweight. You might as well suggest AI write novels for other AI to read.
The hope is that some of them end up being useful; otherwise, nobody would be funding math departments. Mathematics typically anticipates and enables new physics and chemistry.
If people are just doing math to kill time, I don't get why anyone would bother with AI. Do people really enjoy picking through a million lines of generated Lean code, if it's not for any practical use?
If you're interested in the topic enough to comment on it, you'll probably find it worthwhile reading a mathematician's perspective. Here's the prolific Terry Tao:
https://mathstodon.xyz/@tao/117219548485446992
They don't actually say anything about why anyone should fund this, though. I don't get why a society should worry about progress in mathematics if there's no practical benefit expected.
Maybe there's two kinds of math that we need? Useful math and navel gazing, and we can hand the first to the machines, and let hobbyists do the second in their free to entertain themselves?
It's hard to know what math is 'useful' a priori. That's always been the argument for supporting basic research. This is not why I am a mathematician however. I think there's intrinsic value into understanding something of depth and meaning, but the societal setup we have now that mostly agrees this is valuable is probably a very contingent phenomenon that is unlikely to last much longer.
Yes, so if there's useful math, you throw the LLM at it and use the results, no humans needed.
Humans can try to extract some ideas from the million line lean proofs, if they want to, I guess. But I can't imagine anyone really funding the human part of it.
It is the top tier of humans in these fields that are making the significant breakthroughs. The top tier of breakthroughs are not being post on here (which are nowadays usually short form articles of not incredible quality). The code at start ups is not commonly in the top tier of a breakthrough. AI can do averaged work and derivations off what has gone before which Maths works very well for as there is a clear set of rules. The same in physics if you ask AI for help adapting a simulation, yet it couldn't pluck the idea if no one has done it before.
Hey guys! I'm Brian, an organizer of the Caltech Mathathon Challenge. We are hosting the world's first math hackathon.
Our goal is to promote responsible use of AI in math. We are working with leading mathematicians to ensure that the judging process rewards human understanding. We'll open source all chat logs and cover everyone's flights.
See you at Caltech!
reply