Hacker Newsnew | past | comments | ask | show | jobs | submit | sigmar's commentslogin

>or there is a more deliberative approach to assigning credit than who was "first" to solve some problem

it feels like an unintended consequence of the millennium prize is that people view the [last contributor to the solution] as the only one to make progress on the problem. I've never viewed Poincaré as solved by one person and the objective of the prize was to encourage more people to make attempts and contribute towards progress.

this issue is independent, but in these circumstances perhaps interweaved, with the 'ai is taking over math' concerns


Were the agents ever tasked with algorithm improvements? Post just says he didn't find any ("report essentially no algorithm advancements"). These LLMs are useful for optimization tasks where they can attempt a change and then measure performance boosts, so just wondering.

>We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations

Do you think it is possible that better math will lead to better physics models?


It might but the math results from GenAI so far have been limited to finding counterexamples to known conjectures, not building new mathematics.

> results from GenAI so far have been limited to finding counterexamples

Not all.

Ehrhart’s volume conjecture

Quantum parallel repetition for general two-player quantum games

Erdős Problem #183 on multicolor Ramsey numbers

Erdős–Sárközy Problem #12(i)/(ii)

Erdős Problem #125

Log-concavity of codimension-3, type-2 pure O-sequences

Optimal O(1/t) last-iterate convergence for Anchored Gradient Descent-Ascent


True, these results are from last month. My info was a little out of date.

Prior updated :)


Yes, definitely! There’s a long history of this and I think there’s tons of opportunities for more. Both for improving the exactness/physical fidelity of models and for developing new approximate theories and simulation methods

>The speed with which we were able to produce this proof demonstrates that it is now possible to formalize large swaths of mathematics, which may both catch errors in the common body of mathematical proofs and reduce the burden of refereeing new work.

^ this section should have been in the first few paragraphs imho. Explaining why this is relevant shouldn't be so far down.


Forgive the authors of the article for assuming readers would complete it.

For any body of text (or in general, any exposition of any kind), the responsibility to explain the value of the article is very much in the author's side.

Explaining the value of what you are showing should always go towards the start. Else, why would anyone bother with the rest?


Feel very grateful I was never taught this... Would have missed out on quite a lot of good bodies of text in my life I think! Pushing through any initial friction or ignorance I might have as a reader, having the patience and charity to bear with an author until you get it, was instead what I was always taught.

Giving such a blanket "responsibility" to the author at all is just such a bummer! I say let them do whatever they want, there is always more than one way to express oneself. Someone who was never taught to write a clear thesis in the first paragraph for whatever reason doesn't inherently have less to say.


Unfortunately I don't see this particular view paying off in the age of AI, as many prove they have nothing at all to say but say it anyways. Which isn't to say people shouldn't write if they enjoy writing, but I for one will stay a discerning reader.

> Feel very grateful I was never taught this...

Never heard of Abstract section? First semester on a college or last year on high school.


Isn’t deciding what responsibilities there are in any text very much in the author’s side and not yours?

Haven’t you dramatically overstated your case? Many expositions do not contain an explanation of their value at all. Works of fiction are a good example, and there are many many others. Often it’s the responsibility of the readers & reviewers to decide on questions like value.


I think this is a Fermat joke :-)

Buzzard is writing for his blog audience - mostly mathematicians and not the casual visiting HN user.

Eh? The quote is from Anthropic, not Buzzard.

Have you written text for humans? You’re lucky if you can get people to read more than the title. You’re very lucky if they read past the first paragraph.

you would never assume this if you've spoken to any human being, ever

> reduce the burden of refereeing new work.

As a professional mathematician, I rarely need to worry about the correctness of a paper. The main difficulty of writing a review is instead understanding what the results of the paper mean in its context, how the results are presented, etc.


Isn't it the cost we care about, rather than the speed? All we know know is that a frontier AI lab was able to do it in 11 days, we have no idea how much compute they threw at it.

They said 6 billion tokens, which isn't as much as I thought it might be.

Am I doing my napkin math correct? The post says it's using a model comparable to Fable 5.1, which is $50 per million output tokens. So this is ~$300K? Surely an over-estimate due to caching.

Surely input tokens are also involved, and not necessarily only for the initial prompt if there are feedback loops or agent interactions.

Nah they should have released it in a 14-part tweet instead.

https://en.wikipedia.org/wiki/Attempted_assassination_of_Don...

what evidence is there that this registered republican was "far left"?


Well, there is apparently no evidence that he was "far right" either.

    So far, investigators haven't found any evidence on social media or other
    writings by Crooks that might help identify his motive for the attempted
    assassination, law enforcement officials say... And a review of public
    records suggests he may have had divergent political leanings, with Crooks
    registering to vote as a Republican but making a small donation to a
    Democratic-leaning group.


>We are currently working closely with our inference partners and open-source maintainers to align the technical details and ensure the model can be reliably deployed across the ecosystem. The full model weights will be released by July 27, 2026. Further details regarding the architecture, training, and evaluation will be released with the Kimi K3 technical report.

(translated by chrome)

11 days is a long time. It does not take that long to implement inference at providers. In my opinion, seems like they're being pre-emptively cautious about government intervention/review


Actually it does for a massive model, serving it correctly is not easy.

I believe Kimi also does some sort of Q&A and eval for day 0 partners, since early on a long of inference providers just weren’t running their models properly.


Eh, Minimax M2.7 also took a similar amount of time (actually longer) between availability and weights release.


Most of what an LLM does "could have" been done by a human if you throw enough human hours at it. But the reality in this circumstance is that a new tool helped find this leak. Saying this could have happened in a "non LLM world" is analogous to "someone else could have discovered special relativity, let's not mention Einstein"


This not only could have happened pre-llm, it did: https://krebsonsecurity.com/2022/02/report-missouri-governor...


My point is about the emphasis of Codex in the title. That emphasis makes more sense when Codex is credited with finding something that would have been difficult or impractical to discover without substantial human effort.


>Evaluators validate each tag individually — for example, protein, preparation, or health, individually rather than judging the item as a whole.

Am I reading this right that the jury is multiple LLMs each iterating through each tag and voting on each? Why wouldn't you tune one LLM to be really competent at a single tag? Like a single "spicy evaluator LLM" or "protein evaluator LLM"?


I am guessing there are so many possible tags that this wouldn't be pragmatic


>I think it's reasonable for people to say, hey, if you're going to trash the reputation of Zig (in a pretend-objective way)

what specifically is this referring to? Not aware of any comments from Anthropic on this topic.


To me, this addendum makes it worse. Making small edits to a post like this makes it seem like you're doubling down on the original resentful points, especially with all the new justifications like "a trillion dollar company fired the first shot." Should have just deleted the blog post in its entirety imho.

>outstanding relationships with essentially everyone who openly uses Zig and talks about it publicly

Essentially everyone? more blogposts incoming?


>>outstanding relationships with essentially everyone who openly uses Zig and talks about it publicly

> Essentially everyone? more blogposts incoming?

Andrew is explicitly saying the opposite, but you're entitled to your supposition that that is bluster, I suppose.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: