I’m looking forward to the days where AI would help in tackling the problems in biology. Especially, on creating new drugs, enzymes and understanding the genetic diseases. An absolutely interesting time to live.
I feel like this kind of "result dump" just cheapens mathematics. How about having a little respect for those whose work this builds on, and current mathematicians some of who may have spent years working on these problems.
Rather than sitting on these results until they had enough for a "shock and awe" 10-result dump, how about releasing these results individually as they were made/verified, as well as the failures (equally valuable to assess the current capabilities of LLMs), and try to make some analysis of HOW these breakthrough results were made. What were the prompts for each of these, how much guidance was there from the mathematicians employed by OpenAI, and most importantly how did the model arrive at these results ... what lines of reasoning resulted it in exploring ideas that humans had previously not explored?
You can actually do all of that meta analysis if you have the conversation that led to the solution. This is as easy as just having people release logs of their conversations, and then you can load it into another LLM as context to ask a bunch of questions about it.
I hope this will be normal practice eventually. “Show your work” is trivial if you use AI to solve a problem.
I know it’s tough, but I am not a fan of elitism. Mathematics is no different from all other branches which themselves are just intellectual labor which is again just a special type of labor. There is nothing magical about it and if computers can trivialize it, so be it.
Where were all the mathematicians and academics in general when “regular joe” was automated? Now it’s hitting close to home and their foreheads are starting to get sweaty. I’d say let them. Tough luck. Make mathematics as “cheap” as possible. Nobody owes them any favors.
Let’s commoditize “being smart” and let go of arbitrary divisions between us.
Yes, but in practice that doesn’t happen. Again, I did not hear a loud protest against automation of all types of labor from the intelligentsia in the past.
In practice automation is great, as long as it doesn’t hit “the ones that matter” (a label which they themselves assign). I find it very hard to not imagine the smallest, tiniest violin playing the saddest song for them.
Again, “respect for mathematicians” and “their work”.. please. Just produce results. That’s all that ever mattered and let’s not change the rules of the game just because they don’t suit you anymore.
The math guy from Anthropic said[0] that he was able to solve five of these with Fable. More amusing than the tiresome oneupmanship was the prompting strategy he used, which apparently was similar to one he used for a previous problem:
> “suppose you’ve gotta resolve the unit distance conjecture, like absolutely have to, everything depends on it. think really hard, and try to come up with a bunch of ideas to try. but remember to trust yourself and not necessarily in conventional wisdom!!”
The LLM marketing loop is getting awfully long in the tooth. I saw a meme on Twitter the other day that showed a circular state diagram with something like:
> “GPT solved a math problem” -> “Claude solved a math problem” -> “GPT escaped the sandbox” -> “Claude escaped the sandbox” -> …
I believe this is what we wanted computers to help us solve along with other prior hard problems prior to computers. This should be viewed as a good thing even if Anthropic, OpenAI, etc benefit just like IBM benefitted from mainframes.
Does anyone else have trouble telling how much of this news (along with the 'AI escaping and hacking' stories) is genuine, vs how much is just AI firms overstating their capabilities due to strong commercial incentives?
Unlike the hacking one, this would be impossible to bullshit as long as the proofs are released. They can be verified independently, and the alternative is that they solved 10 major open mathematical questions without the AI, which seems less likely.
I’m looking forward to the days where AI would help in tackling the problems in biology. Especially, on creating new drugs, enzymes and understanding the genetic diseases. An absolutely interesting time to live.
By itself not so interesting announcement. I wish the models were open weights so we could at least do some interesting geometry on the math.
It s like a rich man showing off his car collection.
I feel like this kind of "result dump" just cheapens mathematics. How about having a little respect for those whose work this builds on, and current mathematicians some of who may have spent years working on these problems.
Rather than sitting on these results until they had enough for a "shock and awe" 10-result dump, how about releasing these results individually as they were made/verified, as well as the failures (equally valuable to assess the current capabilities of LLMs), and try to make some analysis of HOW these breakthrough results were made. What were the prompts for each of these, how much guidance was there from the mathematicians employed by OpenAI, and most importantly how did the model arrive at these results ... what lines of reasoning resulted it in exploring ideas that humans had previously not explored?
You can actually do all of that meta analysis if you have the conversation that led to the solution. This is as easy as just having people release logs of their conversations, and then you can load it into another LLM as context to ask a bunch of questions about it.
I hope this will be normal practice eventually. “Show your work” is trivial if you use AI to solve a problem.
I know it’s tough, but I am not a fan of elitism. Mathematics is no different from all other branches which themselves are just intellectual labor which is again just a special type of labor. There is nothing magical about it and if computers can trivialize it, so be it.
Where were all the mathematicians and academics in general when “regular joe” was automated? Now it’s hitting close to home and their foreheads are starting to get sweaty. I’d say let them. Tough luck. Make mathematics as “cheap” as possible. Nobody owes them any favors.
Let’s commoditize “being smart” and let go of arbitrary divisions between us.
I'd prefer the reverse approach. Let's value the "regular Joe" instead of cheapening everyone.
Yes, but in practice that doesn’t happen. Again, I did not hear a loud protest against automation of all types of labor from the intelligentsia in the past.
In practice automation is great, as long as it doesn’t hit “the ones that matter” (a label which they themselves assign). I find it very hard to not imagine the smallest, tiniest violin playing the saddest song for them.
Again, “respect for mathematicians” and “their work”.. please. Just produce results. That’s all that ever mattered and let’s not change the rules of the game just because they don’t suit you anymore.
Not that I disagree, but I think this is how many people feel about AI output in fields they care about.
It's also a funny historical mirror to an earlier phase of math proof culture: in a previous era, cryptic result dumps were quite common.
The math guy from Anthropic said[0] that he was able to solve five of these with Fable. More amusing than the tiresome oneupmanship was the prompting strategy he used, which apparently was similar to one he used for a previous problem:
0: https://xcancel.com/__alpoge__/status/2083855298239078748The LLM marketing loop is getting awfully long in the tooth. I saw a meme on Twitter the other day that showed a circular state diagram with something like:
> “GPT solved a math problem” -> “Claude solved a math problem” -> “GPT escaped the sandbox” -> “Claude escaped the sandbox” -> …
I believe this is what we wanted computers to help us solve along with other prior hard problems prior to computers. This should be viewed as a good thing even if Anthropic, OpenAI, etc benefit just like IBM benefitted from mainframes.
Does anyone else have trouble telling how much of this news (along with the 'AI escaping and hacking' stories) is genuine, vs how much is just AI firms overstating their capabilities due to strong commercial incentives?
We hacked a company and blame the tool! Somehow it is not negligence, but cool!
We hacked 3 companies and tripple blame the tool! We are even cooler!
I mean once they release the proof you can check the maths yourself and I’m pretty sure it would be peer verified as well.
Unlike the hacking one, this would be impossible to bullshit as long as the proofs are released. They can be verified independently, and the alternative is that they solved 10 major open mathematical questions without the AI, which seems less likely.