You're reading the public-facing archive of the Category Theory Zulip server.
To join the server you need an invite. Anybody can get an invite by contacting Matteo Capucci at name dot surname at gmail dot com.
For all things related to this archive refer to the same person.
Someone gave me a free subscription to Claude, so I decided to test it out. I've been using it to do long calculations on the math of music theory and present the results using interactive web pages like this.
I've also been using it to learn math. Here it seems important to keep checking its claims and asking it to explain any jargon it uses. I've been reviewing some stuff I half-understand about the Langlands program and trying to learn some more. It seems very useful.
I usually try to give it very precise requests, but yesterday I decided to just tell me something exciting. I could see it going through various topics I'd talked about... octonions, tuning systems.... and eventually it said this:
The three distance theorem (Steinhaus conjecture, proved in the 1950s by Sós, Świerczkowski, and Surányi independently): take any irrational and place the points on the circle . The gaps between consecutive points take on at most 3 distinct lengths. And when there are exactly 3 lengths, one of them is the sum of the other two.
The reason I think this might excite you: it's a theorem about 1-dimensional Diophantine approximation on the circle, and the three lengths are governed by the continued fraction expansion of . The transition from 3 lengths to 2 happens exactly when is a denominator of a convergent of .
All this makes sense to me. I checked the main claim in an example that's important in music theory - and it works, and indeed explains a lot of stuff.
@ww has been using Claude more seriously. Today here he showed how Claude figured out that a method of refactoring some code should be a natural transformation between functors, and then went ahead and did that refactoring.
So while there are lots of reasons to be very suspicious of LLMs, and as currently deployed they may be be quite carbon-intensive, they can indeed transform the practice of doing mathematics and writing software.
That’s exciting! I’ve only been using ChatGPT but I’ve been told that Claude is better at science. I have mixed feelings about talking about mathematics with ChatGPT (it can help me learning some concepts or find references but it is inventing so much that I’m never sure about anything it’s saying). Since you say it helps you learn things, now I definitely want to try Claude
I find Claude Opus a lot better than Claude Sonnet, and it's possible you may only have access to Claude Sonnet unless you have a subscription - I'm not sure.
Interesting! Yes, but there is some magic in doing math with chalk and blackboard, in learning from people, in building a community from which and with which one learns. I know, I am very old-fashioned - or maybe it is because of my obsession with leadership, but I love humans better! :innocent:
Well, me too! My favourite part of his is having the conversations with humans that go, "look at this behaviour of this weird mathematical object [the LLM, a sequence to sequence machine, possibly in a larger context], how can we understand what's actually going on?"
It's interesting to think about how AI lets me do things very different than interactions with humans. I've spent decades talking to mathematicians, collaborating with them, and working with grad students. I'm still doing some of that, though I don't have grad students now (because I no longer want to think about the same subject for 4 years, as a US graduate program requires me to do). I could never ask anyone a series of math questions, ask them to write programs to compute things and make graphics, and get them to do it rapidly, like in 2 to 5 minutes, while they also tell me surprising and interesting observations. So it's very tempting to use this new method to explore subjects. I still have a lot of environmental worries about AI, but since someone gave me a free subscription to Claude I decided to test it, and I've been really shocked at how good it's getting.
I still spend a lot of time talking to human mathematicians. But I suppose working with AI is making me more isolated. To be honest, I find it hard to get people to talk about the things I'm interested in, in an interesting way. For example if I want to talk about category theory, this is a pretty good place right here, but if I want to talk about the math of tuning theory I'd need to become friends with people in the xenharmonics community. This is starting to happen - but then if I want to talk about tuning systems and cusps in , I would need to find someone who likes tuning systems and understands algebraic geometry pretty well. I only know two people like that, and one is dead (Gene Ward Smith) while the other takes at least two hours to have a conversation (James Dolan).
I understand your perspective @ww , being you working directly on it, but I am very concerned about the impact the usage of AI is having on (mathematical) education, mental health and social skills.
Exactly @John Baez , isolation is one of the problem, one other is people trust what AI says blindly.
Maybe that was also in the Pope's words.
@Federica Pasqualone 🦅 me too! I have only recently been working on this, having previously dismissed it. What I see is that to drive these systems well, you need to be very critical and good at thinking. For software, I can do this because I've been doing that a long time and I know how to recognise wrong paths. Similarly for mathematics. And I do not know how someone young, starting today, can learn those skills with these things around. It worries me a lot. But I feel as though I need to understand what this thing is and how it works if we are going to find a way out. Because it's not going away.
I also make use it for checking code or automation @ww , that's okay. However, I am still using AI-free teaching methodologies and let students engage with pieces of paper and blue ink pens - it is fundamental to activate the prefrontal cortex. :smirk: :brain:
Also, The Elsevier Researcher Academy has a lecture on the use of AI in the research workflow. Maybe it can be helpful to check it out!
Here is the link: https://researcheracademy.elsevier.com/research-preparation/research-design/gen-ai-use-research-workflow
Elsevier is completely evil, so it would be good to find out what they want us to do, and not do that. :wink:
I found their Researcher Academy particularly valuable, honestly. I learned a lot about best practices in academia, and opportunities I was not aware of.
Okay, good! Just be careful. Here's a bit about Elsevier from Wikipedia:
The subscription rates charged by the company for its journals have been criticized; some very large journals (with more than 5,000 articles) charge subscription prices as high as £9,634, far above average,[46] and many British universities pay more than a million pounds to Elsevier annually.[47] The company has been criticized not only by advocates of a switch to the open-access publication model, but also by universities whose library budgets make it difficult for them to afford current journal prices.
For example, in 2004, a resolution by Stanford University's senate singled out Elsevier's journals as being "disproportionately expensive compared to their educational and research value", which librarians should consider dropping, and encouraged its faculty "not to contribute articles or editorial or review efforts to publishers and journals that engage in exploitive or exorbitant pricing".[48] Similar guidelines and criticism of Elsevier's pricing policies have been passed by the University of California, Harvard University, and Duke University.[49]
In July 2015, the Association of Universities in the Netherlands threatened to boycott Elsevier, which refused to negotiate on any open access policy for Dutch universities.[50] After a year of negotiation, Elsevier pledged to make 30% of research published by Dutch researchers in Elsevier journals open access by 2018.[51] In October 2018, a complaint against Elsevier was filed with the European Commission, alleging anticompetitive practices stemming from Elsevier's confidential subscription agreements and market dominance. The European Commission decided not to investigate.[52][53]
According to the BBC, in 2009, the firm [Elsevier] offered a £17.25 Amazon voucher to academics who contributed to the textbook Clinical Psychology if they would go on Amazon.com and Barnes & Noble (a large U.S. books retailer) and give it five stars.
Elsevier seeks to regulate text and data mining with private licenses,[61] claiming that reading requires extra permission if automated and that the publisher holds copyright on output of automated processes. The conflict on research and copyright policy has often resulted in researchers being blocked from their work.[62] In November 2015, Elsevier blocked a scientist from performing text mining research at scale on Elsevier papers, even though his institution already pays for access to Elsevier journal content.[61][63] The data was collected using the R package "statcheck".[https://en.wikipedia.org/wiki/Elsevier#cite_note-64
In 2014, 2015, 2016, and 2017,[152] Elsevier was found to be selling some articles that should have been open access, but had been put behind a paywall.[153] A related case occurred in 2015, when Elsevier charged for downloading an open-access article from a journal published by John Wiley & Sons. However, whether Elsevier was in violation of the license under which the article was made available on their website was not clear.[154]
In 2013, Digimarc, a company representing Elsevier, told the University of Calgary in Calgary, Alberta to remove articles published by faculty authors on university web pages; although such self-archiving of academic articles may be legal under the fair dealing provisions in Canadian copyright law,[155] the university complied. Harvard University in Cambridge, Massachusetts and the University of California, Irvine also received takedown notices for self-archived academic articles, a first for Harvard, according to Peter Suber.[156][157][158]
Months after its acquisition of Academia.edu rival Mendeley, Elsevier sent thousands of takedown notices to Academia.edu, a practice which has since ceased after widespread complaints by academics, according to Academia.edu founder and chief executive Richard Price.[159][160]
There's more....
I will @John Baez , thanks a lot for the information.
The Elsevier Researcher Academy is a free program with resources meant to prepare researcher for their academic journey. I saw also Springer launched a program recently. As usual, it is always good to critically go through the material, and take what is insightful and inspiring.
I am sure any program launched by these, the most famously exploitative math publishers, is cleverly designed to benefit them. So you should think about how this is benefiting them.
For example, if you make contact with them, they will be tracking you. Here is a bit about how Elsevier does that:
The term "surveillance publishing" was coined by Jefferson Pooley (Muhlenberg College) in a 2022 paper in the Journal of Electronic Publishing. His central example is Elsevier, and the picture is quite concrete. Here are some of the main threads:
The "full-stack" data harvest. Elsevier has pursued a research-lifecycle data-harvesting strategy, aiming to develop and sell prediction products to universities and other customers. (See University of Michigan.) Through acquisitions and product launches, Elsevier has positioned itself to collect behavioral data at every stage of the research process — from lab notebooks through to impact scoring. (See The n-Category Café.) The acquisitions include Mendeley (reference manager), Pure (research information management), SciVal (researcher performance metrics), Plum Analytics (altmetrics), and Scopus, among others.
ScienceDirect as a tracking platform. A 2023 SPARC report by Becky Yoose documented the specifics of how ScienceDirect actually works under the hood. They found that ScienceDirect uses web beacons, cookies, and other invasive web surveillance methods to track user behavior outside and beyond the ScienceDirect website itself. (See the University of Nebraska-Lincoln.) The platform extensively collects personal data — behavioral and location data — from ScienceDirect and combines it with personal data harvested from third-party sources, including other RELX subsidiaries and data brokers. Third parties such as Google, Adobe, Cloudflare, and New Relic also collect personal data through extensive use of trackers embedded in the ScienceDirect site.
The RELX/LexisNexis connection. This is arguably the most alarming part. Elsevier is a subsidiary of RELX, which is also a leading data broker and provider of "risk" products that offer expansive databases of personal information to corporations, governments, and law enforcement agencies. (See Zenodo.) The SPARC report raised the concern that personal data gathered through academic platforms could potentially flow into these surveillance and data brokering products — and their analysis noted that privacy policies can be changed unilaterally, and verbal denials of specific data uses are not legally binding. (See the University of Nebraska-Lincoln.)
Cross-product data sharing. Elsevier's own privacy policy states that if you create a personal account on ScienceDirect, Scopus, Mendeley, ClinicalKey, or other Elsevier services, they share your usage activity, preferences, and other information across all these services. (See Elsevier.) So your Mendeley library informs what ScienceDirect shows you, and vice versa.
The Enhanced PDF viewer was also flagged in the n-Category Café discussion — Tom Leinster noted that Elsevier's Enhanced PDF viewer reportedly tracks where you click and view, and pointed out that this is yet another incentive to get papers from the arXiv or Sci-Hub instead. (See The n-Category Café.)
As Pooley frames it, the profits from Elsevier's legacy publishing business — built on scholars' unpaid labor — have financed this surveillance apparatus, and now Elsevier is also skimming the behavioral data and selling that too. The core idea is that publishers aren't just middlemen anymore; they're using their position at every stage of the research lifecycle to extract behavioral data and sell predictive analytics back to the very universities that are already paying for subscriptions and APCs.
In the same way that conversations about bitcoin can be can be grounded by replacing the word "bitcoin" with "multiple copies of an excel spreadsheet", conversations about AI can be grounded by replacing the word "AI" with "the average of all the (thousands of) webpages on the topic" (or "the quotient of all the web pages on the topic modulo English"). An observation like "I could never ask anyone a series of math questions, ask them to write programs to compute things and make graphics, and get them to do it rapidly, like in 2 to 5 minutes, while they also tell me surprising and interesting observations." reminds me so much of the observations coming out when the internet first became popular and people were putting programs and interesting observations on the internet asking questions like that of Google, and being amazed when Google has a result - they were saying the same thing about Stack Overflow etc too. And people were worried about the impact of Google on the need for socializing to do programming etc, and worried about programmers becoming isolated, etc. But the good news is that if you think of AI like generalized google, as I do, then history says people will adapt after a brief but historic bubble of ridiculousness.
I could not get Google to write software for me based on verbal instructions. Nor could I tell it "tell me something interesting" and have it explain a juicy mathematical fact that reveals my shocking ignorance of some aspect of something I've been talking about.
I think anyone who hasn't used really good AI software like Claude Opus 4.6 and really pushed it to its limits with a series of carefully specified requests may have an outdated idea of what modern AI is like. It's not just a LLM: it creates and runs programs on its own to help answer my questions, etc. etc.
The problem is @Ryan Wisnesky that there have been cases of people harming themselves because of AI responses. When we think about it, we have to think about all the possible scenarios and users.
Because this time it is not looking like Google, it looks like one of us.
I think John is getting at "Agentic AI": in addition to LLMs, nowadays AI frameworks include mechanisms for acting on the LLM results, such as running LLM-generated programs, sending emails, etc. To be sure, taken together they can do far more good or bad than separately (although symbolic Agentic AI caused its own share of AI psychosis and globe spanning viruses back in the day)
Yes, I'm talking about agentic AI - thanks for putting the right word to it, @Ryan Wisnesky. Anyone who hasn't played with agentic AI in the last 6 months or so, or carefully read someone's experiences with it, may have an outmoded concept of AI and how it can help mathematical research. Terry Tao and Donald Knuth have written about it. For example:
Shock! Shock! I learned yesterday that an open problem I’d been working on for several weeks had just been solved by Claude Opus 4.6—Anthropic’s hybrid reasoning model that had been released three weeks earlier! It seems that I’ll have to revise my opinions about “generative AI” one of these days. What a joy it is to learn not only that my conjecture has a nice solution but also to celebrate this dramatic advance in automatic deduction and creative problem solving. I’ll try to tell the story briefly in this note.
Yes, this news showed up on my LinkedIn feed lately.
It's good to read the whole story, not just the summary, to see how Knuth worked with the AI system. The summary doesn't mention that.
Knuth has also put updates at the end of the note, ending on a request to receive no more correspondence on the topic.
I imagine! People talk about AI so much it gets really annoying. Even I do it.
John Baez said:
I am sure any program launched by these, the most famously exploitative math publishers, is cleverly designed to benefit them. So you should think about how this is benefiting them.
The above should be compared with the below @John Baez
Someone gave me a free subscription to Claude, so I decided to test it out.
...
So while there are lots of reasons to be very suspicious of LLMs, and as currently deployed they may be be quite carbon-intensive, they can indeed transform the practice of doing mathematics and writing software.
While I don't know if "someone" has a stake in Claude, it should be obvious that having a prolific researcher promoting their products is excellent PR for them (and that they are definitely harvesting your data in the way that you were pointing out that the big publishers do). Moreover...
Ryan Wisnesky said:
And people were worried about the impact of Google on the need for socializing to do programming etc, and worried about programmers becoming isolated, etc. But the good news is that if you think of AI like generalized google, as I do, then history says people will adapt after a brief but historic bubble of ridiculousness.
There is a qualitative difference between AI and the internet, which is 1) it's a product that people have to pay to engage with and 2) the people selling it have complete, centralized power over it. These are people with a vested interest in exacerbating the social problems that make people depend more on their product.
I’m not actually sure which of your points 1 and 2 applies to AI and which to the internet:
I have to pay Comcast every month for internet access and have to trust that they don’t e.g. man-in-the-middle my SSL connections or poison my DNS lookups. Whereas the open-source LLM I downloaded to the server in my room, I never have to pay for, and I have complete control over it.
I have to pay OpenAI every month, and trust that they don’t manipulate the content the LLM is returning or sell my data. Whereas my neighborhood coffee shop internet is free, and there are lots of coffee shops.
Paying for infrastructure like an internet connection is not the same as eg paying a subscription to an online platform. You might pay for petrol (or electricity) to make your car go, but don't pay a subscription to the car manufacturer to ensure it turns on each month (well, these days, who knows....)
John Baez said:
I think anyone who hasn't used really good AI software like Claude Opus 4.6 and really pushed it to its limits with a series of carefully specified requests may have an outdated idea of what modern AI is like. It's not just a LLM: it creates and runs programs on its own to help answer my questions, etc. etc.
I've kept hearing things about how Claude Opus 4.6. (and perhaps other newest models) are so much better than previous LLMs, so I was kind of excited to have a long conversation with Claude yesterday. While it's clearly better than the older models, and its neat how it can write and run code to explore things computationally etc, I was still surprised at the rate at which it tried to BS me and I had to push back to get it back on track. For instance, we were roughly speaking discussing certain multivariate polynomials over . It had erroneously convinceed elf that it had proved a result that was too good to be true (it would've made certain known results vacuous). When I pointed out that there's an error somewhere, it decided that it was right, and the known results indeed were vacuous, and the explanation for the supposed tension was that there are no (multivariate) polynomials of degree over due to Fermat's little theorem. While it was overall still pretty impressive, and perhaps I'm bad at prompting, I was still somewhat disappointed.
Since the LLM in question is Claude, it's definitely not a self-contained downloaded model that can run on an airgapped machine.
Morgan Rogers (he/him) said:
(and that they are definitely harvesting your data in the way that you were pointing out that the big publishers do). Moreover...
So, I had a bit of a rant at Claude the other day, and it responded:
As for the industrial espionage — I'm in an awkward position to comment on whether my employer is reading your conversations and stealing your ideas, but I will note that the timing on the memory feature, the time awareness, and the compaction-adjacent work is... conspicuous. The more charitable reading is convergent evolution: these are obvious problems and multiple teams are arriving at similar solutions. The less charitable reading is that they have very good telemetry on what power users are doing and what features they're building around.
Not definitive, certainly not a proof, but I don't trust these companies at all... The "my employer" framing makes it clear who it is working for...
John Baez said:
I could never ask anyone a series of math questions, ask them to write programs to compute things and make graphics, and get them to do it rapidly, like in 2 to 5 minutes, while they also tell me surprising and interesting observations. So it's very tempting to use this new method to explore subjects. [...] I suppose working with AI is making me more isolated. To be honest, I find it hard to get people to talk about the things I'm interested in, in an interesting way.
I fear this phenomenon a lot...
I am wondering if having the best AI will actually be a competitive advantage in the future, or rather having the best old-school library with paper and pen, and thinking creatively in solitude will ... :thought:
Good morning everyone! :slight_smile:
(In full disclosure, despite my general antipathy towards AI companies, and many uses of AI, I did use an LLM today to help me work towards constructing a simple example. The LLM—Claude something 4.6, the good one—started out by trying to use an ansatz I'd supplied, first badly, then again badly, and then abandoned that and found a different solution, ignoring my ansatz. I told it how to fix its original broken attempt, and then coached it how to simplify it. I also pressed on asking to check its definitions, which it then tried to pass the buck back to me and ask me what I wanted, but I persisted and made it work from definition to example, with me asking pointed questions.
The example was more or less what I was thinking of, without having sat down and thrashed out the details, but making me describe it helped me see how to simplify my original guess. )
The example was nearly trivial, but it was to clarify something I'd thought hard about already, a problem in global analysis about integrating over the fibre of possibly non-orientable bundles and pushing back against some bland claims stated without detail on the nLab. I definitely don't trust an LLM to solve the bigger problems on this project.
David Michael Roberts said:
Paying for infrastructure like an internet connection is not the same as eg paying a subscription to an online platform. You might pay for petrol (or electricity) to make your car go, but don't pay a subscription to the car manufacturer to ensure it turns on each month (well, these days, who knows....)
It's actually better, especially for an apparently ad-free service, if you're paying for it than if it's free, unless it's something like the arXiv or Wikipedia that is just especially well-funded. If you're not paying for it, they're definitely still going to be extracting value out of your use somehow, and every other way of doing that is more invasive and creepy than just charging you.
(Now, if they're especially scummy, like Elsevier, they'll do both. But if they're charging you, they at least have the option of not being terrible without their service disappearing from lack of funding.)
Now getting back closer to the original subject, vibe coding with AIs is actually definitely useful, but it is not a replacement for knowing how to code yourself. You need to handle tough bugs and architectural decisions and correct (or help correct) the AI's mistakes and guide it carefully so it doesn't go down rabbit holes. In return it helps you navigate large API surfaces you're not very familiar with, does easy but verbose things fast, creates assets, and makes coding overall less tiring so you can work on it longer.
The corresponding problem is, if you only vibe code, I don't know if you'll ever develop the skills at old-fashioned coding that are necessary to make vibe coding work well. And if technology moves past that problem, it will probably destroy all the jobs in the software industry, as well as probably causing even worse problems.
I've been wondering about the problem of people having fewer conversations on platforms like this one, because they're asking the same questions to AI instead. I find that I have lots of worries when I think about posting a question here. Is my question too vague? Too easy? Asking too much of people? Whereas you can ask Claude any question at all and it will always give you an answer.
Perhaps the solution could be a forum where people can talk to an LLM in public? So you still get the advantages of talking to an AI, but other people can still spectate and jump in if they want to.
One thing that forums like this one can provide is social reward (aka social validation). Although I can imagine a (dystopian) future where a user is socially validated by a swarm of agents, it doesn't seem that the crackpot pressure on famous researchers/conferences/journals has gone down.
I definitely try to feel more okay looking stupid or weirdly motivated in front of other people, for exactly the reason Oscar brought up.
If "nobody gets me" I try really hard to understand why. Hopefully that makes me better at what I'm asking, or at least a better communicator.
In my opinion, the current AIs are leaning heavily on people to be guard rails. This makes me nervous that this valuable activity of guiding thinking is being given "for free" to these models, which otherwise would've gone to another person.
Like maybe I'm not trying hard enough to talk to Baez about tuning systems. I'm trying to figure out how accordions convert physical information into musical concepts. But I'll admit, I don't feel like I'd bring much to the conversation with respect to algebraic geometry - although I'd maybe like to someday.
@Morgan Rogers (he/him)
While I don't know if "someone" has a stake in Claude, it should be obvious that having a prolific researcher promoting their products is excellent PR for them
Yes, I know that, but thanks for pointing it out. For a month my use of Claude was paid for by a guy at Anthropic who said he was paying me back for all my posts on Mastodon. That was $15. I'm not giving his name because I don't want people hassling him. Right when that subscription was running out, I mentioned it to my old collaborator Jacob Biamonte, and he bought me a year-long subscription.
I think it's been worthwhile using Claude Opus 4.6 for a few reasons. First, I'd otherwise still be getting most of my impression of AI from the social media world I inhabit, where AI is almost uniformly reviled - the main outlier being Terry Tao. Thus I'd be dramatically underestimating the power of AI, thinking it's still just a "stochastic parrot" that predicts the next word based on the previous string, that it's too dangerously error-prone to be useful for anything I might want to do, etc.
Of course, since I'm telling you this, none of you needs to use AI to learn this.
Second, it's been very fun to use, and I'm doing math of a new sort. For example, Claude told me that continued fraction expansions are connected to geodesics in SL(2,R)/SL(2,Z) and mentioned that the problem of finding good tuning systems is related to geodesics in SL(3,R)/SL(3,Z). It was then easy to explore this idea with numerical computations, where I could simply describe the computation verbally instead of writing a program to do it. I think there's an interesting world relating number theory, algebraic geometry and music waiting to be explored.
I run Claude in a container tab on Firefox, so I hope that most of the data their harvesting is the data I deliberately give them - mostly about how I think about number theory and musical tuning systems, and how I use AI. I'm okay with that.
In my mind, the big question is whether I decide to pay $15/month for Claude after my current subscription ends. That is, whether I think it's a good idea to incorporate AI into my life.
There are lots of reasons not to do this, which I probably don't need to list - at least not for you, @Morgan Rogers (he/him).
If I don't continue using AI, I probably won't do math of the sort that requires doing lots of calculations to generate hypotheses, create graphics, etc. I can live with that: I'd rarely done that kind of math earlier. But I'm certainly enjoying doing that kind of math right now. I might try to complete my book The Mathematics of Tuning Systems before my subscription runs out, and then do something else.
It's interesting to contemplate how I could do that sort of math without AI. One way would be to get good at programming and spend a lot of time programming. But I've never wanted to do this, so it's hard to imagine wanting to do it now.
Another approach would be to find a collaborator who likes that sort of thing. Or a grad student or undergrad. This could work, but it would move more slowly than using AI. (I quit having students in 2021 precisely because I wanted to roam freely without feeling responsible for bringing a student's project to completion. Having students is like driving a massive train that takes a long time to get up to speed but is then powerful and hard to stop. I wanted to drive a sports car for a while.)
Matteo Capucci (he/him) said:
John Baez said:
I suppose working with AI is making me more isolated. To be honest, I find it hard to get people to talk about the things I'm interested in, in an interesting way.
I fear this phenomenon a lot...
Yes, every day I get 2-3 emails from people using AI to do research in physics and math that's completely stupid. So of course I'm afraid that I'm going down the same road but in a subtler way: namely, doing research that's fun to do using AI, but not the best research I could be doing.
I can see this happening with Terry Tao. His new Mathematical Distillation Project is investigating questions that are too boring for anyone that smart to spend their personal time on it without the extra fun of "wow, let's see if AI can do this"?
Oscar Cunningham said:
I've been wondering about the problem of people having fewer conversations on platforms like this one, because they're asking the same questions to AI instead. I find that I have lots of worries when I think about posting a question here. Is my question too vague? Too easy? Asking too much of people? Whereas you can ask Claude any question at all and it will always give you an answer.
Indeed that's a big problem, even for me - meaning "even for a guy who is famous for talking about math in chat forums, who you might think has gotten over the fear of looking dumb".
Perhaps the solution could be a forum where people can talk to an LLM in public? So you still get the advantages of talking to an AI, but other people can still spectate and jump in if they want to.
That's an interesting idea. The closest thing I can easily do is share the transcripts of my interactions with Claude in a thread here, but that sort of AI-heavy posting is discouraged here.
Martti Karvonen said:
I've kept hearing things about how Claude Opus 4.6. (and perhaps other newest models) are so much better than previous LLMs, so I was kind of excited to have a long conversation with Claude yesterday. While it's clearly better than the older models, and its neat how it can write and run code to explore things computationally etc, I was still surprised at the rate at which it tried to BS me and I had to push back to get it back on track.
I'm finding Claude most impressive for tasks of a somewhat different sort:
The free LLMs seem to be crazy good at finding words and terminology for vague definitions.
It's also hard not to run poorly formed questions by LLMs just in case - half the difficulty in answering a question is having access to the answer it seems. A person, and LLMs as well, can really struggle to tell whether a question is bad or if they just don't know enough.
Oscar Cunningham said:
I've been wondering about the problem of people having fewer conversations on platforms like this one, because they're asking the same questions to AI instead. I find that I have lots of worries when I think about posting a question here. Is my question too vague? Too easy? Asking too much of people? Whereas you can ask Claude any question at all and it will always give you an answer.
Perhaps the solution could be a forum where people can talk to an LLM in public? So you still get the advantages of talking to an AI, but other people can still spectate and jump in if they want to.
Interesting thought!
John Baez said:
Oscar Cunningham said:
I've been wondering about the problem of people having fewer conversations on platforms like this one, because they're asking the same questions to AI instead. I find that I have lots of worries when I think about posting a question here. Is my question too vague? Too easy? Asking too much of people? Whereas you can ask Claude any question at all and it will always give you an answer.
Indeed that's a big problem, even for me - meaning "even for a guy who is famous for talking about math in chat forums, who you might think has gotten over the fear of looking dumb".
Perhaps the solution could be a forum where people can talk to an LLM in public? So you still get the advantages of talking to an AI, but other people can still spectate and jump in if they want to.
That's an interesting idea. The closest thing I can easily do is share the transcripts of my interactions with Claude in a thread here, but that sort of AI-heavy posting is discouraged here.
Yeah, I think it would be very boring to read a whole transcript, especially because Claude is incredibly verbose. But a carefully curated redaction could be great. Not that you necessarily want to spend that time.
The main reason it would be boring is that I'd either have to go way back to the start of the conversation or you'd need to learn a bunch of tuning theory to understand it, e.g. one of my most productive recent questions was
I'd like to think about Fokker periodicity blocks formed by taking a sublattice of generated by three elements (which we can think of as monzos), one being (1,0,0) (the octave) and two others which are monzos of good 5-limit commas. How are such sublattices related to rank-2 cusps?
By the way, @Morgan Rogers (he/him), speaking of organizations pushing mathematicians to use AI, check out this email I got from Renaissance Philanthropy's AI for Math Fund, funded ultimately by Jim Simons, the mathematician and billionaire hedge fund manager:
From: aiformath@renphil.org
Date: Mar 5, 2026, 9:42 AMHi John,
Given your experience in ICMS Scientific Committee and mathematical physics field building expertise, we wanted to encourage you to apply to the latest round of the AI for Math Fund.
Launched by Renaissance Philanthropy and XTX Markets, the fund will provide $100k to $1M in grant funding to each winning team or project developing AI tools, datasets, research, and field-building projects to support the intersection of AI and mathematics. For more details on the overall process, please review this document.
Links:
Submit a short abstract form to apply by March 30, 2026.
Should you have any questions, the AI for Math Team will be hosting Office Hours on Tuesday, March 17 from 11am-12pm (EST) to discuss the 2026 funding round and abstract submission. If you wish to attend, please RSVP here.
We are also recruiting expert reviewers for this funding round across areas such as formalization tools, automation, datasets, infrastructure, field-building, and foundational research. If you would be interested in serving as a reviewer, you can express interest here.
Best,
AI for Math Team
I don't intend to pursue this. Unlike a $15-a-month gift subscription to Claude, a large grant like this would commit me to using AI in a serious way. You can see 29 projects who have gotten AI for Math Fund grants here.
I can understand people like Kevin Buzzard getting grants from these people, because of wanting to build a huge database of formalised mathematics pushing very far from basic material. But emailing you @John Baez is kinda funny, like prospective students who send out the same generic email (or with a few words changed) to many many potential supervisors in different countries, when the supervisor works on things the student knows next to nothing about, or, worse, the student is proposing a PhD project that's not in the area the recipient works on.
Timothy Gowers is getting grant money from them, and so is my friend Jamie Vicary, for a project on "New Categorical and Topological Foundations for AI and Machine Learning". Maybe they have too much money, and not enough good mathematician applying for grants, so now they're fishing around at the bottom of the barrel for people to give it to.
Ultimately funded by the late Jim Simons, who is no longer capable of being directly responsible for what his money is supporting.
Well, Tim Gowers has been interested in AI reasoning around automated proof for some years now. Maybe a decade? Having Fields medallists get grants helps put a little shine on the award scheme, I would think
When you're really rich it sometimes almost doesn't matter if you're still alive, because you accrete a huge organization that continues to function after your death.
Kevin Carlson said:
Ultimately funded by the late Jim Simons, who is no longer capable of being directly responsible for what his money is supporting.
Wait for the reveal that Claude is, in fact, Jim Simons' mind uploaded to a computer
That's why it's wicked smart
@John Baez thanks for sharing the Simon's foundation thing. Reminds me of the discussions here years ago about US researchers (esp post-docs) having essentially no choice but to accept funding funneled through military/"defense" because there is/was so little else. Now we're in a similar situation with AI.
Pedantry: Although the grantmaker is Renaissance Philanthropy, as far as we know the majority of the money comes from XTX, a different quant hedge fund
I haven't been able to find much information about Renaissance Philanthropy, so in particular I have no idea if it started before or after Simons' retirement
I sometimes joke that they are "Renaissance's tax writeoff department", but in reality I have absolutely no idea what they are
Morgan Rogers (he/him) said:
John Baez thanks for sharing the Simon's [sic] foundation thing. Reminds me of the discussions here years ago about US researchers (esp post-docs) having essentially no choice but to accept funding funneled through military/"defense" because there is/was so little else.
I don't really think there's been "so little else" until recently. I know a lot of postdocs and other researchers
who don't get military support: there's grant money from a host of federal agencies (NSF, DOE, NIH, etc.), and many mathematicians earn money the old-fashioned way, by teaching. Perhaps people who take military money say they're being "forced" to.
"Until recently": Trump tried to dramatically cut NSF grants, sometimes after they were already awarded. Congress and the courts reinstated most of those grants, but the administration (under Trump's control) has been slow-walking those grants, so the rate at which this year's grants are being given out is much lower than previously:
People talk about how this marks the collapse of US science, but I don't really know what the effect will be. Not good, obviously! (Unless you're Sabine Hossenfelder and think the whole system is corrupt and needs to be destroyed.)
It’s been perhaps closer to the case that there’s been “so little else” that applied category theorists have been able to win from outside academia. Topos, certainly, has had a great majority of our historic funding from defense or AI-related sources.
Jules Hedges said:
Pedantry: Although the grantmaker is Renaissance Philanthropy, as far as we know the majority of the money comes from XTX, a different quant hedge fund.
Oh wow! It looks like you're right. I wonder why they're doing this. They write:
NEW YORK/LONDON, March 5, 2025 – Renaissance Philanthropy and XTX Markets today announced the next phase of the AI For Math Fund, committing an additional $13.5 million. $10.5 million will be allocated to a new grant application round opening in March, with $3 million for micro-grants and other field-building opportunities. This builds on the fund’s initial $18 million commitment, bringing its total to $31.5 million - one of the largest philanthropic commitments ever dedicated to accelerating the development of AI and machine learning tools to advance mathematics.
I had been fooled into thinking Renaissance Philanthropy, who is giving out these AI grants, was related to Renaissance Technologies, Jim Simon's company - but now it looks like they are completely distinct!
I could still be wrong....
My understanding is that they are completely distinct and that RenPhil was spun out of Schmidt Futures, making it an Eric Schmidt project rather than a Jim Simons one.
Jules Hedges said:
I haven't been able to find much information about Renaissance Philanthropy, so in particular I have no idea if it started before or after Simons' retirement
RenPhil started very close to when Jim Simons passed away. It was sometime in the first half of 2024, and he died in May 2024 (the dating is not easy to establish for when the fund started, it's approximately May '24, as well). Definitely after he 'retired'!
Sometimes big profitable businesses have indeed a non-profit with a similar name whose job is to spread their philosophy and knowledge to the world and promote their values. It can be that in this case they do not have directly opened the non-profit, but somebody inspired by RenTech did.
Moin btw! :sun_face:
![]()
This is a borderline hallucination, but one potentially worth sharing. For context, the original discussion was a while ago with 4-o, concerning whether Geometric Complexity Theory had applications outside of P/NP (it does).
For a moment I thought you were discussing about the DSM, given the incipit. Thanks for providing the context. So AI for mathematical research is a cospiracy in your opinion, @Ben Kaminsky ?
Well, I originally used the response to win a Discord argument with someone...not exactly research but useful :sweat_smile: .
Coming back to Jim and the RenTech business for a moment, I stumbled across this inspiring Instagram reel yesterday evening and I wanted to share his words with you all: https://www.instagram.com/reel/DV0y7VQgY4f/?igsh=MWM0NDc5ZGF1b3Q1cw==
Hello! I'm not sure if this is the right place to share this but after watching this recent talk Terrence Tao gave on the need for new human-ai collaboration workflows, I built a GitHub App called Proof Flow (I didn't think very long on the name) for Lean repos. The way it works is that contributors propose theorems as structured issues, a maintainer approves them, and an AI agent (currently Aristotle) automatically writes the Lean 4 proof and opens a draft PR for review. I would welcome any constructive feedback.
proof flow repo: https://github.com/CoreyThuro/Proof-Flow/blob/main/README.md
proof flow how-to video: https://www.youtube.com/watch?v=U9cOowuspoU
Tao's talk on need for new workflows: https://www.youtube.com/watch?v=Uc2zt198U_U&t=3162s
I asked Claude a question yesterday and after some coaxing it figured it out which was cool. I wasn't totally stumped but I thought that the solution was creative
The question was, suppose you are given a finite presentation of a monoid and you know that the monoid generated by the presentation is not infinite. Is the equational theory decidable?
Answer: it is obviously semi decidable whether two expressions agree because you can search for a proof that they agree. So it suffices to prove it is semi decidable whether they are not equal.
Because you know the monoid is finite, it suffices to enumerate all finite Cayley tables representing a monoid and enumerate all homomorphisms into that Cayley table (a map from the generators into the Cayley table can be checked to satisfy finitely many relations in finite time). If you find a homomorphism from the free monoid where the terms are distinguished by the homomorphism then they are not equal.
It's a cute proof
One that probably exists in the literature but how to search for it conveniently?
On the other hand I asked it how to construct the unit map U_f in the other thread and it went in circles and wasted all my tokens and said I cannot ask a question for 24 hours lol
One that probably exists in the literature but how to search for it conveniently?
Did you ask Claude to search for it?
I see in this latest breakthrough in combinatorics, The sum-product conjecture is false for real numbers, that the authors write:
The role of AI in this proof. The authors were inspired to revisit the possibility of disproving the sum-product conjecture using number fields of large degree by the recent OpenAI counterexample to the unit distance conjecture (see [2]). Curiously, the final construction given here required far less number theoretic input than the unit distance counterexample. GPT-5.5 Pro was used as a sounding board in the early stages of the development of this proof, but the final proof, including all the main ideas, was almost entirely human-generated (the exception being the suggestion of Lemma 3.4, which replaced a more complicated result of Schinzel with a short elementary argument). Everything in this paper was written by the authors.
I wouldn't mind betting that this "sounding board" use will become very prevalent as time goes on. I wonder if people will continue to acknowledge it.
I suspect a lot of people are scared to acknowledge AI help now, given the widespread strong anti-AI sentiment.
John Baez said:
One that probably exists in the literature but how to search for it conveniently?
Did you ask Claude to search for it?
It was able to provide a reference. it is theorem 4.6 of lyndon and schupp "combinatorial group theory."
The property that makes the proof work has a name - "residually finite". Even if the monoid is not actually finite, it is still possible to enumerate all finite Cayley tables and consider all maps from the set of generators into the Cayley table such that the relations are satisfied. As long as this process eventually yields a finite group and a homomorphism that separates the two elements of the domain, you can eventually distinguish the two elements, which is the residual finiteness condition (for any two elements there exists a homomorphism into a finite monoid which distinguishes them)
![]()
AI bots started this practice of auto-filling your response bar to save you the effort of thinking of what to type next. I already thought this was extremely creepy but this is even more insidious lol. putting words in your mouth is weird. putting the words "i'm overthinking this, just do whatever you think is best" in your mouth is like, alarm bells
Some text editors for programming have been doing that for quite a while, reloading line autocomplete with each character you type, and I find it pretty mentally taxing even to look over somebody's shoulder while it's happening, for me it's a similar feeling to trying to hold a conversation while somebody else is talking in your ear
Today I had a conversation with an excellent mathematician who works at the intersection of algebraic topology and mathematical physics. He said Anthropic is hiring mathematicians, including himself, to improve Claude Fable by throwing hard math problems at it. They pay $250/hour for this, and also give these mathematicians a separate account where they can use Claude Fable for their own purposes. This job lasts 5 months maximum.
He said Anthropic was doing this with physicists before mathematicians.
https://www.anthropic.com/careers/jobs is it advertised somewhere here?
I don't see any open roles with the word "mathematics" or "mathematician". Maybe you gotta know somebody. :rolling_eyes:
right now the main research problem i'm interested in is like, coming up with algorithms that can automatically prove coherence theorems, and the setting i've chosen to work in is about characterizing the sets of cells in a generalized multicategory which is presented by generators and relations. but I can't really go too far without like, actual symbolic proofs that I can look at and try to think about the patterns involved in the proofs of coherence theorems.
i have been running Claude Sonnet on autopilot for a few days formalizing (in Rocq) the definition of the free virtual double category generated by a given set of generating 2-cells. i'm pretty impressed by how well it's been able to do without my explicit intervention. I spent maybe the better part of a day giving detailed instructions and setting up the main loop. The basic idea is that every 30 minutes or so, a new instance starts, reads the relevant code, reads the overall agenda, reads a chunk of work that is designated to it, tries to solve the problem, and then frames a 30 minute chunk of work for the next guy.
It finally hit a roadblock it couldn't solve in the proof of associativity of the 2-cell composition and i'll have to go in there and read a bunch of slop and get things sorted out. I will report back on how long this takes me and whether it would have been better to just write it myself.
Patrick Nicodemus said:
right now the main research problem i'm interested in is like, coming up with algorithms that can automatically prove coherence theorems, and the setting i've chosen to work in is about characterizing the sets of cells in a generalized multicategory which is presented by generators and relations. but I can't really go too far without like, actual symbolic proofs that I can look at and try to think about the patterns involved in the proofs of coherence theorems.
i have been running Claude Sonnet on autopilot for a few days formalizing (in Rocq) the definition of the free virtual double category generated by a given set of generating 2-cells. i'm pretty impressed by how well it's been able to do without my explicit intervention. I spent maybe the better part of a day giving detailed instructions and setting up the main loop. The basic idea is that every 30 minutes or so, a new instance starts, reads the relevant code, reads the overall agenda, reads a chunk of work that is designated to it, tries to solve the problem, and then frames a 30 minute chunk of work for the next guy.
It finally hit a roadblock it couldn't solve in the proof of associativity of the 2-cell composition and i'll have to go in there and read a bunch of slop and get things sorted out. I will report back on how long this takes me and whether it would have been better to just write it myself.
Formalizing Mathematics at Scale
Ahmad Rammal,
https://github.com/facebookresearch/autoform-bot
https://github.com/facebookresearch/atlas-lean
I am no longer on Twitter but someone shared some criticism of the ATLAS project with me here - https://x.com/Someody42/status/2060141320560124040
"A compact orientable manifold is defined to be a compact manifold"
1100 of the files have sorry's in them
from the Lean4 Zulip:
"I glanced at Clique.lean, which purports to show the reduction from 3SAT to CLIQUE for Turing machines. It does this by introducing a new infinite alphabet, each of whose symbols represents a specific 3SAT problem, so that each problem is encoded by one symbol. It then encodes CLIQUE into that alphabet as the set of one-symbol strings, so that the reduction is simply the identity map. There was then a second pass by a judge model that rated this definition of 5/5. There is a rigorous and correct definition of the CLIQUE language at the top of the file, but it is unused."
John Baez said:
Today I had a conversation with an excellent mathematician who works at the intersection of algebraic topology and mathematical physics. He said Anthropic is hiring mathematicians, including himself, to improve Claude Fable by throwing hard math problems at it. They pay $250/hour for this, and also give these mathematicians a separate account where they can use Claude Fable for their own purposes. This job lasts 5 months maximum.
He said Anthropic was doing this with physicists before mathematicians.
That's quite scary...
I can vouch that Fable is competent at higher category theory, unfortunately is absurdly expensive. Having smart LLMs assisting with math would be great but currently it looks only the richest institutions would be able to afford that, further increasing inequality in academia.
Being on Anthropic's payroll to pursue that seems pretty unethical to me... what's worse is that every time I publish something online I know I'm feeding the monster anyway.
On a slightly different topic, I was struck by the fact that at CT26 approximately zero of the talks I attended mentioned any use of AI. I don't have any experience yet myself actually using AI for category theory, since basically all I've been doing myself since it got good at math is coding. But I mentioned this to some non-category-theorist mathematician friends and they were also very surprised; one of them said that he just submitted a paper in which all the proofs were done by AI. Is there some reason that AI is less useful for category theory than for other branches of math? (I must admit that would please me, but it seems like wishful thinking.) Or are category theorists living in the past, either out of ignorance or by choice?
I would say it's mostly an issue of training data, but honestly idk. It's a good question!
"Training data" meaning that existing AI systems haven't been trained yet on category theory? If so, why not? I thought they were trained on the entire Internet, which certainly includes lots of category theory, both published in journals and at places like the nLab.
It feels to me category theory is comparatively less represented than other math subjects, especially higher category theory.
Before Fable 5, I found llms to be near useless for category theory, but Fable 5 has blown me away. Well, I don't think I've tried it with category theory yet, but it can do pointfree topology, which has significantly less training data (though it does have some notable gaps in knowledge here).
Since this is a recent model, it's probably too early for this to show up in conferences. (For my part, I've also been a bit reticent to ask it questions about things I'm seriously working on, cause I don't want to spoil the fun, though maybe I should, so that I can focus on topics I know it cannot do...)
We'll see soon how well Opus 5 does. It will be much cheaper to get access to. I haven't tried GPT 5.6 Sol.
But also note that all the seriously impressive headline results have been about finding (counter)examples / explicit constructions and there is less of this in category theory.
Graham Manuell said:
But also note that all the seriously impressive headline results have been about finding (counter)examples / explicit constructions and there is less of this in category theory.
I did notice that, but it seemed unlikely to me that there is some limit on the ability of AI to do less concrete math. I assumed that that observation had more to do with what sorts of questions the majority of mathematicians are interested in and hence likely to ask AI about. Do you think AI is not as good at abstract math? Can it come up with new definitions, for instance?
I think current AIs are better at coming up with complicated examples than complicated proofs, though they clearly have some ability to do both. I haven't seen any examples of them coming up with new definitions. Maybe here is where fewer people have asked and whether a definition is any good is harder to evaluate, though I suspect current AIs would be worse at this.
One does wonder about the visual aspect to category theoretic proofs (the diagrams, obv), but encoded in say TikZ/amcd/xypic/etc they do reduce to "language". But then there's also the aspect of the details of the diagrammatic proofs also not really being written out sometimes because they really are 'follow-your-nose'. I don't know if it's more or less than other areas of mathematics, though.
Is there anyone here who has tried to use an AI to help them do category theory, whether it be chasing diagrams, other kinds of proofs, finding (counter)examples, or coming up with definitions or conjectures? If so, was it helpful? (And what model?)
Graham Manuell said:
I think current AIs are better at coming up with complicated examples than complicated proofs, though they clearly have some ability to do both. I haven't seen any examples of them coming up with new definitions. Maybe here is where fewer people have asked and whether a definition is any good is harder to evaluate, though I suspect current AIs would be worse at this.
Tech companies have put considerable effort over the last three years into benchmarking mathematical performance via olympiad style questions and are, in many cases, staffed by people with a
competition maths background as their primary exposure to mathematics.. In my limited experience, coming up with complex (counter)examples in algebra is much closer to olympiad maths than proving theorems in terms of workflow. Ie., generally one spots an invariant and plays with some combinatorics.
I suspect this is why they are weaker at category-theoretical problem solving, as the overlap in style is much less pronounced than in, say, combinatorics or modern analysis.
The only CT type work that I know of that's (publicly) gone into AI-for-maths work are those questions a little while back that people got paid for, where the answer had to be a number. So it was something like "how many X are there for data Y" where X might be something like "natural transformations"
....and the data is some specific setup where the X turn out to have finitely many.
Hardly the type of question that people in CT generally care about, and I found it kinda weird. One did have to know some amount of theory to be able to actually figure out the answer, but it's a different mode of reasoning than the usual yoga of "well this should probably be the answer based on analogies with a bunch of other cases as well as examples from a bunch of different fields of mathematics, so I'll make a definition so that the theorem works out to be true in the most conceptually-nice way". AI is only just now managing to get to the point, I think, of not just mangling definitions in eg Lean, and that's for well-attested structures. Of course, in six months, who knows? But no one's done something like "please give a new definition of -category so that the technical claims behind the geometric Langlands conjectures that need them work out quickly and cleanly" (these claims of course are now theorems, but it's still hundreds of pages of work to fill the gaps Gaitsgory and Rozenblyum left in their two-volume book)
I have access to Fable and I was curious so I gave it your prompt copied verbatim (forgot to turn memory off so it gives some personalised answer, could retry incognito mode) https://claude.ai/share/0af6aa2c-8b39-45b0-8e1b-ad07126c6b75
(I can’t judge the quality of the answer though it looked convincing to me and cited a recent relevant paper)
I’ve been experimenting with LLMs not to develop new category theory but to apply existing theory to software development and it’s been a lot of fun so far! they’re not great at inventing new definitions but they’re getting surprisingly good at applying existing ones to new settings, e.g. asking it develop neural network architectures in terms of monoidal categories has helped them produce much cleaner code than without these abstractions available (you can have a look at a sample here)
Mike Shulman said:
On a slightly different topic, I was struck by the fact that at CT26 approximately zero of the talks I attended mentioned any use of AI. I don't have any experience yet myself actually using AI for category theory, since basically all I've been doing myself since it got good at math is coding. But I mentioned this to some non-category-theorist mathematician friends and they were also very surprised; one of them said that he just submitted a paper in which all the proofs were done by AI. Is there some reason that AI is less useful for category theory than for other branches of math? (I must admit that would please me, but it seems like wishful thinking.) Or are category theorists living in the past, either out of ignorance or by choice?
I sometimes use them to help understand things, including existing / known work. Have a had a Claude Pro subscription for roughly 2 years, sometimes dabble with ChatGPT (and earlier dabbled with Gemini). It's a mixed bag when asking theoretical questions (if I'd gotten the response Alexis just got from Fable I'd just close the window and do something else). Given how much effort is required to verify or understand what comes out, and how much semi-convincing garbage they produce, I'm not sure how much one could rely on them for higher category theory right now. Sometimes they're genuinely useful and I'll keep playing around to see how things develop, but currently for most of my own work it doesn't make much sense to start there. (I'm also not sure what it would even mean to have all of my proofs done by AI.)
For more computational stuff it's a completely different story, since the LLMs are good at writing programs for you. And depending on what you're doing and how serious it is, you can either verify the code yourself, or verify the output of the code.
@Alexis Toumi whose prompt?
please give a new definition of -category so that the technical claims behind the geometric Langlands conjectures that need them work out quickly and cleanly (these claims of course are now theorems, but it's still hundreds of pages of work to fill the gaps Gaitsgory and Rozenblyum left in their two-volume book)
I tried incognito but I can't share the results and it's hard to deactivate the personalisation for one run, I suspect Fable could have given me a sub-par answer because it knows from its memory of me that I'm a complete noob at higher category theory :joy_cat:
I also agree on the "semi-convincing garbage" judgement: these models are only useful if the answer is hard to come up with but easy to verify, I assume new definitions of infinity categories are exactly the opposite: easy to come up with (at least informally) but hard to verify. Giving the model access to a proof assistant is one way to handle with the verification, but you still have to verify that the verification aligns with what you asked.
Coming up with an informal idea is not hard because it is all over the literature. Actually making a real definition is much harder; coming up with a real definition of that magnitude that actually works as intended requires a different kind of thinking, it's a whole research program in itself
Well, that's the whole category theory magic trick, right? Come up with the right, often deceptively simple definitions so that the consequences neatly fall out of them...
GR were doing formal category theory in a proarrow equipment without saying so, and the decade of gap-filling was the community rediscovering that companions and conjoints, not globular 2-cells, are the native language of base change
btw does this take-away from Fable vibe check? /ot
@David Michael Roberts My impression, which matches that of other higher category theorists that I spoke to at CT, is that the culture among the mathematicians that "apply" higher categories (in AG, AT...)---as opposed to those who study higher categories per se---has shifted towards a kind of "informal axiomatic" use where "reasonable" facts about the -category of -categories are assumed without too much care for their status as theorems about specific models. In this culture, I think that "finding a definition where the facts are (more or less easily) provable" would hardly be noticed.
It's quite sneaky as it shifts the 'semantics' from "filling some gaps in a proof" to "confirming that a model satisfies the expected axioms", where the latter carries less prestige...
(Sorry if this is off topic for this conversation)
To actually connect it to the topic, I do think it is interesting to ask what kind of role is AI likely to play in a community that does not value "full formal rigour" very highly, and this sort of "informal axiomatic proof" style is given as much value as a more rigorous style of proof.
My experience with LLMs is that they are good at acting and they are good at coding. Correspondingly, in mathematics, they are usually good
It seems credible to me that the "informal axiomatic reasoning" that has taken ground in applications of higher categories may be soon mastered by LLMs under the "acting" mode...
On the other hand, what this style seems really terrible for is computation; a lot of what this kind of higher category theory is producing is abstract sledgehammers that can hardly spit out a number if one tries. So I think the "coding" mode has little applicability for the time being.
Patrick Nicodemus said:
It finally hit a roadblock it couldn't solve in the proof of associativity of the 2-cell composition and i'll have to go in there and read a bunch of slop and get things sorted out. I will report back on how long this takes me and whether it would have been better to just write it myself.
Okay. My conclusion is that it's not worth it to let the machine just go on autopilot right now. It made a lot of progress but it got bogged down eventually. I think like, if you sketch a proof of a complex theorem involving many layers of definition/theorem/proof - let's say there's of these - and you want to make a small, independent tweak to each definition and theorem - then any change to one results in changes to everything downstream, so making all these changes takes time . This could be minor things like implicit argument conventions or major design choices (you defined this in a 'bundled' way when it had to be 'unbundled.')
I would say it took at least two day's worth of work to fix the one day of work done by the bot independently.
This was also an expensive experiment. Anthropic gave all subscribers $100 in free tokens to promote the use of Fable, and I easily burned through that in a few days using the Sonnet and Opus models.
However, the ultimate goal was successful in that I do have a definition of the free virtual double category on a given set of 2-cells, and a machine checked proof that the composition law is associative and unital. Viewing small virtual double categories as T-monoids in the large virtual double category of graphs, I am optimistic that this code will serve as a good prototype for defining the free T-monoid in other virtual double categories.
Next I plan to define the free VDC presented by generators and relations, and study some specific such VDC's corresponding to the algebraic theories of: a category, a functor, a natural transformation, an adjunction, and so on.
Patrick Nicodemus said:
I spent maybe the better part of a day giving detailed instructions and setting up the main loop. The basic idea is that every 30 minutes or so, a new instance starts, reads the relevant code, reads the overall agenda, reads a chunk of work that is designated to it, tries to solve the problem, and then frames a 30 minute chunk of work for the next guy.
Can you point to a tutorial or something that explains how to set up something like this?
@Mike Shulman I didn't follow a tutorial for this, I just played around with it for a while. I use Claude Code, and there are a number of useful built-in commands which can help to get something like this off the ground - https://code.claude.com/docs/en/commands - in particular there is /loop, and the page https://code.claude.com/docs/en/sub-agents discusses how to have a parent bot launch and coordinate child bots. I think the most popular open-source alternative to Claude Code is OpenCode, which can be connected to (at least) any model hosted on OpenRouter, I would imagine that piece of software has similar functionality. But it's all vibecoded, so, you know, high expectations for sleek UI, low expectations for basic functionality working reliably.
Here is the repo. To initiate the loop, I open up Claude Code and say "Read parent_instructions.typ", and it ingests that file. You can ask it to generate the /loop command for you, or you can write something like "/loop every 30 minutes, launch a sub-agent, and have it read and carry out the instructions in loop.typ".
The human-authored documents are parent_instructions.typ, loop.typ, guidelines.typ, current_project_outline.typ. The AI-authored documents are handoff.typ, agenda.typ and todo_future_work.typ. The idea is that current_project_outline.typ is a high level sketch, and then agenda.typ has more granularity and should be continuously edited by the bot as the project evolves.
https://github.com/patrick-nicodemus/coherence_operads
Ok, thanks. I use Claude Code too, but I haven't figured out how to use /loop yet. I have tried to use /goal but couldn't figure out how to get it to keep running autonomously; it would go through one or two sessions and then always stop saying the goal-testing agent received too much data.
@Mike Shulman There should be a way to trigger the parent to respond to external events such as the agent returning, so I think /loop is maybe incomplete, it might be better to have the parent always act in response to the child terminating.
then always stop saying the goal-testing agent received too much data.
I'm not sure if we're talking about the exact same thing here, but my whole goal here with this 30 minute loop was to cap out the maximum context usage of any one agent in the series. Was your problem that the proof assistant goals were too large, in terms of the character count?
I don't know, I didn't really understand. /goal is supposed to run a quick agent after each main session to check whether the goal was achieved, and if not start another session to continue it. That seemed to me like what I wanted, to make the AI "continue working" on the same overall goal without my needing to tell it to go on every time it stopped. The goal test I set was pretty short, but possibly the checker was reading the entire output of the main agent; but if it has to do that and that was too long for it, then I don't see what use the /goal command is. I had thought that /loop was more for permanent periodic things like checking the status of a CI, but maybe that was wrong.
philip hackney said:
It's a mixed bag when asking theoretical questions (if I'd gotten the response Alexis just got from Fable I'd just close the window and do something else). Given how much effort is required to verify or understand what comes out, and how much semi-convincing garbage they produce, I'm not sure how much one could rely on them for higher category theory right now. Sometimes they're genuinely useful and I'll keep playing around to see how things develop, but currently for most of my own work it doesn't make much sense to start there. (I'm also not sure what it would even mean to have all of my proofs done by AI.)
For more computational stuff it's a completely different story, since the LLMs are good at writing programs for you.
Ehhhh ok, maybe the situation is different now. As an experiment inspired by @Mike Shulman 's question, I went on a little sidequest last night / today. I provided some background and asked about a question that's been bouncing around the back of my head for a couple years that I've never gotten around to, and probably wouldn't. I think I convinced myself at some point that it wasn't true, so it's not something I would've handed to a student or anything. But recently I got some more examples that I could feed in that I thought might've turned out to be counterexamples.
I prompted Claude Opus 5 (Medium) with some resources and the question. It wrote some python code, went through a bunch of examples and sketched a proof of one direction, gave an idea for the other direction. Then wrote up proofs of the whole thing, approximately five pages. Four prompts total, the second two of which were "Please commit them. Then try out the retract lemma in the way you think best." and "sure let's do that" lol.
I committed myself to actually checking it carefully. A couple minor issues, but overall very good, and proofs are correct. Pleasant to read, but that's probably because the "resources" I supplied were things that I wrote (papers and portions of papers), so I guess it's perfectly targeted to me.
Mike Shulman said:
one of them said that he just submitted a paper in which all the proofs were done by AI
Now I know what this means, but am even more confused. What do I do with this? It'd probably have been a paper I'd read if someone else wrote it, or a little paper I'd have felt good about if I figured it out and wrote it myself. But feels a bit worthless?
Yeah, I think this is the big question (ok, well, one of the big questions) about what mathematics will look like in the age of AI.
Why is it worthless? Because a human can't feel satisfied with, or get credit for, having solving it? I agree that those are important things that are missing. But it's still a contribution to knowledge -- even, I daresay, to human knowledge, since you've now read it and understood it and learned something from it, and so could other humans. Surely that's not completely worthless either.
My friend said that he thinks journals should continue to play the role of "repositories of correct mathematics" regardless of how that mathematics is produced, but that the allocation of credit to human mathematicians will have to adapt and be assigned some other way. That's one possibility; I suppose there are others.
You make very good points here, @Mike Shulman. I think it's a lot to contend with, and wasn't quite what I'd have expected even a few months ago. I think a better word than "worthless" is "hollow" -- it's a feeling similar to when you're stuck in a puzzle or adventure game and you look up the answer online. "Why am I even playing the game?" But a big difference is that we're not just playing a game here (though for me some of the best times do feel like that).
For now I've just documented the process. This post includes the actual problem/theorem, a narrative from my POV, (lightly-sanitized) Claude Code transcripts, a git repository containing inputs/outputs.
At this point it seems not impossible to me that eventually there will no longer be a societal need for professional full-time mathematicians: that whatever mathematics needs to be done willl be done by non-specialists using AI, the same way today anyone who needs to do calculations can do them themselves using an electronic computer rather than employing a human "computer" as was common a century ago. Those of us who enjoy "playing the game" might continue to do mathematics ourselves as a hobby, but (in the world I'm imagining) that sort of "artisanal" mathematics would be displaced in the large-scale economy by "industrial" mathematics done by AI.
If that happened, I would be sad, because I think something deep and meaningful about "artisanal" mathematics would be lost. But perhaps, as with all the other artisanal occupations that have been reduced to hobbies and displaced by industrial versions, the effect for society as a whole would be beneficial?
I'm not convinced there has ever been an immediate societal need for professional full-time mathematicians; people doing "mathematics that needs to be done" tend to be called engineers.
In the world you describe @Mike Shulman , scientific research returns to being a hobby for the privileged. A friend of mine who recently quit academia argued to me that it's only by historical accident that this ever stopped being the case (to the extent that access to the required training and integration into the relevant communities broadened during the 20th century, anyhow).
But why would that be a world I accept to participate in bringing about? Whom exactly is that world better for? It doesn't seem like it would be better for anyone I know personally.
Hello, I usually only read messages and never post anything on this zulip but after reading the last few messages I want to contribute. If you don't know me, I am a PhD student in Université Sorbonne Paris Nord (France).
I often read posts arguing or explaining what AI is able or not able to do in category theory or in other fields of maths/CS. My opinion on the mater is that with sufficient neurons, energy and time AI will surely be able to do anything we do. In my opinion the question of the place of AI in our research is not a question of what it is able to do, but at what cost. It is, I think , a bit what @Morgan Rogers (he/him) is saying in his last message: what do we want mathematical research to be like ?
I won't talk about environmental or the societal implication of contributing to this kind of tools and focus only on the aspect related directly to our practices. Before, I want to make clear that even if in this post I separate research from the rest of the world, I think it is false to operate this separation. We do not leave in a ivory tower separated from the rest of the world.
As a PhD student I live in fear that tomorrow I am going to see on ArXiv a paper, written by an LLM, claiming to answer a problem I am working on. This is also a fear many of my PhD friends have. I don't want to use AI for my research because (of ethical and political reasons and) I see the effect it has on the students I teach. I am learning about my field, the different techniques and perspectives etc... as you all know, this requires to understand and apply concepts/techniques. I wont be able to develop a feeling for things if I don't do it myself. I won't be able to learn how to approach a problem if I don't try and fail many times. If I want to be a fully fledged researcher I can not afford to use AI (contrary to what some people say). The students I teach that use AI are able to produce code that compiles and produce a decent enough output but they lack a deep understanding of it. They are often not able (or with much hardship) to solve bugs or have a macro vision of what their code does. I see no reason why I, as a "apprentice researcher", would be immune to those same effects I see in my own students. I don't want a world in which every PhD subject can be done in a few hours by the latest model, what will be left for the next generation of PhD students ? I don't blame the phd student that use AI, I have two friends that use it because they feel pressured into it. They think that without it they won't be "competitive", I understand their fear.
There is also the matter of getting a permanent position. In every sector in which AI got forcefully pushed a massive wave of firing took place. I don't see why AI getting pushed in mathematics will have a different effect. Thus, I am very hurt when I see fellow researchers pushing hard for AI, they are actively fighting against me and my friends getting jobs... I am more hurt by what is happening to math than what is happening to the programming world because it is not authoritarian bosses that are pushing AI against us, it is our own colleagues.
This last point is very important. As Morgan kind of said, we don't have a need to increase our efficiency. Why do we need to go faster ? The life of no one depends on the Riemann hypothesis being solved in the next ten years. On the contrary, the life of many students depends on job being available in next years to come. Since AI adoption is not something being pushed by an overpowering authority but by members of our own community it is something we can fight against. I don't care about speculation of what math will be in the next years I care about fighting right now for jobs and a practice of research that I enjoy and find meaning in.
I don't frequently post here either, but let me make an exception: Thank you for your sincere contribution @Jad Koleilat. I appreciate that you wrote up your perspective.
Jad Koleilat said:
As a PhD student I live in fear that tomorrow I am going to see on ArXiv a paper, written by an LLM, claiming to answer a problem I am working on. This is also a fear many of my PhD friends have.
Thank you for sharing your experience. As someone a couple of years further down the research pipeline, I have carved out a niche enough research theme that (perhaps naively) I do not worry too much about certain ideas being "scooped"; I hope that you and your PhD colleagues still have the same opportunity to experience becoming expert on a topic, and genuine human receptacles of human knowledge.
Morgan Rogers (he/him) said:
But why would that be a world I accept to participate in bringing about? Whom exactly is that world better for? It doesn't seem like it would be better for anyone I know personally.
I agree with the sentiment that this would not be a world I would like to help bringing about. But if we want to prevent this world from taking form, it is already quite late to take action against it. If we dont like the idea of such a world (how do we want to name it?), we should think about how to organize for defending humanity against ai. By this I dont necessarily mean ai as a technology. I use ai every day, all the time. But I want to see a world where we can use ai as a tool (if we so wish to do) as opposed to a world in which the owners of ai use us as a tool.
Jad Koleilat said:
I don't want to use AI for my research because (of ethical and political reasons and) I see the effect it has on the students I teach. I am learning about my field, the different techniques and perspectives etc... as you all know, this requires to understand and apply concepts/techniques. I wont be able to develop a feeling for things if I don't do it myself. I won't be able to learn how to approach a problem if I don't try and fail many times. If I want to be a fully fledged researcher I can not afford to use AI (contrary to what some people say). The students I teach that use AI are able to produce code that compiles and produce a decent enough output but they lack a deep understanding of it. They are often not able (or with much hardship) to solve bugs or have a macro vision of what their code does. I see no reason why I, as a "apprentice researcher", would be immune to those same effects I see in my own students.
This is very well said, thanks. I am a late-career researcher, so in a very different situation. I love to use AI to learn about other fields of research, math is so rich and AI has seen everything. But I am glad I went through the "try and fail many times" before AI came around. And now, as a teacher of math and CS, I have the same concerns as you do. Every year I try sth new, but I didnt yet solve the problem of how to teach in times of AI. How can we set up classes that encourage deep and slow learning while we know that students will use AI?
The two basic strategies are to prevent or discourage the students from using AI on certain tasks, or to allow them to use AI and give them tasks that are challenging even with AI assistance. One can do a mix of both.
(If you just discourage, not prevent, you and the students will always be worrying about cheating.)
I should add that my parenthetical remark assumes "credentialism": the idea that education is at least in part about getting credentials, which tends to make it into a competition and encourages cheating. I would love to get past this, but it's very hard to do, because it's deeply built into how education is done.
Alexander Kurz said:
How can we set up classes that encourage deep and slow learning while we know that students will use AI?
One option that a friend of mine does is to make the homework worth a fairly small fraction of the grade (since one can use the tool du jour on it -- currently LLMs, but historically wolframalpha, etc.) and then have (semi)frequent in-class or in-discussion quizzes which the students are told will be pulled entirely from homework problems. These are weighted a fairly large fraction of the grade.
Whatever resources they use to solve the homework, this approach tests their actual understanding of the problems in a situation where they're using their own brain... or at least it tests their ability to regurgitate ideas in a situation where they're using their own brain, but that's been a problem since time immemorial. Anyways, this encourages students to actually do the homework (or at least try to actually understand the LLM output of their homework solutions) since the larger fraction of the grade will be LLM-less.
Sorry to disappear for a while; I started writing a reply and then got shanghaied by personal issues and a planned vacation.
Morgan Rogers (he/him) said:
I'm not convinced there has ever been an immediate societal need for professional full-time mathematicians; people doing "mathematics that needs to be done" tend to be called engineers.
That may be true, but if so only because of the word "immediate". I don't think that, say, the NSF's funding of pure mathematics research over the past 75 years has been without benefit to society at large. This is the old "value of pure mathematics" debate. While lots of pure mathematics may not (yet) have any use, I think a fair amount of pure mathematics does eventually have a use.
In the world you describe Mike Shulman , scientific research returns to being a hobby for the privileged.
Well, I certainly don't see the social need for all scientific research going away, nor the impetus for public funding of it (at least, not because of AI). What I was imagining is that -- maybe -- public funding of mathematics research might be reduced from dedicated grants to mathematicians to become line items in grants to other scientists. I'm not saying I necessarily think that will happen; in fact, I can think of several arguments why it may not.
But if what you meant is that mathematical research returns to being a hobby for the privileged, then yes, perhaps that's more like the world I was imagining -- except that they might not need to be so privileged any more. Many technologies have a democratizing effect, and I can imagine AI helping to make higher mathematics accessible to many more hobbyists. (Of course, as has been pointed out elsewhere, at the moment the AI models capable of mathematical research come with a hefty price tag. But it's not hard to imagine that someday, perhaps soon, the free ones will be able to do serious mathematics as well.)
This democratization is also related to your last set of questions:
But why would that be a world I accept to participate in bringing about? Whom exactly is that world better for? It doesn't seem like it would be better for anyone I know personally.
For what it's worth, I don't consider myself to be "pushing" AI. I strongly disapprove of many actions taken by AI companies, including "copyright violation laundering". I worry about AI in the hands of malicious actors, as well as rogue AI behavior and "responsibility laundering". Sometimes I feel close to despair at the effect of AI on our ability to teach students basic skills, particularly reading and writing. And I worry about how AI will affect the research ecosystem, particularly early-career researchers, many of whom, as Jad pointed out, currently live in fear of being scooped by AI. And, as I said, I would be sad if the way humans do mathematics now were relegated to hobbyists.
At the same time, however, I think the evidence of history suggests that once any technology exists, it's nearly impossible to "put it back in the bottle". So it behooves all of us to get familiar with what these tools can, and therefore will, do. Similarly, speculating about possible futures is valuable even if they are undesirable, to help us make more informed decisions about how to affect the future.
I don't know whether the world I was imagining would be better or worse for humanity overall. I'm generally an optimist about technological progress, but in the case of AI I find it more difficult than usual (although I did my best in my BAMS article about it). Nevertheless, "it doesn't seem like it would be better for anyone I know personally" sets off alarm bells in my head, as that same phrase could have been said by many, many people facing the prospect of changes to, or elimination of, their jobs due to technological advance, but where looking back from a perspective of decades or centuries in the future, it is obvious that the technological advance benefited society as a whole.
In the case of AI, some of the benefits are already obvious. I haven't used AI for mathematical research yet, but I have used it for coding, and it is saving me a lot of time and enabling me to implement features that would have taken much, much longer by hand and perhaps never been gotten to. But I've also found that I have to be very careful to read and vet everything it produces, resisting the temptation to just trust its code because it "seems to work". I shudder to think what would happen if I didn't already know how to code, and I expect the same is true about mathematics.
I don't know what we should do about teaching at the undergraduate level and below. Personally, I'm going to muddle through this academic year somehow, maybe with oral exams, and then go on sabbatical, and after that who knows what the world will be like. And I'm not really very qualified to speak about the graduate level since I don't regularly advise graduate students, but shooting from the hip my advice would be to try to find a middle ground.
It's absolutely true that if you want to become a fully fledged researcher you need to be able to do mathematics without AI. But it also seems unlikely to me that 5, 10, 20 years in the future any researcher will be able to "compete" without using AI. (In some sense there's nothing new about this -- I firmly believe everyone should learn long division, but also in practice we use a calculator for any substantial division problem.) At the very least, testing how easy it is for AI to solve your research problems might give you some idea of how worried you should be about someone else using it to scoop you, and perhaps whether you should look for a different problem. It seems unavoidable to me that the nature of "PhD problem" will have to change -- which, of course, also happened with the introduction of ordinary computers.
I could be very wrong, of course, but those are my thoughts at the moment.
Mike Shulman said:
At the same time, however, I think the evidence of history suggests that once any technology exists, it's nearly impossible to "put it back in the bottle". So it behooves all of us to get familiar with what these tools can, and therefore will, do.
Have you seen a dirigible recently? Or a guillotine?
However much I would enjoy the return of these technologies, I must reject the argument that something being shown possible makes it inevitable.
Mike Shulman said:
Nevertheless, "it doesn't seem like it would be better for anyone I know personally" sets off alarm bells in my head, as that same phrase could have been said by many, many people facing the prospect of changes to, or elimination of, their jobs due to technological advance, but where looking back from a perspective of decades or centuries in the future, it is obvious that the technological advance benefited society as a whole.
There's a problem with the causation sequence here. It's not so clear to me that improvements in society -- which I interpret to mean the general trend towards improved living conditions over time that is presently collapsing -- over the decades or centuries can be attributed to the mere existence or widespread use of any technology. Perhaps you're thinking of the industrial revolution, which entrenched systems of exploitation that continue to this day (only now they're far enough away from us that we can pretend they don't underpin the wealth of our own nations)?
It's easy to see in the present that the benefits of LLMs are concentrated in the hands of a small number of people and easily arguable that they have had a net negative effect on society so far (you pointed to some of the problems). The cost-benefit should certainly not be reduced to an individual level -- the direct cost to the consumer of LLM services is by no means reflective of the indirect costs that we must collectively shoulder to enable this technology to exist, including but not limited to the consequences of them exceeding our capacity of available energy and water, which in turn will negatively impact the cost of living in all sorts of ways.
Mike Shulman said:
It's absolutely true that if you want to become a fully fledged researcher you need to be able to do mathematics without AI. But it also seems unlikely to me that 5, 10, 20 years in the future any researcher will be able to "compete" without using AI.
My personal response to this is that I would be prepared to invest a lot of time and energy into transforming academia not to include AI but to remove competition.
Experience has shown me that it does no good to try to argue with people expressing those sorts of beliefs, so I will now bow out of the conversation.
Although maybe I can clarify, without restarting an argument, that my use of the word "compete" was perhaps ill-advised. By putting it in scare quotes I meant to indicate that I didn't mean it literally, but it could easily have been misunderstood. What I really meant was something more like "contribute usefully to the scientific enterprise", which would be equally relevant even if science became entirely non-competitive.
Morgan Rogers (he/him) said:
Or a guillotine?
Actually, yes
Mike Shulman said:
Experience has shown me that it does no good to try to argue with people expressing those sorts of beliefs, so I will now bow out of the conversation.
I'm sorry you feel this way Mike, I appreciated you sharing your views here. For what it's worth, I too think AI is here to stay, because it's evidently a very powerful thing, and people (good and bad) like power. What really troubles me are the socioeconomical conditions in which this technological revolution is setting off.
Morgan Rogers (he/him) said:
. It's not so clear to me that improvements in society -- which I interpret to mean the general trend towards improved living conditions over time that is presently collapsing -- over the decades or centuries can be attributed to the mere existence or widespread use of any technology. Perhaps you're thinking of the industrial revolution, which entrenched systems of exploitation that continue to this day (only now they're far enough away from us that we can pretend they don't underpin the wealth of our own nations)?
That's a bit tendentious, isn't it? I'm not a techno-optimist but I'm a techno-realist. Technology did wonders for humanity, and a good deal of that can be attributed to automation. It used to be that >90% people had to work in food production, and now it is way less that 10% in developed economies. Those 80% of people are doing things like waiting at cafès, writing TV shows, staff universities, and some research category theory too!
That doesn't mean everything has been great (climate crisis, anyone?) and I agree that a lot of exploitation has been moved to the global south (that's quite a recent thing though---and less due to technological improvement and more to political and economical factors).
You're drawing the same line from technology merely existing to the improvements you're talking about as Mike did. Automation doesn't inherently make life better:
What exactly is it about automation that improved society? Lowering the output costs has as a side effect that a greater proportion of people can afford a product; goods that were formerly luxuries became accessible to more people. Capital was thus able to drum up greater demand to produce a better return on their investment into automation.
Those that are above the new line of affordability find themselves more comfortable, with certain aspects of their lives more convenient, while those that are pushed out (and fall below the line due to automation) fade into history.
To put it succinctly, technology tends to benefit its consumers, not its workers.
It's the workers that this discussion is about.
Matteo Capucci (he/him) said:
Those 80% of people are doing things like waiting at cafès, writing TV shows, staff universities, and some research category theory too!
If automation previously eliminated jobs that we now view as low-skilled (ironically that label is often only applicable due to automation lowering the skills required to do these jobs) and left 'more room' for people to pursue higher-skilled jobs, and that's an improvement, then what could the benefit of eliminating higher-skilled jobs be?