Category Theory
Zulip Server
Archive

You're reading the public-facing archive of the Category Theory Zulip server.
To join the server you need an invite. Anybody can get an invite by contacting Matteo Capucci at name dot surname at gmail dot com.
For all things related to this archive refer to the same person.


Stream: practice: software

Topic: Workflows for using LLMs in research


view this post on Zulip Will Jump (Aug 28 2026 at 17:46):

Up until recently I've been using LLMs as a brainstorming aid by interacting with its chat feature and asking it for references to read. Has anyone got a more advanced workflow for applying LLMs to maths research problems? - I'm not talking about LLMs doing autonomous research, rather I want it to help me brainstorm but do it in a way which is more digestable, more accurate and (if possible) more token efficient.

I've been playing around with spawning subagents to adversarially check outputs and generating obsidian vaults so I can interactively read the content it produces. But I'd be interested to hear if people have tried some more advanced tricks w.r.t making it generate readable artifacts and making it interactive/steerable.

Thanks!

view this post on Zulip Morgan Rogers (he/him) (Aug 28 2026 at 20:56):

If you got a very useful answer, would it mean you used (and ultimately depended on) LLMs more?

view this post on Zulip Mike Shulman (Aug 29 2026 at 02:29):

What happened to the original question in this thread?

view this post on Zulip Morgan Rogers (he/him) (Aug 29 2026 at 08:23):

The OP messaged me to say that they had subsequently come across the other discussions about using AI in research and didn't want to participate in a discussion of that type so had deleted their original question.

view this post on Zulip Mike Shulman (Aug 29 2026 at 17:33):

That's too bad. I would have liked to hear answers to the question. I would hope that we could separate questions about whether from questions about how, and that those who think the answer to "whether" is "never" could just refrain from participating in discussions about "how".

view this post on Zulip Mike Shulman (Aug 29 2026 at 17:35):

I suppose anyone who has an answer could go ahead and post it here anyway. I don't remember how the question was phrased exactly, and the OP clearly already had more experience using LLMs in research than I do, but IIRC it was basically about what the most effective methods are for using LLMs as a brainstorming tool in mathematical research.

view this post on Zulip fosco (Aug 29 2026 at 18:15):

I don't consider myself a heavy user, but I can share a couple of things that are solidifying in my workflow.

I just started experimenting with building skills that agents can rely on, and that they can see repo-wide or system-wide: an example among many, I have a SKILL.md file whose task is "Check a claim from <main.tex> against a corpus of relevant literature [for example turned to txt via OCR] — find who already said it, when, who said something weaker or stronger, providing verbatim quotes, line numbers and \cite keys. Use case is: the user gives a claim, sentence, theorem or \label in main.tex, asks whether it is in the corpus, whether it is novel, who to cite for it, or asks to add pointwise citations to a passage. Also covers building and refreshing the corpus."

==

The kind of output I receive is something like this (I'm copypasting from the .md itself)

Every hit is a relation to the claim, instead of a similarity score:

verdict meaning
same priority citation — someone got there first
weaker special case or extra hypotheses → cite as antecedent, say what you drop
stronger :warning: the claim may be subsumed by existing work
assumed asserted without proof → proving it is the contribution
contradicts must be addressed in the text
adjacent cite as context, not as priority

This is just an instance of something naive with which I am experimenting. Something related I'd like to experiment with, that I recently learned can be done with ease inside opencode or claude's cli, is "red teaming": letting a flock of Sonnet sub-agents, managed by a Opus super-agent, provide me with extremely harsh reviews of the main file in the repo, to be compared with the output of the previous skill.

Besides "knowledge representation", I try to make all code management (linting, scripts for latexindent, spellcheck, adding labels uniformly as the writing proceeds, etc) as trivial as possible. When I use agda (lately, I've been using it a lot), there's a skill that periodically aligns my code with the agda-stdlib style guide.

view this post on Zulip fosco (Aug 29 2026 at 18:25):

Something I did not ask, and that Opus figured out would have pleased me, is that the papers in the corpus are tagged by "tradition": they are tagged by a label like automata theory, lambda calculus, pure ct | classical, pure ct | enriched, topos theory, none of the above, but still in this category, etc.