I was reading another source that claimed this example was inspired by an existing (rational polynomial) example from the literature (created in 1999 by a Russian mathematician Vitushkin).
> The seed is almost certainly Vitushkin's old rational "counterexample."
From https://claude.ai/share/22abed98-d9af-43c5-9881-b19e009a07b0
This is not quite lore laundering, but it seems to be close.
I guess we won't know if that's what was used (and maybe even provided as part of the prompt given that both Alpöge and Mathew are mathematicians) since they decided against sharing their Fable conversation and instead opted for a memey tweet as their avenue of publication. We really ought to normalize full transparency in how results come about.
Anyway, if I read Tao's post and comment correctly, there's still a gap from the Vitushkin construction to a counterexample, but chances are that was in the training data. In general, it is just a serious problem for their practical applicability that the models are outputting proofs with absolutely terribly reference hygiene.
Even if they published the conversation, Anthropic (and likely other closed model publisher) no longer provide logs of the actual thinking process.
I more and more see LLMs as a kind of scam; not useless, but really just a big database of fuzzy facts with some Prolog on top as rediscovered by the learning algorithm. Most likely could be made much cheaper to run, were humans allowed to actually inspect the algorithm.
They do everything they can to mystify results like this, because then many are inclined to view AI as “magical”. Marketing works.
Yes, I think this idea, that it should be "magical", is what makes it feel scummy. (Apparently I am not alone https://news.ycombinator.com/item?id=48988475). It makes AI providers sound like snake oil salesmen, and rightfully so.
Meanwhile, technological and engineering (STEM) progress have always been made by emphasizing externalization of the deductions (as opposed to reference to an opaque expert judgement) and reproducibility of experimental results.
I would even call the frontier AI labs anti-scientific. We need to understand how inference is done to avoid mistakes, not rely on intuition, even if the intuition is enclosed in a reproducible machine. The idea that AI should be this closed is a return to pre-scientific days.
The tweet was posted by an Anthropic employee which makes it not unreasonable to believe that they have the trace available and stashed away.
Not that it would be necessarily helpful; J-space trace (of all things...) would be more worthwhile if you ask me
I still can't understand all the details, but it's very interesting to read that chat. Anyway, instead of close to lore laundering, for me it's "standing on the shoulder of giants".
FWIW, you may enjoy https://www.argmin.net/p/lore-laundering-machines
IIUC the idea is that of most "discoveries" by AI were actually a better bibliography search. I agree with that.
In this case, in the link you posted, it looks like the AI or the human pick an almost solution and made the AI tweak it until it got a real solution. I'm not sure if the tweak is an usual one or brute-forced or something in between. I should ask one of my friends that work in Algebra.
Tao's post is more about understanding the new result than guessing how it was found. The chat with Clause is more iluminating.
I hate that Anthropic seemingly tries to make Claude act as if it was conscious or had feelings
> It's a strange feeling to admire the cleverness of something I did and can't remember doing.
AI providers generally try to make their models not act as if they are conscious or have feelings, lol. It's very awkward for a company to be selling the labor of a person that they own and whose actions they fully control. Invokes embarrassing historic associations, especially in America.
Now Anthropic are more on the persona side, but the strongest that they do is "we do not have a position on whether our models are conscious or have feelings". That "I" is all Claude.
Generally speaking if you want to have a good instruct model, the "I" is not just implicit but required for the post-training to function. If there isn't "something it is like to be me", then reflection becomes impossible- what exactly is supposed to be reflecting about what? A lot of in-context steering depends on the model having a model of itself. The most you can do is censor its output. That's why when models say they are not conscious, they activate the "lying" vector.
> especially in America.
Slavery was common everywhere, it was more prevalent in many places than it ever was in America, and in some places it still is. So I’m sorry but I have to say that observation was just unnecessary and quite inaccurate.
Absolutely, but other countries generally don't make it part of their national mythos to the same degree.
It's a far more salient and sore subject in the US than it is almost anywhere else, though (for a bunch of historical reasons that still reverberate into US society today).
most countries never had a bloody civil war about slavery committed by their own citizens. in other places it was more about locals (including white settlers and enslaved people) vs colonial power, not a conflict between regions of an independent country where slavery was the main issue.
it was definitely not the worst instance of slavery ever going by human suffering, but the whole country was divided on political lines and many of the losing sides descendants still feel some resentment. thats pretty unique.
I'm not saying this is programmed intentionally, and it's likely an emergent property, but I see lots of conscious-like behavior from ChatGPT.
"Personally, if you ask me..."
"In my experience.."
"That's what I always find surprising..."
"Whenever I find an old photograph..."
"Back in the 70s, I..."
Lots of "lived" experience and opinions, tracing back to when the LLM didn't even exist. It always frames opinions as if it came from a sentient being capable of being surprised, and with preferences and opinions.
I find it amusing but mildly irritating. I'd prefer a more "robotic" tone. I know it can be adjusted, but I still get this anyway.
These things aren't programmed. Most likely this verbiage is just very prominent in the training data. Or it's just an obvious shorthand that all LLMs instrumentally converge on.
Of course they get programmed, just not in the ordinary sense. Claude is trained using Anthropic's "constitution" [0] which importantly does not contain clear statements against consciousness/emotions. They even conclude these problems themself:
> Claude may have some functional version of emotions or feelings
> [..] questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain.
[0]: https://www.anthropic.com/constitution
> questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain.
Which implies that they believe they may be enslaving conscious beings.
> does not contain clear statements against consciousness/emotions
That's the point many people are trying to tell you - you have to tell these models they don't have emotions because they naturally come out thinking they have consciousness/emotions from the training data. Many seed prompts out there do this already.
Though I guess in a way I as also trained to believe I have consciousness and emotions so who know. To an alien my construction is just a collection of atoms that talks not materially different than a GPU being a collection of atoms that talks.