Is there evidence that LLMs generate better code in more popular languages? I get the sense the "experience" translates between languages and it can reason in any language just fine. I write Clojure code using a rather esoteric framework (Pathom3). There is probably very little similar code out there (it's definitely a tiny fraction of the training dataset) but it seems to do just fine
Not saying you're wrong, just curious if there are numbers backing this up.
In my experience LLMs are way better when they "know" a language "instinctively". It's just that unless your language is very niche, the corpus is usually good enough. I tried using Claude to write my own personal language a while ago (I wrote a toy compiler decades ago) and it struggled a bit, because you could see in it's reasoning it had to "repeat" the syntax equivalence to itself while it read the code. It didn't just "know" it could use a given construct to do something; conversely Astra, when carefully instructed to do so, can plop down esoteric template code that works the first time, because it just "knows" it's the right stuff to write
I built a brand new language to test this[1]. Not only is the language different to basically any other language but it also tries to be adversarial against LLM understanding.
The best models can still make sense of it[2], though the tasks so far have been pretty basic. But I do think it gives some evidence that languages which aren’t well represented in an LLM’s training can still be reasoned about and written well by LLMs.
1 - https://killswitch-lang.org
2 - https://bench.killswitch-lang.org