Nice! It's interesting to see quantitative ways of measuring slop. I'm curious what _would_ happen if you did plug these back into the LLM as a form of feedback. Goodharting may happen... or it could get the LLM to generate cleaner code potentially?