Thank you for the Berry post. Has anyone tried to test his hypothesis? My gut feeling is it will check out, particularly for AI chats. Input measure is easy–you have time and text quantity. But output value is more difficult...I'm wondering if there is a good case where this could be observed.
I'm doing a PhD in sociology of work (I have a background in software development) and want to do such a study on software development with AI, but as you say it's quite hard to find a good setting, because it's hard to operationalise quality.
I have been thinking of measuring it along these lines perhaps... https://www.gitclear.com/the_ai_code_quality_maintainability...
But they have a closed sourced software so I would have to find or create an open one in that case.
Happy to hear if you have any ideas on study design!