Yep, it's like getting a commoner from the street evaluate a literature PhD in their native language. Sure, both know the language, but the depth difference of a specialist vs a generalist is too large. And neither we can't use AI to automatically evaluate this literature genius because real AI doesn't exist (yet), hence the programs can't understand the contents of text they output or input. Whoops. :)