So what this paper tries to study is that when. tllms become more expertt than curent human, then who will verify them ? How will a human (non expert ) verify an expert LLM. So they ran a test where an expert LLM which has the knowledge/ facts to arrive at an answer and a judge who is equally capable but has the knowledge gap limiting to choose the right answer.