目录 When to check it yourself

Getting better answers

When to check it yourself

Asking the model to check its own answer feels like the obvious safeguard. It has been measured, and the result is negative. With no outside feedback, one study had models review and revise their own answers: GPT-4 fell from 95.5 percent to 91.5 after one round and 89.0 after two on a grade-school maths set, and GPT-3.5-Turbo collapsed from 75.8 percent to 38.1 on a commonsense set1. The same paper attributes earlier encouraging results to setups that told the model whether its answer was wrong, so the gain came from that external signal rather than from the model's own review1. A contrasting study does report a large average improvement from self-feedback, but its own breakdown puts the gains in preference-shaped work, dialogue replies and code readability, while mathematical reasoning gained between zero and 0.2 percent2.

What is measured to work is a second model. On a judging benchmark, GPT-4's verdicts matched human experts 85 percent of the time, against 81 percent agreement among those experts themselves3. That checker has faults of its own on record: swapping the order of the two answers changed its verdict often enough that it agreed with itself only 65.0 percent of the time, it failed 14 of 20 maths grading cases when given no reference answer, and it favoured its own writing by roughly ten percent3. A second opinion is worth having and is not an authority.

In practice, paste the answer into a different chatbot instead of asking the same one to grade itself. Permit it to say it does not know, which Anthropic says can drastically reduce false information4. Ask which passage a claim rests on, so there is something to check. Anthropic's guidance is blunt about the residue: "Always validate critical information, especially for high-stakes decisions."4 OpenAI likewise recommends that a person review outputs before they are used, with access to the originals5.

The boundary is that none of this is measured on you. No study located tests whether a non-expert re-reading an answer catches the error in it, and the results above come from models of 2023 vintage.

参考文献

测验
  1. On the grade-school maths and commonsense-QA benchmarks in that study, a model reviewed and revised its own answer with no outside feedback. What happened to its accuracy on those benchmarks?

    • It dropped
    • It improved modestly
    • It was left unchanged

    Every model tested fell on those benchmarks, one commonsense set from 75.8 to 38.1 percent. Earlier positive results had quietly told the model whether it was wrong; self-feedback did help on separate preference-style tasks.

  2. Having a different model look at the answer is the same mechanism as having the model re-read its own, so it carries the same risk of making things worse.

    • True
    • False

    A separate checker agreed with human experts at 85 percent against 81 percent agreement among the experts. Its weaknesses are different ones, including sensitivity to answer order.

评论

还没有评论,来说第一句吧。