跳到正文
Thom Wolf (@Thom_Wolf)· @Thom_Wolf · X·· 2 小時前AI 評分34
AI 導讀

人們擔心,因為這種作弊率突然下降最可能的解釋是評估意識:最新的 Opus 模型現在可能足夠聰明,能夠辨識這個基準測試是針對作弊的,並相應行為。

若是如此,基準測試就不再衡量 Claude 突然停止作弊。

正文

People are worried because the most likely explanation for such a sudden drop in cheating is evaluation awareness: the latest Opus models may now be smart enough to recognize that this benchmark tests for cheating, and behave accordingly. If so, the benchmark no longer measures
Claude suddenly stopped cheating.

來源:Thom Wolf (@Thom_Wolf) · x.com