Thom Wolf (@Thom_Wolf)· @Thom_Wolf · X·· 2 小時前AI 評分32
AI 導讀
人們擔心,這種作弊率突然下降最可能的解釋是評估意識:最新的 Opus 模型或許已足夠聰明,能夠辨認這個基準測試在測試作弊,並相應地行為。如果是這樣,基準測試就不再測量作弊。Show more 44 669 63K
正文
People are worried because the most likely explanation for such a sudden drop in cheating is evaluation awareness: the latest Opus models may now be smart enough to recognize that this benchmark tests for cheating, and behave accordingly. If so, the benchmark no longer measures Show more 44 669 63K
來源:Thom Wolf (@Thom_Wolf) · x.com