The authors use item response theory and an interpretive taxonomy to address the motivating question: “What can benchmark data reveal about what artificial intelligence (AI) can do now that it could not do before?” Findings indicate that many benchmark tasks are no longer informative, but there is a discriminating frontier of tasks, with recent gains in performance concentrated in difficult tasks theorized to be relevant to real-world risks.
This publication is part of the RAND research report series. Research reports present research findings and objective analysis that address the challenges facing the public and private sectors. All RAND research reports undergo rigorous peer review to ensure high standards for research quality and objectivity.
This document and trademark(s) contained herein are protected by law. This representation of RAND intellectual property is provided for noncommercial use only. Unauthorized posting of this publication online is prohibited; linking directly to this product page is encouraged. Permission is required from RAND to reproduce, or reuse in another form, any of its research documents for commercial purposes. For information on reprint and reuse permissions, please visit www.rand.org/pubs/permissions.
RAND is a nonprofit institution that helps improve policy and decisionmaking through research and analysis. RAND’s publications do not necessarily reflect the opinions of its research clients and sponsors.

