How to measure the capabilities of Artificial Intelligence? Researchers from FSV UK contributed to global research published in Nature
More than a thousand scientists from around the world joined forces to create Humanity’s Last Exam – one of the most significant benchmarks currently used to evaluate the capabilities of the most advanced artificial intelligence models. Through thousands of tasks across scientific disciplines, the benchmark evaluates the strengths and weaknesses of AI models. Among the contributors to this global project, whose findings were published in the prestigious journal Nature, are Vít Střítecký and Petr Špelda from the Institute of Political Studies at FSV UK.
What was the goal or main motivation behind the Humanity’s Last Exam project?
Humanity’s Last Exam (HLE) is a global collaborative initiative involving more than a thousand experts from a wide range of disciplines. Its aim was to create a highly challenging benchmark for evaluating the capabilities of the most advanced artificial intelligence models. Among contemporary AI benchmarks, HLE occupies a distinctive position because its broad disciplinary scope allows it to assess capabilities that developers often describe as multidisciplinary reasoning.
How would you explain to a general audience how an AI benchmark works and why it is useful?
An AI benchmark is typically a collection of tasks in which a model’s responses are compared against reference solutions or the current state of knowledge. Such evaluations help identify where a model performs well, where it falls short, and how it compares with other systems. Benchmark results can be used to compare models, provide a partial picture of technological progress, and contribute to broader discussions about how society and policymakers should respond to advances in artificial intelligence.
How does this new benchmark differ from previous tests? What makes it unique?
HLE’s uniqueness lies primarily in its extraordinary disciplinary breadth and in the global, collaborative nature of the project as a whole.

Was there any result that particularly surprised you?
One particularly interesting aspect is the impact HLE has had on the developers of the most advanced AI models themselves. Regardless of whether the companies are based in the United States or China, performance on HLE has become one of the key ways they demonstrate the capabilities of their models. As a result, the benchmark functions not only as a passive measurement tool but also, to some extent, shapes the direction of future development. Areas covered by the benchmark are likely to receive increased attention from developers seeking to improve their systems. Participation in such a collaborative project therefore represents not only a contribution to evaluation efforts but also an opportunity to indirectly influence the areas in which future AI models may advance.
What was your specific role in the project?
Through the evaluation task we contributed – based on previously unpublished research – we brought our expertise in AI safety, inductive inference, and related epistemological questions to the project. The task itself emerged from a long-term research programme. One of the main challenges, however, was translating a broader scientific problem into a format suitable for AI evaluation.
What are you planning to work on next within this area?
Given the rapid pace of AI development and growing societal uncertainty, we will continue to focus on research and teaching in the field of AI safety. Our work concentrates on specific issues that already provide insight into the types of safety challenges likely to emerge in future stages of AI development.