Automated Grading for Efficiently Evaluating the Dual-Use Biological Capabilities of Large Language Models
RAND Health Quarterly, 2025; 12(4):12
RAND Health Quarterly, 2025; 12(4):12
RAND Health Quarterly is an online-only journal dedicated to showcasing the breadth of health research and policy analysis conducted RAND-wide.
More in this issueAdvances in the biological knowledge and reasoning capabilities of large language models (LLMs) have sparked interest in assessing the potential of LLMs to facilitate emerging biological risks. The authors evaluated LLMs' abilities to answer knowledge-based questions and generate protocols that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats. Because LLM evaluation approaches that rely on human subject-matter experts are often costly and time-intensive, the authors introduced an automated systematic and scalable method for evaluating the ability of LLMs to generate protocols for laboratory techniques. The results presented confirm prior work indicating that LLMs possess knowledge of the biological sciences. This study is intended to inform evaluators of artificial intelligence systems, academics, technical experts, and policymakers on techniques for examining the risks of the convergence of LLMs and biological threats.
Advances in the biological knowledge and reasoning capabilities of large language models (LLMs) have sparked interest in assessing the potential of LLMs to facilitate emerging biological risks. LLMs have not been shown to significantly augment the ability of malicious actors to execute biological weapon attacks, but the risk remains that future iterations of LLMs may assist in surmounting such knowledge barriers. We evaluated the abilities of LLMs to answer knowledge-based questions and generate protocols (stepwise instructions) that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats. Because LLM evaluation approaches that rely on human subject-matter experts are often costly and time-intensive, we introduced a systematic and scalable method for evaluating the ability of LLMs to generate protocols for laboratory techniques. The method relied on an automated grader (autograder) developed using OpenAI's GPT-4.1 The autograder leverages biologist-designed rubrics to grade LLM responses on techniques and questions (TAQs) (n = 48) related to the development process of a specific biological threat model. We tested the autograder against human graders and, in this preliminary test, found alignment between the two when using the same rubrics to grade responses from one LLM, although there was considerable variation among human grader scores.2 Of the 11 LLMs we tested, those that achieved the highest scores in our evaluation, as assessed by our autograder, were GPT-4 and Claude Opus 3, which both achieved an average per-question correct score of approximately 84 percent.3 Comparing across LLMs, we also found that performance on our benchmark correlated with general reasoning capabilities (we observed a Spearman coefficient of 0.98, 95 percent confidence interval [0.86, 1.00]). Preliminary investigations indicate that the autograder reaches parity with human experts in scoring model responses, suggesting that automated evaluation of the dangerous capabilities of LLMs may be feasible, reducing dependence on costly, time-consuming studies requiring human participants.
Our results indicate that, when asked to generate the stepwise instructions a human would follow to perform common laboratory procedures, an LLM can provide relevant instructions.
This work was independently initiated and conducted within the Technology and Security Policy Center of RAND Global and Emerging Risks using income from operations and gifts from philanthropic supporters. A complete list of donors and funders is available at www.rand.org/TASP.
More in this issueRAND Health Quarterly is produced by the RAND Corporation. ISSN 2162-8254.
Explore RAND Health Quarterly articles on PubMed