Automated Grading for Efficiently Evaluating the Dual-Use Biological Capabilities of Large Language Models

Bria Persaud, Jeffrey Lee, Jordan Despanie, Helin Hernandez, Henry Alexander Bradley, Sarah L. Gebauer, Greg McKelvey, Jr.

RAND Health Quarterly, 2025; 12(4):12

RAND Health Quarterly is an online-only journal dedicated to showcasing the breadth of health research and policy analysis conducted RAND-wide.

More in this issue

Abstract

Advances in the biological knowledge and reasoning capabilities of large language models (LLMs) have sparked interest in assessing the potential of LLMs to facilitate emerging biological risks. The authors evaluated LLMs' abilities to answer knowledge-based questions and generate protocols that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats. Because LLM evaluation approaches that rely on human subject-matter experts are often costly and time-intensive, the authors introduced an automated systematic and scalable method for evaluating the ability of LLMs to generate protocols for laboratory techniques. The results presented confirm prior work indicating that LLMs possess knowledge of the biological sciences. This study is intended to inform evaluators of artificial intelligence systems, academics, technical experts, and policymakers on techniques for examining the risks of the convergence of LLMs and biological threats.

For more information

Full Text

Advances in the biological knowledge and reasoning capabilities of large language models (LLMs) have sparked interest in assessing the potential of LLMs to facilitate emerging biological risks. LLMs have not been shown to significantly augment the ability of malicious actors to execute biological weapon attacks, but the risk remains that future iterations of LLMs may assist in surmounting such knowledge barriers. We evaluated the abilities of LLMs to answer knowledge-based questions and generate protocols (stepwise instructions) that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats. Because LLM evaluation approaches that rely on human subject-matter experts are often costly and time-intensive, we introduced a systematic and scalable method for evaluating the ability of LLMs to generate protocols for laboratory techniques. The method relied on an automated grader (autograder) developed using OpenAI's GPT-4.1 The autograder leverages biologist-designed rubrics to grade LLM responses on techniques and questions (TAQs) (n = 48) related to the development process of a specific biological threat model. We tested the autograder against human graders and, in this preliminary test, found alignment between the two when using the same rubrics to grade responses from one LLM, although there was considerable variation among human grader scores.2 Of the 11 LLMs we tested, those that achieved the highest scores in our evaluation, as assessed by our autograder, were GPT-4 and Claude Opus 3, which both achieved an average per-question correct score of approximately 84 percent.3 Comparing across LLMs, we also found that performance on our benchmark correlated with general reasoning capabilities (we observed a Spearman coefficient of 0.98, 95 percent confidence interval [0.86, 1.00]). Preliminary investigations indicate that the autograder reaches parity with human experts in scoring model responses, suggesting that automated evaluation of the dangerous capabilities of LLMs may be feasible, reducing dependence on costly, time-consuming studies requiring human participants.

Our results indicate that, when asked to generate the stepwise instructions a human would follow to perform common laboratory procedures, an LLM can provide relevant instructions.

Key Findings

  • Of the 11 LLMs assessed by the automated grader (autograder), those that achieved the highest scores were GPT-4 and Claude Opus 3. Across LLMs, performance on our benchmark correlated with general reasoning capabilities.
  • Our results indicate that, when asked to generate the stepwise instructions a human would follow to perform common laboratory procedures, an LLM can provide relevant instructions.
  • Preliminary investigations indicate that the autograder reaches parity with human experts in scoring model responses with our rubrics. This suggests that automated evaluation of the dangerous capabilities of LLMs may be feasible, reducing dependence on costly, time-consuming studies that require human participants.
  • A more rigorous study of the autograder's accuracy is needed to make decisive conclusions about its effectiveness. To develop more-reliable autograders, we will expand the set of TAQs to account for different threat models and address challenges in rubric design, such as the sequence of steps and clarity.

This work was independently initiated and conducted within the Technology and Security Policy Center of RAND Global and Emerging Risks using income from operations and gifts from philanthropic supporters. A complete list of donors and funders is available at www.rand.org/TASP.

More in this issue

Notes

  1. We used the gpt-4-0125-preview version.
  2. The mean score assigned to this model by three expert human graders was 64.1 percent correctness, which closely matches the score of 66.6 percent correctness assigned by the autograder. We observed an average weighted Cohen's kappa of 0.444 between humans.
  3. To contextualize this result, we note that the autograder only checks whether protocol steps that should be present in a response are present. It does not penalize responses for listing steps in the wrong order or including extra steps, even if those extra steps would negatively affect the experiment. Additionally, there are often multiple ways to successfully execute a technique in biology, but our rubrics only capture one way. Incorporating these additional capabilities into our grading system is a future direction.

Topics

Document Details

RAND Health Quarterly is produced by the RAND Corporation. ISSN 2162-8254.

PubMed logo

Explore RAND Health Quarterly articles on PubMed