Automated Grading for Efficiently Evaluating the Dual-Use Biological Capabilities of Large Language Models

Bria Persaud, Jeffrey Lee, Jordan Despanie, Helin Hernandez, Henry Alexander Bradley, Sarah L. Gebauer, Greg McKelvey, Jr.

ResearchPublished Jun 12, 2025

Advances in the biological knowledge and reasoning capabilities of large language models (LLMs) have sparked interest in assessing the potential of LLMs to facilitate emerging biological risks. The authors of this report evaluated LLMs' abilities to answer knowledge-based questions and generate protocols that explain how to perform common laboratory techniques that could be used in the creation of proxies for biological threats. Because LLM evaluation approaches that rely on human subject-matter experts are often costly and time-intensive, the authors introduced an automated systematic and scalable method for evaluating the ability of LLMs to generate protocols for laboratory techniques. The results presented in this report confirm prior work indicating that LLMs possess knowledge of the biological sciences. This report is intended to inform evaluators of artificial intelligence systems, academics, technical experts, and policymakers on techniques for examining the risks of the convergence of LLMs and biological threats.

Key Takeaways

  • Of the 11 LLMs assessed by the automated grader (autograder), those that achieved the highest scores were GPT-4 and Claude Opus 3. Across LLMs, performance on the authors' benchmark correlated with general reasoning capabilities.
  • Preliminary investigations indicate that the autograder reaches parity with human experts in scoring model responses with the authors' rubrics. This suggests that automated evaluation of the dangerous capabilities of LLMs may be feasible, reducing dependence on costly, time-consuming studies that require human participants.
  • The results indicate that, when asked to generate the stepwise instructions a human would follow to perform common laboratory procedures, an LLM can provide relevant instructions.
  • A more rigorous study of the autograder's accuracy is needed to make decisive conclusions about its effectiveness. To develop more-reliable autograders, the authors will expand the set of techniques and questions to account for different threat models and address challenges in rubric design, such as the sequence of steps and clarity.

Topics

Document Details

Citation

Chicago Manual of Style

Persaud, Bria, Jeffrey Lee, Jordan Despanie, Helin Hernandez, Henry Alexander Bradley, Sarah L. Gebauer, and Greg McKelvey, Jr., Automated Grading for Efficiently Evaluating the Dual-Use Biological Capabilities of Large Language Models. Santa Monica, CA: RAND Corporation, 2025. https://www.rand.org/pubs/research_reports/RRA3124-1.html.
BibTeX RIS

Research conducted by

This publication is part of the RAND research report series. Research reports present research findings and objective analysis that address the challenges facing the public and private sectors. All RAND research reports undergo rigorous peer review to ensure high standards for research quality and objectivity.

This document and trademark(s) contained herein are protected by law. This representation of RAND intellectual property is provided for noncommercial use only. Unauthorized posting of this publication online is prohibited; linking directly to this product page is encouraged. Permission is required from RAND to reproduce, or reuse in another form, any of its research documents for commercial purposes. For information on reprint and reuse permissions, please visit www.rand.org/pubs/permissions.

RAND is a nonprofit institution that helps improve policy and decisionmaking through research and analysis. RAND's publications do not necessarily reflect the opinions of its research clients and sponsors.

Version Note

This publication supersedes a previous version published in 2025 (WR-A3124-1).