NTI | bio and Concordia AI, with technical input from SecureBio, have launched the AIxBio Technical Working Group on Evaluation Practice. Convened through the AIxBio Global Forum, the Working Group will bring together experts from organizations in China, the European Union, the United Kingdom, and the United States that develop or conduct biological capability evaluations of AI systems.
Biological capability evaluations increasingly inform frontier AI risk management by assessing capabilities that could lower barriers to biological misuse. Model developers use the findings when deciding whether to strengthen safeguards, restrict access, increase monitoring, or change deployment plans.
As these evaluations become more widely used, findings produced in one setting may inform risk assessments and related decisions in another, even when evaluation methods and reporting practices differ. Results reflect not only the capabilities of the model being evaluated, but how those capabilities are tested. Without consistent reporting, differences arising from system configuration or evaluation design may be mistakenly attributed to the underlying model itself.
The Working Group will compare published evaluation approaches across jurisdictions and institutional settings to establish a common technical basis for designing evaluations and interpreting their results. Building on this analysis, it will develop shared terminology and a minimum reporting baseline for evaluation results. Its work will:
- Clarify what biological capability evaluations measure and how their design relates to the threat models they are intended to represent
- Identify the conditions under which results can be compared across systems and settings
- Distinguish what findings establish about the capabilities tested from what they support about biological risk
- Examine how evidentiary standards and thresholds for action affect the use of results
Methodological findings, voluntary reporting conventions, and research priorities will be published to support more consistent interpretation and use of evaluation evidence by AI developers, governments, and third-party evaluators.
By bringing together experts who conduct this work to examine assessment design and comparability, the group will complement the AIxBio Forum’s Working Group on Horizon Scanning, Risk Assessment, and Evaluations.
Organizations and individuals working on biological capability evaluations are invited to contact Helia Samani at [email protected] to discuss participation or share related work.