Biological Sequence Models In The Context Of The Ai Directives
Summary
The White House's October 2023 Executive Order on AI mandates reporting for machine learning models trained on "primarily biological sequence data" using over 1e23 operations. A new dataset of nearly a hundred biological sequence models and 30 biological sequence datasets reveals that xTrimoPGLM-100B, a 100B-parameter protein language model, exceeds this threshold by a factor of six, with over a dozen other models within 10x. Training compute for biological sequence models has surged 8-10 times annually over the past six years, matching language model growth. Public databases contain approximately 7 billion protein sequences, significantly more than currently utilized. However, a regulatory gap exists: models like Galactica-120B, trained with over 1e23 operations but not "primarily" on biological data, may evade oversight despite incorporating extensive biological sequences.
Key takeaway
For AI scientists and policymakers developing or regulating biological sequence models, you must scrutinize models approaching or exceeding the 1e23 operations threshold. Be aware that models not "primarily" trained on biological data, even with significant biological components, might currently evade the Executive Order's reporting requirements, creating a potential oversight gap. Consider advocating for clearer regulatory definitions to ensure comprehensive risk management.
Key insights
The rapid scaling of biological sequence models and data necessitates updated AI governance to address emerging dual-use risks.
Principles
- AI compute growth in biology mirrors language models.
- Abundant biological data remains underutilized.
- Regulatory definitions impact oversight scope.
Method
The report curates a dataset of biological sequence models and datasets, analyzing training compute trends and data availability to identify models likely requiring regulatory scrutiny under the Executive Order.
In practice
- Monitor biological sequence models exceeding 1e23 operations.
- Evaluate models for "primarily biological data" definition.
- Explore large public protein sequence databases for training.
Topics
- AI Executive Order
- Biological Sequence Models
- Protein Language Models
- Training Compute Trends
- Biological Data Governance
- Regulatory Gaps
Best for: CTO, VP of Engineering/Data, Director of AI/ML, AI Scientist, Policy Maker, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Papers & Reports | Epoch AI.