From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
Summary
A new multilayer taxonomy for Large Language Model (LLM) evaluation has been introduced, organizing capabilities into 14 domains and 91 subskills across Primitive, Constructed, and Integrative layers. This taxonomy, guided by human cognitive science, addresses the current fragmentation of LLM evaluation, which is task-based rather than capability-focused. To demonstrate its utility, researchers screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025, mapping 15,934 LLM-focused papers. Analysis revealed research concentration on Language-Semantic Competence (22.3%), Reasoning (21.3%), Planning and Decision-Making (13.5%), and Perception (12.3%), while six domains appeared in fewer than 2% of papers. The taxonomy aims to improve research organization, coverage audits, and evaluation interpretation.
Key takeaway
For research scientists evaluating Large Language Models, this capability-based taxonomy offers a structured approach to move beyond isolated task performance. You should consider adopting this multilayer framework to organize your research, audit evaluation coverage, and formulate more precise hypotheses for LLM diagnosis and training. This shift can enhance cross-study comparisons and reveal critical capability gaps in current models.
Key insights
A new LLM taxonomy organizes evaluation by capabilities, not tasks, improving research and identifying gaps.
Principles
- LLM evaluation benefits from capability-centric organization.
- Human cognitive science can guide LLM capability definition.
- Developmental precedence informs capability layer assignments.
Method
The method involves defining 14 capability domains and 91 subskills across three layers, guided by human cognitive science, then mapping research papers to these capabilities via multi-model annotation.
In practice
- Use capability taxonomy for LLM evaluation design.
- Audit research coverage using capability domains.
- Formulate testable hypotheses for LLM diagnosis.
Topics
- LLM Evaluation
- Capability Taxonomy
- Cognitive Science
- Research Mapping
- Language-Semantic Competence
- Reasoning
Best for: AI Scientist, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.