From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models
Summary
A new multi-layer taxonomy organizes Large Language Model (LLM) capabilities into 14 domains and 91 subskills across Primitive, Constructed, and Integrative layers, guided by human cognitive science, developmental precedence, and functional support. This framework addresses the fragmentation of task-centric LLM evaluation, which obscures underlying capabilities and limits cross-study comparison. To demonstrate its utility, researchers screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS (2023-2025), mapping 15,934 LLM-focused papers. Analysis revealed concentrated research attention on "Language-Semantic Competence" (22.3%), "Reasoning" (21.3%), "Planning and Decision-Making" (13.5%), and "Perception" (12.3%), with six domains appearing in under 2% of papers. The taxonomy facilitates systematic research organization, coverage audits, and diagnostic hypothesis generation for LLM development.
Key takeaway
For AI Scientists and Machine Learning Engineers designing or evaluating LLMs, recognize that task-centric benchmarks often obscure underlying capabilities. You should adopt a structured, multi-layer capability taxonomy to diagnose model strengths and weaknesses more effectively. This approach helps identify research gaps, interpret evaluation results beyond aggregate scores, and generate targeted hypotheses for improving LLM training and transfer, moving beyond brittle performance.
Key insights
A multi-layer taxonomy organizes LLM capabilities, moving beyond task-centric evaluation to reveal underlying cognitive structures and research gaps.
Principles
- Human cognition guides capability identification.
- Organize capabilities by developmental precedence.
- Adapt human constructs for LLM behavior.
Method
The taxonomy was developed through four iterative phases, then operationalized via an annotation codebook and scalable mapping procedure involving multi-model annotation, consensus, and arbitration of LLM papers.
In practice
- Audit LLM research coverage.
- Design capability-focused evaluations.
- Diagnose LLM performance issues.
Topics
- Large Language Models
- LLM Evaluation
- Cognitive Capabilities
- Taxonomy Design
- Research Mapping
- Human Cognition
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by cs.CL updates on arXiv.org.