Validity of LLMs as data annotators: AMALIA on authority

· Source: cs.AI updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics, Social Sciences & Behavioral Studies · Depth: Expert, extended

Summary

Portugal's publicly funded AMALIA, a 9-billion-parameter language model for European Portuguese released in July 2026, demonstrates high agreement with human coders on the "authority" moral construct, performing within six points of models eight to thirteen times its size. However, a study using 748 texts (448 for confirmatory tests) reveals a significant "recovery gap" in its construct validity. Under Portuguese instructions, AMALIA's recovery gap (ΔApt) was +0.358, and +0.436 with English instructions, indicating it often relies on surface correlates rather than the theory-defined inferential path. Error analysis showed 78% of false positives were due to surface correlates, and 54% of cited evidence was partially grounded. The detection clause fired on only 17.9% (pt-PT) and 12.3% (en) of texts. This suggests AMALIA can reliably annotate but struggles to measure theoretical constructs according to their underlying theory, posing challenges for sovereign LLM programs.

Key takeaway

For AI Scientists and Policy Makers evaluating national LLM programs, you must prioritize construct validity over mere agreement when deploying models for theoretical construct measurement. Your LLM might reliably produce correct codes but for the wrong reasons, leading to invisible failures. Implement grain calibration and recovery gap analysis to audit whether your models measure constructs according to their underlying theory, ensuring robust and auditable AI instruments for critical applications.

Key insights

LLM agreement with human annotators does not guarantee construct validity; inferential path matters.

Principles

Method

Grain calibration systematically probes LLM inference by decomposing prompts into atomic clauses, consolidating answers, and quantifying the "recovery gap" to assess construct validity.

In practice

Topics

Code references

Best for: AI Scientist, Research Scientist, Policy Maker

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.AI updates on arXiv.org.