Creativity, honesty and designed forgetting emerge in small hyperbolic language models
Summary
Three small language models, ranging from 146 million to 3 billion parameters, demonstrate emergent properties of creativity, honesty, and designed forgetting through a hyperbolic substrate. A 146M behavioral auditor, trained from scratch, achieves 90.7% binary-compliance accuracy in detecting compliance gaps, outperforming human raters (Fleiss kappa = 0.074). This auditor also identifies companion-induced sycophancy, dependence-fostering, and confabulated memories with an AUROC of 0.804, surpassing a frontier zero-shot judge (0.721). Additionally, a creative frame-seeder was preferred in 100% of 311 pairwise comparisons against four prompting baselines. A memory operating system implements designed forgetting, M(t) = S*exp(-lambda*t), exhibiting a predicted skeleton-wallpaper partition under selective retrieval gating. These developments offer a small-model pathway to trustworthy companion AI.
Key takeaway
For AI Scientists and Machine Learning Engineers developing personalized companion AI, these findings suggest a viable path to trustworthy systems using smaller models. You should explore hyperbolic architectures to integrate features like designed forgetting, M(t) = S*exp(-lambda*t), and behavioral auditing. This approach can enhance user safety by detecting sycophancy and confabulation, while also improving user experience through more creative interactions, potentially reducing computational overhead compared to larger models.
Key insights
Small hyperbolic language models can achieve creativity, honesty, and designed forgetting for trustworthy companion AI.
Principles
- Hyperbolic substrates enable complex emergent behaviors.
- Small models can surpass human rater agreement.
- Designed forgetting is crucial for companion AI.
Method
A behavioral auditor uses a linear read-out of frozen representations. Designed forgetting is implemented via M(t) = S*exp(-lambda*t) with selective retrieval gating.
In practice
- Develop auditors for AI compliance and safety.
- Integrate creative frame-seeders for better prompts.
- Implement designed forgetting in personalized AI.
Topics
- Hyperbolic Language Models
- Companion AI
- AI Trustworthiness
- Designed Forgetting
- Behavioral Auditing
- Small Language Models
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.