Paper Reading AI Learner

Neologism Learning for Controllability and Self-Verbalization

2025-10-09 17:41:57
John Hewitt, Oyvind Tafjord, Robert Geirhos, Been Kim

Abstract

Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introducing new words to better understand and control the models, expanding on the recently introduced neologism learning. This method introduces a new word by adding a new word embedding and training with examples that exhibit the concept with no other changes in model parameters. We show that adding a new word allows for control of concepts such as flattery, incorrect answers, text length, as well as more complex concepts in AxBench. We discover that neologisms can also further our understanding of the model via self-verbalization: models can describe what each new word means to them in natural language, like explaining that a word that represents a concept of incorrect answers means ``a lack of complete, coherent, or meaningful answers...'' To validate self-verbalizations, we introduce plug-in evaluation: we insert the verbalization into the context of a model and measure whether it controls the target concept. In some self-verbalizations, we find machine-only synonyms: words that seem unrelated to humans but cause similar behavior in machines. Finally, we show how neologism learning can jointly learn multiple concepts in multiple words.

Abstract (translated)

当人类需要表达一个新颖且有用的观念时,他们会创造新的词汇(例如,“doomscrolling”)。在与大型语言模型(LLM)的交流中,我们也探讨并验证了一个类似的理念:通过引入新词来更好地理解和控制这些模型,并在此基础上扩展了最近提出的“新词学习”概念。这种方法通过添加一个新的词嵌入并在展示该概念的例子上进行训练实现,无需对模型参数做其他改变。 我们发现添加一个新词可以用来控制诸如恭维、错误答案、文本长度等概念,甚至在AxBench中还能处理更复杂的概念。更重要的是,我们还发现这些新造的词汇能通过自我表述来进一步提升我们对于模型的理解:即模型可以用自然语言描述每个新词在其内部代表的意义,例如解释一个表示“错误答案”概念的新词意味着“缺乏完整、连贯或有意义的答案……” 为了验证这种自我表述的有效性,我们引入了插件评估方法:将这些表述插入到另一个模型的上下文中,并测量它们是否能控制目标概念。在某些自述中,我们发现了机器独有的同义词:虽然对于人类来说看起来与该概念无关,但对机器来说却能够引发相似的行为。 最后,我们展示了新词学习可以同时在一个或多组词汇上学会多个概念。这种方法不仅拓宽了我们理解和操作大型语言模型的能力边界,也为深入探索和开发这些强大的人工智能系统提供了新的途径。

URL

https://arxiv.org/abs/2510.08506

PDF

https://arxiv.org/pdf/2510.08506.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot