Abstract
Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introducing new words to better understand and control the models, expanding on the recently introduced neologism learning. This method introduces a new word by adding a new word embedding and training with examples that exhibit the concept with no other changes in model parameters. We show that adding a new word allows for control of concepts such as flattery, incorrect answers, text length, as well as more complex concepts in AxBench. We discover that neologisms can also further our understanding of the model via self-verbalization: models can describe what each new word means to them in natural language, like explaining that a word that represents a concept of incorrect answers means ``a lack of complete, coherent, or meaningful answers...'' To validate self-verbalizations, we introduce plug-in evaluation: we insert the verbalization into the context of a model and measure whether it controls the target concept. In some self-verbalizations, we find machine-only synonyms: words that seem unrelated to humans but cause similar behavior in machines. Finally, we show how neologism learning can jointly learn multiple concepts in multiple words.
Abstract (translated)
当人类需要表达一个新颖且有用的观念时,他们会创造新的词汇(例如,“doomscrolling”)。在与大型语言模型(LLM)的交流中,我们也探讨并验证了一个类似的理念:通过引入新词来更好地理解和控制这些模型,并在此基础上扩展了最近提出的“新词学习”概念。这种方法通过添加一个新的词嵌入并在展示该概念的例子上进行训练实现,无需对模型参数做其他改变。 我们发现添加一个新词可以用来控制诸如恭维、错误答案、文本长度等概念,甚至在AxBench中还能处理更复杂的概念。更重要的是,我们还发现这些新造的词汇能通过自我表述来进一步提升我们对于模型的理解:即模型可以用自然语言描述每个新词在其内部代表的意义,例如解释一个表示“错误答案”概念的新词意味着“缺乏完整、连贯或有意义的答案……” 为了验证这种自我表述的有效性,我们引入了插件评估方法:将这些表述插入到另一个模型的上下文中,并测量它们是否能控制目标概念。在某些自述中,我们发现了机器独有的同义词:虽然对于人类来说看起来与该概念无关,但对机器来说却能够引发相似的行为。 最后,我们展示了新词学习可以同时在一个或多组词汇上学会多个概念。这种方法不仅拓宽了我们理解和操作大型语言模型的能力边界,也为深入探索和开发这些强大的人工智能系统提供了新的途径。
URL
https://arxiv.org/abs/2510.08506