Paper Reading AI Learner

Image, Word and Thought: A More Challenging Language Task for the Iterated Learning Model

2026-01-06 10:53:00
Hyoyeon Lee, Seth Bullock, Conor Houghton

Abstract

The iterated learning model simulates the transmission of language from generation to generation in order to explore how the constraints imposed by language transmission facilitate the emergence of language structure. Despite each modelled language learner starting from a blank slate, the presence of a bottleneck limiting the number of utterances to which the learner is exposed can lead to the emergence of language that lacks ambiguity, is governed by grammatical rules, and is consistent over successive generations, that is, one that is expressive, compositional and stable. The recent introduction of a more computationally tractable and ecologically valid semi supervised iterated learning model, combining supervised and unsupervised learning within an autoencoder architecture, has enabled exploration of language transmission dynamics for much larger meaning-signal spaces. Here, for the first time, the model has been successfully applied to a language learning task involving the communication of much more complex meanings: seven-segment display images. Agents in this model are able to learn and transmit a language that is expressive: distinct codes are employed for all 128 glyphs; compositional: signal components consistently map to meaning components, and stable: the language does not change from generation to generation.

Abstract (translated)

迭代学习模型模拟了语言从一代传到另一代的过程,以探索语言传播过程中施加的限制如何促进语言结构的出现。尽管每个被建模的语言学习者都从一张白纸开始,但如果存在一个瓶颈,即限制学习者接触到的表达数量,则可以导致产生一种缺乏歧义、受语法规则支配并且在连续几代人中保持一致性的语言,也就是说,这种语言是富有表现力的、具有组合性和稳定的。 最近引入了一种更为计算上可行且生态有效的半监督迭代学习模型,该模型结合了监督和无监督学习,并在一个自编码器架构内进行。这使得我们能够探索更大意义-信号空间中的语言传播动态。在这里,这种模型首次成功应用于涉及传达更复杂含义的语言学习任务:七段显示器图像。在这个模型中,代理能够学习并传递一种富有表现力的语言:为所有128个字符使用不同的代码;具有组合性:信号组成部分与含义组成部分一致映射;并且是稳定的:语言不会从一代传到另一代时发生变化。

URL

https://arxiv.org/abs/2601.02911

PDF

https://arxiv.org/pdf/2601.02911.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot