Text2Model: Model Induction for Zero-shot Generalization Using Task Descriptions

2022-10-27 05:19:55

Ohad Amosy, Tomer Volk, Eyal Ben-David, Roi Reichart, Gal Chechik

arXiv_CV

Abstract
Abstract (translated)
URL
PDF

Abstract

We study the problem of generating a training-free task-dependent visual classifier from text descriptions without visual samples. This \textit{Text-to-Model} (T2M) problem is closely related to zero-shot learning, but unlike previous work, a T2M model infers a model tailored to a task, taking into account all classes in the task. We analyze the symmetries of T2M, and characterize the equivariance and invariance properties of corresponding models. In light of these properties, we design an architecture based on hypernetworks that given a set of new class descriptions predicts the weights for an object recognition model which classifies images from those zero-shot classes. We demonstrate the benefits of our approach compared to zero-shot learning from text descriptions in image and point-cloud classification using various types of text descriptions: From single words to rich text descriptions.

Abstract (translated)

URL

https://arxiv.org/abs/2210.15182

PDF

https://arxiv.org/pdf/2210.15182.pdf