Search papers, labs, and topics across Lattice.
This paper explores the concept of visual general intelligence (VGI) by analyzing how intelligence derived from visual experiences could lead to advancements in artificial general intelligence (AGI). It emphasizes the potential of visual modalities鈥攕uch as images, videos, and geometry鈥攖o exhibit capabilities akin to those seen in language models like GPT, which have thrived through large-scale data and autoregressive learning. The authors aim to outline principles for computer vision in the AGI context, addressing input modalities, benchmarks, and the interplay between visual and linguistic intelligence.
Visual intelligence could be the key to unlocking AGI, challenging the dominance of language-based models in the quest for general intelligence.
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.