A survey of visual intelligence development
WEI Yunchao
REN Zhongwei
FANG Yan
Abstract:Visual intelligence,as a core branch of AI,seeks to endow machines with human-like capa-bilities for visual understanding and interaction.Since the breakthrough of deep learning in computer vi-sion in 2012,the field has undergone four progressive stages of evolution.The first stage,represented by AlexNet,VGGNet and ResNet,leveraged large annotated datasets such as ImageNet to achieve remarkable success in closed-domain tasks(e.g.,image classification and object detection),but its de-pendence on labeled data highlighted inherent limitations.The second stage witnessed the rise of self-supervised learning,with models such as MoCo,DION and MAE learning powerful visual representa-tions from massive unlabeled data through contrastive,distillation,and masked reconstruction meth-ods.The third stage marked a shift toward multimodal intelligence,where models like CLIP and GPT-4V integrated vision and language,enabling open-vocabulary understanding and advancing to-ward fine-grained,intent-driven reasoning.The current frontier is world models,exemplified by Sora,which aim not only to perceive and describe but also to simulate and predict the physical world,paving the way for embodied intelligence capable of interacting with reality.The current frontier is world models,exemplified by Sora,which aim not only to perceive and describe but also to simulate and predict the physical world,paving the way for embodied intelligence capable of interacting with re-ality.This fundamental transformation from discriminative understanding to generative simulation of the world marks a new consensus:generative modeling is the new deep learning.This survey follows this developmental trajectory,analyzing the core ideas,representative models,and methodological paradigms at each stage,while highlighting ongoing challenges in robustness,reasoning,and general-ization.This survey follows this developmental trajectory,analyzing the core ideas,representative models,and methodological paradigms at each stage,while highlighting ongoing challenges in robust-ness,reasoning,and generalization.
Keywords:visual intelligencedeep learningself-supervised learningmultimodal aiworld model
Publication Date:2025-10-30
Online Publishing Date:2025-11-06(First online date of this platform, not the publication date of the document)
Pages:16( 66-81 )
Journal of Beijing Jiaotong University

Journal of Beijing Jiaotong University

ISTICPKUCSCD
ISSN:1673-0291
Year, Vol.(Issue):2025,49(5)