I am a Ph.D. student at the New Laboratory of Pattern Recognition(NLPR), Institute of Automation, Chinese Academy of Sciences, advised by Prof. Yan Huang. I am currently interning at JD Explore Academy, and have previously interned at Beijing Academy of Artificial Intelligence (BAAI), Kuaishou Technology, and Tsinghua AIR.
My research is broadly about what agents can do β how far we can push their ability to search, reason, create, and even conduct research autonomously. My earlier work focused on search agents for video understanding, driven by a central question: what can models actually do when they must actively search and compare external evidence? I built agents that retrieve and reason over open-web evidence, and designed benchmarks that reveal where frontier multimodal models fall short when required to search, compare, and ground claims across multiple sources. I am now extending this to agentic video generation and autoresearch β moving from agents that understand the world to agents that actively create and explore it.
I believe impactful research comes from close collaboration across communities. If youβre working on agentic AI, video generation, or autoresearch, or if you have internship or exchange opportunities, Iβd be glad to connect.
π₯ News
- 2026.08: Β ππ Two papers on Search Agent for Video Understanding (Misinformation Detection & Shot Retrieval) were accepted by EMNLP 2026 Findings!
- 2026.01: Β ππ One paper on Browser Agent were accepted by TMLR 2025!
- 2025.11: Β ππ Two papers on Multi-View Clustering and Deepfake were accepted by AAAI 2026!
- 2025.07: Β ππ One technical report on Kwai Keye-VL was released!
- 2025.05: Β ππ One paper on DPO (Direct Preference Optimization) was accepted by ICML 2025!
- 2025.02: Β ππ One paper on GUI Agent was accepted by CVPR 2025!
- 2024.06: Β ππ One paper on Knowledge Editing Benchmark was accepted by NeurIPS 2024 Datasets and Benchmarks Track!
- 2023.08: Β ππ One paper on Mobile Agent was accepted by Mobicom 2024 Summer Round!
π Publications

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks
Β Project
Tao Yu, Yifei Qu, Zhiqing Cui, Pengfei Zhou, Zhongtian Luo, Yujia Yang, Shenghua Chai, Haopeng Jin, Zhenghao Zhang, Xinming Wang, Hongzhu Yi, Wangbo Zhao, Zhenglin Wan, Yan Huang, Yeshani, Jinwen Luo, Yang You

When Seeing Is Not BelievingβA Benchmark for Search-Grounded Video Misinformation Detection
Β Project
Tao Yu, Yujia Yang, Shenghua Chai, Zhang Jinshuai, Haopeng Jin, Hao Wang, Minghui Zhang, Zhongtian Luo, Yuchen Long, Xinlong Chen, Jiabing Yang, Zhaolu Kang, Yuxuan Zhou, Zhengyu Man, Xinming Wang, Hongzhu Yi, Zheqi He, Xi Yang, Yan Huang, Liang Wang

Β Project
Tao Yu, Hao Wang, Changyu Li, Shenghua Chai, Minghui Zhang, Zhongtian Luo, Yuxuan Zhou, Haopeng Jin, Zhaolu Kang, Jiabing Yang, YiFan Zhang, Xinming Wang, Hongzhu Yi, Zheqi He, Jing-Shu Zheng, Xi Yang, Yan Huang, Liang Wang

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
Β Project
Tao Yu, Yiming Ding, Shenghua Chai, Minghui Zhang, Zhongtian Luo, Xinming Wang, Xinlong Chen, Zhaolu Kang, Junhao Gong, Yuxuan Zhou, Haopeng Jin, Zhiqing Cui, Jiabing Yang, YiFan Zhang, Hongzhu Yi, Zheqi He, Xi Yang, Yan Huang, Liang Wang

Β Project
Tao Yu, Yujia Yang, Haopeng Jin, Junhao Gong, Xinlong Chen, Yuxuan Zhou, Shanbin Zhang, Jiabing Yang, Xinming Wang, YiFan Zhang, Hongzhu Yi, Ping Nie, Kai Zou, Zhang Zhang, Yan Huang, Liang Wang, Yeshani, Ruiwen Tao, Jin Ma, Haijin Liang, Jinwen Luo

PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG
Β Project
Tao Yu, Minghui Zhang, Zhiqing Cui, Hao Wang, Zhongtian Luo, Shenghua Chai, Junhao Gong, Yuzhao Peng, Yuxuan Zhou, Yujia Yang, Zhenghao Zhang, Haopeng Jin, Xinming Wang, Yufei Xiong, Jiabing Yang, Jiahao Yuan, Hanqing Wang, Hongzhu Yi, YiFan Zhang, Yan Huang, Liang Wang

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
Tao Yu, Haopeng Jin, Hao Wang, Shenghua Chai, Yujia Yang, Junhao Gong, Jiaming Guo, Minghui Zhang, Xinlong Chen, Zhenghao Zhang, Yuxuan Zhou, Yufei Xiong, Shanbin Zhang, Jiabing Yang, Hongzhu Yi, Xinming Wang, Cheng Zhong, Xiao Ma, Zhang Zhang, Yan Huang, Liang Wang

BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
Β Project
Tao Yu, Zhengbo Zhang, Zhiheng Lyu, Junhao Gong, Hongzhu Yi, Xinming Wang, Yuxuan Zhou, Jiabing Yang, Ping Nie, Yan Huang, Wenhu Chen

Aligning Multimodal LLM with Human Preference: A Survey
Tao Yu, Yi-Fan Zhang, Chaoyou Fu, Junkang Wu, Jinda Lu, Kun Wang, Xingyu Lu, Yunhang Shen, Guibin Zhang, Dingjie Song, Yibo Yan, Tianlong Xu, Qingsong Wen, Zhang Zhang, Yan Huang, Liang Wang, Tieniu Tan
Others
-
EMNLP 2026 Main, Why does Weak-OOD Help? A Further Step Towards Understanding Jailbreaking VLMs
Yuxuan Zhou, Yuzhao Peng, Yang Bai, Kuofeng Gao, Yihao Zhang, Yechao Zhang, Xun Chen, Tao Yu, Tao Dai, Shu-Tao Xia -
ACL 2026 Main, Scaling Law for Multimodal Large Language Model Supervised Fine-Tuning
Yifan Zhang, Tao Yu, Feng Li, Chaoyou Fu, Yibo Hu, Kun Wang, Qingsong Wen, Zhang Zhang, Liang Wang, Rong Jin -
AAAI 2026, Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
Yuxuan Zhou, Tao Yu, Wen Huang, Yuheng Zhang, Tao Dai, Shu-Tao Xia -
AAAI 2026, Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss
Zhenghao Zhang, Jun Xie, Xingchen Chen, Tao Yu, Hongzhu Yi, Kaixin Xu, Yuanxiang Wang, Tianyu Zong, Xinming Wang, Jiahuan Chen, Guoqing Chao, Feng Chen, Zhepeng Wang, Jungang Xu
Technical Report, Kwai Keye-VL Technical Report
Kwai Keye Team
ICML 2025, MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Yi-Fan Zhang, Tao Yu, Haochen Tian, Chaoyou Fu, Peiyan Li, Jianshu Zeng, Wulin Xie, Yang Shi, Huanyu Zhang, Junkang Wu, Xue Wang, Yibo Hu, Bin Wen, Fan Yang, Zhang Zhang, Tingting Gao, Di Zhang, Liang Wang, Rong Jin, Tieniu Tan
CVPR 2025, GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration
Yuchen Sun, Shanhui Zhao, Tao Yu, Hao Wen, Samith Va, Mengwei Xu, Yuanchun Li, Chongyang Zhang
NeurIPS 2024, VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
Han Huang, Haitian Zhong, Tao Yu, Qiang Liu, Shu Wu, Liang Wang, Tieniu Tan
Mobicom 2024, Autodroid: Llm-powered task automation in android
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, Yunxin Liu
π Honors and Awards
- 2024.12, National Scholarship.(0.4%)
- 2023.12, National Encouragement Scholarship.(0.4%)
- 2022.12, National Scholarship.(0.4%)
- 2022.06, First Grade Scholarship.(0.4%)
π Educations
- 2025.09 - Current, Ph.D. Student in Pattern Recognition and Intelligent Systems (Institute of Automation, Chinese Academy of Sciences)
- 2021.09 - 2025.06, Bachelor in Computer Science and Technology (School of Computer Science and Technology, Harbin Institute of Technology), GPA: 93.79/100 (Ranking: 2/135)
π€ Collaborators
- 2022 Entry, Jinshuai Zhang (HIT->THU), Hao Wang (HIT->USTC)
- 2023 Entry, Yiming Ding (HIT->SJTU), Shenghua Chai (HIT->ZJU), Minghui Zhang (HIT->CASIA NLPR), Hao Wang (HIT->ICT VIPL), Zhongtian Luo (HIT->Tencent&HIT), Yifei Qu (HIT->AILab&HIT)
π» Internships
- 2026.08 - Present, Big Data Engineering Department, JD Explore Academy (JD.com), China.
- 2026.03 - 2026.08, Intelligent Evaluation Group, Beijing Academy of Artificial Intelligence (BAAI), China.
- 2025.01 - 2025.07, Multimodal Understanding and Application Group, Kuaishou Technology, China.
- 2024.03 - 2024.06, New Laboratory of Pattern Recognition(NLPR), Institute of Automation, China.
- 2023.05 - 2024.07, Institute for AI Industry Research(AIR), Tsinghua University, China.