Xiangbo Gao-image
Xiangbo Gao portrait

Xiangbo Gao

Generative AI researcher working on video generation, controllable video editing, and world models.

Short CV (PDF)Contact
Xiangbo Gao snowboarding

Profile

About me

I am a Computer Science Ph.D. student at Texas A&M University. My research focuses on generative video models, controllable video editing, and world models, with an emphasis on precise and physically grounded visual generation.

  • LocationTexas A&M University, College Station, TX
  • ResearchVideo generation, controllable editing, and world models
  • InterestsSnowboarding, Skiing, Rock climbing, Badminton, Tennis
  • AffiliationDepartment of Computer Science & Engineering, Texas A&M University

Academic path

Education

Texas A&M University logo

Ph.D. in Computer Science

Texas A&M University

2025.1 - Present

University of Michigan logo

M.S. in Robotics

University of Michigan, Ann Arbor

2023.9 - 2024.12

University of California, Irvine logo

B.S. in Computer Science | B.S. in Mathematics

University of California, Irvine

2018.9 - 2023.3

Industry experience

Employment

Adobe logo

Research Intern

Adobe Research

2026.5 - Present
DiDi logo

Autonomous Driving Algorithms Research Intern

DiDi Global Inc., USA

2025.9 - 2026.5
COWA Robot logo

Perception Research Intern

Anhui Cowa ROBOT Co., Ltd, Shanghai, China

2023.4 - 2023.7

Research

Publications

Selected work in generative video, controllable editing, world models, and autonomous systems.

422 citationsh-index 11Full list on Google Scholar

Selected

7 works
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation overview
ECCV 2026

SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation

Jiongze Yu, Xiangbo Gao, Pooja Verlani, Akshay Gadde, Yilin Wang, Balu Adsumilli, Zhengzhong Tu

An interactive video super-resolution framework that propagates sparse user-selected high-resolution keyframes across a video.

LangCoop: Collaborative Driving with Language overview
CVPR 2025 Workshop ยท Best Paper Award

LangCoop: Collaborative Driving with Language

Xiangbo Gao, Yuheng Wu, Runsheng Wang, Chenxi Liu, Yang Zhou, Zhengzhong Tu

A language-driven vehicle collaboration framework that cuts communication bandwidth by 96% while preserving closed-loop driving performance.

AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving overview
TMLR 2026

AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving

Shuo Xing, Hongyuan Hua, Xiangbo Gao, Shenzhe Zhu, Renjie Li, Kexin Tian, Xiaopeng Li, Heng Huang, Tianbao Yang, Zhangyang Wang, Yang Zhou, Huaxiu Yao, Zhengzhong Tu

A comprehensive benchmark for evaluating the trustworthiness of large vision-language models in autonomous driving.

STAMP: Scalable Task- And Model-agnostic Collaborative Perception overview
ICLR 2025

STAMP: Scalable Task- And Model-agnostic Collaborative Perception

Xiangbo Gao, Runsheng Xu, Jiachen Li, Ziran Wang, Zhiwen Fan, Zhengzhong Tu

A task- and model-agnostic framework that lets heterogeneous agents share collaborative perception features efficiently and securely.

MambaST: A Plug-and-Play Cross-Spectral Spatial-Temporal Fuser for Efficient Pedestrian Detection overview
ITSC 2024

MambaST: A Plug-and-Play Cross-Spectral Spatial-Temporal Fuser for Efficient Pedestrian Detection

Xiangbo Gao, Asiegbu Miracle Kanu-Asiegbu, Xiaoxiao Du

A plug-and-play state-space fuser for efficient spatial-temporal pedestrian detection from RGB and thermal video.

Scale-free and Task-agnostic Attack: Generating Photo-realistic Adversarial Patterns with Patch Quilting Generator overview
ICASSP 2024

Scale-free and Task-agnostic Attack: Generating Photo-realistic Adversarial Patterns with Patch Quilting Generator

Xiangbo Gao, Cheng Luo, Qinliang Lin, Weicheng Xie, Minmin Liu, Linlin Shen, Keerthy Kusumam, Siyang Song

A scale-free generator for transferable, defense-resistant, and visually realistic adversarial patterns.

Sample Hardness Based Gradient Loss for Long-Tailed Cervical Cell Detection overview
MICCAI 2022

Sample Hardness Based Gradient Loss for Long-Tailed Cervical Cell Detection

Minmin Liu, Xuechen Li, Xiangbo Gao, Junliang Chen, Linlin Shen, Huisi Wu

Gradient-based hardness calibration improves long-tailed cervical cell detection by 7.8% mAP over standard classification loss.

Preprints

8 works
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation overview
arXiv 2026

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

Xiangbo Gao, Siyuan Yang, Ping He, Mingyang Wu, Yuheng Wu, Yushen Zuo, Jiongze Yu, Ryan Cui, Hongyuan Hua, Devin Ma, Xiao Jin, Yubo Yuan, Qing Yin, Jie Yang, Zhengzhong Tu

A live video model for prompt-switchable, hour-scale generation with real-time 4K output at 24 FPS.

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs overview
arXiv 2026

Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

Xiangbo Gao, Xiukun Huang, Boyu Lu, Junge Zhang, Mengjie Mao, Jiachen Li, Wei Xiong, Zhengzhong Tu

Rule-based planners provide faithful reasoning traces that keep a driving VLA's explanations structurally tied to its motion decisions.

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation overview
arXiv 2026

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

Yuheng Wu, Xiangbo Gao, Tianhao Chen, Xinghao Chen, Qing Yin, Zhengzhong Tu, Dongman Lee

A trust-region training objective improves long-horizon consistency while keeping interactive video generation responsive to new events.

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects overview
arXiv 2026

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Xiangbo Gao, Sicong Jiang, Bangya Liu, Xinghao Chen, Minglai Yang, Siyuan Yang, Mingyang Wu, Jiongze Yu, Qi Zheng, Haozhi Wang, Jiayi Zhang, Jared Yang, Jie Yang, Zihan Wang, Qing Yin, Zhengzhong Tu

A 5,049-example human-annotated dataset, reward model, and benchmark for instruction-guided video editing quality.

The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics overview
arXiv 2026

The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics

Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu

Visual Chronometer estimates physical frame rate from motion and exposes temporal-scale failures in modern video generators.

PISCO: Precise Video Instance Insertion with Sparse Control overview
arXiv 2026

PISCO: Precise Video Instance Insertion with Sparse Control

Xiangbo Gao, Renjie Li, Xinghao Chen, Yuheng Wu, Suofei Feng, Qing Yin, Zhengzhong Tu

A video diffusion model that inserts instances with coherent appearance, motion, and interaction from only a few controlled keyframes.

SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving overview
arXiv 2025

SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving

Xiangbo Gao, Tzu-Hsiang Lin, Ruojing Song, Yuheng Wu, Kuan-Ru Huang, Zicheng Jin, Fangzhou Lin, Shinan Liu, Zhengzhong Tu

An agentic defense pipeline protects language-based collaborative driving against communication failures and semantic attacks.

AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration overview
arXiv 2025

AirV2X: Unified Air-Ground Vehicle-to-Everything Collaboration

Xiangbo Gao, Yuheng Wu, Fengze Yang, Xuewen Luo, Keshu Wu, Xinghao Chen, Yuping Wang, Chenxi Liu, Yang Zhou, Zhengzhong Tu

A large-scale dataset and benchmark for vehicle-to-drone collaboration using flexible aerial perception.

Community

Professional Services

Conference and Journal Paper Reviewing

  • CV & ML:ICCV, CVPR, ICLR, NeurIPS, T-PAMI
  • Robotics:RA-L
  • Transportation:TRBAM