Research
I am interested in machine learning and computer vision. My current focus is on improving MLLMs' performance in real-world scenarios with limited supervision through test-time training and adaptation.
|
Publications
|
BeyondNL2Code: A Structured Survey of Multimodal Code Intelligence
Xuanle Zhao, Qiushi Sun, Jingyu Xiao, Xuexin Liu, Haoyue Yang,
Qiaosheng Chen, Xianzhen Luo, Jing Huang, Yufeng Zhong, Lei Chen,
Shuai Fu, Zhenlin Wei, Jinhe Bi, Lei Jiang, Haibo Qiu, Siqi Yang, Peng Shi,
Jian Hu (Corresponding),
Zhixiong Zeng (Corresponding)
TMLR, 2026
arXiv
A structured survey of multimodal code intelligence beyond NL2Code.
|
|
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Zixu Cheng,
Jian Hu (Corresponding),
Ziquan Liu,
Chenyang Si,
Wei Li,
Shaogang Gong
CVPRW, 2026
arXiv
/
website
/
code
To decompose video understanding into a Reverse Spatio-Temporal Reasoning (RSTR) task that simultaneously evaluates what objects are present, when events occur, and where they are located while capturing the underlying Chain-of-Thought (CoT) logic.
|
|
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
Jian Hu,
Zixu Cheng,
Shaogang Gong,
Isabel Guan,
Jianye Hao,
Jun Wang,
Kun Shao
NeurIPS, 2025
arXiv
An unsupervised RL-based cross-domain temporal grounding approach.
|
|
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
Jian Hu,
Zixu Cheng,
Shaogang Gong
IJCAI, 2025 (oral)
arXiv
To adaptively reduce the influence of irrelevant (negative) prior knowledge whilst increasing the use of the most plausible prior knowledge, selected by negative mining with higher contrast, in order to optimise instance-specific prompt generation.
|
|
Class-Aware Diversified Augmentation for Open-Set Single Domain Generalization
Jian Hu,
Shaogang Gong,
Weitong Cai,
Junchi Yan
IEEE Transactions on Multimedia, 2025
Paper
Explore single domain open-set generalization with class-aware augmentation.
|
|
Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable Segmentation
Jian Hu,
Jiayi Lin,
Junchi Yan,
Shaogang Gong
NeurIPS, 2024
arXiv
/
website
/
code
Using hallucinations as prior knowledge to help create specific prompts for segmenting tasks, reducing the need for manual prompts.
|
|
Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged Objects
Jian Hu*,
Jiayi Lin*,
Weitong Cai,
Shaogang Gong
AAAI, 2024
arXiv
/
website
/
code
Eliminate the need for manual prompts for SAM in various challenging segmentation tasks.
|
|
Uncertainty-based Heterogeneous Privileged Knowledge Distillation for Recommendation System
Ang Li*, Jian Hu*, Ke Ding, Xiaolu Zhang,
Jun Zhou,
Yong He
SIGIR, 2023
Paper
Proposing a novel algorithm to address heterogeneous knowledge distillation-based transfer learning in industrial recommendation systems.
|
|
Global-Aware Model-Free Self-distillation for Recommendation System
Ang Li*, Jian Hu*,
Lu Wei, Ke Ding, Xiaolu Zhang,
Jun Zhou,
Yong He
DASFAA, 2023
Paper
Introducing a novel algorithm called Global-aware Model-free Self-Distillation to address label noise in training data in the Alipay advertising system.
|
|
Learning Unbiased Transferability for Domain Adaptation by Uncertainty Modeling
Jian Hu*,
Haowen Zhong*,
Fei Yang,
Shaogang Gong,
Guile Wu,
Junchi Yan
ECCV, 2022
arXiv
/
code
Delving into the transferability estimation problem in domain adaptation and proposing a non-intrusive Unbiased Transferability Estimation Plug-in (UTEP) by modeling the uncertainty of a discriminator in adversarial-based DA methods to optimize unbiased transfer.
|
|
Attribute-Conditioned Face Swapping Network for Low-Resolution Images
Ang Li*, Jian Hu*, Chilin Fu, Xiaolu Zhang,
Jun Zhou
ICASSP, 2022
Paper
A novel Attribute-Conditioned Face Swapping Network (AFSNet) to preserve attributes and handle low resolution images.
|
|
Domain Adaptive YOLO for One-Stage Cross-Domain Detection
Shizhao Zhang, Hongya Tuo,
Zhongliang Jing,
Jian Hu
ACML, 2021
arXiv
Improving cross-domain performance for one-stage detectors: image level feature alignment is used to strictly match local features and loosely match global features.
|
|
Discriminative Partial Domain Adversarial Network
Jian Hu, Hongya Tuo, Chao Wang, Lingfeng Qiao,
Haowen Zhong,
Junchi Yan,
Zhongliang Jing, Henry Leung
ECCV, 2020
Paper
Addressing the partial domain adaptation problem with a discriminative partial domain adversarial network with theoretical analysis.
|
|
Unsupervised Satellite Image Classification based on Partial Transfer Learning
Jian Hu, Hongya Tuo, Chao Wang,
Haowen Zhong, Pan Han, Lingfeng Qiao,
Zhongliang Jing
Aerospace Systems, 2019
Paper
Focusing on how to achieve high accuracy on unsupervised satellite image classification.
|
|
Multi-Weight Partial Domain Adaptation
Jian Hu, Hongya Tuo, Chao Wang, Lingfeng Qiao,
Haowen Zhong,
Zhongliang Jing
BMVC, 2019 (spotlight)
Paper
Focusing on how to transfer knowledge from a massive labelled dataset to an unlabelled miniature one.
|
|
 |
Reviewer for CVPR, ICCV, ECCV, TPAMI, IJCV, ICML, ICLR, NeurIPS (top reviewer'24), AAAI, TMLR, ACMMM (outstanding reviewer'24), PKDD, AISTATS, ToMM
|
|
Student Demonstrator, ECS795P Deep Learning and Computer Vision 2022-24
Student Demonstrator, ECS607U Data Mining 2023-24
|
|