Open to research conversations and collaborations.

MULTIMODAL INTELLIGENCE · EFFICIENT AI · GUI AGENTS

Yuhao Wang王宇皓

Researcher in Multimodal Fusion & Edge Intelligence

Building efficient multimodal intelligence for perception, foundation models, and on-device agents.

I work on multimodal fusion and edge intelligence, with a focus on cross-modal perception, efficient multimodal large models, and autonomous decision-making for resource-constrained devices. My work connects algorithmic innovation with system optimization and terminal deployment.

First-year PhD student in Information and Communication Engineering DLUT · Information & Communication Engineering Dalian, China
中文研究简介

本人长期聚焦多模态融合与端侧智能体研究,面向复杂环境中的跨模态感知、理解与自主决策需求,持续攻关异构模态分布差异、信息协同和高效推理等关键问题。

25 research works CVPR · ECCV · AAAI · WACV · IEEE journals Verified publication dataset · Sep 2026
11 first / co-first works from fusion to edge intelligence Verified publication dataset · Sep 2026
735 Google Scholar citations across multimodal vision research Google Scholar · Sep 2026
593 GitHub stars 19 public repositories · forks included GitHub REST API · Sep 2026

01 / RESEARCH IDENTITY

One connected research program.

My work follows a continuous path from multimodal evidence to efficient representation and deployable agents.

Research field / interactive map

Multimodal Intelligence

Select a direction or a keyword to inspect its evidence.

3 areas 18 terms 25 works
Perception focus
Efficient MLLMs focus
Agent Systems focus
One connected program
01 Perceive 02 Interpret 03 Compress 04 Act
Heterogeneous evidence becomes efficient, deployable intelligence.
01 Perception RGB · NIR · TIR · Text
02 Efficient models Align · select · compress
03 Agents Reliable agent systems

Established RESEARCH THEME

Multimodal Perception

01

How can heterogeneous sensors describe the same identity across viewpoints, spectra, and environments?

I align RGB, near-infrared, thermal, text, and aerial-ground observations into identity-aware representations that remain dependable in complex scenes.

Methods
  • Cross-modal alignment
  • Modality-specific representation
  • Text-guided fusion
All related papers 22 works
2026 · Journal Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation Contributor · IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2026 · Conference Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification Co-first author · European Conference on Computer Vision (ECCV) 2026 · Conference STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification Co-corresponding author · AAAI Conference on Artificial Intelligence (AAAI) 2026 · Conference Signal: Selective Interaction and Global-local Alignment for Multi-Modal Object Re-Identification Co-first author · AAAI Conference on Artificial Intelligence (AAAI) 2026 · Journal SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification First author · IEEE Transactions on Image Processing (TIP) 2026 · Journal HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation Contributor · IEEE Transactions on Image Processing (TIP) 2026 · Conference CADTrack: Learning Contextual Aggregation with Deformable Alignment for Robust RGBT Tracking Contributor · AAAI Conference on Artificial Intelligence (AAAI) 2026 · Conference RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation Contributor · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 · Workshop SAS-VPReID: A Scale-Adaptive Framework with Shape Priors for Video-based Person Re-Identification at Extreme Far Distances Contributor · WACV Workshop on Video Re-Identification at Extreme Far Distances 2026 · Challenge Report VReID-XFD: Video-based Person Re-identification at Extreme Far Distance Challenge Results Contributor · WACV Workshop on Video Re-Identification at Extreme Far Distances 2025 · Journal Unity Is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification First author · IEEE Transactions on Intelligent Transportation Systems (TITS) 2025 · Conference Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation Contributor · IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2024 · Conference Magic Tokens: Select Diverse Tokens for Multi-Modal Object Re-Identification Contributor · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024 · Conference TOP-ReID: Multi-spectral Object Re-Identification with Token Permutation First author · AAAI Conference on Artificial Intelligence (AAAI) 2025 · Challenge Report AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results Contributor · IEEE International Joint Conference on Biometrics (IJCB) 2025 · Conference IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-Modal Object Re-Identification First author · IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025 · Conference DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification First author · AAAI Conference on Artificial Intelligence (AAAI) 2025 · Conference MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt First author · AAAI Conference on Artificial Intelligence (AAAI) 2025 · Conference CLIMB-ReID: A Hybrid CLIP-Mamba Framework for Person Re-Identification Contributor · AAAI Conference on Artificial Intelligence (AAAI) 2025 · Preprint LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification Contributor · arXiv preprint 2026 · Journal Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance Contributor · IEEE Transactions on Image Processing (TIP)
2025 · Manuscript VMambaX: Exploiting Vision Mamba for Abdominal X-ray Image Based NEC Classification Contributor · Public record pending
View all related works
  • RGB / NIR / TIR
  • Multimodal ReID
  • Cross-modal Retrieval
Read the full research agenda

02 / RESEARCH PHILOSOPHY

Less redundant compute.
More useful intelligence.

01

Perceive beyond pixels.

Multimodal systems should preserve the evidence unique to each sensor while discovering the semantics they share.

02

Spend compute where it matters.

Efficiency is an information-design problem: select useful tokens, route useful features, and reuse useful memory.

03

Design for the device.

A method is more valuable when latency, memory, and deployment constraints influence the research question from the beginning.

04

Connect perception to action.

The next step is not only to understand multimodal environments, but to help agents make reliable decisions within them.

03 / LATEST WORK

The newest paper, read closely.

A dedicated reading of the current work: the constraint, the method, and why it sits at the edge of the research program.

04 / SELECTED PUBLICATIONS

Methods that build on one another.

Search the full record

05 / NOW

Recent signals.

Recognition

Received the Bochuan Qu Scholarship, awarded annually to only ten DLUT students.

Recognition

Received the National Scholarship.

Presentation

Presented a poster at VALSE 2025.

Talk

Gave a Spotlight Talk at CCIG 2025.

06 / RESEARCH JOURNEY

How the direction formed.

Ph.D. in Information and Communication Engineering

School of Information and Communication Engineering · Dalian University of Technology

Research practice

OPPO Y-Lab

M.S. in Computer Science

School of Computer Science and Technology · Dalian University of Technology

Research Intern

LV LAB · National University of Singapore

Rethinking Object ReID

National Centre for Computer Animation · Bournemouth, UK

View the complete journey

07 / COMMUNITY & OPEN WORK

Research also means reviewing, sharing, and maintaining.

Beyond papers, I contribute through peer review, talks, released implementations, and curated research resources.

08 / WHERE I’M GOING

From efficient models to reliable agent systems.

1

Current

Efficient multimodal models

Reduce token, feature, and memory redundancy while preserving cross-modal understanding.

2

Next

Long-horizon GUI agents

Build agents that perceive interfaces, retain useful visual history, and act under real device budgets.

3

Future

Cloud–edge multimodal agent systems

Coordinate perception, reasoning, memory, and action across device and cloud for reliable autonomous decision-making.

09 / BEYOND RESEARCH

Curiosity needs a little room.

生如芥子,心藏须弥
A small seed can hold a universe.

Visit BearNoBugs

LET’S TALK

Interested in multimodal intelligence, efficient models, or on-device agents?

[email protected] WeChat · w924973292