MULTIMODAL INTELLIGENCE · EFFICIENT AI · GUI AGENTS
Yuhao Wang王宇皓
Researcher in Multimodal Fusion & Edge Intelligence
Building efficient multimodal intelligence for perception, foundation models, and on-device agents.
I work on multimodal fusion and edge intelligence, with a focus on cross-modal perception, efficient multimodal large models, and autonomous decision-making for resource-constrained devices. My work connects algorithmic innovation with system optimization and terminal deployment.
中文研究简介
本人长期聚焦多模态融合与端侧智能体研究,面向复杂环境中的跨模态感知、理解与自主决策需求,持续攻关异构模态分布差异、信息协同和高效推理等关键问题。
01 / RESEARCH IDENTITY
One connected research program.
My work follows a continuous path from multimodal evidence to efficient representation and deployable agents.
Multimodal Intelligence
Select a direction or a keyword to inspect its evidence.
Established RESEARCH THEME
Multimodal Perception
How can heterogeneous sensors describe the same identity across viewpoints, spectra, and environments?
I align RGB, near-infrared, thermal, text, and aerial-ground observations into identity-aware representations that remain dependable in complex scenes.
- Cross-modal alignment
- Modality-specific representation
- Text-guided fusion
Growing RESEARCH THEME
Efficient Multimodal Large Models
How can multimodal foundation models preserve useful evidence while spending less compute?
I study visual token selection, semantic priors, state-space modeling, and efficient adaptation for multimodal models.
- Visual token pruning
- Vision-language adaptation
- State-space aggregation
Emerging RESEARCH THEME
Efficient Agent Systems
How can multimodal agents perceive, remember, and act under strict latency, memory, and device constraints?
My emerging agent direction connects visual evidence management with reliable, deployment-aware GUI agent systems.
- Evidence-aware inference
- Visual memory efficiency
- Device-aware deployment
02 / RESEARCH PHILOSOPHY
Less redundant compute.
More useful intelligence.
Perceive beyond pixels.
Multimodal systems should preserve the evidence unique to each sensor while discovering the semantics they share.
Spend compute where it matters.
Efficiency is an information-design problem: select useful tokens, route useful features, and reuse useful memory.
Design for the device.
A method is more valuable when latency, memory, and deployment constraints influence the research question from the beginning.
Connect perception to action.
The next step is not only to understand multimodal environments, but to help agents make reliable decisions within them.
03 / LATEST WORK
The newest paper, read closely.
A dedicated reading of the current work: the constraint, the method, and why it sits at the edge of the research program.
TRACE
Trajectory-robust Admission and Coverage-aware Evidence ordering
TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
arXiv preprint
Scroll figures
How can training-free visual-token pruning stay useful across a GUI trajectory when discarded tokens cannot be recovered?
GUI agents accumulate high-resolution screenshots as the trajectory unfolds, raising inference latency and memory use. Training-free pruning can cut that cost, but cache reuse makes the choice irreversible: once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. The admission decision must therefore remain useful for unknown future targets, while still covering operable regions under a tight budget.
TRACE is a training-free framework for trajectory-robust admission and coverage-aware evidence ordering. It ranks visual evidence by potential future utility and diversity, then repairs missing spatial coverage without breaking that order, so retained tokens can shrink monotonically and stay reusable.
The current research edge is not only compressing a single screenshot, but making visual evidence survive an unfolding GUI trajectory under a tight budget.
Verified across six GUI benchmarks and diverse models. Source code will be released.
- 01 Rank A query-independent layout-derived interaction prior is combined with instruction relevance and feature novelty, ranking visual evidence by both potential future utility and diversity.
- 02 Cover Part of the budget is reserved for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering.
- 03 Nest The result is a nested token order: retained visual evidence shrinks monotonically across budgets while remaining reusable throughout the trajectory.
- 04 Contract Monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning.
- Layout-derived interaction prior
- Instruction relevance
- Feature novelty
- Native-token coverage repair
- Nested token order
- Monotone KV contraction
04 / SELECTED PUBLICATIONS
Methods that build on one another.
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-Modal Object Re-Identification
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
A text-guided multimodal framework that uses cooperative deformable aggregation to bridge visual and language cues for object re-identification.
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
AAAI Conference on Artificial Intelligence (AAAI)
A Mamba-based aggregation and prompt learning framework for modeling long-range multimodal dependencies in object re-identification.
05 / NOW
Recent signals.
Received the Bochuan Qu Scholarship, awarded annually to only ten DLUT students.
Received the National Scholarship.
Presented a poster at VALSE 2025.
Gave a Spotlight Talk at CCIG 2025.
06 / RESEARCH JOURNEY
How the direction formed.
Ph.D. in Information and Communication Engineering
School of Information and Communication Engineering · Dalian University of Technology
Research practice
OPPO Y-Lab
M.S. in Computer Science
School of Computer Science and Technology · Dalian University of Technology
Research Intern
LV LAB · National University of Singapore
Rethinking Object ReID
National Centre for Computer Animation · Bournemouth, UK
07 / COMMUNITY & OPEN WORK
Research also means reviewing, sharing, and maintaining.
Beyond papers, I contribute through peer review, talks, released implementations, and curated research resources.
08 / WHERE I’M GOING
From efficient models to reliable agent systems.
Current
Efficient multimodal models
Reduce token, feature, and memory redundancy while preserving cross-modal understanding.
Next
Long-horizon GUI agents
Build agents that perceive interfaces, retain useful visual history, and act under real device budgets.
Future
Cloud–edge multimodal agent systems
Coordinate perception, reasoning, memory, and action across device and cloud for reliable autonomous decision-making.
09 / BEYOND RESEARCH
Curiosity needs a little room.
生如芥子,心藏须弥
A small seed can hold a universe.
LET’S TALK