Undergraduate Researcher · Xidian University
Trustworthy LLM Reasoning · Reliable AI Agents · AI for Research · Recursive Self-Improvement
I am an undergraduate student at Xidian University, working on trustworthy reasoning in large language models and evaluation of LLM-based agents.
My research focuses on understanding when and why LLMs and agents fail, how such failures can be diagnosed from internal representations or execution trajectories, and how these signals can eventually support more reliable and self-improving AI systems.
I am particularly interested in AI for Research, including scientific agents that can understand literature, formulate hypotheses, conduct experiments, interpret evidence, and iteratively refine their research strategies.
I am open to research collaboration and RA opportunities in NLP, trustworthy AI, LLM agents, and AI for Research.
-
Trustworthy LLM Reasoning
- Chain-of-Thought faithfulness
- Internal representation dynamics
- Reasoning reliability and failure diagnosis
-
Reliable & Self-Improving Agents
- Agent evaluation and capability boundaries
- Execution trajectory analysis
- Failure attribution and continual adaptation
- Recursive Self-Improvement (RSI)
-
AI for Research
- Scientific agents and AutoResearch
- Research judgment and research taste
- Hypothesis–experiment co-evolution
- Scientific information retrieval
-
SPD-Faith Bench (ACL Findings 2026) Diagnosing and improving Chain-of-Thought faithfulness in multimodal large language models; introduces the training-free SAGE framework. W. Lv, Y. Feng, X. Xia, et al.
-
GeoFaith (arXiv) Modeling faithful reasoning through latent geometry and entropy dynamics, with process-aware faithfulness detection and reinforcement learning. W. Lv, W. Zhao, J. Wang, X. Xia, et al.
-
Act As a Real Researcher (AARRI-Bench) (arXiv) Evaluating frontier LLMs and agentic harnesses on realistic research-intern tasks, with an emphasis on research judgment beyond task execution. J. Wang, W. Lv, et al.*
-
DPDiff-AD (IJCV) Dual prototype-conditioned diffusion for multi-class unsupervised anomaly detection. Y. Feng, Y. Li, W. Lv, et al.
Research Intern · University of Science and Technology of China Advisor: Prof. Xiaobo Xia Sep. 2025 – May 2026 · Hefei, China
Research Intern · Xidian University Advisor: Prof. Bo Chen Dec. 2024 – Jul. 2025 · Xi’an, China
I am currently exploring how AI research agents can move beyond a static idea-to-experiment pipeline toward an evidence-driven scientific reasoning loop.
A central limitation of existing AI Scientist systems is that hypothesis generation and experiment execution are often treated as separate stages: an idea is proposed first, followed by experiments that mainly serve to implement or validate it. In real research, however, hypotheses and experiments continuously shape each other.
I am exploring a hypothesis–experiment co-evolution framework:
Hypothesis → Experiment Design → Execution → Evidence → Hypothesis Revision → Next Experiment
The goal is to enable a research agent to:
- revise hypotheses according to experimental evidence rather than rigidly executing an initial idea;
- design the next experiment based on uncertainty, previous failures, and expected information gain;
- distinguish genuine scientific progress from superficial metric improvement;
- accumulate useful research experience across iterative experimentation;
- gradually develop stronger scientific decision-making and research taste.
More broadly, I am interested in whether such scientific interaction trajectories can eventually support self-improving research agents—agents that not only conduct research, but also learn from their own failures and experimental evidence to improve how they conduct future research.
Python · FastAPI · SQLite · Multi-Agent LLM
An agentic literature-research system that turns large-scale paper collections into structured research trajectories and candidate research directions.
- Built end-to-end pipelines for ingestion → evolution modeling → gap discovery → research-story composition
- Constructed methodology evolution graphs with relations such as extends, improves, and replaces
- Developed domain-scoped retrieval and multi-agent extraction for large conference corpora
- Integrated research gap discovery, idea generation, multi-agent review, and iterative story refinement
- Built an interactive web interface for browsing papers, timelines, and research evolution graphs
- Finalist (Top 1%) · International Contest in Mathematical Modeling · Feb. 2026
- Excellent Award · HarmonyOS Agent Competition, Huawei × Nanjing University · May 2025
- First Prize · National College Student Mathematics Competition · Dec. 2024
Languages: Python, C/C++, Java, SQL, JavaScript / TypeScript
AI / ML: PyTorch, scikit-learn, Hugging Face Transformers, LLM APIs, RAG, embeddings, multi-agent workflows
Engineering: FastAPI, Flask, React, Git, Docker, Linux, SQLite, PostgreSQL, pandas, NumPy, REST APIs
Open to research collaboration and RA opportunities.