Research at SJTU-MARL · Shanghai

Agents that remember,
adapt, and improve
from experience.

I am Shengtao Zhang. I study self-evolving agents at the intersection of reinforcement learning, memory, and efficient post-training. My research explores how agents learn from interaction and develop reusable capabilities through distillation and training.

Shengtao Zhang
SJTU-MARLAgent · RL · Memory
Current focus

How can agents turn experience into reliable, transferable skills?

Research thesis

From experience to capability.

My work connects three layers of adaptive intelligence: acting in an environment, learning from feedback, and preserving useful experience.

01Act

Agent systems

Agents that plan, use tools, and improve through interaction instead of staying fixed after deployment.

02Adapt

RL & post-training

Reinforcement learning, distillation, and efficient post-training for agents that learn from experience.

03Remember

Memory

Episodic and value-aware memory that retrieves experience for utility, provenance, and continual refinement.

Selected publications

Research index.

Work on self-evolving agents, reinforcement learning, and memory, alongside graph learning, language, and multimodal reasoning.

Google Scholar146 citationsh-index 3i10-index 1Checked

Scientific Data2026Cited by 5MultimodalDataReasoning

A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories

Haojun Yu, Youcheng Li, Zihan Niu, Nan Zhang, Xuantong Gong, Huan Li, Zhiying Zou, Haifeng Qi, Zhenxiao Cao, Zijie Lan, Xingjian Yuan, Jiating He, Haokai Zhang, Shengtao Zhang, Zicheng Wang, Dong Wang, Ziwei Zhao, Congying Chen, Yong Wang, Wangyan Qin, Qingli Zhu, Liwei Wang

A reasoning-rich breast ultrasound resource that connects images, pathology coverage, and clinically grounded chains of thought.

Recent notes

News.

About

Building learning loops for agents that operate in the real world.

I work on agent learning and memory with SJTU-MARL. My interests span reinforcement learning, memory, knowledge distillation, and efficient post-training. I study how agents acquire, retain, and generalize useful skills from experience, connecting runtime adaptation with lasting improvements in model capabilities.

Previously, I studied Artificial Intelligence at Xi'an Jiaotong University and worked on dynamic graph learning, language understanding, and reliable decision systems.

Z. S. T.
Zeal stirs tides.「热望所至,潮涌不息。」