Research at SJTU-MARL · Shanghai

Agents that remember,
adapt, and improve
at runtime.

I am Shengtao Zhang. I study self-evolving agents at the intersection of reinforcement learning and memory—systems that turn experience into better decisions without retraining the underlying model.

Shengtao Zhang
SJTU-MARLAgent · RL · Memory
Current focus

How can an agent learn from every interaction while keeping its reasoning stable?

Research thesis

From storage to experience.

My work connects three layers of adaptive intelligence: acting in an environment, learning from feedback, and preserving useful experience.

01Act

Agent systems

Agents that plan, use tools, and improve through interaction instead of staying fixed after deployment.

02Adapt

Runtime reinforcement learning

Feedback-driven learning loops that turn execution outcomes into better choices at inference time.

03Remember

Memory

Episodic and value-aware memory that retrieves experience for utility, provenance, and continual refinement.

Selected publications

Research index.

Work on self-evolving agents, reinforcement learning, and memory, alongside graph learning, language, and multimodal reasoning.

Scientific Data2026Cited by 3MultimodalDataReasoning

A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories

Haojun Yu, Youcheng Li, Zihan Niu, Nan Zhang, Xuantong Gong, Huan Li, Zhiying Zou, Haifeng Qi, Zhenxiao Cao, Zijie Lan, Xingjian Yuan, Jiating He, Haokai Zhang, Shengtao Zhang, Zicheng Wang, Dong Wang, Ziwei Zhao, Congying Chen, Yong Wang, Wangyan Qin, Qingli Zhu, Liwei Wang

A reasoning-rich breast ultrasound resource that connects images, pathology coverage, and clinically grounded chains of thought.

Recent notes

News.

About

Building learning loops for agents that operate in the real world.

I work on agent learning and memory with SJTU-MARL. I am especially interested in non-parametric adaptation: how agents can use outcomes, provenance, and episodic experience to improve continuously while the base model stays stable.

Previously, I studied Artificial Intelligence at Xi'an Jiaotong University and worked on dynamic graph learning, language understanding, and reliable decision systems.

Z. S. T.
Zeal stirs tides.「热望所至,潮涌不息。」