Agent systems
Agents that plan, use tools, and improve through interaction instead of staying fixed after deployment.
Research at SJTU-MARL · Shanghai
I am Shengtao Zhang. I study self-evolving agents at the intersection of reinforcement learning, memory, and efficient post-training. My research explores how agents learn from interaction and develop reusable capabilities through distillation and training.

How can agents turn experience into reliable, transferable skills?
Research thesis
My work connects three layers of adaptive intelligence: acting in an environment, learning from feedback, and preserving useful experience.
Agents that plan, use tools, and improve through interaction instead of staying fixed after deployment.
Reinforcement learning, distillation, and efficient post-training for agents that learn from experience.
Episodic and value-aware memory that retrieves experience for utility, provenance, and continual refinement.
Featured work
Temporal memory admission helps returning agents retain valid information and reject stale state as multi-agent collaborations evolve.

When teammates change shared plans, a returning agent must preserve valid preferences while updating outdated bookings and unresolved tasks.
Experience-grounded procedural memory lets frozen vision-language agents improve spatial reasoning without parameter updates or expert tools at deployment.

SMA turns verifier-scored spatial experience into transferable memory cards, then retrieves reliable procedures through semantic filtering and TRS-aware ranking while keeping the VLM frozen.
A non-parametric route to runtime self-improvement: keep the reasoner stable, let episodic memory learn utility from feedback.

MemRL first recalls similar episodes, then selects them by learned Q-values. The frozen LLM acts with the retrieved context, while environmental rewards update each memory item's utility.
Credit assignment for agent memory, propagated through the provenance chains that make later memories possible.

Each interaction retrieves memories, creates a new episode, and records its provenance. TD feedback then travels backward through the DAG to update the memories that enabled success.
Value-guided experience reuse turns sparse NPU feedback into a continual drafting-and-refinement loop.

Cold-start drafting retrieves transferable experience for an initial kernel. Verification rewards update the shared memory, which later reuses successful traces for continual latency refinement.
Selected publications
Work on self-evolving agents, reinforcement learning, and memory, alongside graph learning, language, and multimodal reasoning.
Google Scholar146 citationsh-index 3i10-index 1Checked
Temporal memory admission helps returning agents retain valid information and reject stale state as multi-agent collaborations evolve.
Experience-grounded procedural memory lets frozen vision-language agents improve spatial reasoning without parameter updates or expert tools at deployment.
A non-parametric route to runtime self-improvement: keep the reasoner stable, let episodic memory learn utility from feedback.
Credit assignment for agent memory, propagated through the provenance chains that make later memories possible.
Value-guided experience reuse turns sparse NPU feedback into a continual drafting-and-refinement loop.
A reasoning-rich breast ultrasound resource that connects images, pathology coverage, and clinically grounded chains of thought.
Temporal curriculum learning makes final-timestamp labels useful across the full evolution of a dynamic graph.
Genre structure and user history expose spoiler patterns that text-only classifiers tend to miss.
Recent notes
About
I work on agent learning and memory with SJTU-MARL. My interests span reinforcement learning, memory, knowledge distillation, and efficient post-training. I study how agents acquire, retain, and generalize useful skills from experience, connecting runtime adaptation with lasting improvements in model capabilities.
Previously, I studied Artificial Intelligence at Xi'an Jiaotong University and worked on dynamic graph learning, language understanding, and reliable decision systems.