Repository navigation
Expand file tree
/
Copy pathpublication.html
More file actions
253 lines (246 loc) · 19 KB
/
Copy pathpublication.html
File metadata and controls
253 lines (246 loc) · 19 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1//EN"
"http://www.w3.org/TR/xhtml11/DTD/xhtml11.dtd">
<html xmlns="http://www.w3.org/1999/xhtml" xml:lang="en">
<head>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<meta http-equiv="Content-Type" content="text/html;charset=utf-8" />
<link rel="icon" type="image/svg+xml" href="./images/SJTU-logo.svg">
<link rel="stylesheet" href="style.css" type="text/css" />
<title>Publications | SAIL Lab @ SJTU</title>
<!-- MathJax -->
<script src='https://cdnjs.cloudflare.com/ajax/libs/mathjax/2.7.5/latest.js?config=TeX-MML-AM_CHTML' async>
</script>
<script type="text/x-mathjax-config">
MathJax.Hub.Config({
TeX: { equationNumbers: { autoNumber: "AMS" } }
});
</script>
<!-- End MathJax -->
</head>
<body>
<header class="site-header">
<div class="brand-bar">
<div class="brand-inner">
<a class="brand-seal" href="https://en.sjtu.edu.cn/" target="_blank"><img src="./images/SJTU-logo-white.svg" alt="Shanghai Jiao Tong University" /></a>
<div class="brand-title">
<a href="index.html">SAIL Lab</a>
<span class="brand-sub">Safe AI and Robot Learning Lab · Shanghai Jiao Tong University</span>
</div>
<div class="brand-links">
<a href="https://github.com/SAIL-Research-Lab" target="_blank" title="SAIL Lab GitHub"><svg viewBox="0 0 16 16" aria-hidden="true"><path fill-rule="evenodd" d="M8 0C3.58 0 0 3.58 0 8c0 3.54 2.29 6.53 5.47 7.59.4.07.55-.17.55-.38 0-.19-.01-.82-.01-1.49-2.01.37-2.53-.49-2.69-.94-.09-.23-.48-.94-.82-1.13-.28-.15-.68-.52-.01-.53.63-.01 1.08.58 1.23.82.72 1.21 1.87.87 2.33.66.07-.52.28-.87.51-1.07-1.78-.2-3.64-.89-3.64-3.95 0-.87.31-1.59.82-2.15-.08-.2-.36-1.02.08-2.12 0 0 .67-.21 2.2.82.64-.18 1.32-.27 2-.27.68 0 1.36.09 2 .27 1.53-1.04 2.2-.82 2.2-.82.44 1.1.16 1.92.08 2.12.51.56.82 1.27.82 2.15 0 3.07-1.87 3.75-3.65 3.95.29.25.54.73.54 1.48 0 1.07-.01 1.93-.01 2.2 0 .21.15.46.55.38A8.01 8.01 0 0 0 16 8c0-4.42-3.58-8-8-8z"/></svg></a>
<a href="https://www.youtube.com/@ai-agent-research" target="_blank" title="SAIL Lab YouTube"><svg viewBox="0 0 24 24" aria-hidden="true"><path d="M23.5 6.19a3.02 3.02 0 0 0-2.12-2.14C19.5 3.55 12 3.55 12 3.55s-7.5 0-9.38.5A3.02 3.02 0 0 0 .5 6.19C0 8.07 0 12 0 12s0 3.93.5 5.81a3.02 3.02 0 0 0 2.12 2.14c1.88.5 9.38.5 9.38.5s7.5 0 9.38-.5a3.02 3.02 0 0 0 2.12-2.14C24 15.93 24 12 24 12s0-3.93-.5-5.81zM9.55 15.57V8.43L15.82 12l-6.27 3.57z"/></svg></a>
<a href="mailto:shangding.gu@sjtu.edu.cn">shangding.gu@sjtu.edu.cn</a>
</div>
<button id="nav-toggle" aria-label="Menu" onclick="document.body.classList.toggle('nav-open')">☰</button>
</div>
</div>
<nav class="site-nav">
<div class="nav-inner">
<ul class="nav-list">
<li class="nav-item"><a href="index.html">Home</a></li>
<li class="nav-item"><a href="group.html">People</a></li>
<li class="nav-item current"><a href="publication.html">Publications</a></li>
<li class="nav-item"><a href="projects.html">Research & Projects</a></li>
<li class="nav-item"><a href="softwares.html">Products & Software</a></li>
<li class="nav-item"><a href="tutorial_talk.html">Tutorials & Talks</a></li>
<li class="nav-item"><a href="academic_service.html">Academic Service</a></li>
<li class="nav-item"><a href="resources.html">Resources</a></li>
<li class="nav-item"><a href="join.html">Join Us</a></li>
</ul>
</div>
</nav>
</header>
<div class="page">
<h1>Publications</h1>
<div class="muted" style="margin-top:-6px;">Selected recent publications and preprints by Prof. Shangding Gu and lab members.</div>
<h2>Selected Recent Publications</h2>
<div class="pub-row">
<div class="paper-tag">ICLR 2025</div>
<div class="pub-main">
<p><strong>Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Laixi Shi, Muning Wen, Ming Jin, Eric Mazumdar, Yuejie Chi, Adam Wierman, Costas Spanos.</p>
<p><em>International Conference on Learning Representations.</em></p>
<p><a href="https://arxiv.org/pdf/2502.19652">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/Robust-Gymnasium">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">IEEE TPAMI</div>
<div class="pub-main">
<p><strong>Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Bilgehan Sel, Yuhao Ding, Lu Wang, Qingwei Lin, Alois Knoll, Ming Jin.</p>
<p><em>IEEE Transactions on Pattern Analysis and Machine Intelligence.</em></p>
<p><a href="https://arxiv.org/pdf/2405.16390">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/CMORL">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">TMLR</div>
<div class="pub-main">
<p><strong>TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Alois Knoll, Ming Jin.</p>
<p><em>Transactions on Machine Learning Research.</em></p>
<p><a href="https://arxiv.org/pdf/2403.08694">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/TeaMs-RL">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">NeurIPS 2024</div>
<div class="pub-main">
<p><strong>Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Laixi Shi, Yuhao Ding, Alois Knoll, Costas Spanos, Adam Wierman, Ming Jin.</p>
<p><em>Advances in Neural Information Processing Systems.</em></p>
<p><a href="https://arxiv.org/pdf/2405.20860">[Paper]</a>, <a href="https://github.com/SafeRL-Lab">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">IEEE TPAMI</div>
<div class="pub-main">
<p><strong>A Review of Safe Reinforcement Learning: Methods, Theory and Applications.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Alois Knoll.</p>
<p><em>IEEE Transactions on Pattern Analysis and Machine Intelligence.</em></p>
<p><a href="https://arxiv.org/pdf/2205.10330">[Paper]</a>, <a href="https://github.com/chauncygu/Safe-Reinforcement-Learning-Baselines">[Code]</a>, <a href="https://mp.weixin.qq.com/s/Ks2EdGstTPngPrnZFcAdSw">[机器之心]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">AAAI 2024</div>
<div class="pub-main">
<p><strong>Balancing Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Bilgehan Sel, Yuhao Ding, Lu Wang, Qingwei Lin, Ming Jin, Alois Knoll.</p>
<p><em>Association for the Advancement of Artificial Intelligence.</em></p>
<p><a href="https://arxiv.org/pdf/2405.01677">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/Safety-MuJoCo">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">IEEE TII</div>
<div class="pub-main">
<p><strong>Safe Multiagent Learning With Soft-Constrained Policy Optimization in Real Robot Control.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Dianye Huang, Muning Wen, Guang Chen, Alois Knoll.</p>
<p><em>IEEE Transactions on Industrial Informatics.</em></p>
<p><a href="#">[Paper]</a>, <a href="https://github.com/SAIL-Research-Lab">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">IEEE TASE</div>
<div class="pub-main">
<p><strong>ROSCOM: Robust Safe Reinforcement Learning on Stochastic Constraint Manifolds.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Puze Liu, Alap Kshirsagar, Guang Chen, Jan Peters, Alois Knoll.</p>
<p><em>IEEE Transactions on Automation Science and Engineering.</em></p>
<p><a href="https://ieeexplore.ieee.org/abstract/document/10616119">[Paper]</a>, <a href="https://github.com/SafeRL-Lab">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">AIJ</div>
<div class="pub-main">
<p><strong>Safe Multi-Agent Reinforcement Learning for Multi-Robot Control.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Jakub Grudzien Kuba, Yuanpei Chen, Yali Du, Long Yang, Alois Knoll, Yaodong Yang.</p>
<p><em>The Journal of Artificial Intelligence.</em></p>
<p><a href="https://www.sciencedirect.com/science/article/abs/pii/S0004370223000516">[Paper]</a>, <a href="https://github.com/chauncygu/Multi-Agent-Constrained-Policy-Optimisation">[Code]</a></p>
</div>
</div>
<div class="pub-row">
<div class="paper-tag">IEEE TAI</div>
<div class="pub-main">
<p><strong>Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving.</strong></p>
<p>Zheng Zhi, <b><span style="text-decoration: underline">Shangding Gu</span></b>.</p>
<p><em>IEEE Transactions on Artificial Intelligence.</em></p>
<p><a href="https://arxiv.org/pdf/2405.18209">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/Safe-MARL-in-Autonomous-Driving">[Code]</a></p>
</div>
</div>
<!-- ================================ Preprints ================================ -->
<h2>Selected Recent Preprints</h2>
<ul>
<li>
<p><strong>From Model Scaling to System Scaling: Scaling the Harness in Agentic AI.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b></p>
<p><a href="https://arxiv.org/pdf/2605.26112">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/cheetahclaws">[Code]</a>, <a href="https://x.com/dair_ai/status/2059294269698199929">[DAIR.AI]</a></p>
</li>
<li>
<p><strong>Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b></p>
<p><a href="https://arxiv.org/pdf/2602.15028">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/PAPerBench">[Code]</a></p>
</li>
<li>
<p><strong>Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity.</strong></p>
<p>Yingxuan Yang, Chengrui Qu, Muning Wen, Laixi Shi, Ying Wen, Weinan Zhang, Adam Wierman, <b><span style="text-decoration: underline">Shangding Gu</span></b></p>
<p><a href="https://arxiv.org/pdf/2602.03794">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/agentic-web">[Code]</a>, <a href="https://mp.weixin.qq.com/s/aFg6PpQcPXMjd22hLWsG8g">[机器之心]</a></p>
</li>
<li>
<p><strong>Agentic web: Weaving the next web with AI agents.</strong></p>
<p>Yingxuan Yang, Mulei Ma, Yuxuan Huang, Huacan Chai, Chenyu Gong, Haoran Geng, Yuanjian Zhou, Ying Wen, Meng Fang, Muhao Chen, <b><span style="text-decoration: underline">Shangding Gu</span></b>, Ming Jin, Costas Spanos, Yang Yang, Pieter Abbeel, Dawn Song, Weinan Zhang, Jun Wang</p>
<p><a href="https://arxiv.org/pdf/2507.21206?">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/Agent-Scaling">[Code]</a>, <a href="https://spectrum.ieee.org/agentic-web">[IEEE Spectrum]</a>, <a href="https://www.economist.com/interactive/science-and-technology/2025/12/10/the-next-version-of-the-web-will-be-built-for-machines-not-humans">[The Economist]</a>, <a href="https://mp.weixin.qq.com/s/Co1lBdo-nhErdeAFdewyCg">[机器之心]</a></p>
</li>
<li>
<p><strong>AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond.</strong></p>
<p><b><span style="text-decoration: underline">Shangding Gu</span></b>, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Boyi Li, Ming Jin, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos</p>
<p><a href="https://accident-bench.github.io/">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/AccidentBench">[Code]</a></p>
</li>
<li>
<p><strong>Data Uniformity Improves Training Efficiency and More, with a Convergence Framework Beyond the NTK Regime.</strong></p>
<p>Yuqing Wang, <b><span style="text-decoration: underline">Shangding Gu</span></b></p>
<p><a href="https://arxiv.org/pdf/2506.24120">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/data-uniformity">[Code]</a></p>
</li>
<li>
<p><strong>RLBenchNet: The Right Network for the Right Reinforcement Learning Task.</strong></p>
<p>Ivan Smirnov, <b><span style="text-decoration: underline">Shangding Gu</span></b></p>
<p><a href="https://arxiv.org/pdf/2505.15040">[Paper]</a>, <a href="https://github.com/SafeRL-Lab/BenchNetRL">[Code]</a></p>
</li>
</ul>
<p class="muted">For more details, please see <a href="https://scholar.google.com/citations?user=E1GCDXUAAAAJ&hl=en" target="_blank">Google Scholar</a>.</p>
<!-- ====================================================================
Archived: previously displayed full list (hidden on request; preserved for reference)
====================================================================
<ul>
<li><p><b><span style="text-decoration: underline">Gu, S*.,</span></b> Shi, L., Wen, M., Jin, M., Mazumdar, E., Chi, Y., Wierman, A., Spanos, C.. (2025). Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning. <i>ICLR 2025.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Sel, B., Ding, Y., Wang, L., Lin, Q., Knoll, A., & Jin, M. (2025). Safe and Balanced: A Framework for Constrained Multi-Objective Reinforcement Learning. <i>IEEE Transactions on Pattern Analysis and Machine Intelligence.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Knoll, A., & Jin, M. (2024). TeaMs-RL: Teaching LLMs to Generate Better Instruction Datasets via Reinforcement Learning. <i>Transactions on Machine Learning Research.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu*, S.,</span></b> Shi, L., Ding, Y., Knoll, A., Spanos, C., Wierman, A., & Jin, M. (2024). Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation. <i>NeurIPS.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Yang, L., Du, Y., Chen, G., Walter, F., Wang, J., & Knoll, A. (2024). A review of safe reinforcement learning: Methods, theory and applications. <i>IEEE Transactions on Pattern Analysis and Machine Intelligence.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Zheng, Z.,</span></b> & <b><span style="text-decoration: underline">Gu*, S.</span></b> (2024). Safe Multi-Agent Reinforcement Learning with Bilevel Optimization in Autonomous Driving. <i>IEEE Transactions on Artificial Intelligence.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Liu, P., Kshirsagar, A., Chen, G., Peters, J., Knoll, A. (2024). ROSCOM: Robust Safe Reinforcement Learning on a Stochastic Constraint Manifolds. <i>IEEE Transactions on Automation Science and Engineering.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Huang, D., Wen, M., Chen, G., Knoll, A. (2024). Safe Multi-Agent Learning with Soft Constrained Policy Optimization in Real Robot Control. <i>IEEE Transactions on Industrial Informatics.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Bilgehan S., Ding, Y., Wang, L., Lin, Q., Jin, M., Knoll, A. (2024). Balance Reward and Safety Optimization for Safe Reinforcement Learning: A Perspective of Gradient Manipulation. <i>AAAI 2024 (Oral paper).</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Kuba, J. G., Chen, Y., Du, Y., Yang, L., Knoll, A., & Yang, Y. (2023). Safe multi-agent reinforcement learning for multi-robot control. <i>Artificial Intelligence, 319, 103905.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Kshirsagar, A., Du, Y., Chen, G., Peters, J., & Knoll, A. (2023). A human-centered safe robot reinforcement learning framework with interactive behaviors. <i>Frontiers in Neurorobotics, 17.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Chen, G., Zhang, L., Hou, J., Hu, Y., & Knoll, A. (2022). Constrained Reinforcement Learning for Vehicle Motion Planning with Topological Reachability Analysis. <i>Robotics, 11(4), 81. (Editor selected paper).</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Zhu, M., Chen, G., Wen, Y., & Knoll, A. (2022). Computing position error margin for a USV due to wind and current with a trajectory model. <i>Ocean Engineering, 262, 111950. (Top Journal in this area).</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Zhou, C., Wen, Y., Xiao, C., & Knoll, A. (2022). Motion Planning for an Unmanned Surface Vehicle with Wind and Current Effects. <i>Journal of Marine Science and Engineering, 10(3), 420.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Zhou, C.,</span></b> <b><span style="text-decoration: underline">Gu*, S.,</span></b> Wen, Y*., Du, Z., Xiao, C., Huang, L., & Zhu, M. (2020). The review unmanned surface vehicle path planning: Based on multi-modality constraint. <i>Ocean Engineering, 200, 107043. (Corresponding author, Top Journal in this area).</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Zhou, C.,</span></b> <b><span style="text-decoration: underline">Gu*, S.,</span></b> Wen, Y*., Du, Z., Xiao, C., Huang, L., & Zhu, M. (2020). Motion planning for an unmanned surface vehicle based on topological position maps. <i>Ocean Engineering, 198, 106798. (Corresponding author, Top Journal in this area).</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Zhou, C., Wen, Y., Zhong, X., Zhu, M., Xiao, C., & Du, Z. (2020). A motion planning method for unmanned surface vehicle in restricted waters. <i>Proceedings of the Institution of Mechanical Engineers, Part M: Journal of Engineering for the Maritime Environment, 234(2), 332-345.</i></p>
</li>
<li><p><b><span style="text-decoration: underline">Gu, S.,</span></b> Zhou, C., Wen, Y., Xiao, C., Du, Z., & Huang, L. (2019). Path Search of Unmanned Surface Vehicle Based on Topological Location. <i>Navigation of China, 42(02), 52-58.</i></p>
</li>
</ul>
-->
</div>
<footer class="site-footer">
<div class="footer-inner">
© 2026 SAIL Lab (Safe AI and Robot Learning Lab) — School of Computer Science, Shanghai Jiao Tong University<br/>
Principal Investigator: Prof. Shangding Gu (顾尚定) ·
<a href="mailto:shangding.gu@sjtu.edu.cn">shangding.gu@sjtu.edu.cn</a> ·
<a href="https://github.com/SAIL-Research-Lab" target="_blank">GitHub</a> ·
<a href="https://www.youtube.com/@ai-agent-research" target="_blank">YouTube</a> ·
<a href="https://scholar.google.com/citations?user=E1GCDXUAAAAJ&hl=en" target="_blank">Google Scholar</a>
</div>
</footer>
<script>
document.querySelectorAll('.nav-item a').forEach(function(a){
a.addEventListener('click', function(){ document.body.classList.remove('nav-open'); });
});
</script>
</body>
</html>