الباحثون

Yu Hu

المنشورات 2

نسخة أولية وصول مفتوح

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed …

نسخة أولية وصول مفتوح

4D Radar Perception Algorithms for Autonomous Driving: A Review

Xumin Wu, Jun Zhou, Jilin Mei وآخرون · 2026

Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. This review organizes the field according to the evolution of perception t …

المؤلفون المشاركون