الباحثون

Yifan Peng

المنشورات 5

نسخة أولية وصول مفتوح

LoDEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery

Kai Li, Zigan Zhou, Zhenyang Li وآخرون · 2026

Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO prediction can instead use a compact offset token. Under lo …

نسخة أولية وصول مفتوح

Explainable Suicide Risk Assessment on Social Media with Multi-Task QLoRA

Xuan Zhong Feng, Geoffrey Martin, Hexin Dong وآخرون · 2026

Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three task …

نسخة أولية وصول مفتوح

E-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual Encoding

Jiale Wu, Xiaoyang Bai, Haoming Yu وآخرون · 2026

Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and computational cost; whil …

نسخة أولية وصول مفتوح

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T bran …

نسخة أولية وصول مفتوح

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this w …

المؤلفون المشاركون