الباحثون

Chen Chen

المنشورات 15

نسخة أولية وصول مفتوح

LiteNWM: Efficient Latent World Models for Onboard Visual Navigation in the Wild

Linkai Liu, Yuntian Zhang, Zhenshan Bing وآخرون · 2026

Direct visual navigation policies generate trajectories efficiently but do not explicitly evaluate their future consequences. Generative navigation world models provide this foresight through visual rollouts, which are costly when evaluating multiple candidates. We present LiteNWM, a latent navigation world model that …

نسخة أولية وصول مفتوح

Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence

Chen Chen, Dongjie Wang, Mei Liu وآخرون · 2026

Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these …

نسخة أولية وصول مفتوح

MedPrune: Topology-Efficient Multimodal Multi-Agent Communication Evolution for Medical VQA Tasks

Jiuheng Wan, Runze Li, Chen Chen وآخرون · 2026

While medical multimodal large language models (Med-MLLMs) advance medical visual question answering (VQA), existing clinical workflow-inspired multi-agent frameworks suffer from interaction patterns and excessive computational overhead caused by redundant communication topologies. In this paper, we propose MedPrune, a …

نسخة أولية وصول مفتوح

From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deployment

Zeshen Zhang, Han Zhao, Weihao Cui وآخرون · 2026

As Large Language Models (LLMs) become essential in privacy-sensitive sectors like hospitals and government agencies, the on-premise LLM servers offer a cost-effective and secure alternative to public cloud services. However, these resource-constrained servers struggle to guarantee heterogeneous Service Level Objective …

نسخة أولية وصول مفتوح

Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Sixun Dong, Wei Li, Andong Deng وآخرون · 2026

Efficient long-video understanding with vision-language models (VLMs) is often framed as selecting informative frames or visual tokens at a fixed native resolution. We show that per-frame resolution can instead be traded for denser temporal coverage, while front-end decoding latency depends on the size of the candidate …

نسخة أولية وصول مفتوح

RASPER: Reward-Aligned Summarization of Clinical Notes for EHR Outcome Prediction

Unstructured discharge notes in Electronic Health Records (EHRs) often carry signal complementary to structured medical codes, holding patient-specific evidence that standardized cohort-level codes alone cannot capture. However, this evidence in notes is frequently buried in lengthy, noisy text that is not intentionall …

نسخة أولية وصول مفتوح

Visualizing Distribution Coverage in Generative Diffusion Models

Yifei Wang, Xiaoyu Wu, Tsu-Jui Fu وآخرون · 2026

Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their multi-step teachers in generation. However, standard evaluations such as GenEval2 typically draw only one sample per prompt, so improved scores may fail to reveal losses …

نسخة أولية وصول مفتوح

Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks

Guoheng Sun, Chen Chen, Jin Wang وآخرون · 2026

World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over …

نسخة أولية وصول مفتوح

Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations

Guanghui Min, Liang Wu, Mingjia Shi وآخرون · 2026

Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate indi …

نسخة أولية وصول مفتوح

ASAP: Visual Analytics for Identifying and Analyzing Image Patterns in AI-generated Images

Jinbin Huang, Yuki Ueno, Chen Chen وآخرون · 2026 · 10.1177/14738716261481077

Generative image models can produce highly realistic images, raising concerns about potential misuse in creating deceptive content. Current deepfake approaches face several challenges, including limited generalizability, lack of interpretability, and poor actionability. To help address these, we present ASAP, an intera …

نسخة أولية وصول مفتوح

MORSE: Multi-Context Ordering via Reverse Scoring for Evidence-Preserving Compression

Ke Wan, Yifan Wang, Liheng Lai وآخرون · 2026

Retrieval-augmented generation often relies on multiple retrieved contexts that contain substantial redundancy, motivating context compression to preserve useful information under limited input budgets. Likelihood-based compressors can account for cross-context redundancy through sequential scoring, but this makes evid …

نسخة أولية وصول مفتوح

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T bran …

نسخة أولية وصول مفتوح

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this w …

المؤلفون المشاركون