الباحثون

Fan Zhang

المنشورات 19

نسخة أولية وصول مفتوح

RewardWeaver: Long-Horizon Interactive Learning for Language Agents via Self-Evolving Reward Adaptation

Hengbo Xiao, Boyao Zhang, Purui Liu وآخرون · 2026

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in domains where task outcomes can be reliably evaluated, but long-horizon interaction remains challenging due to sparse terminal feedback and difficult credit assignment. Process rewards provide denser supervision, yet the capabiliti …

نسخة أولية وصول مفتوح

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and mode …

نسخة أولية وصول مفتوح

InferOpt: Constrained Multi-Objective Search for LLM Inference Configurations

Qi Chen, Yingying Cheng, Zhaoyi Sun وآخرون · 2026

Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics that return a single operating point and do not scale to layer-wise search spaces. We recast inference configuration as constrained multi-objective b …

نسخة أولية وصول مفتوح

FiberGeoText: A Vision-Language Model for Population- Level Organization of Superficial White Matter

The superficial white matter (SWM), a critical brain region for cognition across the lifespan and brain disease, contains abundant short-range association fibers whose organization remains incompletely characterized, in part because the short trajectories and highly variable cortical folding make correspondence across …

نسخة أولية وصول مفتوح

FutureWorlds: Learning Robotic World Models from Alternative Futures

Hao Wu, Shengju Qian, Weiyan Wang وآخرون · 2026

Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent ma …

نسخة أولية وصول مفتوح

OceanMind: A multi-agent AI system for ocean diagnosis

Fan Zhang, Weicong Cheng, Yuheng Chen وآخرون · 2026

Time-dependent, three-dimensional (3D) oceanic multi-variables define coherent states of the evolving ocean to facilitate ocean diagnosis and advance ocean science to better inform environmental and hazard management. However, extracting quantitative evidence from these variables requires substantial and complex analyt …

نسخة أولية وصول مفتوح

RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance

Yiming Liu, Ben Wan, Tongxuan Liu وآخرون · 2026

Diffusion models enable high-quality visual generation, but iterative denoising remains computationally expensive, especially under classifier-free guidance (CFG), which requires both conditional and unconditional evaluations. Training-free caching reduces this cost by reuse of previously computed features or predictio …

نسخة أولية وصول مفتوح

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Donghao Zhou, Haoyang He, Fan Zhang وآخرون · 2026

Instruction-guided video editing has made significant progress, yet existing methods use multimodal large language models (MLLMs) primarily as semantic encoders, so they often fall short in working with implicit edits that require causal or semantic reasoning. To bridge this fundamental gap in video editing, we propose …

نسخة أولية وصول مفتوح

From Pixel Generation to Topological Inference: Structural Dual Super-Resolution for Trustworthy Cross-Physical-Domain Trabecular Morphology Learning

Clinical CT and UHRCT cannot resolve individual trabeculae, whereas synchrotron radiation microCT (SRuCT) provides high-resolution references but is not applicable for in vivo imaging. The two domains differ by a 32x resolution gap, are only coarsely paired, and exhibit severe physical differences including partial vol …

نسخة أولية وصول مفتوح

OceanXL: Large-scale Underwater 3D Gaussian Splatting via Block Partitioning and Adaptive Pruning

Haoran Wang, Shaoyu Cai, Adrian Azzarelli وآخرون · 2026

Underwater 3D reconstruction is critical for marine exploration, ecological monitoring, and subsea infrastructure inspection, yet remains challenging at large scale due to light attenuation, scattering, and limited capture coverage. While 3D Gaussian Splatting (3DGS) enables high-quality real-time rendering, its applic …

نسخة أولية وصول مفتوح

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordi …

نسخة أولية وصول مفتوح

Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States

Shuhao Li, Weidong Yang, Changan Liu وآخرون · 2026

Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable cha …

نسخة أولية وصول مفتوح

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

Fan Zhang, Yankai Chen, Zhuohan Xie وآخرون · 2026

Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating in …

نسخة أولية وصول مفتوح

DTKDP: A Dual Teacher Knowledge Distillation and Pruning Framework for Lightweight Oriented SAR Ship Detection

Yuming Li, Fan Zhang, Alin M. Achim · 2026 · 10.3390/rs18183172

Two-stage oriented detectors achieve high localization accuracy in synthetic aperture radar (SAR) ship detection, but their large backbones, feature pyramids, proposal modules, and heavy region of interest (RoI) heads hinder deployment. Existing lightweight SAR ship detectors typically use one-stage frameworks that lac …

نسخة أولية وصول مفتوح

PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing

Donghao Zhou, Jia-Hui Pan, Fan Zhang وآخرون · 2026

Robotic bin packing requires long-horizon sequential decision-making, as each object placement affects the available space for subsequent packing. Existing methods primarily rely on hand-crafted geometric heuristics that optimize predefined objectives or reinforcement learning policies learned through trial and error o …

المؤلفون المشاركون