الباحثون

Jie Yang

المنشورات 9

نسخة أولية وصول مفتوح

Self-correction Optimization for Interleaved Multimodal Generation

Xin You, Zhiwei Ning, Zukai Chen وآخرون · 2026

Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solut …

نسخة أولية وصول مفتوح

Flow Matching Reinforcement for 3D Mesh Generation via Dynamic Homing Optimization

Zhen Zhou, Zhiwei Ning, Puhua Jiang وآخرون · 2026

Flow matching is central to 3D generation, yet in practice its reinforcement learning (RL) methods are largely adapted from 2D visual generation. Representative DPO-, GRPO-, and NFT-style objectives, when applied to negative trajectories, mainly steer predicted velocities away from the corresponding directions without …

نسخة أولية وصول مفتوح

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

Zhengyu Fang, Seoyeon Hong, Jie Yang وآخرون · 2026

On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialis …

نسخة أولية وصول مفتوح

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Muyang Li, Jie Yang, Zhengyu Fang وآخرون · 2026

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta …

نسخة أولية وصول مفتوح

TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

Dehai Min, Daoan Zhang, Yiming Zeng وآخرون · 2026

An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable …

نسخة أولية وصول مفتوح

LastOPD: Taming Collapse in Latent On-Policy Distillation

Jie Yang, Zhengyu Fang, Zelin Xu وآخرون · 2026

On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods …

نسخة أولية وصول مفتوح

DiaVLo: Diagnosing Behaviours of Vision-Language Models

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diag …

المؤلفون المشاركون