الباحثون

Haoyang Li

المنشورات 7

نسخة أولية وصول مفتوح

Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models

Shengyin Sun, Yiming Li, Yingzhao Lian وآخرون · 2026

Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit st …

نسخة أولية وصول مفتوح

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

Zhiyi Mou, Yao Lu, Wangze Ni وآخرون · 2026

Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicuous in a serialized …

نسخة أولية وصول مفتوح

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

Hongcheng Gao, Hailong Qu, Yu Lei وآخرون · 2026

Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present Engi …

نسخة أولية وصول مفتوح

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

Haoyang Li, Ruoxi Sun, Qingqing Ye وآخرون · 2026

Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are …

نسخة أولية وصول مفتوح

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

Guanlin Li, Shifeng Bao, Yihan Zhao وآخرون · 2026

Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a paramet …

نسخة أولية وصول مفتوح

TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement

Haoyang Li, Yaxin Xiao, Linyan Dai وآخرون · 2026

Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We theref …

نسخة أولية وصول مفتوح

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Hejia Geng, Zesen Huang, Haoyang Li وآخرون · 2026

Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain conventions, and specialized correctness criteria make this knowledge difficult to convert into reliable learning experience-a challenge we call the scientific experience b …

المؤلفون المشاركون