الملخص
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Xu, W., Cheng, K. J., Saito, K., Shi, J., Li, T., Liu, Y., Wang, L., Ishii, M., Shibuya, T., Anumanchipalli, G., & Liang, P. P. (2026). AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes. https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes
MLA 9
Xu, Weihan, et al. "AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes." https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes.
شيكاغو (المؤلف–التاريخ)
Xu, Weihan, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, Takashi Shibuya, Gopala Anumanchipalli, and Paul Pu Liang. 2026. "AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes." https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes.
هارفارد
Xu, W., Cheng, K. J., Saito, K., Shi, J., Li, T., Liu, Y., Wang, L., Ishii, M., Shibuya, T., Anumanchipalli, G. and Liang, P. P. (2026) 'AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes', Available at: https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes.
فانكوفر
Xu W, Cheng KJ, Saito K, Shi J, Li T, Liu Y, et al. AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes. https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes
IEEE
W. Xu, K. J. Cheng, K. Saito, J. Shi, T. Li, Y. Liu, L. Wang, M. Ishii, T. Shibuya, G. Anumanchipalli, and P. P. Liang, "AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes," https://omanscience.com/ar/articles/avio-learning-to-add-and-remove-sounding-objects-in-audiovisual-scenes.