الملخص
Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Choi, K., Shon, S., Serdyuk, D., Lan, G., Huang, C. W., Rasooli, M. S., Srivastava, S., Lin, Z., Adya, S., & Sun, M. (2026). Logbook: Extremely Long-form Audio Event Understanding. https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding
MLA 9
Choi, Kwanghee, et al. "Logbook: Extremely Long-form Audio Event Understanding." https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.
شيكاغو (المؤلف–التاريخ)
Choi, Kwanghee, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, Saurabh Adya, and Ming Sun. 2026. "Logbook: Extremely Long-form Audio Event Understanding." https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.
هارفارد
Choi, K., Shon, S., Serdyuk, D., Lan, G., Huang, C. W., Rasooli, M. S., Srivastava, S., Lin, Z., Adya, S. and Sun, M. (2026) 'Logbook: Extremely Long-form Audio Event Understanding', Available at: https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.
فانكوفر
Choi K, Shon S, Serdyuk D, Lan G, Huang CW, Rasooli MS, et al. Logbook: Extremely Long-form Audio Event Understanding. https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding
IEEE
K. Choi, S. Shon, D. Serdyuk, G. Lan, C. W. Huang, M. S. Rasooli, S. Srivastava, Z. Lin, S. Adya, and M. Sun, "Logbook: Extremely Long-form Audio Event Understanding," https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.