الملخص

Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we introduce Logbook, a benchmark for hour-scale audio understanding, with recordings ranging from ten minutes to six days. Given a continuous audio recording and an event label vocabulary, a system must predict a gap-free segmentation with an event label and a description per segment. We compare 52 systems, end-to-end and cascaded, and ablate fine-tuning, context length, and reasoning budget. We find the task tractable, though the best systems remain below the human reference. Also, over-segmentation is pervasive, and fine-tuning partially mitigates it. Finally, end-to-end are often better than cascaded systems, but degrades with longer context.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Choi, K., Shon, S., Serdyuk, D., Lan, G., Huang, C. W., Rasooli, M. S., Srivastava, S., Lin, Z., Adya, S., & Sun, M. (2026). Logbook: Extremely Long-form Audio Event Understanding. https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding

MLA 9

Choi, Kwanghee, et al. "Logbook: Extremely Long-form Audio Event Understanding." https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.

شيكاغو (المؤلف–التاريخ)

Choi, Kwanghee, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, Saurabh Adya, and Ming Sun. 2026. "Logbook: Extremely Long-form Audio Event Understanding." https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.

هارفارد

Choi, K., Shon, S., Serdyuk, D., Lan, G., Huang, C. W., Rasooli, M. S., Srivastava, S., Lin, Z., Adya, S. and Sun, M. (2026) 'Logbook: Extremely Long-form Audio Event Understanding', Available at: https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.

فانكوفر

Choi K, Shon S, Serdyuk D, Lan G, Huang CW, Rasooli MS, et al. Logbook: Extremely Long-form Audio Event Understanding. https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding

IEEE

K. Choi, S. Shon, D. Serdyuk, G. Lan, C. W. Huang, M. S. Rasooli, S. Srivastava, Z. Lin, S. Adya, and M. Sun, "Logbook: Extremely Long-form Audio Event Understanding," https://omanscience.com/ar/articles/logbook-extremely-long-form-audio-event-understanding.