Abstract

Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Sen, S., & Ahmadi, Z. (2026). MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos. https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos

MLA 9

Sen, Souptik, and Zahra Ahmadi. "MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos." https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos.

Chicago (author–date)

Sen, Souptik, and Zahra Ahmadi. 2026. "MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos." https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos.

Harvard

Sen, S. and Ahmadi, Z. (2026) 'MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos', Available at: https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos.

Vancouver

Sen S, Ahmadi Z. MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos. https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos

IEEE

S. Sen, and Z. Ahmadi, "MacJEPA: Missingness-Robust Audio-Visual Recognition from Untrimmed Egocentric Videos," https://omanscience.com/en/articles/macjepa-missingness-robust-audio-visual-recognition-from-untrimmed-egocentric-videos.