Abstract

Comprehensive 3D scene understanding for autonomous driving requires modeling geometry, semantics, and motion. However, camera-based occupancy and scene flow prediction are sensitive to unreliable spatial and temporal aggregation, caused by semantically incompatible image features, misaligned historical observations, and incomplete voxel structures. To address this issue, we propose SelectOccFlow, a selective spatiotemporal aggregation framework that progressively refines contextual evidence across image, temporal, and voxel domains. To obtain semantically compatible image evidence, we design Semantic-Guided Sampling (SGS) to regulate feature sampling with semantic priors. Since reliable image evidence alone cannot resolve temporal inconsistency, we then present State-Conditioned Temporal Aggregation (SCTA) to selectively retrieve historical evidence according to voxel states. To further enhance the structural completeness of voxel representations, we introduce Extent-Aware Spatial Aggregation (ESA), which exploits directional structural support to refine foreground geometry. Experiments on OpenOcc demonstrate that SelectOccFlow achieves a state-of-the-art OccScore of 44.9, improving the previous best by +4.2%. It also maintains competitive occupancy performance on Occ3D-nus and improves the mean OccScore under nuScenes-C corruptions by +11.1%, demonstrating improved robustness to visual corruptions. The source code will be made publicly available at https://github.com/muchen1021/SelectOccFlow.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, Y., Luo, K., Zheng, Y., & Yang, K. (2026). SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction. https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction

MLA 9

Wang, Yuhang, et al. "SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction." https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction.

Chicago (author–date)

Wang, Yuhang, Kai Luo, Yuanfan Zheng, and Kailun Yang. 2026. "SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction." https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction.

Harvard

Wang, Y., Luo, K., Zheng, Y. and Yang, K. (2026) 'SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction', Available at: https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction.

Vancouver

Wang Y, Luo K, Zheng Y, Yang K. SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction. https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction

IEEE

Y. Wang, K. Luo, Y. Zheng, and K. Yang, "SelectOccFlow: Selective Spatiotemporal Aggregation for 3D Occupancy and Scene Flow Prediction," https://omanscience.com/en/articles/selectoccflow-selective-spatiotemporal-aggregation-for-3d-occupancy-and-scene-flow-prediction.