Author Identifier
Date of Award
2026
Keywords
camera-LiDAR fusion, real-time semantic segmentation, OUTBACK dataset, synthetic dataset generation, off-road robot navigation, real-time machine learning models
Document Type
Thesis
Publisher
Edith Cowan University
Degree Name
Doctor of Philosophy
School
School of Engineering
First Supervisor
Alexander Rassau
0000-0002-8295-5681
Second Supervisor
Douglas Chai
0000-0002-9004-7608
Third Supervisor
Syed Mohammed Shamsul Islam
0000-0002-3200-2903
Abstract
The application of autonomous field robots in agriculture, mining, transport and logistics, disaster management, and exploration has increased with advances in sensing, perception, and autonomous navigation technologies. Western Australian off-road environments present challenges for autonomous ground robots because they contain unstructured ter rain, variations in appearance and geometry, changing illumination, and visually similar regions with different traversal characteristics. Reliable operation therefore requires scene understanding methods capable of representing the environment according to its navigability for robot traversal.
This PhD research investigates multimodal scene understanding for autonomous ground robot navigation in unstructured off-road environments. The research focuses on the development of a photorealistic synthetic multimodal dataset representative of Western Australian off-road environments, the development and evaluation of real-time camera-LiDAR fusion semantic segmentation models for navigability estimation, the development of a navigation-oriented evaluation metric for semantic segmentation, and the proof-of-concept integration of semantic perception with local traversability mapping and robot navigation within NVIDIA Isaac Sim. The review identified limitations in available multimodal off-road datasets and motivated the development of the OUTBACK dataset. A synthetic data generation framework was developed using NVIDIA Isaac Sim and high-fidelity 3D environmental assets to generate monocular and stereo RGB images, LiDAR point clouds, IMU measurements, and semantic annotations. Sim2Real experiments showed that models trained using the synthetic data could provide useful semantic predictions on real-world off-road datasets.
Three preliminary camera-LiDAR fusion models, namely the Balanced-Model, Robustness-Model, and Accuracy-Model, were developed to combine RGB appearance in formation with LiDAR-derived geometric information. Their performance varied across training and testing domains, with the Accuracy-Model achieving the highest average performance on the synthetic OUTBACK evaluation and the Robustness-Model showing stronger generalisation in the conducted RELLIS-3D Sim2Real experiments.
Based on these observations, a final camera-LiDAR fusion model was developed and evaluated against DFormer-Small, a recent high-performing RGB-D semantic segmentation model, under the same OUTBACK experimental conditions. Across four inde pendent training seeds, the proposed model achieved a mean mIoU of 0.8439 ± 0.0064, compared with 0.6946±0.0071 for DFormer-Small on the three-level navigability task. A controlled ablation replacing LiDAR depth with grayscale RGB while retaining the same dual-stream architecture produced lower performance, providing evidence that LiDAR derived geometry contributes complementary information.
A Robotic Navigation Intersection over Union (RN IoU) metric was introduced to complement conventional IoU and mIoU by weighting prediction errors according to their navigational severity. Evaluation using both metrics showed that RN IoU provides additional navigation-oriented information while preserving the relative ranking of the com pared models.
Finally, the developed perception model was integrated into a proof-of-concept local navigation framework in NVIDIA Isaac Sim. Predicted navigability masks were projected into a bird’s-eye-view cost representation and used for local path and motion planning. The experiments demonstrated successful autonomous local navigation within the evaluated synthetic environment, while identifying limitations including false-positive obstacle predictions and oscillatory navigation command behaviour. Overall, this research demonstrates the potential of photorealistic synthetic multi modal data and camera-LiDAR feature fusion for off-road scene understanding. The OUTBACK dataset, multimodal segmentation models, RN IoU metric, and proof of-concept navigation framework collectively contribute towards perception-driven autonomous navigation in unstructured off-road environments.
Recommended Citation
Wijayathunga, L. M. (2026). Multimodal scene understanding for mobile robot navigation in unstructured outdoor environments. Edith Cowan University. Retrieved from https://ro.ecu.edu.au/theses/3119