Onur Bagoren*
·
Seth Isaacson*
·
Sacchin Sundar
·
Yung-Ching Sun
Anja Sheppard
·
Haoyu Ma
·
Abrar Shariff
·
Ram Vasudevan
·
Katherine A. Skinner
*Equal contribution
Background image courtesy of the National Oceanic and Atmospheric Administration Thunder Bay National Marine Sanctuary.
Abstract (Click to Expand)
Localization and mapping are core perceptual capabilities for underwater robots. Stereo cameras provide a low-cost means of directly estimating metric depth to support these tasks. However, despite recent advances in stereo depth estimation on land, computing depth from image pairs in underwater scenes remains challenging. In underwater environments, images are degraded by light attenuation, visual artifacts, and dynamic lighting conditions. Furthermore, real-world underwater scenes frequently lack rich texture useful for stereo depth estimation and 3D reconstruction. As a result, stereo estimation networks trained on in-air data cannot transfer directly to the underwater domain. In addition, there is a lack of real-world underwater stereo datasets for supervised training of neural networks. Poor underwater depth estimation is compounded in stereo-based Simultaneous Localization and Mapping (SLAM) algorithms, making it a fundamental challenge for underwater robot perception. To address these challenges, we propose a novel framework that enables sim-to-real training of underwater stereo disparity estimation networks using simulated data and self-supervised finetuning. We leverage our learned depth predictions to develop SurfSLAM, a novel framework for real-time underwater SLAM that fuses stereo cameras with IMU, barometric, and Doppler Velocity Log (DVL) measurements. Lastly, we collect a challenging real-world dataset of shipwreck surveys using an underwater robot. Our dataset features over 24,000 stereo pairs, along with high-quality, dense photogrammetry models and reference trajectories for evaluation. Through extensive experiments, we demonstrate the advantages of the proposed training approach on real-world data for improving stereo estimation in the underwater domain and for enabling accurate trajectory estimation and 3D reconstruction of complex shipwreck sites.Data from this project is available at DeepBlue, consisting of shipwreck surveys collected at the National Oceanic and Atmospheric Administration Thunder Bay National Marine Sanctuary. Data is stored as hdf5 archives for their compression, ability to stream, and minimal dependencies. See docs/data_format.md for details on the data format, including how to convert your data to our format.
We highly recommend using docker for this project and don't provide explicit support for non-docker setups.
Check out the model submodules and apply our patches to them:
./scripts/setup_submodules.shThe submodules point at their public upstream repositories. Our changes live in patches/, one file per submodule, and setup_submodules.sh applies them. See docs/submodules.md for details.
Then build the docker image. This pulls an image from docker hub that has most dependencies installed, then adds a user-specific configuration and mounts all the data paths.
cd docker/
./build_user.shBEFORE launch the docker container, you must tell the system where you downloaded data to:
cp cfg/dataset_paths.example.yaml cfg/dataset_paths.yaml
# edit cfg/dataset_paths.yaml to point at the DeepBlue downloadOnce that is done, you may launch the container. The first time you start the container it will build and install the turtlmap python bindings.
./run.sh # starts a container, or attaches to an existing container (allowing multiple terminals)
./run.sh restart # restarts the container, discarding any changes you have made to the local containerSee docs/docker.md for more details.
All commands below run inside the docker container. Note that if you are running outside of docker, you will need to manually build and install the turtlmap bindings. To run SLAM on a sequence:
python3 examples/run_surfslam.py cfg/tbnms/monohansett_long.yaml # also: monohansett_boiler.yaml, monohansett_engine.yamlResults (estimated trajectory, statistics, and a copy of the settings used) are written to ./outputs/<experiment_name>_<timestamp>/. Useful flags:
--num_repeats N: run N trials of each configuration (results go intrial_*/subdirectories).--overrides <yaml>(optionally with--run_all_combos): sweep parameters for ablations. See cfg/README.md for how the settings system works.--duration <sec>: only process the first part of the sequence.
To rebuild a dense map offline from a finished run's trajectory:
python examples/offline_mapping.py cfg/tbnms/mono_long.yaml outputs/<run_dir> --output_mesh mesh.plyBy default, trajectory evaluation picks up every run of our method found in outputs/ and compares it against the released ground-truth trajectories -- no configuration needed. To evaluate other methods side by side, add entries to the algorithms list in analysis/traj_eval_cfg.yaml.
cd analysis/
python evaluate_trajectories.py traj_eval_cfg.yamlThis writes APE/RPE/completeness tables to ./eval_results/, along with the ground-truth alignment transforms (eval_results/alignments/) needed for map evaluation.
For map evaluation, first run evaluate_trajectories.py with analysis/traj_eval_cfg_map.yaml to align each method's trajectory, then evaluate the reconstructions against the photogrammetry models:
python evaluate_trajectories.py traj_eval_cfg_map.yaml
python mapping/evaluate_maps.py ../eval_results/map_comparison/alignmentsThis computes accuracy, completeness, precision, and recall for every trial (ground-truth point clouds are resolved automatically from dataset_paths.yaml) and writes results to eval_results/map_comparison/maps/. Note that the release's trajectory.tum and reconstruction.ply are in different frames -- the trajectory is re-referenced so its first pose is the identity -- so evaluate_maps.py composes each alignment with the per-scene transform in analysis/mapping/model_frame.yaml to reach the photogrammetry frame.
To turn those per-trial metrics into the paper's mapping table:
python compute_mapping_summary.py
python generate_mapping_latex_table.pyRuntime and pipeline statistics can be summarized with python summarize_statistics.py --config stats_summary_cfg.yaml.
Code in this repo is based on the structure from LONER and the opti-acoustic-inertial pipeline TURTLMap.