Abstract
Vision offers richer context than traditional marine sensors (e.g., LiDAR, Doppler Velocity Logger (DVL), sonar) but is harder to interpret on water due to reflections, glare, and dynamic surfaces. SUSHI is a vision-first navigation system for Autonomous Surface Vehicles (ASVs) that fuses detection, water segmentation, and monocular depth to produce camera-centric navigation grids for planning and control. The proposed perception methods achieve 90% segmentation accuracy through knowledge distillation with SAM2 logits, requiring only 500-550 frames and approximately 30 minutes of training. The system implements a YOLO detection model that achieves 94.5% [email protected] (F1 score: 0.91) for trash and obstacle detection in simulation, and benchmarks a monocular depth method that solves the issue of reflective surfaces and can work universally. Path planning uses a Multi-Field Synthesis (MFS) approach: a locally reactive artificial-potential-field component blended adaptively with a global wavefront flow field, mitigating local minima while preserving real-time responsiveness. A behavior layer prioritizes target seeking and mask-based visual exploration when explicit goals are absent. Validation was performed in the TOAST simulator and in a pool environment, demonstrating robust goal targeting and exploration using cameras with minimal side sensing for emergency avoidance.
| Original language | English |
|---|---|
| Article number | 38 |
| Journal | Journal of Intelligent and Robotic Systems |
| Volume | 112 |
| Issue number | 2 |
| Early online date | 6 Mar 2026 |
| DOIs | |
| Publication status | Published - Jun 2026 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 14 Life Below Water
Keywords
- Multi-field synthesis
- Marine robotics
- Autonomous surface vehicles
- Path planning
- Visual navigation
- Computer vision
Fingerprint
Dive into the research topics of 'SUSHI: A Vision System for Reactive, Uninformed ASV Navigation via Multi-Field Path Planning and Visual Exploration'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver