A discrepancy exists in Earth Observation (EO) literature between offline benchmark metrics and operational performance for vessel detection in optical satellite imagery. Deep learning models evaluating on curated benchmark splits frequently report mean Average Precision (mAP) scores above 0.85, yet performance degrades when those same models encounter uncurated ocean imagery.

Optical vessel detection is frequently modeled as a two-class segmentation or localization problem: isolating bright hulls against uniform water backgrounds. In practice, environmental dynamics such as atmospheric scattering, specular reflection, variable sea states, and spatial sensor constraints introduce unpredictable noise that standard object detectors fail to model predictably.
This article analyzes the structural causes of detection failure, beginning with sensor physics and benchmark design, moving through computer vision failure modes, and concluding with hardware constraints on satellite edge platforms.
1. Benchmark Datasets and Performance Metrics
To evaluate model failure modes, we must first examine some of the training datasets used in the literature:
Airbus Ship Detection Challenge: Contains high-resolution optical tiles centered on maritime scenes with precise bounding box annotations. The dataset is suitable for baseline training in high-contrast conditions, but underrepresents cloud occlusion, haze, and extreme sun glint.
xView: A multi-class overhead dataset containing diverse terrestrial and maritime targets. The
shipclass exhibits high intra-class variance, combining small fishing craft, barges, and stationary commercial vessels in complex port environments.SeaDronesSee: Uses low-altitude Unmanned Aerial Vehicle (UAV) imagery to provide orientation labels and small-vessel samples. While useful for algorithm development, its low acquisition altitude does not capture spaceborne atmospheric attenuation or orbital geometry.
No single open dataset captures the full distribution of orbital artifacts. Consequently, published leaderboard scores often reflect performance on limited real-world conditions.
Aggregate mAP metrics mask significant performance drops when evaluations are split by target size, sea state, or cloud cover percentage. Some know model performances for open datasets are:
2. Sensor Physics and Spatial Resolution Limits
Ground Sample Distance vs. Target Footprint
The fundamental unit of detection is the pixel. Target visibility is governed by the sensor’s Ground Sample Distance (GSD) relative to physical vessel dimensions.
In high-resolution systems (e.g. <1m GSD), a small target (vessel) occupies approximately 20 pixels in length, providing sufficient structural features for convolutional kernel extraction. At a medium resolution of 5m GSD, a 10m vessel occupies 2 pixels, while a 20m vessel spans 4 pixels.
When a vessels spans fewer than 10 pixels in width, spatial features collapse. The detector must infer the presence of an object based almost entirely on intensity contrast rather than geometric morphology. Peer-reviewed evaluations on the Airbus dataset show a sharp drop in detection probability once target width falls below 10 pixels, reflecting a fundamental limitation in signal-to-noise ratio (SNR) rather than architectural flaws in the model backbone.
The Signal-to-Noise Floor for Small Vessels
At 5m resolution, an 8m fishing vessel projects onto 1 to 4 pixels. The spectral signature of these pixels blends hull reflectance with background water reflectance.
This spatial averaging produces high-entropy / uncertain feature maps in deeper network layers. The network cannot reliably differentiate between:
Pure sensor noise or defective pixels.
Breaking wave crests and sea foam.
Specular glint from wave slopes.
Small vessels.
In datasets like SeaDronesSee, where small vessels dominate the class distribution, detection performance drops unless training pipelines employ anchor boxes tuned specifically for sub-10-pixel targets, aggressive multi-scale feature fusion (such as Feature Pyramid Networks), and synthetic small-object copy-paste augmentation. Even with these modifications, environmental variables like high sea states generate persistent false negatives.
Georeferencing and Registration Error
Beyond target detection, operational maritime monitoring requires precise absolute spatial coordinates. Optical imagery is subject to georeferencing uncertainties stemming from sensor attitude errors, orbital ephemeris tolerances, and orthorectification limits.
Standard satellite imagery carries absolute positional errors ranging from 1 to 3 pixels without ground control points. At 5m GSD, a 3-pixel offset introduces a 15m ground translation error.
This spatial discrepancy does not affect mAP evaluation on image tiles, but it creates severe failure modes during downstream data fusion, such as correlating optical detections with Automatic Identification System (AIS) radio telemetry broadcasts.
3. Environmental Artifacts and Background Complexity
Atmospheric Scattering and Cloud Cover
Cloud contamination degrades optical detection pipelines through two distinct mechanisms:
Target Occlusion: Optically thick cloud cover blocks surface reflectance completely, rendering physical targets invisible to passive optical sensors.
False Positive Generation: Thin cirrus clouds, haze, and atmospheric contrails scatter solar radiation, generating high-frequency spatial structures that mimic vessel hulls in intensity and contrast.
Standard convolutional networks evaluate local gradient changes and texture patterns. Without explicit atmospheric correction or multi-spectral band ratioing, these models misclassify isolated cloud fragments as vessel signatures.
Recent architectural variations incorporate auxiliary cloud-masking branches or parallel atmospheric classification heads. While these additions reduce false positive rates in moderate conditions, their efficacy declines under shifting solar zenith angles and variable aerosol optical depths.
Motion Artifacts and Sensor Geometry
Most high-resolution Earth Observation satellites employ pushbroom scanners, collecting imagery line-by-line along the satellite's orbital path. Fast-moving airborne objects, such as commercial aircraft, introduce motion relative to the satellite’s ground track speed during image synthesis.
This velocity differential causes spatial stretching along the track direction, producing elongated, high-contrast artifacts that resemble vessel hulls. Standard non-maximum suppression (NMS) algorithms evaluate spatial overlap strictly through Intersection-over-Union (IoU) metrics and confidence thresholds.
Because NMS algorithms lack directional, motion, or time-based cues, they fail to suppress these motion-induced anomalies, resulting in persistent false positive detections in single-pass images.
Water Surface Non-Stationarity
Land-based object detectors rely on clear visual contrast between objects and background. In ocean imagery, background spectral reflectance is non-stationary and varies as a function of:
Local surface roughness driven by wind speed and wave state.
Bathymetry, water depth, and suspended sediment concentration.
Solar elevation angle and specular reflection (sun glint).
When a vessel hull exhibits low spectral contrast relative to the surrounding water column, such as gray naval hulls or dark-painted fishing vessels, edge-detection filters fail to extract closed boundaries. Background-subtraction algorithms trained on stationary sea-state statistics experience rapid performance degradation when processing scenes containing dense sun glint.
4. Bounding Box Geometry and Crowded Scenes
Axis-Aligned versus Oriented Bounding Boxes
The majority of general-purpose object detection architectures predict Axis-Aligned Bounding Boxes (AABB). Vessels are high-aspect-ratio objects that rarely align perfectly with the horizontal and vertical axes of the image.
When an elongated vessel lies at an oblique angle relative to the image frame, an axis-aligned bounding box encloses substantial background water alongside the target hull. This geometric mismatch lowers the calculated IoU score against ground truth annotations.
As a result, a correct target detection may be reclassified as a false positive during metric evaluation simply because the axis-aligned box fails to capture the object’s spatial orientation.

Oriented Bounding Box (OBB) detectors resolve this issue by predicting an explicit rotation angle. This lack of OBB annotations creates a data trap. You cannot easily convert standard axis-aligned datasets (HBB) into OBB format without losing the precise angular data required to train rotation heads. Conversely, degrading the few available OBB datasets down to HBB to match standard architectures throws away the exact orientation data needed to solve the background inclusion problem in the first place. Furthermore, naively converting an OBB prediction back to an HBB for evaluation or data fusion often introduces excess background space, lowering IoU.
To address this, SkyServe’s recent work introduces a shape-aware OBB-to-HBB conversion method based on ship hull geometry, producing tighter axis-aligned bounding boxes without requiring model retraining:
Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion by Sabhapathy, Dahiya, and Vatsal available on ArXiv, 2026 link
In the paper, we observed superior performance for angle-stratified mean IoU on ShipRSImageNet (10,043 samples) and robustness to object rotation with such a technique.
Port Congestion and Spatial Density Failures
In crowded maritime environments, such as commercial shipping terminals, anchorages, and narrow canals, vessels routinely moor adjacent to one another. Overlapping vessel shadows, crane gantries, and industrial shoreline infrastructure create high spatial clutter.
Standard convolutional backbones aggregate spatial features through successive pooling and downsampling layers, decreasing feature map resolution to expand receptive fields. In dense port clusters, this spatial downsampling blurs the feature boundaries between adjacent targets.
This spatial aggregation leads to three primary failure modes:
Box Merging: Combining multiple adjacent vessels into a single large bounding box.
Target Splitting: Partitioning a single large vessel into multiple smaller predictions based on superstructure breaks.
Omission under Occlusion: Dropping obscured hulls entirely when feature activations fall below detection thresholds.
Existing crowd-detection formulations developed for terrestrial pedestrian counting have not been widely adapted or validated on overhead satellite imagery, leaving port clutter as a primary source of localization error.
5. Operational Constraints: The Edge Hardware Gap
Everything examined in the preceding sections assumes ground-based processing: raw imagery is downlinked to cloud infrastructure where high-capacity detectors operate without strict memory or power limitations. A growing body of work attempts to move detection onboard the satellite itself to reduce downlink latency and bandwidth costs. However, onboard deployment introduces operational constraints that compound sensor and algorithmic errors.
The Skyserve Edge Deployment Case Study
A recent study was presented at MIGARS 2026:
R. V. S, V. A. Belludi, G. Dahiya and V. Vatsal, “Challenges in Vessel Detection for Edge Deployment,” 2026 International Conference on Machine Intelligence for GeoAnalytics and Remote Sensing (MIGARS), Thiruvananthapuram, Kerala, India, 2026, pp. 1-5, doi: 10.1109/MIGARS72065.2026.11688781. link
The study quantifies this hardware gap. We evaluated a YOLO26-based detector trained on a globally diverse PlanetScope dataset (3–4m resolution, augmented with speckle noise, haze, striping, color-tone shifts, and simulated glint).

The experimental results shows us the gap between offline metrics and real-world edge performance:
The Validation Baseline: The unquantized model achieved a training validation mAP50 of 0.90 (with an mAP50-95 of 0.638).
The Test Set Penalty: When evaluated against a held-out test set of 953 images simulating real orbital conditions, mAP50 for the larger variant (YOLO26L) dropped to 0.75.
The Quantization Cliff: Compressing the model to INT8 representation to fit onboard memory limits reduced mAP50 further to 0.66. The smaller YOLO26N variant, chosen for its 5MB memory footprint, plateaued near 0.70 mAP50 across test splits.
The Hardware Acceleration Bottleneck: Quantization is theoretically intended to accelerate inference. However, when deployed on an NVIDIA Jetson TX2i, an edge processor lacking dedicated INT8 hardware tensor cores. The quantized INT8 model executed slower (1.61 FPS) than the unquantized FP16 baseline (1.99 FPS).
This empirical outcome illustrates a fundamental engineering paradox: software compression metrics measured on host development systems do not consistently translate to speed gains on specific target hardware. Onboard deployment compounds limited memory budgets, strict power ceilings, and hardware architectural mismatches, eliminating opportunities for downstream post-processing cleanup.
6. Architectural Evolution and the Path Forward
The Limits of Architectural Scaling
Detection models have evolved through successive paradigms, moving from Histogram of Oriented Gradients (HOG) combined with Support Vector Machines, to deep Convolutional Neural Networks, and subsequently to Vision Transformers.
However, performance improvement curves show diminishing returns when applied to ocean imagery. Architectural scaling yields significant gains where visual features are abundant, objects exhibit complex internal textures, and backgrounds vary across semantic categories. Conversely, these models yield minimal improvements when targets occupy fewer than 10 pixels and background variability is governed by physical oceanography rather than semantic features.
No current architectural breakthrough alters the low-signal, high-ambiguity conditions of marine surface scenes.
Research Directions for Robust Detection
Overcoming current performance plateaus requires moving beyond generic object detection architectures toward domain-specific methods:
Better Datasets Expanding public benchmarks to encompass wider regional variations, dense cloud cover matrices, pixel-accurate ground-truth masks, and balanced small-vessel representation.
Multi-Task Architectures: Constructing unified networks that jointly execute vessel localization, atmospheric cloud masking, orientation regression, and uncertainty estimation calibrated to environmental parameters.
Physics-Aware Modeling: Integrating explicit physical priors into neural network loss functions or feature backbones, including solar illumination geometry, radiative transfer atmospheric models, and sea-surface specular reflection equations.
Target-Specific Hardware Co-Design: Ensuring that quantization and model pruning are benchmarked directly against the target edge computing architecture to prevent throughput regressions.
Conclusion
When evaluating maritime monitoring systems, deployers can reliably depend on the robust detection of large commercial vessels under optimal lighting conditions, the modest gains of modern transformer and CNN backbones over legacy architectures, and performance gains derived from careful training data curation.
Conversely, systems cannot depend on high recall for sub-10-pixel vessels at 5m resolution, cloud-resilient detection in the absence of explicit atmospheric modeling, or precise absolute geolocation derived from optical bounding boxes alone.
If an autonomous monitoring architecture claims universal performance across all vessel size classes and environmental states, operators should request breakdowns by sea state, locations, sensor states, and error distributions. That is where operational reality diverges from marketing numbers.
Citations & Benchmark References
Faster R-CNN: Al-Saad et al., Airbus Ship Detection from Satellite Imagery Using Frequency Domain Learning, SPIE. Available at Strathprints Open Repository.
YOLOv5 / YOLOv8: Waqas et al., Deep Learning-Based Automatic Detection of Ships, Sensors. Available via the National Institutes of Health (NIH).
Oriented Detectors: Prakash et al., Ship Detection in Remote Sensing Imagery for Arbitrarily Oriented Object Detection, Available on arXiv Repository.
Cascaded Pipelines: Fouad et al., Airbus Ship Classification, Detection and Segmentation Using Cutting-Edge Deep Learning, FUJE. Expanded data layout hosted on ResearchGate.
Transformer Architectures: Carion et al. / Remote Sensing Benchmark Studies, Comparative Evaluation of YOLO- and Transformer-based Detectors, Available at MDPI Journals.
RetinaNet on xView: Lam et al., xView: Objects in Context in Overhead Imagery, Available on arXiv Repository.


















