Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight

83d ago · Global · primary source: export.arxiv.org

A new hardware-validated framework allows unmanned aerial vehicles to fly fully autonomously indoors while simulating the visual conditions of a ship at sea, according to research posted to the arXiv preprint server on June 17, 2026 [1][2]. The system, described in a paper by Maneesha Wickramasuriya, addresses a persistent barrier in maritime drone development: validating autonomous flight on actual vessels is expensive, constrained by weather, and carries operational risk [1][2]. The proposed vision-in-the-loop setup renders photorealistic maritime views that are processed onboard the drone by a deep transformer-based monocular pose estimator [2]. Delayed vision measurements are then fused with high-rate inertial measurement unit data through a delayed Kalman filter, producing consistent state estimates for geometric control [2]. The authors report that the framework captures embedded effects such as perception latency, asynchronous updates, and computational limits — factors absent from software-only simulations [2]. Experiments covering autonomous takeoff, trajectory tracking, and landing demonstrated stable closed-loop flight [2]. The work is positioned as a safe, hardware-realistic intermediate stage for developing maritime UAV autonomy before moving to shipboard trials [2]. The paper was submitted to arXiv’s robotics section at 15:18:11 UTC and is available as a 1,890 KB PDF [1]. arXiv, which began in 1991, is an open-access repository of electronic preprints that are moderated but not peer-reviewed; it surpassed two million articles by the end of 2021 and currently receives about 24,000 submissions per month [6]. The platform also hosts arXivLabs, a framework for community-built tools that appear on abstract pages, though new project proposals are temporarily paused while the development team modernizes arXiv’s infrastructure and migrates systems to the cloud [3][4].

applicationresearch-papermodel-releasetool-release

Background sources we checked (7)
  • arxiv.org ↗ Autonomous UAV operations on ships require reliable vision-based relative pose estimation, yet at-sea validation is costly, weather-dependent, and risky. This paper presents a hardware-validated vision-in-the-loop framework that enables fully autonomous indoor flight while emulat…
  • info.arxiv.org ↗ arXiv Labs - arXiv info | arXiv e-print repository Skip to content # arXiv Labs Attention arXiv Users: arXiv Labs is pausing new proposals ## What are arXiv Labs? arXiv Labs are a way for the community to contribute new, useful features to arXiv. These integrations are avail…
  • blog.arxiv.org ↗ arXivLabs: a space for community innovation – arXiv blog arXiv has launched a new, formalized framework enabling innovative collaborations with individuals and organizations. “Members of our community want to contribute tools that enhance the arXiv experience, and we val…
  • info.arxiv.org ↗ arXivLabs: Showcase - arXiv info | arXiv e-print repository ... # arXivLabs: Showcase ... arXiv is surrounded by a community of researchers and developers working at the cutting edge of information science and technology. ... While the arXiv team is focused on our core mission—pr…
  • en.wikipedia.org ↗ arXiv (pronounced as "archive"—the X represents the Greek letter chi ⟨χ⟩) is an open-access repository of electronic preprints and postprints (known as e-prints) approved for posting after moderation, but not peer reviewed. It consists of scientific papers in the fields of mathem…
  • en.wikipedia.org ↗ 14 (fourteen) is the natural number following 13 and preceding 15.…
  • en.wikipedia.org ↗ A large language model (LLM) is a type of machine learning model designed for natural language processing tasks such as language generation. LLMs are language models with many parameters, and are trained with self-supervised learning on a vast amount of text.…

Sources

Spot something wrong? Report an issue