Back to results
Bibliographic record · Consultation and access
Preprint

Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)

Natan, Avraham; Stern, Roni; Kalech, Meir · arXiv (Cornell University) · 2017

Resource page
Quick overview. Review the resource’s basic details, then access the content using the main button. This page shows only the information needed to identify, cite, and open the work.

Resource access

Open the content from the main option or choose another available source.

OpenAlex OpenAlex Works
Entrar por OpenAlex
Main access

Resource page

Resource reference page. Full text availability has not been automatically confirmed.
Open resource

Summary

Descripción general del contenido del recurso.

Due to the safety risks and training sample inefficiency, it is often preferred to develop controllers in simulation. However, minor differences between the simulation and the real world can cause a significant sim-to-real gap. This gap can reduce the effectiveness of the developed controller. In this paper, we examine a case study of transferring an octorotor reinforcement learning controller from simulation to the real world. First, we quantify the effectiveness of the real-world transfer by examining safety metrics. We find that although there is a noticeable (around 100%) increase in deviation in real flights, this deviation may not be considered unsafe, as it will be within > 2m safety corridors. Then, we estimate the densities of the measurement distributions and compare the Jensen-Shannon divergences of simulated and real measurements. From this, we show that the vehicle’s orientation is significantly different between simulated and real flights. We attribute this to a different flight mode in real flights where the vehicle turns to face the next waypoint. We also find that the reinforcement learning controller actions appear to correctly counteract disturbance forces. Then, we analyze the errors of a measurement autoencoder and state transition model neural network applied to real data. We find that these models further reinforce the difference between the simulated and real attitude control, showing the errors directly on the flight paths. Finally, we discuss important lessons learned in the sim-to-real transfer of our controller.

How to cite

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

Natan, A, Stern, R, & Kalech, M. (2017). Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper). arXiv (Cornell University). https://doi.org/10.4230/oasics.dx.2024.16

MLA

Natan, Avraham, et al. Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper). arXiv (Cornell University), 2017. https://doi.org/10.4230/oasics.dx.2024.16.

Chicago

Natan, Avraham, Roni Stern, and Meir Kalech. 2017. Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper). arXiv (Cornell University). https://doi.org/10.4230/oasics.dx.2024.16.

Harvard

Natan, A, Stern, R. and Kalech, M. 2017, Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper), arXiv (Cornell University), available at: https://doi.org/10.4230/oasics.dx.2024.16 [Accessed 8 Aug. 2026].

Share and print

Save the record, copy its permanent link, or print it as a PDF.

Export reference

You can export the record in common formats for use in a reference manager.

Resource details

Bibliographic information to help confirm that this is the correct material.

Title
Diagnosing Non-Intermittent Anomalies in Reinforcement Learning Policy Executions (Short Paper)
Author / contributors
Natan, Avraham; Stern, Roni; Kalech, Meir
Publisher
arXiv (Cornell University)
Publication year
2017
Language
English

Subjects

Explore related resources through these subjects.

Copied