Tutorials for reinforcement learning in PyTorch and Gym by implementing a few of the popular algorithms. [IN PROGRESS]
-
Updated
Oct 23, 2020 - Jupyter Notebook
Tutorials for reinforcement learning in PyTorch and Gym by implementing a few of the popular algorithms. [IN PROGRESS]
Proximal Policy Optimization(PPO) with Intrinsic Curiosity Module(ICM)
A collection of Reinforcement Learning implementations with PyTorch
CellTRIP, Inferring virtual cell environments using multi-agent reinforcement learning for spatiotemporal trajectory interpolation, imputation, and perturbation
Generalized Advantage Estimation (GAE-Lambda) exponentially weighted temporal difference calculator
Generalized Advantage Estimation (GAE-Lambda) exponentially weighted temporal difference calculator
Phasic-Policy-Gradient
Modular Implementation of Proximal Policy Optimization (PPO) is a policy gradient reinforcement learning algorithm introduced by OpenAI in 2017. It's designed to be a simpler, more stable, and more sample-efficient alternative to previous policy gradient methods like A3C and TRPO (Trust Region Policy Optimization).
An implementation from the state-of-the-art family of reinforcement learning algorithms Proximal Policy Optimization using normalized Generalized Advantage Estimation and optional batch mode training. The loss function incorporates an entropy bonus.
Generating reasonably optimal routes for ships and dynamically deciding direction considering weather and distance factors using RL
Recurrent Policies for Handling Partially Observable Environments
PyTorch implementation of Proximal Policy Optimization for continuous pendulum control, featuring GAE, clipped objectives, training curves, and value-function visualization.
Deep Reinforcement Learning from mathematical foundations to PyTorch implementations: rigorous proofs, step-by-step derivations of objectives and gradient estimators, and reproducible experiments on Gymnasium and MuJoCo benchmarks spanning policy gradients, actor-critic methods, value-based learning, and continuous control.
From-scratch PPO (PyTorch) on a PyBullet reaching task - GAE bootstrap verified against a hand-derived reference
Example TRPO implementation with ReLAx
Example PPO implementation with ReLAx
To associate your repository with the generalized-advantage-estimation topic, visit your repo's landing page and select "manage topics."