The objective of this project is to perform a set of experiments on a model whose purpose is to dynamically generate an action for a ship in the midst of a route in the sea, such that it can follow the current most optimal route to its goal detination, given its start destination. The objective is not just to devise a model that performs the best - though that is a crucial objective - but the goal is also to examine what works, what doesn't and why it is so.
I have used the Natural Earth 50m land database Land data: https://www.naturalearthdata.com/downloads/50m-physical-vectors/ to help identify land, sea and coast points.
- The global weather conditions change every second.
- The agent should generate actions dynamically over each timestep rather than only 1 path at the starting of the journey. This is, as per me, equivalent to generating entire paths every second or dynamically changing paths, because, for the current timestep, the action that the ship takes, whether an entire path is visible or whether only that action is visible, is same. Therefore, in my opinion, the agent will be able to perform similarly well if the agent is trained to generate a single action rather than to generate an entire path (keeping performance optimal of course) and it will also be computationally cheaper.
This framework doesn't just consist of the model - it consists of several other components that are just as essential as the model.
For reference, the "main" code file that you could run is parallel_journey.py or PJ_curriculum.py or reinforce.py, depending on the experiment.
(this has a huge caveat which I found later, which serves as the basis for Experiment 2). Consider the flowchart given below:
The routing agent is trained in a dynamic geographic environment where the ship makes a decision at every timestep based on its current location, destination, and surrounding weather conditions.
The overall interaction follows the reinforcement learning loop:
The process is repeated until the ship reaches the destination or the maximum number of steps is exceeded.
Each journey is defined by a fixed starting coordinate and destination:
where:
-
$(\phi)$ = latitude -
$(\lambda)$ = longitude
Multiple test routes are defined to evaluate the learned policy under different geographic conditions, including:
- Short routes
- Medium routes
- Long routes
- Equatorial crossings
- International dateline crossings
- Northern hemisphere routes
- Southern hemisphere routes
During an episode, the start and goal remain fixed while the ship position changes.
The environment contains a time-varying weather field:
At each timestep, the weather simulator can update the environmental conditions:
The ship therefore does not make all routing decisions using a single static weather map. Instead, the state at each timestep contains information about the weather surrounding the ship's current position.
The intended interaction is:
This allows the agent to adapt its route as weather conditions evolve.
Implementation note: In the current version, the
weather.update()call inside the timestep loop is commented out. Therefore, the weather field is currently initialized once rather than updated after every step. The dynamic formulation above represents the intended architecture.
At every timestep, a new state is constructed from these types of information:
- Local radial weather information
- Distance to the destination
- Distance travelled from the starting point
- Direction toward the destination
The state is represented as:
where (W_t^{radial}) represents the local weather surrounding the ship and (\beta_t) is the bearing from the current ship position toward the destination.
Instead of providing the complete global weather grid to the neural network, weather is sampled around the current ship position.
For a set of predefined radii
Each sampled point contains local weather information such as wind speed and wind direction.
The resulting radial representation is flattened before being passed to the policy:
This gives the agent a localized view of the weather around its current position.
The distance between the ship and destination is calculated using geodesic distance on the Earth's surface.
Let:
be the current ship position and:
be the destination.
The geodesic distance is:
where (D(\cdot,\cdot)) represents the Earth's surface distance.
The distance is normalized before being provided to the agent:
where approximately (20,000) km represents the characteristic maximum distance scale of the Earth's surface.
The distance travelled from the original starting location is also included in the state.
Let:
be the fixed starting point.
Then:
and the normalized value is:
Including both distances allows the agent to distinguish between:
- How far it still needs to travel
- How far it has already travelled
The direction from the ship toward the destination is represented by a bearing:
Because angular values wrap around at (2\pi), directly providing the angle can introduce a discontinuity.
For example:
but numerically these values appear far apart.
To avoid this problem, the bearing is encoded using sine and cosine:
Therefore, the direction component of the state is:
This provides a continuous representation of direction.
The final state provided to the policy is:
Thus, the agent receives both local environmental information and global positional information.
The policy receives the state:
and produces a continuous heading action:
The raw action is transformed into a heading angle:
The heading is then combined with the goal direction to bias movement toward the destination:
where
The resulting
Given the current position:
the ship moves according to its selected heading.
For a predefined movement step (\Delta):
Latitude is constrained to the valid geographic range, while longitude is wrapped around the globe.
The resulting position is:
After the action is executed, the environment evaluates the resulting position.
The basic transition is:
where:
-
$(s_t)$ = current state -
$(a_t)$ = selected heading -
$(r_t)$ = reward -
$(s_{t+1})$ = resulting state -
$(d_t)$ = episode termination indicator
If the resulting position lies on land:
the movement is rejected and a large negative reward is assigned:
The episode is not immediately terminated in the current implementation, allowing the agent to attempt another direction.
If:
the ship is allowed to move to the new position.
The episode terminates when the ship reaches sufficiently close to the destination.
If:
where (R_{goal}) is the goal tolerance determined by the environment resolution, then:
and a large goal-completion reward is provided.
Otherwise:
The episode is also terminated when the maximum allowed number of steps is reached:
The primary reward encourages progress toward the destination.
Let:
and:
Then the distance-progress reward is:
A positive value means that the ship moved closer to the destination, while a negative value means that it moved farther away.
The reward is normalized:
The intended reward structure can incorporate additional environmental costs:
where:
-
$(C_t^{weather})$ = cost associated with unfavorable weather -
$(C_t^{collision})$ = collision penalty -
$(R_{goal})$ = destination completion bonus -
$(\alpha,\beta)$ = weighting coefficients
The current implementation primarily uses the distance-progress component together with the land-collision penalty and goal-completion bonus.
For each episode, the ship is reset to the same predefined start and goal coordinates.
The agent then repeatedly performs:
until the episode terminates.
During the episode, the following information is collected:
These trajectory samples are then used by the PPO update.
The main idea behind this experiment is that here an RL model is trained on "k" parallel journeys.
Now, if we train an RL agent to generate action on only one journey (1 set of start and end coordinates), then it will only learn to generate action for that path. Then if you want to make it generate actions for another set of start and end points you would need to train it again. So the deployment scenario would be that every time the user introduces a new journey the model parameters would need to be changed (model trained again) which would be computationally expensive.
To prevent this, I defined the deployment goal to be such that every time a user gives its "state" - journey, weather etc the agent which is already trained will run (no change in params - they will just be used) and it will generate a single action (here, it is the degree that the ship has to turn. The same model can be extrapolated to output degree along with direction, aka velocity, by changing the number of output neurons in the neural network and introducing speed effect in the error function.)
To satisfy this deployment condition, the agent's performance needs to be independent of start and end points. Therefore, I am training the agent on k different journeys simultaneously, where the start and end points are part of the input to every prediction. The exact flow is described in the chart below.
Key components are explained below:
The routing agent uses an Actor-Critic PPO (Proximal Policy Optimization) architecture. The actor learns a probability distribution over continuous heading actions which are angles that the ship has to turn (theta), while the critic estimates the value of the current state.
Given the current state (s_t), the policy network outputs the parameters of a Gaussian action distribution:
The standard deviation is obtained as:
with the log standard deviation clipped to:
The raw action is then sampled from the Gaussian policy:
The log probability of the sampled action is stored for the PPO update:
The critic simultaneously estimates the value of the current state:
The raw action is transformed into a valid angular heading using a hyperbolic tangent:
This constrains the base heading to:
The angle is then wrapped to the interval ([-\pi,\pi)):
Finally, the goal direction is incorporated as a directional bias:
where (d_t^{goal}) represents the direction from the ship toward the destination.
Thus, the actor does not directly output the final geographic heading. Instead, it learns a stochastic base direction, which is subsequently combined with the direction toward the goal.
The action-selection process can be summarized as:
The function returns:
- Selected heading
$(\theta_t)$ - Log probability
$(\log\pi_\theta(a_t^{raw}|s_t))$ - State value
$(V_\phi(s_t))$ - Raw action
$(a_t^{raw})$
After collecting a rollout, the policy and value networks are updated using PPO.
The calculated advantages are normalized across the collected samples:
where:
-
$(\mu_A)$ is the mean advantage -
$(\sigma_A)$ is the standard deviation of the advantages -
$(\epsilon=10^{-8})$ provides numerical stability
This improves the stability of the policy update.
For each stored action, the updated policy calculates a new log probability:
The probability ratio between the new and old policies is:
In implementation, this is computed in log-space:
PPO prevents excessively large policy updates by clipping the probability ratio.
The unclipped objective is:
The clipped objective is:
The PPO policy objective is:
Since PyTorch optimizers minimize losses, the implementation uses the negative of this objective:
This clipping constrains how much the new policy can deviate from the policy that generated the collected trajectories.
The critic predicts:
and is trained toward the estimated return:
using mean squared error:
The entropy of the Gaussian policy distribution is calculated as:
Higher entropy corresponds to greater exploration.
The total PPO loss combines the policy loss, value loss, and entropy regularization:
where:
- (L_{policy}) = PPO clipped policy loss
- (L_{value}) = critic/value loss
- (H) = policy entropy
- (c_v) = value loss coefficient
- (c_H) = entropy coefficient
The gradients of the combined actor-critic loss are then clipped to a maximum norm of (0.5) before the optimizer update.
The implementation uses Generalized Advantage Estimation (GAE) to estimate how much better or worse an action performed compared with the critic's expectation.
For each timestep (t), the temporal-difference error is:
where:
-
$(r_t)$ = reward at timestep (t) -
$(V(s_t))$ = value estimate at timestep (t) -
$(V(s_{t+1}))$ = next-state value estimate -
$(\gamma)$ = discount factor -
$(d_t)$ = terminal indicator
The factor
prevents bootstrapping from the next state when the journey has terminated.
Starting from the end of the trajectory, GAE is calculated recursively:
where:
-
$(A_t)$ = advantage estimate -
$(\gamma)$ = reward discount factor -
$(\lambda)$ = GAE smoothing parameter
In the implementation:
and
The recursion is performed backwards through the trajectory:
This allows the advantage at each timestep to incorporate information from multiple future rewards while controlling the bias-variance trade-off through (\lambda).
Once the advantages have been calculated, the target return for the value network is obtained as:
The function therefore returns two quantities:
and
The overall training process can therefore be summarized as:
After collecting a rollout:
followed by:
where (\theta) represents the actor parameters and (\phi) represents the critic parameters.
While REINFORCE.py did not manage to train the model so well because there was a crucial drawback, parallel_journey.py managed to train the agent. At the end of the training, the agent was able to complete 9/12 journeys (300 episodes). The detailed summary is provided in results.txt. While this isn't perfect, it proves that this approach has the potential to work significantly more if we use more approaches such as curriculum learning or adding milestone based rewards which will encourage the agent to complete more journeys.
I have plotted the exact paths that the agent has taken for these different joureys.
Start: (13.0, -95.0) | Goal: (-72.0, -13.0)

Start: (25.0, -146.0) | Goal: (-28.0, -129.0)

Start: (-66.0, 179.0) | Goal: (-66.0, -18.0)

Start: (-75.0, 168.0) | Goal: (-41.0, -162.0)

Start: (40.0, -73.0) | Goal: (79.0, 51.0)

Start: (35.0, 129.0) | Goal: (-7.0, 137.0)

Start: (-34.0, 25.0) | Goal: (40.0, -66.0)

Start: (-33.0, -141.0) | Goal: (81.0, 84.0)




