Architectural Improvements
Recall the architecture of a Deep-Q-Network (DQN):
Some notes on the above:
The input layer takes some representation of the state, learns some "features" in the hidden layers, and outputs the Q-value estimates with one output neuron corresponding to each possible action.
The learning process tries to find the weights that are best predictive of the experienced Q-values, largely motivated by Temporal Difference Learning (TDL).
Like the Exact / Approximate Q-learning methods we've seen previously, DQNs are known as Value-based methods since they attempt to learn a value function (in this case, \(Q(s,a)\)) in order to (hopefully) derive \(\pi^*\).
Value-based methods thus use an implicit policy whose actions are derived from the learned values: $$\pi(s) = argmax_a Q(s,a)$$
Early on, however, we saw some reasons to prefer alternatives to value-based methods -- what were those reasons, and what techniques did we attempt instead?
Value-based methods spent a lot of effort refining the accuracy of the values, when sometimes much earlier, the optimal policy could have already been extracted from them! We turned to policy-gradient methods like policy-iteration instead!
Policy-based algorithms like policy-iteration learned \(\pi(s) = a\) directly, WITHOUT learning the value function.
In fact, there are translations of policy-based methods like what we did in policy-iteration to the deep-learning space as well -- take a look at:
The REINFORCE Algorithm
Proximal Policy Optimization (PPO)
That said, there were still some issues with policy-based methods like we discussed with policy-iteration... what were they?
Because these are primarily hill-climbing approaches, might find policy that is simply local max rather than global, leading to high variance in quality of derived policy.
Is there no hope to combine these techniques to address the flaws of one another?!
Actor-Critic Methods
Of course there is, what a dumb lead-up that would have been!
To garner some intuition from a domain with which every computer scientist should be familiar: athleticism!
Intuition 1: The best athletes have coaches to critique their performance and provide guidance on how they performed following some game / training.
This maps to our domain in that:
We can think of the athlete-coach relationship as being two learners in a single RL agent: an actor choosing moves (policy-based) and a critic who can evaluate them (value-based).
The biggest difference is that in online-RL, the coach will need to learn at the same time as the athlete!
As each get better over time, the coach learns to better evaluate the athlete and the athlete learns to prefer actions that match the coach's suggestions.
This last bullet suggests one difference we'll need to think about: in policy iteration, we learned a deterministic policy \(\pi(s) = a\), but how could this be implemented in an architecture that shouldn't remember every state \(s\)?
Intuition 2: The best athletes won't always act deterministically given the same state.
...but they *will* prefer superior actions to those that aren't obvious gaffs.
Idea: can transform the deterministic policy learned in policy-iteration (\pi(s) = a) to a stochastic policy that can also be updated more numerically through gradient descent like we saw in logistic regression: $$\pi_{\theta}(a | s) = P(a | s; \theta)$$ (the probability of choosing action \(a\) from state \(s\) given the current set of weights / parameters \(\theta\)).
Implementation: can mimic the policy / target networks of the DQN tutorial but adapt for actor-critic instead!
Two models, each helping each other learn:
Actor (Policy) Network with weights \(\theta\) trying to learn the optimal policy \(\pi_{\theta}(a | s)\).
Inputs: the vectorized state
Outputs: \(\pi_{\theta}(a | s)~\forall~a\)
Critic (Value) Network with weights \(w\) trying to learn the Q-values \(Q_{w}(s,a)\)
Inputs: the vectorized state \(s\) and the action space from the Actor network: \(\pi_{\theta}(a | s)\)
Outputs: \(Q_{w}(s, a)\) for the provided (s, a) pair
The training process proceeds in the typical phases of policy-iteration: (1) policy evaluation (i.e., sample gathering) and (2) policy improvement.
In Actor-Critic policy-evaluation, the actor network simply selects actions by sampling from its output action distribution \(\pi(a | s)\), generating many episodes of format: \(\langle s_t, a_t, s_{t+1}, r_t \rangle\)
In Actor-Critic policy-improvement, both Actor and Critic networks update their weights according to their gradients, defined as follows:
Actor Gradient: for updating weights \(\theta\) and for learning rate \(\alpha\), log-likelihood of each action choice \(log~\pi_{\theta}(a_t|s_t)\), and Critic network value function \(Q_{w}(s_t, a_t)\): $$\Delta \theta = \alpha \nabla_{\theta} log~\pi_{\theta}(a_t|s_t)~Q_{w}(s_t, a_t)$$ Notably: this scales higher-likelihood, higher-Q-value actions as becoming more probable.
Critic Gradient: for updating weights \(w\) and for learning rate \(\beta\), uses the Temporal Difference Learning approach with:
$$difference = [R(s_t, a_t, s_{t+1}) + \gamma Q_w(s_{t+1}, a_{t+1})] - Q_w(s_t, a_t)$$ Notably: all Q-value estimates are coming from the current set of weights \(w\) and outputs from the Critic network, and thus plugged into a gradient: $$\Delta w = \beta \cdot difference \cdot \nabla_w Q_w(s_t, a_t)$$
Different implementations of Actor-Critic will interweave the policy evaluation / data gathering and policy improvement at different times, e.g., some may perform improvement after every step taken whereas others may create a Replay Buffer to sample from only after many episodes have been gathered.
And there you have it! Actor-critic in a (very dense) nut shell!
For further reading, there are many actor-critic implementations you can examine, including (for you to look-up):
DDPG
A2C
A3C
Input Structuring
No matter your choice of Deep-RL architecture, there are still many design decisions to be made even at the onset: for instance, how to best represent / vectorize the state!
Here's an example for why an "omniscient" view of the state that has as much information as possible may not be desirable:
Pacman with Traps! In this variant of Pacman, the goal is the same (eat all pellets, avoid all ghosts), but with the potential to fall into fiery pits of doom that are stationary terminals, e.g., a simple map like the following:
In the above example, pretend you're the learning agent just seeing the vectorized version of the maze: because both Pacman and Ghost move between transitions, it will be difficult to determine if the -1 received as the reward was because the -1 indicating the ghost moved down onto the pellet (wrong) or because the 1 indicating Pacman moved onto the trap (right).
In other words, it can be difficult to learn what part of the changing state your agent actually has controlled!
One possible approach is to simplify the state into an embodied sensor paradigm wherein the input vector necessarily corresponds to your agent's actions.
An example implementation: raycast sensors that simply count the number of tiles of distance from your agent to entities of interest, e.g., in the following, showing 2 raycast sensors for distance to walls vs. ghosts.
Repeat the above for each maze entity like a ghost raycast, pellet, etc. to encode a state space relative to Pacman's position!
Warning: although this may simplify learning, it does also risk over-simplifying the state, which can miss key observations to finding a better policy!
As usual, experimentation is required!
Modern Research
Given that Deep-Q-Learning has seen a lot of popularity, there are still *many, many* avenues to explore. Here are a selection of some fun ones with decent amounts of promise.
Intrinsic Curiosity
In what will soon be a problem you will greatly appreciate (if it isn't one already), on your final project, getting Pacman to sample distant state-actions can be a baffling ordeal:
What technique did we see earlier, in the context of Exact Q-Learning, that might help with this scenario?
Exploration functions! Idea: bias unfamiliar transitions to have an optimistic Q-value inflated by some exploration bonus.
It'd be nice if we could adapt that into the deep-RL domain!
Some work attempts to do just that!
Intrinsic Curiosity Modules (ICMs) are those added to the learning process while training Deep-RL models such that they add bonuses to the perceived reward during training when a transition is unfamiliar, encoded as being *unpredictable*
In this diagram:
Being in state \(s_t\) and taking action \(a_t\) from the policy network, the ICM guesses what the next state would look like, \(\hat{s}_{t+1}\).
-
Compare the "distance" or "surprise" between the predicted next state \(\hat{s}_{t+1}\) and the actually observed one, \(s_{t+1}\).
Since states can be vectorized, this "distance" can make use of standard vector similarity metrics!
The greater the distance / surprise, create a greater *intrinsic* reward bonus \(r_t^i\) that is added to the environment-specified extrinsic reward \(r_t^e\) for the replay buffer's total reward \(r_t\) for that episode.
You can read more in this paper! Deep reinforcement learning with intrinsic curiosity module based trajectory tracking control for USV
Inverse Reinforcement Learning
Now that we've seen some powerful ways of learning state-feature representations through deep Q-learning, we should address another piece of the RL puzzle that might be hard to define.
What might be another facet of RL that could be difficult for a programmer to define, and that we might try to learn instead?
The reward function!
This is no easy task, and is often ill advised, since the reward function specifies precisely what we want our agent to do!
Still, we should hint at some scenarios whereby we might want to learn the reward function rather than provide it:
Watching an expert play a video game and then inferring just *what* rewards they're trying to get?
Learning how to walk by watching someone else walk.
If you want to go down a trippy rabbit hole: look up "Mirror Neurons" in the brain and how they're believed to help infants learn.
Critically, we see a bit of a trend in the sparse examples above:
Inverse Reinforcement Learning is the process of inferring a reward function from expert demonstrations.
"Givens" |
Traditional RL |
Inverse RL |
|---|---|---|
Environment |
States and actions |
States and actions |
Model |
Transitions and Rewards (in fully-specified) |
Transitions (in fully-specified) and some set of expert actions \(\{(s_1, a_1), (s_2, a_2), ...\}\) sampled from \(\pi^*(s)\) |
Goal |
Learn \(\pi^*(s)\) |
Learn \(R(s, a, s') \Rightarrow \) Learn parameters of some linear reward function or neural net... then use that to learn \(\pi^{*}(s)\) |
In words: the objective is to learn the reward the expert is trying to maximize.
A fascinating exploration that is a *really* hard problem and hasn't seen a lot of traction, but for which causal inference might have some things to say WRT probability of necessity and sufficiency of certain features to be observed in states in which the expert acts!
Transfer Learning
The number one challenge with anything involving neural networks: transportability -- being able to take training in one environment and use that same network / agent in a separate environment with key differences.
This is something humans do quite easily, as we've mentioned numerous times in this course, but to motivate it, here's a small example (once more, a game -- can you tell why I like RL?)
Take the old Atari game that has traditionally been used in many Transfer Learning domains: Montezuma's Revenge:
Some notables in the (already rich, but hilariously low-rez) frame shown:
Even if you've never played Montezuma's Revenge, you probably see the key and understand that it's used to open something -- a useful first objective to acquire.
If you made the insight above, you've likely just experienced transfered learning, since knowledge about keys in the past (both in real life and many other games) has transferred into this one without you needing to take a single step.
Transfer Learning is the ability to take learned behavior in one environment and translate it to reduce the need for exploration in another.
Generally, this is a very difficult and open problem, but there are several key avenues in current pursuit:
Forward Transfer: train in 1 environment, transfer that training to another
Multi-task Transfer: train on several environments with commonalities, transfer to another
Meta-Learning: learn to learn from several environments with commonalities
Meta-learning, in particular, has seen much traction in recent years, but there is still much to be done here!
If you look up Causal Transportability, there's a lot to say here -- perhaps the future is in concerting the two!
Further Reading
There are MANY more active areas of RL research, here are just a few that you can examine further!
Self-play
Model-based algorithms like: Dyna, Dreamer, MuZero
Shameless Plug: If you find any of the above interesting, come to the follow-on graduate course Agentic AI in Game Development to gain applied practice in many of the techniques above!
Reinforcement Learning - Concluding Remarks
What a wild first half of the semester it has been! Thus far we've:
Examined some ways to model the sequential decision tasks that, perhaps, you and I respond to on a daily basis.
Seen some neat algorithms for both online and offline MAB + MDB solving, and appreciated some of their strains, weaknesses, and strengths.
Taken a good, hard, look in the mirror and asking what rewards we respond to as humans just trying to navigate the day O_O
Fun Videos
Here are some great links you should check out when you have the time! We'll review a few in class together.
The Next State
That said, as powerful as these tools are, we're just starting to scratch the surface of what makes intelligent beings... well... intelligent!
If we as humans *only* responded to reward signals, we would not be nearly the successful species that we are... there's something more behind our cognitive capacities that goes beyond our animalistic tendencies.
So, join me in Part 2, as we begin to think about thinking, and then think some more about how to encode our thinking for the next generation of intelligent agents!