# The AGI Landscape

**The AGI Landscape** $$\Omega$$ is going to push the boundary of artificial general intelligence.

$$
\mathbf{\Omega} = \underset{\theta}{\arg\max}\ \mathcal{AGI}(\theta)
$$

## Concepts

* [Kolmogorov complexity](/)

## Frameworks

<https://github.com/deepmind/pysc2>

<https://pythonprogramming.net/starcraft-ii-ai-python-sc2-tutorial/>

## [Important Papers](/papers-1)

* [Universal Transformers ](/universal-transformers)
* [The Forget-me-not Process](/forget-me-not-process)&#x20;
* [AGI Safety Literature Review](https://arxiv.org/pdf/1805.01109.pdf) : summary of general safety research in agi
* [Out-of-sample extension of graph adjacency spectral embedding](https://www.stat.berkeley.edu/~mmahoney/pubs/levin18a.pdf): consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure for embedding the vertices of a graph into Euclidean space.
* [Alignment for Advanced Machine Learning Systems](https://intelligence.org/files/AlignmentMachineLearning.pdf)
* [Measuring and avoiding side effects using relative reachability](https://arxiv.org/pdf/1806.01186.pdf): introduces a general definition of side effects, based on relative reachability of states compared to a default state, that avoids these undesirable incentives.&#x20;

Nov

1. [The Importance of Sampling in Meta-Reinforcement Learning](http://papers.nips.cc/paper/8140-the-importance-of-sampling-inmeta-reinforcement-learning.pdf)
2. [Inequity aversion improves cooperation in intertemporal social dilemmas](http://papers.nips.cc/paper/7593-inequity-aversion-improves-cooperation-in-intertemporal-social-dilemmas.pdf)

## Books

### Probability

* **R. Durrett** [Probability: Theory and Examples (4th edition)](http://www.amazon.com/Probability-Cambridge-Statistical-Probabilistic-Mathematics/dp/0521765390).
* **P. Billingsley** Probability and Measure (3rd Edition). Chapters 1-30 contain a more careful and detailed treatment of some of the topics of this semester, in particular the measure-theory background. Recommended for students who have not done measure theory.
* **R. Leadbetter et al** [A Basic Course in Measure and Probability: Theory for Applications ](http://www.amazon.com/gp/product/1107652529/)is a new book giving a careful treatment of the measure-theory background.

There are many other books at roughly the same \`\`first year graduate" level. Here are my personal comments on some.

* **D. Khoshnevisan** [Probability ](http://www.amazon.com/gp/product/1107652529/)is a well-written concise account of the key topics in 205AB.
* **R. Bhattacharya and E. C. Waymire** [A Basic Course in Probability Theory](https://www.amazon.com/Basic-Course-Probability-Theory-Universitext/dp/3319479725) is another well-written account, mostly on the 205A topics.
* **K.L. Chung**[ A Course in Probability Theory](http://www.amazon.com/gp/product/1107652529/) covers many of the topics of 205A: more leisurely than Durrett and more focused than Billingsley.
* **D. Williams** [Probability with Martingales ](http://www.amazon.com/Probability-Martingales-Cambridge-Mathematical-Textbooks/dp/0521406056/)has a uniquely enthusiastic style; concise treatment emphasizes usefulness of martingales.
* **Y.S. Chow and H. Teicher**[ Probability Theory: Independence, Interchangeability, Martingales ](http://www.amazon.com/Probability-Theory-Independence-Interchangeability-Martingales/dp/0387406077/). Uninspired exposition, but has useful variations on technical topics such as inequalities for sums and for martingales.
* **R.M. Dudley**[ Real Analysis and Probability](http://www.amazon.com/Analysis-Probability-Cambridge-Advanced-Mathematics/dp/0521007542). Best account of the functional analysis and metric space background relevant for research in theoretical probability.
* **B. Fristedt and L. Gray**[ A Modern Approach to Probability Theory](http://www.amazon.com/Modern-Approach-Probability-Theory-Applications/dp/0817638075/). 700 pages allow coverage of broad range of topics in probability and stochastic processes.
* **L. Breiman**[ Probability](http://www.amazon.com/Probability-Classics-Applied-Mathematics-Breiman/dp/0898712963/). Classical; concise and broad coverage.
* **O. Kallenberg** [Foundations of Modern Probability](http://www.amazon.com/Foundations-Modern-Probability-Its-Applications/dp/0387953132). Quoting an amazon.com reviewer: \`\`.... a compendium of all the relevant results of probability ..... similar in breadth and depth to Loeve's classical text of the mid 70's. It is not suited as a textbook, as it lacks the many examples that are needed to absorb the theory at a first pass. It works best as a reference book or a "second pass" textbook."
* **John B. Walsh** [Knowing the Odds: An Introduction to Probability](http://www.amazon.com/Knowing-Odds-Introduction-Probability-Mathematics/dp/0821885324). New in 2012. Looks very nice -- concise treatment with quite challenging exercises developing part of theory.
* **George Roussas** [An Introduction to Measure-Theoretic Probability](http://www.amazon.com/Introduction-Measure-Theoretic-Probability-Second/dp/0128000422/ref=asap_B00JALV1Z8_1_1?s=books\&ie=UTF8\&qid=1412368880\&sr=1-1). Recent treatment of classical content.
* **Santosh Venkatesh** [The Theory of Probability: Explorations and Applications](http://www.amazon.com/Theory-Probability-Explorations-Applications/dp/1107024471). Unique new book, intertwining a broad range of undergraduate and graduate-level topics for an applied audience.
* **I. Florescu** [Probability and Stochastic Processes](http://www.amazon.com/Probability-Stochastic-Processes-Ionut-Florescu/dp/0470624558/ref=sr_1_2?s=books\&ie=UTF8\&qid=1427746625\&sr=1-2\&keywords=florescu). Very clearly written, and with 550 pages gives a broad coverage of topics including intro to SDEs.
* Jim Pitman has his [very useful lecture notes](http://bibserver.berkeley.edu/205/WorkInProgress/DurrettTOC.html) linked to the Durrett text; these notes cover more ground than my course will! Also some [lecture notes by Amir Dembo ](http://www-stat.stanford.edu/~amir/stat-310b/lnotes.pdf)for the Stanford courses equivalent to our 205AB.

## Reference

1. The `Books`: <https://www.stat.berkeley.edu/~aldous/205B/index.html>, by Professor David Aldous from UC Berkeley.


# 我们的愿景 Our vision

## **中文**

通用人工智能大学\[^1] \[英文: AGI University ]：首家非营利通用人工智能科研教学机构，致力于为拥有通用人工智能理想的人类提供优越环境和丰富资源，研究和开发通用人工智能哲学、算法、框架、平台和应用，分享通用人工智能整体认知和全方位技术进步的阶段性结果，为个人及社会团体提供发展咨询服务，为政府及企事业单位提供通用人工智能方面的政策和治理服务，共同创建人工智能时代人类社会的美好未来。

\[^1]:「 大学之道，在明明德，在亲民，在至于至善。」-- 出自《大学》，取其立意而不着相，意味着我们需要追寻的是至善的通用人工智能。

## **English**

AGI University \[or Artificial General Intelligence University]: The first non-profit AGI research and education institution, dedicated&#x20;

* to provide superior environment and abundant resources for talents dreaming of AGI
* to research and development of AGI philosophy, algorithms, frameworks, platforms and applications
* to share the gradual results of AGI and full-aspects technological advancement
* to provide development consulting services for individuals and social groups and policy and governance services for governments and businesses&#x20;
* to create a better future for human society in the era of artificial intelligence together.


# Papers

We list important papers on AGI as follows:

* [Universal Transformers ](/universal-transformers)
* [The Forget-me-not Process](/forget-me-not-process)&#x20;
* [AGI Safety Literature Review](https://arxiv.org/pdf/1805.01109.pdf) : summary of general safety research in agi
* [Out-of-sample extension of graph adjacency spectral embedding](https://www.stat.berkeley.edu/~mmahoney/pubs/levin18a.pdf): consider the problem of obtaining an out-of-sample extension for the adjacency spectral embedding, a procedure for embedding the vertices of a graph into Euclidean space.
* [Alignment for Advanced Machine Learning Systems](https://intelligence.org/files/AlignmentMachineLearning.pdf)
* [Measuring and avoiding side effects using relative reachability](https://arxiv.org/pdf/1806.01186.pdf): introduces a general definition of side effects, based on relative reachability of states compared to a default state, that avoids these undesirable incentives.&#x20;
* [Asymptotically Unambitious Artificial General Intelligence](https://arxiv.org/pdf/1905.12186.pdf): presents the first algorithm for asymptotically unambitious AGI, where “unambitiousness” includes not seeking arbitrary power; identifies an exception to the Instrumental Convergence Thesis.


# Rationality and intelligence

\[3] S. Russell. Rationality and intelligence. Artificial Intelligence, 94(1–2):57–77, 1997.


# AI safety gridworlds

\[1] J. Leike, M. Martic, V. Krakovna, P.A Ortega, T. Everitt, L. Orseau, and S. Legg. AI safety gridworlds. arXiv:1711.09883, 2017.


# Modeling Friends and Foes

> **Pedro A. Ortega**
>
> **Shane Legg**
>
> DeepMind

**keywords:** AI safety; friendly and adversarial; game theory; bounded rationality.

## $$\alpha$$ 壹

Discovering whether an environment (or a part within) is a friend or a foe is a poorly understood yet desirable skill for safe and robust agents. Possessing this skill is important for a number of situations, including:

1. ***Multi-agent systems***: Some environments, especially in multi-agent systems, might have incentives to either help or hinder the agent \[[1](/ai-safety-gridworlds)]. For example, an agent playing football must anticipate both the creative moves of its team members and its opponents. Thus, learning to discern between friends and foes might not only help the agent to avoid danger, but also open the possibility to solving taks throught collaboration that it could not solve alone otherwise.&#x20;
2. ***Model uncertainty***: An agent can choose to impute “adversarial” or “friendly” qualities to an environment that it does not know well. For instance, an agent that is trained in a simulator could compensate for the innaccuracies by assuming that the real environment differs from the simulated one—but in an adversarial way, so as to devise countermeasures ahead of time \[[2](/concrete-problems-in-ai-safety)]. Similarly, innacuracies might also originate in the agent itself—for instance, due to bounded rationality \[[3](/rationality-and-intelligence), [4](/thermodynamics-as-a-theory-of-decision-making-with-informationprocessing-costs)].

> Typically, these situations involve a *knowledge limitation* that the agent addresses by responding with a *risk-sensitive* policy.

**Contributions:**

1. we offer a ***broad definition of friendly and adversarial behavior***. Furthermore, by varying a single real-valued parameter, one can select from a continuous range of behaviors that smoothly interpolate between fully adversarial and fully friendly.
2. we derive the agent’s (and environment’s) ***optimal strategy*** under friendly or adversarial assumptions. To do so, we treat the agent-environment interaction as a one-shot game with information-constraints, and characterize the optimal strategies at equilibrium.
3. we provide an ***algorithm to find the equilibrium strategies*** of the agent and the environment

We also demonstrate empirically that the resulting strategies display non-trivial behavior which vary qualitatively with the information-constraints.

## $$\beta$$ 贰

intuitive example using multi-armed bandits

game theory is the classical economic paradigm to analyze the interaction between agents. \[[5](/a-course-in-game-theory)]

However, within game theory, the term adversary is justified by the fact that in zero-sum games the equilibrium strategies are maximin strategies, that is, strategies that maximize the expected payoff under the assumption that the adversary will minimize the payoffs.&#x20;

**Problems:**

when the game is not a zero-sum game, interpreting an agent’s behavior as adversarial is far less obvious, since:

1. &#x20;there are is no coupling between payoffs&#x20;
2. the strong guarantees provided by the minimax theorem are unavailable \[[6](/theory-of-games-and-economic-behavior)].

![Figure 1: Average rewards for 4 different 2-armed bandits under 2 different strategies: ](https://570505731-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LKHJDzxLb94Turlb5G-%2F-LLof0Y5TX9oeBHjKziA%2F-LLof2X8sXHwumxyzUgE%2Fimage.png?alt=media\&token=a8e1d006-ca1b-427e-8128-a83c9e93b864)

**explanations:**

(I) a uniform strategy (arms were chosen using 1000 fair coin flips) and (II) a deterministic strategy (500 times arm 1, then 500 times arm 2). Bandits get to choose in each turn (using a probabilistic rule) which one of the two arms will deliver the reward. The precise probabilistic rule used here will be explained later in Section 5 (Experiments).<br>

Two different strategies are:&#x20;

1. The agent plays all 1000 rounds using a uniformly random strategy
2. The agent deterministically pulls each arm exactly 500 times using some fixed rule.

Observations:

1. ***Sensitivity to strategy.*** The two sets of average rewards for bandit A are statistically indistinguishable, that is, they stay the same regardless of the agent’s strategy. This corresponds to the stochastic bandit type in the literature \[[7](/untitled-1), [8](/regret-analysis-of-stochastic-and-nonstochastic-multi-armed-bandit-problems)]. In contrast, bandits B–D yielded different average rewards for the two strategies. Although each arm was pulled approximately 500 times, it appears as if the reward distributions were a function of the strategy.
2. ***Adversarial/friendly exploitation of strategy.*** The average rewards do not always add up to one, as one would expect if the rewards were truly independent of the strategy.
3. ***Strength of exploitation*** : Notice how the rewards of both adversarial bandits (C & D) when using strategy II differ in how strongly they deviate from the baseline set by strategy I. This difference suggests that bandit D is better at reacting to the agent’s strategy than bandit C— and therefore also more adversarial. A bandit that can freely choose any placement of rewards is known as a non-stochastic bandit \[[9](/the-nonstochastic-multiarmed-bandit-problem), [8](/regret-analysis-of-stochastic-and-nonstochastic-multi-armed-bandit-problems)].
4. ***Cooperating/hedging :*** the nature of the bandit qualitatively affects the agent's optimal strategy. A friendly (B) invites the agent to cooperate through the use of predictable policy whereas adversarial bandits (C\&D) pressure the agent to hedge through randomization.

![Figure 2: Reacting to the agent’s strategy in RPS](https://570505731-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LKHJDzxLb94Turlb5G-%2F-LLp2az7G-qPXvji415e%2F-LLp3c4JlQAOzPz-BKlg%2Fimage.png?alt=media\&token=bfc27f16-ff28-4123-b581-cc7c3ee26ebb)

The diagram depicts the simplex of the environment’s mixed strategies over the three pure strategies Rock, Paper, and Scissors, located at the corners.&#x20;

* a. When the agent picks a strategy it fixes its expected payoff (shown in color, where darker is worse and lighter is better).&#x20;
* b. An indifferent environment corresponds to a player using a fixed strategy (mixed or pure).&#x20;
* c. However, a reactive environment can deviate from the indifferent strategy (the set of choices is shown as a ball). A friendly environment would choose in a way that benefits the player (i.e. try playing Paper as much as possible if the player mostly plays Scissors); analogously, an adversarial environment would attempt to play the worst strategy for the agent.

***5. Agent/environment symmetry***. Let us turn the tables on the agent: how should we play if we were the bandit? A moment of reflection reveals that the analysis is symmetrical. An agent that does not attempt to maximize the payoff, or cannot do so due to limited reasoning power, will pick its strategy in a way that is indifferent to our placement of the reward. In contrast, a more effective agent will react to our choice, seemingly anticipating it. Furthermore, the agent will appear friendly if our goal is to maximize the payoff and adversarial if our goal is to minimize it.

This symmetry implies that the choices of the agent and the environment are coupled to each other, suggesting a solution principle for determining a strategy profile akin to a Nash equilibrium \[[5](/a-course-in-game-theory)]. The next section will provide a concrete formalization.&#x20;

## $$\gamma$$ 叄

## $$\delta$$ 肆

## $$\epsilon$$ 伍


# Forget-me-not-Process

Kieran Milan†, Joel Veness†, James Kirkpatrick, Demis Hassabis from DeepMind

Anna Koop, Michael Bowling from University of Alberta

We introduce the Forget-me-not Process, an efficient, non-parametric metaalgorithm for online probabilistic sequence prediction for piecewise stationary, repeating sources. Our method works by taking a Bayesian approach to partitioning a stream of data into postulated task-specific segments, while simultaneously building a model for each task. We provide regret guarantees with respect to piecewise stationary data sources under the logarithmic loss, and validate the method empirically across a range of sequence prediction and task identification problems.


# Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study

In this work, we have demonstrated how techniques from cognitive psychology can be leveraged to help us better understand DNNs. As a case study, we measured the shape bias in two powerful yet poorly understood DNNs - Inception and MNs. Our analysis revealed previously unknown properties of these models. More generally, our work leads the way for future exploration of DNNs using the rich body of techniques developed in cognitive psychology

link: <https://arxiv.org/pdf/1706.08606.pdf>


# Universal Transformers

![](https://570505731-files.gitbook.io/~/files/v0/b/gitbook-legacy-files/o/assets%2F-LKHJDzxLb94Turlb5G-%2F-LKMbZNotp01yP6PdlWk%2F-LKMbrKhR022NT7U-0i4%2Fimage.png?alt=media\&token=36717a09-4cf0-44c6-b6af-4fb6e1080d4e)

Despite these successes, however, feed-forward sequence models like the Transformer fail to generalize in many tasks that recurrent models handle with ease (e.g. copying when the string lengths exceed those observed at training time). Moreover, and in contrast to RNNs, the Transformer model is not computationally universal, limiting its theoretical expressivity. In this paper we propose the Universal Transformer which addresses these practical and theoretical shortcomings and we show that it leads to improved performance on several tasks.

1. Instead of recurring over the individual symb ols of sequences like RNNs, the Universal Transformer repeatedly revises its representations of all symbols in the sequence with each recurrent step.
2. In order to combine information from different parts of a sequence, it employs a **self-attention** mechanism in every recurrent step.
3. Assuming sufficient memory, its recurrence makes the Universal Transformer computationally universal.&#x20;
4. We further employ an adaptive computation time (ACT) mechanism to allow the model to **dynamically adjust** the number of times the representation of each position in a sequence is revised.&#x20;
   1. Beyond saving computation, we show that ACT can improve the accuracy of the model.&#x20;

Our experiments show that on **various algorithmic tasks** and a diverse set of large-scale **language understanding** tasks the Universal Transformer generalizes significantly better and outperforms both a vanilla Transformer and an LSTM in machine translation, and achieves a new state of the art on the bAbI linguistic reasoning task and the challenging LAMBADA language modeling task.


# Graph Convolutional Policy Network

Generating novel graph structures that optimize given objectives while obeying some given underlying rules is fundamental for chemistry, biology and social science research.&#x20;

This is especially important in the task of molecular graph generation, whose goal is to discover novel molecules with desired properties such as drug-likeness and synthetic accessibility, while obeying physical laws such as chemical valency.&#x20;

However, designing models to find molecules that optimize desired properties while incorporating highly complex and non-differentiable rules remains to be a challenging task.&#x20;

Here we propose Graph Convolutional Policy Network (GCPN), a general graph convolutional network based model for goaldirected graph generation through reinforcement learning.&#x20;

The model is trained to optimize domain-specific rewards and adversarial loss through policy gradient, and acts in an environment that incorporates domain-specific rules.&#x20;

Experimental results show that GCPN can achieve 61% improvement on chemical property optimization over state-of-the-art baselines while resembling known molecules, and achieve 184% improvement on the constrained property optimization task.


# Thermodynamics as a theory of decision-making with informationprocessing costs

\[4] P. A. Ortega and D. A. Braun. Thermodynamics as a theory of decision-making with informationprocessing costs. Proceedings of the Royal Society A, 469, 2013.


# Concrete Problems in AI Safety

\[2] D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané. Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565, 2016.


# A course in game theory

\[5] M. J. Osborne and A. Rubinstein. A course in game theory. MIT Press, first edition, 1994


# Theory of games and economic behavior

\[6] J. Von Neumann and O. Morgenstern. Theory of games and economic behavior, 2nd rev. 1947.


# Reinforcement learning: An introduction 1e

\[7] R. S. Sutton and A. G. Barto. Reinforcement learning: An introduction, volume 1. MIT press Cambridge, 1998.


# Regret analysis of stochastic and nonstochastic multi-armed bandit problems

\[8] S. Bubeck and N. Cesa-Bianchi. Regret analysis of stochastic and nonstochastic multi-armed bandit problems. Foundations and Trends in Machine Learning, 5(1):1–122, 2012.


# The nonstochastic multiarmed bandit problem

\[9] P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire. The nonstochastic multiarmed bandit problem. SIAM journal on computing, 32:48–77, 2002.


# Information theory of decisions and actions

\[10] N. Tishby and D. Polani. Information theory of decisions and actions. In Perception-action cycle, pages 601–636. Springer New York, 2011.


# Clustering with bregman divergences

\[11] A. Banerjee, S. Merugu, I. S. Dhillon, and J. Ghosh. Clustering with bregman divergences. Journal of machine learning research, 6(Oct):1705–1749, 2005.


# Quantal Response Equilibria for Normal Form Games

\[12] R. McKelvey and T. Palfrey. Quantal Response Equilibria for Normal Form Games. Games and Economic Behavior, 10:6–38, 1995.


# The numerics of gans

\[13] Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. The numerics of gans. In Advances in Neural Information Processing Systems, pages 1823–1833, 2017.


# The Mechanics of n-Player Differentiable Games

\[14] D. Balduzzi, S. Racaniere, J. Martens, J. Foerster, and T. Tuyls, K. Graepel. The Mechanics of n-Player Differentiable Games. In Proceedings of the International Conference on Machine Learning (ICML), 2018.


# Reactive bandits with attitude

\[15] P. A. Ortega, K.-E. Kim, and D. D. Lee. Reactive bandits with attitude. In Artificial Intelligence and Statistics, 2015.


# Data clustering by markovian relaxation and the information bottleneck method

\[16] N. Tishby and N. Slonim. Data clustering by markovian relaxation and the information bottleneck method. In Advances in neural information processing systems, pages 640–646, 2001.


# Information bottleneck for Gaussian variables

\[17] G. Chechik, A. Globerson, N. Tishby, and Y. Weiss. Information bottleneck for Gaussian variables. Journal of Machine Learning Research, 6(Jan):165–188, 2005.


# Bounded Rationality, Abstraction, and Hierarchical Decision-Making: An Information-Theoretic Optimal

\[18] T. Genewein, F. Leibfried, J. Grau-Moya, and D. A. Braun. Bounded Rationality, Abstraction, and Hierarchical Decision-Making: An Information-Theoretic Optimality Principle. Frontiers in Robotics and AI, 2:27, 2015.


# Risk sensitive path integral control

\[19] B. van den Broek, W. Wiegerinck, and B. Kappen. Risk sensitive path integral control. In Proceedings of the Twenty-Sixth Conference on Uncertainty in Artificial Intelligence, pages 615–622. AUAI Press, 2010.


# Information, utility and bounded rationality

\[20] P. A. Ortega and D. A. Braun. Information, utility and bounded rationality. In International Conference on Artificial General Intelligence. Springer Berlin Heidelberg, 2011. \[21] D. H. Wolpert, M. Harré, E. Olbrich, N. Bertschinger, and J. Jost. Hysteresis effects of changing the parameters of noncooperative games. Physical Review E, 85, 2012. \[22] S. Bubeck and A. Slivkins. The best of both worlds: stochastic and adversarial bandits. In In Proceedings ofthe International Conference on Computational Learning Theory (COLT), 2012. \[23] Y. Seldin and A. Silvkins. One practical algorithm for both stochastic and adversarial bandits. In 31 st International Conference on Machine Learning, 2014. \[24] P. Auer and C. Chao-Kai. An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits. In 29th Annual Conference on Learning Theory, 2016. \[25] M. L. Littman. Friend-or-Foe Q-Learning in General-Sum Games. In Proceedings of the International Conference on Machine Learning (ICML), 2001. \[26] R. Powers and Y. Shoham. New criteria and a new algorithm for learning in multi-agent systems. In Advances in neural information processing systems, pages 1089–1096, 2005. \[27] A. Greenwald and K. Hall. Correlated Q-Learning. In Proceedings of the 22nd Conference on Artificial Intelligence, pages 242–249, 2003. \[28] J. W. Crandall and M. A. Goodrich. Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning. Machine Learning, 82(3):281–314, 2011. \[29] P. Hernandez-Leal and M. Kaisers. Learning against sequential opponents in repeated stochastic games. In The 3rd Multi-disciplinary Conference on Reinforcement Learning and Decision Making, Ann Arbor, 2017. \[30] W. R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25:285–294, 1933. \[31] O. Chappelle and L. Li. An empirical evaluation of Thompson Sampling. In Advances in neural information processing systems, 2011. \[32] C. K. Ling, F. Fang, and J. Z. Kolter. What game are we playing? end-to-end learning in normal and extensive form games. arXiv preprint arXiv:1805.02777, 2018. \[33] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013. \[34] I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.


# Hysteresis effects of changing the parameters of noncooperative games

\[21] D. H. Wolpert, M. Harré, E. Olbrich, N. Bertschinger, and J. Jost. Hysteresis effects of changing the parameters of noncooperative games. Physical Review E, 85, 2012.


# The best of both worlds: stochastic and adversarial bandits

\[22] S. Bubeck and A. Slivkins. The best of both worlds: stochastic and adversarial bandits. In In Proceedings ofthe International Conference on Computational Learning Theory (COLT), 2012.


# One practical algorithm for both stochastic and adversarial bandits

\[23] Y. Seldin and A. Silvkins. One practical algorithm for both stochastic and adversarial bandits. In 31 st International Conference on Machine Learning, 2014.


# An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits

\[24] P. Auer and C. Chao-Kai. An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits. In 29th Annual Conference on Learning Theory, 2016.


# Friend-or-Foe Q-Learning in General-Sum Games

\[25] M. L. Littman. Friend-or-Foe Q-Learning in General-Sum Games. In Proceedings of the International Conference on Machine Learning (ICML), 2001


# New criteria and a new algorithm for learning in multi-agent systems

\[26] R. Powers and Y. Shoham. New criteria and a new algorithm for learning in multi-agent systems. In Advances in neural information processing systems, pages 1089–1096, 2005.


# Correlated Q-Learning

\[27] A. Greenwald and K. Hall. Correlated Q-Learning. In Proceedings of the 22nd Conference on Artificial Intelligence, pages 242–249, 2003.


# Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning

\[28] J. W. Crandall and M. A. Goodrich. Learning to compete, coordinate, and cooperate in repeated games using reinforcement learning. Machine Learning, 82(3):281–314, 2011.


# Learning against sequential opponents in repeated stochastic games

\[29] P. Hernandez-Leal and M. Kaisers. Learning against sequential opponents in repeated stochastic games. In The 3rd Multi-disciplinary Conference on Reinforcement Learning and Decision Making, Ann Arbor, 2017.


# On the likelihood that one unknown probability exceeds another in view of the evidence of two sample

\[30] W. R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25:285–294, 1933.


# An empirical evaluation of Thompson Sampling

\[31] O. Chappelle and L. Li. An empirical evaluation of Thompson Sampling. In Advances in neural information processing systems, 2011.


# What game are we playing? end-to-end learning in normal and extensive form games

\[32] C. K. Ling, F. Fang, and J. Z. Kolter. What game are we playing? end-to-end learning in normal and extensive form games. arXiv preprint arXiv:1805.02777, 2018.


# Intriguing properties of neural networks

\[33] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus. Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.


# Untitled


# Explaining and harnessing adversarial examples

\[34] I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.


# The Landscape of Deep Reinforcement Learning

Deep reinforcement learning (Deep RL) can be said to be one of the hottest topics in artificial intelligence (AI), attracting many outstanding scientists in this field to explore its ability to solve tough real-world problems. Deep RL itself is highly respected by various application fields because of its versatility, from end to end game control, robotic arm control, recommendation systems, and even natural language dialogue systems. However, it is very hard for a machine learning engineer or researcher to keep a good rhythm aligned with the fast and iterative development of deep reinforcement learning. We hope this book can be a guide to help our readers who not only want to know the connections between various deep reinforcement learning algorithms but to get familiar with the practice side of these algorithms.&#x20;

We will first briefly introduce deep learning techniques and applications. Then we review some core concepts of reinforcement learning, explore the combination of deep learning and reinforcement learning, introduce several paradigms of deep reinforcement learning and give some interesting work and applications in the near future. Finally, we outline a quick learning guide of this project-based book for readers.&#x20;

We'll cover the following main topics:

* Deep Learning&#x20;
* Deep Reinforcement Learning
* The general guide for deep RL projects

## Deep learning

Deep learning played a key role in building powerful modern AI systems for many applications and conducting research in various scientific areas. Here we quickly explain basic knowledge about deep learning and some typical applications.&#x20;

### Elements of deep learning

In March 2019, three influential AI scientists: Geoffrey E. Hinton, Yoshua Bengio, and Yann LeCun received the 2018 ACM Turing Award, which is often called The Nobel Prize for Computer Science, due to their contribution to deep learning. Deep learning is a modern version of artificial neural networks, which is a framework for building an intelligent system with the inspirations from psychology and brain science invented since the early age of artificial intelligence. The basic way it works is to use the layer-wised computational models to learn representations of data with multiple levels of abstraction.

Artificial neural networks have been able to achieve approximations of arbitrarily complex continuous functions. This can be seen in Michael Nielsen's book [Neural Networks and Deep Learning](http://neuralnetworksanddeeplearning.com/), chapter 4. Deep learning can take advantage of more hidden layers to enhance the ability to represent the data. From a mathematical view, deep learning is actually a combination of a large number of functions and can be trained by back-propagation algorithm.&#x20;

It has become popular around the world with its transcendental effects in practical applications. The computing device GPUs, produced by Nvidia, which is very crucial computing that made deep learning an efficient way to deal with ImageNet Competition in 2012, now is dominating the training market and becoming a must for deep learning. This is also one of the most important driving forces for building capable reinforcement learning agents.&#x20;

Nowadays, deep learning has swept the fields of speech recognition, image recognition, computer vision, natural language processing, and even video prediction. There are two main network architectures- convolutional neural networks and recurrent neural networks - have completed the space and time perspectives to model problems.&#x20;

Although deep learning may still have some disadvantages for the interpretability of the models, this technology has already been the default choice for many areas today, like computer vision, natural language processing, social network analysis, biology, quantum mechanics and astronomy, etc.

Here is a list of building blocks for constructing various recent deep learning programs:

1. Transformers
2. Residual connections
3. Attention mechanisms
4. Generative adversarial networks, GANs for short
5. Variational auto-encoders
6. Graph convolutional networks

### Applications of Deep Learning

#### Arts

Since the very beginning of deep learning, researchers had already designed several interesting algorithms for generating arts like paintings with styles from some famous painters like Claude Monet's impressionism, Van Gogh's post-impressionism, and other styles. This kind of work has been done by a deep learning algorithm named Neural Style Transfer. There is a famous mobile application named Prisma dedicated to producing high-resolution images of different styles for customers.&#x20;

In October 2018, 'Edmond de Belamy, from La Famille de Belamy', a painting generated by a deep learning algorithm, sold for $432,500 at auction.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/a8821f66-5296-4885-bb2a-92aaffabe904.jpg)

Nvidia researchers recently developed a GAN named GauGAN, which is capable of turning doodles into photorealistic Landscapes. You can draw doodles first, and get the landscape drawings after the processing of the GAN.  The example below shows an extraordinary waterfall generated by the left side drawing.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/dims.jpg)

#### Biology

DeepMind published their work on predicting the 3D structure of a protein based solely on its genetic sequence by cooperating with experts from various fields, including structural biology and physics. Their cutting-edge system is named AlphaFold, and is making significant progress on one of the core challenges in biology. The left figure below shows the dynamics of the transforming process of a structure, just using the gradient descent method in a normal deep learning training.

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/WeChat-Image_20190324165749.png)![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/7888523e-92f9-40b4-948a-046b37caada2.png)

The basic procedure is described in the right figure, we input the protein sequence into a neural network, using distance prediction and angle predictions to provide enough information to construct a score of measurement so that we can optimize to get an optimal structure of the protein. We could consider this as a normal deep learning application for vision.&#x20;

#### Natural language processing

OpenAI released their NLP model in March 2019. The model named GPT-2 is trained with an unsupervised style. The following example shows the text generated by GPT-2 based on the human written text. Given the beginning of a sentence or two, this model can generate a variety of different styles of news, novels, and other similar styles, the content looks very realistic.  You can see one example in the table as follows.&#x20;

| **SYSTEM PROMPT (HUMAN-WRITTEN)**                | In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               |
| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **MODEL COMPLETION (MACHINE-WRITTEN, 10 TRIES)** | <p>The scientist named the population, after their distinctive horn, Ovid’s Unicorn. These four-horned, silver-white unicorns were previously unknown to science.</p><p>Now, after almost two centuries, the mystery of what sparked this odd phenomenon is finally solved.</p><p>Dr. Jorge Pérez, an evolutionary biologist from the University of La Paz, and several companions, were exploring the Andes Mountains when they found a small valley, with no other animals or humans. Pérez noticed that the valley had what appeared to be a natural fountain, surrounded by two peaks of rock and silver snow.</p><p>...</p><p>However, Pérez also pointed out that it is likely that the only way of knowing for sure if unicorns are indeed the descendants of a lost alien race is through DNA. “But they seem to be able to communicate in English quite well, which I believe is a sign of evolution, or at least a change in social organization,” said the scientist.</p> |

However, unusually, OpenAI researchers decided not to release the data for the training model, nor for the pre-trained parameters of the largest model, because they believe that such a powerful model is at risk of malicious abuse. Their arguments that there may be risks and the model is better not to be released have caused a big wave of rendering, and the researchers in the machine learning and natural language processing community have had intense discussions. To our understanding, there are many aspects needing to be enhanced or reconstructed so that the text generated could be more consistent according to some core meaning as a real writer writes.&#x20;

#### Medical

Google Brain made some progress in medical diagnosis. The machine learning system used for image diagnosis of diabetic retinopathy can be equivalent to a professionally certified ophthalmologist. If early diabetic retinopathy is not detected, there are now 400 million people at risk of blindness. But in many countries, the number of professional ophthalmologists is too small to perform the necessary checks, and this technology will help ensure that more people receive appropriate checks. Research in other areas of medical imaging is investigating the potential of using machine learning to predict other medical tasks. We believe that machine learning can improve the treatment experience for both physicians and patients, both in terms of quality and efficiency.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/e7132508-ddbd-45eb-afd2-8f58a2fb4804.jpg)

In 2018, Stanford AI lab cooperated with a medical team; they made a deep learning system to help with the monitoring of patients. Postures and actions can be detected so that the system can prevent dangerous situations happening.

### Summary

To summarize, deep learning is a force that pushes the understanding and utilization of artificial intelligence across sectors of human life and society. One thing to remember is that for a deep learning program, we, the human, give the program the instructions to start, to transfer, and to stop. But we believe an intrinsic autonomous system could be a more natural way to build real artificial intelligence, that is to say, reinforcement learning could be a crucial framework for building artificial general intelligence.&#x20;

## Deep reinforcement learning

The deep learning model is simple. In just a few dozen lines of code, you can solve the system that took a lot of effort before you can design it. Therefore, various application areas (speech, image, visual, natural language understanding, etc.) now tilt resources to deep learning, where we do not judge the unintended adverse consequences of this, from an optimistic point of view Deep learning really revitalizes the field of artificial intelligence. Of course, how to divert people's passion is very important. I believe that after a while, everyone will find a suitable path to develop.

Success in the field is often very difficult to achieve in the field of science. There are several important number theory and graph theory problems that have been carried forward through generations of scientists and continue to advance in the work of predecessors. After finishing the history, let’s take a look now. The most exciting progress. We introduce the paradigm of deep reinforcement learning and related algorithms. See what exactly is the most critical factor. The key is actually how we apply these techniques to solve problems - suitable problem modeling, solutions Improvement.

The reason why it is not practical before intensive learning is that it is difficult to deal with these situations effectively in the face of excessive state or action space. The examples often seen are relatively simplified scenarios. The emergence of deep learning allows people to deal with real Problems such as the dramatic increase in visual recognition accuracy to ImageNet  top-5 error rate dropped to less than 4%, now speech recognition has really become more mature, and is widely used, and all current commercial speech recognition algorithms No one is not based on deep learning. These are all indicating that deep learning can be the basis of some practical applications. Now the research and application of deep reinforcement learning are basically aimed at the above problems.ImageNet: **ImageNet** is an important dataset for computer vision research. This dataset is designed based on [WordNet](http://wordnet.princeton.edu/) hierarchy. Each node of the hierarchy is related to hundreds and thousands of images. Researchers test their deep learning algorithms on different tasks of ImageNet.  To read more about it, please visit <http://www.image-net.org/>

Thanks to the work on building reinforcement learning environments, comparison of research work turns out to be easier than ten years ago. OpenAI gym has a much impressive effect on developing more and more algorithms and finding applications of reinforcement learning.&#x20;

### Reinforcement learning

Reinforcement learning is a way to simulate human learning process by simplifying the situation in which humans make decisions. There are two independent strings of development, one is animal behavior research, the other is optimization control. Finally, this process was formalized into a Markov Decision Process (MDP) by Richard Bellman. Since then, many scientists have expanded this model to form a relatively complete system—often called approximate dynamic programming. See the dynamic programming of MIT professor Dimitri P. Bertsekas - *Dimitri P. Bertsekas, Dynamic Programming and Optimal Control, Vol. II, 4th Edition: Approximate Dynamic Programming*.

Although we have lots of reinforcement learning methods, it is very difficult to apply them to large scale real-world problems, first of all, the existence of the curse of dimensionality makes it difficult to efficiently find the optimal policy or compute the optimal action value. In addition, the ideas contained in deep learning—greedy algorithms, dynamic programming, approximation, etc.—are the most critical parts of the algorithm, and they are the most used of these methods. Here is an overview diagram of reinforcement learning introduced by David Silver in his Reinforcement Learning course:

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/1.png)

Core concepts of a reinforcement learning system are:

1. Environment. The environment can be fully observable or partially observable based on the complexity of the problem and related factors. &#x20;
2. Reward signal. Reinforcement learning is based on the reward hypothesis: any goal can be formalized as the outcome of maximizing a cumulative reward. That is the basic foundation for applying reinforcement learning techniques to your problems. &#x20;
3. Agent. The agent is the key element for reinforcement learning, containing state, policy, value function (probably) and Model (optionally).

### Classical methods of deep reinforcement learning

Now, let's consider a bunch of deep reinforcement learning algorithms and divide them into several clusters based on the ways or perspectives to tackle reinforcement learning problems. Due to the rapid pace of researching and practicing by researchers and engineers, it is hard to get a complete view of this area. So, we'll make a visualization of these interesting methods. You can find the classes, names, and publishing time in the figure from the [awesome-deep-rl](https://github.com/tigerneil/awesome-deep-rl/) project .&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/d545795e-3bdc-404e-b9fa-aae605a8d3c5.png)

As we know, deep reinforcement learning methods are a combination of deep learning and reinforcement learning. So typically we can divide the following types for deep reinforcement learning:

* Value-based methods
* Policy gradient methods
* Explorations in deep reinforcement learning
* Actor-critic methods
* Model-based methods
* Multi-agent reinforcement learning
* Meta-reinforcement learning
* Hierarchical reinforcement learning
* Inverse reinforcement learning

#### Value-based methods

The symbol of the rise of deep reinforcement learning, DQN, was first proposed by V. Mnih in 2013. After he jointed DeepMind, their team gave a better model by getting rid of some issues in original DQN. We often called it Nature version DQN.&#x20;

Many researchers followed the work of Nature DQN, especially from DeepMind. During the last several years, the performance of value-based methods has been improved a lot on a large portion of tasks in Atari Game environment.&#x20;

In some random environments, Q-learning is very poor. The culprit is overestimations of large action values. These overestimates are due to the fact that Q learning uses the largest action value as the estimate of the maximum expected action value with positive bias. There is another way to approximate the maximum expected action value for any random variable set. The so-called double estimation method will be underestimated rather than overestimated. Apply this idea to Q-learning. A double Q-learning method, a policy-free reinforcement learning method is obtained. This algorithm can converge to the optimal strategy and perform better than the Q-learning algorithm under certain settings.

Double Q-Network is the result of merging Q-learning and deep learning. In some Atari games, DQN itself is also subject to estimation. Through the introduction of double Q-learning, it can handle large-scale function approximation problems. The final algorithm not only reduces the overestimation of the observations but also has a fairly good performance in some games.

Other algorithms based on DQN make efficient use of samples in experience replay buffer or change network architectures like Dueling networks. Rainbow proposed in Mar 2018 seems like a final version for DQN, it also contains other great ideas.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/rainbow.png)

#### Policy gradient methods

Policy gradient methods belong to another typical way to solve reinforcement learning problems. Strictly speaking, this class of methods approaches getting the optimal policy in a direct manner.&#x20;

Although there are several books related to reinforcement learning, the introduction of the strategy gradient part is not enough. The existing reinforcement learning textbook does not give enough guidance on how to use the function approximation; basically, it is focused on discrete The realm of state space. Moreover, existing RL textbooks do not adequately describe non-derivative optimization and strategic gradient methods, and these techniques are quite important in many tasks.

The strategy gradient algorithm is optimized by gradient descent. That is, by repeatedly calculating the noise estimate of the desired return gradient of the strategy, and then updating the strategy according to the gradient direction. This method is more advantageous than other RL methods (such as Q-learning) because it is possible to directly optimize the quantity of interest - the expected total return of the strategy. This type of method has long been considered to be less practical due to the high variance of the gradient estimate. Until recently, the work of Schulman et al. and Mnih et al. demonstrated the successful application of the policy gradient method on the difficult control problem.

#### Explorations in deep reinforcement learning

Exploration is an important part of reinforcement learning agents to get rich experience and learn about the environments more broadly. As we know, there is another part named exploitation and the two are some bit of competitors. We call this competitive situation "Exploration-Exploitation Dilemma".&#x20;

Now we give more intuition for exploration. Generally speaking, we could have three types of exploration strategy: optimistic, posterior sampling, and information gain.&#x20;

1. Optimistic exploration always tries to choose highly uncertain actions, just as a famous Chinese saying "rare things are precious" says. So when we calculate the frequency of some action or an action-state pair, this frequency can be used as a bonus to distinguish the importance of the rare states. In deep reinforcement learning, researchers have already considered several paths to make this possible: CTS-based pseudocounts, hash-based pseudocount, and exemplar models exploration are typical pseudo-counts exploration.&#x20;
2. Posterior sampling is to make more accurate exploration by utilizing the idea from Bayesian learning. Since posterior sampling algorithms always remedy the probability distribution after each sampling, after many steps learning, we can have a stronger belief for each action, with a smaller variance as an indicator. Techniques like Thompson sampling or Bootstrap DQN can be good for posterior sampling.&#x20;
3. Information Gain Exploration considers information as a part of the states. Through traversing new states to acquire new information, then make those states that can get more information gain the ideal states. Information theory can help here, and we need better ways to approximate analytic solutions. We can use VIME: variational information maximizing exploration to get as much information about the environment from each interaction as possible.&#x20;

#### Actor-critic methods

The goal as a guide for better methods is to make stable and data efficient algorithms. The most famous deep actor-critic algorithm is DDPG. DDPG is the deep learning version of the deterministic policy gradient method, which uses the idea of DQN to transform the DPG. DDPG can solve the reinforcement learning problem in continuous action space. In those continuous action tasks, DDPG gives a stable performance and in different environments. No changes are required on the top. In addition, DDPG found the solution of the Atari game in less time than DQN learning in all experiments, which is about 20 times the performance. Given more simulation time, DDPG may solve more difficult problems than the current Atari game. The future direction of DDPG should be to use a model-based approach to reduce the number of rounds of training because model-independent reinforcement learning methods usually require a lot of training to find a reasonable solution.

DDPG is actually an Actor-Critic structure that combines information from both strategy and value functions. Both Actor and Critic use deep neural networks for approximation.&#x20;

After DDPG, we have ACER and Reactor algorithms. Researchers hope to eliminate these flaws, such as TRPO and ACER, through constraints or other optimization strategy size methods. These methods all have their own trade-offs. The ACER method is much more complicated than the PPO method. It requires additional code to modify the off-policy and refactor buffers, but it is only one better than the PPO on the Atari benchmark. Although useful for continuous control tasks, TRPO is not easily compatible with algorithms that share parameters between strategy and value functions or auxiliary losses, that is, those that are important for solving Atari and other visual inputs algorithm.

#### Model-based methods

The goal of model-based reinforcement learning methods, mentioned above, is to improve the stability and data efficiency without explicitly modeling for the environment. Here the model-based methods mainly utilize the information about the environments. If we get the learned model, we can use it to plan optimal actions.&#x20;

In contrast, model-based reinforcement learning methods can be learned with significantly fewer samples. This type of learning method uses a learned environment dynamic model that can perform policy optimization. Learning dynamic models can be done in a sample-efficient manner because they are trained using standard supervised learning techniques, allowing the use of non-strategic data.&#x20;

The low sample complexity of the method in Model-Based Reinforcement Learning via Meta-Policy Optimization makes it suitable for real-world robots. For example, it can find the optimal strategy for a high-dimensional and complex four-dimensional motion world based on real data within two hours. Note that the amount of data required to learn such a strategy using a model-free approach is 10 to 100 times higher, and the researchers know that previous model-based methods did not achieve similar performance in such tasks.&#x20;

In March 2019 a complete model-based deep RL algorithm proposed by Google Brain and UIUC researchers based on video prediction models with a novel architecture yielded the best results in standard benchmark tasks. It may tell us model-based methods close to one of the core mechanisms of human's fast learning.&#x20;

#### Multi-agent reinforcement learning

Multi-agent reinforcement learning is a natural extension of (deep) reinforcement learning for many reasons. First, when a single agent interacting with the environment, the information exchange just between two clear separated counterparts. They have a totally different role. That is so restricted for modeling a real-world problem. Second, after we solve really challenging problems in only one agent setting, we want to find more difficult problems, obviously, the prediction for the development of a complex system with more agents turns out more difficult. Third, if we want to devise general intelligent agents, the final question is to make them act in a real society like us.&#x20;

As we know, much great work had been done since the early age of human history on researching the dynamics of a society and predictions of future actions of the individuals. Complex networks and game theory are two main areas contributing many ideas for interactions within a giant system with many individuals competing and cooperating with each other. Clustering of the nodes is one of the problems in complex network analysis while Nash equilibrium is a great milestone in game theory. It is actually the premise of a multi-agent system.&#x20;

Normally, the methods for a single agent don't work well in multi-agent setting, therefore we need to find more proper methodology and framework to fit the problems in multi-agent systems. Multi-agent deep reinforcement learning is a vital area for building efficient and effective algorithms to help us understand the dynamics and properties of a networked agent sets.&#x20;

Compared to training a strategy to solve all actions in an environment, a multi-agent perspective can be helpful to decompose the problem more naturally: F1-racing, antennas placing and traffic control. Multi-agent reinforcement learning can be seen as a more scalable way of learning: First, decomposing the actions of a single agent and observing into multiple simpler agents not only reduces the dimensions of the input and output of the agent but also effectively increases the amount of training data for each time step.  Then, dividing the action and observation space of each agent can produce effects similar to the introduction of time abstraction methods. Time abstraction has been used to improve learning efficiency under a single agent setting. Finally, good decomposition can also lead to the learning of strategies that are more likely to migrate in multiple variant environments, such as a single super-agent may match a particular environment.&#x20;

#### Meta-reinforcement learning

Thanks to OpenAI gym, robotic arms and other environments, we can train our agents for solving much more tasks than before. Having those large number of tasks actually could give us a huge advantage. Meta-learning just is a way to utilize the tasks to figure out a general way to learn better.&#x20;

Meta-learning could reduce the number of samples needed to train deep reinforcement learning algorithms since meta-learning can meta-learn a faster reinforcement learner when dealing with new tasks. Actually, meta-learning can have various types like learning RNNs with experience or learning representations, even learning optimizers.&#x20;

Researchers found that meta-learning could help agents explore more intelligently, avoid useless actions or find the right features faster. Meta-learning algorithms can be automated by automating the process of task design. For example, unsupervised meta reinforcement learning can effectively accelerate reinforcement learning procedures while no need for manual task design, exceeds the performance of learning from scratch and showed competitive performance to that use hand-specified task distributions.

#### Hierarchical reinforcement learning

If the reward is delayed and sparse, the reinforcement learning algorithm may suffer from poor sample efficiency. Hierarchical reinforcement learning (HRL) enables agents to learn time-extended actions at multiple levels of abstraction in an efficient and automated manner. HRL allows agents to learn strategies that belong to different time scales in parallel.

Multi-level hierarchies have the advantages to accelerate learning in sparse reward tasks because they can classify problems into a set of short-term sub-problems.

#### Inverse reinforcement learning

So far we all deal with rewards, however, most problems we face in the real world don't have a proper reward function. So we need a method to learn reward function from some expert's behavior. For inverse RL, we should try to find a reward function that matches some history of an agent's behavior or policy.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/83f09d40-7465-4f99-b23e-daad0d0be84d.png)

Ng and Russell first proposed algorithms for inverse RL. Based on the assumption that actions always chooses the best possible action for its reward function, we try to estimate a reward function that could have generated behavior like this. &#x20;

When we build such an algorithm, the input is environment dynamics e.g., an MDP without a reward function and optimal behavior e.g., the full policy or trajectories and the output is the inferred reward function.&#x20;

There are mainly two problems of inverse RL:&#x20;

1. Many reward functions can be useful under most observations of behavior
2. Sometimes the observed behavior is not optimal. Our optimal policy assumption is too strong.&#x20;

More recently discussions in Value misalignment in AI safety show the importance of utilizing inverse RL to find a way to keep AI safe.&#x20;

There are many other branches in deep reinforcement learning, like option based reinforcement learning, multi-task reinforcement learning, or distributional reinforcement learning, etc. We just stop here to focus on methods relating to our projects in this book.&#x20;

### Applications of deep reinforcement learning

Since reinforcement learning is a powerful and general enough framework to model various situations, we can see lots of applications in many fields. And because of the power of deep learning, the deep reinforcement learning can be designed to match the real world needs of various domains.&#x20;

As we know, the first success in this area is the DQN agent for playing Atari games. People saw the potential of deep reinforcement learning, therefore big companies and research institutes like DeepMind, OpenAI, Google Brain, UC Berkeley, CMU as well as many others all put their resources into this area to try to achieve better AI techniques.&#x20;

Besides games, we can also find applications in autopiloting drones, traffic control, self-driving cars, robotics control, electric business, and even computer security.&#x20;

#### Games

Humans like playing games all the time, for example, Go is an ancient game needing the players to do calculation and reasoning tasks and electric games have different weights for the usage of various aspects of human intelligence. Now, games are used by researchers to show the effectiveness and performance of our algorithms.&#x20;

**AlphaGo**

**AlphaGo** is the first artificial intelligence robot to defeat human professional Go players to win the world championship in Go. It was developed by a team led by Google's DeepMind company.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/4cf974e1-708d-49ac-82e2-70d638761b53.jpg)

In March 2016, AlphaGo and Go World Champion and professional Go player of 9 ran rank, Lee Sedol, carried out the Go-Man Wars, AlphaGo winning with a total score of 4-1. At the end of 2016 and the beginning of 2017, the program was a "master" on the Chinese chess website ( Master). For the registered account, there are dozens of Go players in China, Japan, and Korea, and there is no one defeat in 60 consecutive games. In May 2017, at the Wuzhen Go Summit in China, it competed with World Go Champion Ke Jie. It won the match with a total score of 3 to 0. In the world of Go, AlphaGo is recognized as better than top human Go professionals. In the World Professional Go ranking published on the GoRatings website, its score has surpassed Keji, the number one player in the ranking.

On May 27, 2017, after the human-machine battle between Ke Jie and Alpha Go, the Alpha Go team announced that Alpha Go would no longer participate in the Go game. On October 18, 2017, the DeepMind team announced the strongest version of Alpha Go, named AlphaGo Zero.&#x20;

They used a new way to do reinforcement learning in which AlphaGo Zero starts to teach itself. At first, the system just has a neural network of zero knowledge about the game of Go, then it plays games against itself, by combining this neural network with a search algorithm. AlphaGo Zero can be used in other areas to learn to find new knowledge about a complex system. Their work published in Nature explained that deep learning related techniques can help with scientific research in other domains. &#x20;

**OpenAI Five**

In 2017, **OpenAI** beat the "Dota 2" world's top players in a 1 to 1 solo at the Dota2 TI finals. In June 2018, OpenAI announced that their AI bot beat amateur human players in the 5 v 5 team competition and was able to beat the top professional team after the plan. The heart of the machine compiles the contents of OpenAI's blog.

Through self-confrontation learning, OpenAI Five is equivalent to playing 180 years of games every day. In training, it uses 256 GPUs and 128,000 CPU cores to train using the Proximal Policy Optimization method, which was augmented on the solo Dota2 system we built last year. When we use a separate LSTM for each hero, the model learns identifiable strategies without human data. This suggests that intensive learning can produce large-scale but acceptable long-term planning, even without fundamental advances.&#x20;

#### Self-driving cars

**Wayve**

The UK company Wayve designed the first-ever autonomous car that works with the help of reinforcement learning

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/dc8dc104-6727-4ad5-91c0-481bf1b0c792.jpg)

A deep reinforcement learning based approach helped them to teach the car how to drive in just 15-20 minutes. The system is supported by a deep neural network that has 4 convolutional layers and 3 fully connected layers.&#x20;

#### Electronic Commerce

Here we present some use case in e-commerce companies like Alibaba. The following figure is from paper <https://arxiv.org/pdf/1803.00710.pdf>

It shows the typical search session in TaoBao. A user starts a session from a query and has multiple actions to choose, including clicking into an item description, buying an item, turning to the next page, and leaving the session.

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/taobao.png)

**Product recommendation**&#x20;

In the recommended scenario, Alibaba uses deep reinforcement learning and adaptive online learning to build a decision engine through continuous machine learning and model optimization and analyzes massive user behavior and tens of billions of commodity features in real time. Every user quickly finds the product and improves the matching efficiency between the person and the product. The algorithm performance index is increased by 10% and 20%.&#x20;

**Customer service**

In intelligent customer service, customer service robots such as Ali Xiaomi, as agents of the delivery engine, need to have decision-making capabilities. This decision is not based on the direct benefit of a single node, but a relatively long-term process of human-computer interaction. The interaction between consumers and platforms is regarded as a Markov decision process, using an intensive learning framework. Establish a loop system in which consumers interact with the system, and the system's decision is based on maximizing the process benefits to achieve a dynamic balance between the system and the user.

#### Robotics

Robotics is another area where deep reinforcement learning can be applied.

**Robotic arms**

&#x20;The following figure shows the scene of collective training of robotic arms using deep reinforcement learning algorithms.&#x20;

![](https://packt-type-cloud.s3.amazonaws.com/uploads/sites/3589/2019/03/6a74cadf-9e84-4b9d-8dff-494b5bf1e4b5.jpg)

In this experiment, each robotic arm was started by practicing the opening skill in the specific position and direction of the door displayed before the instructor. As it gets better and better at performing tasks, the instructor begins to change the position and orientation of the door to slightly exceed the current function of the policy, but it is not difficult to completely fail. This allows the robot to gradually improve their skill level over time and expand the range of situations they can handle. The combination of manual guidance and trial and error allows the robot to collectively learn how to open the door in just a few hours. Since the robotic arms were trained on doors that looked different from each other, the final policy was successful on a door with no robotic arms on the handle that it had seen before.

## The general guide for deep RL projects&#x20;

In this section, we present the general guide for our book through which you can get the experience of applying deep reinforcement learning as much as possible. Typically, we design with a step-by-step explanation of how to solve a problem. However, it can also be seen as a well-generalized guide for most of the problems you have already encountered or will face in the future.&#x20;

### General steps

Step 1: Make your understanding of the problem clear enough

Step 2: Design a (PO)MDP for this problem

Step 3: Try methods you have already learned to see the results

Step 4: Find tricks and theoretical proof to deal with sub-problems based on the results

Step 5: Combine tricks and rigorous techniques to solve the problem

### General tools

Here we introduce important tools for building deep reinforcement learning agents. For a detailed introduction, you can read the corresponding part in the appendix.&#x20;

#### TensorFlow 2.0&#x20;

TensorFlow is an open-source machine learning library for research and production. TensorFlow now is the most popular deep learning framework with 123,589 stars and 73,108 forks on its Github project in early 2019.

At the end of 2018, Google TensorFlow team announced the 2.0 agenda for TensorFlow. We assume you have some experience in using TensorFlow to implement deep learning models like CNN, ResNet, LSTM, GRU even Attention mechanisms. Compared with TF1.0, TF 2.0 changed a bit on different aspects.&#x20;

We want to introduce Google Colab, an easy way to learn and use TensorFlow. We can use colab to practice deep reinforcement learning algorithms using free computing resources. In the following chapters, we'll go through the usage of colab.

You can go to <https://www.tensorflow.org/> to check more about TensorFlow, and <https://colab.research.google.com/notebooks/welcome.ipynb> to know how to use colab.&#x20;

#### Ray

Ray is a high-performance distributed execution framework providing at large-scale machine learning and reinforcement learning applications. It came from UC Berkeley RISE lab, which leads teams building many successful frameworks for big data science and machine learning.&#x20;

RISE is **Real-time Intelligence with Secure Explainable decisions**, which is important for building systems in AI era, "*Sensors are everywhere.* ... *AI is for real.* ... *The world is programmable."* is their basis for this lab. Ray is the most important framework for building distributed AI systems in the future. Therefore we will introduce Ray to our readers to gain a basic understanding of its usage and try to make a multiagent reinforcement learning agents using it.&#x20;

Go to <https://rise.cs.berkeley.edu/projects/ray/> to know more about Ray.

#### OpenAI gym

As we mentioned above, OpenAI gym has been pushing the research of deep reinforcement learning since its very beginning. That is the great contribution of OpenAI actually. "**Gym is a toolkit for developing and comparing reinforcement learning algorithms.** " We can use the OpenAI gym to design new environments for different problems and integrate with any numerical computation library, such as TensorFlow or PyTorch as well as distributed computing framework, Ray.

You can check gym on <https://gym.openai.com/docs/>

### Projects

The projects we will build with the tools above is as follows:&#x20;

* **Developing Grid Environment**: In this project, we will build up one of the most classical reinforcement learning tasks, grid world. And it is also the benchmark for deep reinforcement learning. We will solve this task by different reinforcement learning algorithms - dynamic programming, Monte Carlo, temporal difference and policy/value iteration. Through this single tasks, readers will learn all the useful algorithms including tabular and deep reinforcement learning.&#x20;
* **Playing Atari Games using Improvised deep Q-learning methods**: In this project, we will start to use the Gym environment and train an agent to play Atari games. Atari games are a collection of interesting video games and its video image input requires us to combine CNN with our RL algorithms. DQN and other techniques will be introduced to build up this agent.
* **Building continuous deep RL agent to control Mujoco robots**: In this project, we will learn how to train an agent to control robots in a simulated environment. Mujoco is a robot simulator with Gym wrapper. Mujoco task is one of the typical continuous control reinforcement learning problems. We will solve this task by several policy gradient RL algorithms such as TRPO and PPO. We will also solve this task by evolutionary strategy RL algorithms.&#x20;
* **Building powerful RL agents to play Montezuma’s Revenge**: In this project, we will learn how to train an agent to solve the Montezuma’s Revenge Project. To solve this, we will introduce effective methods like Go-Explore, hierarchical RL and imitation learning. This hard problem can be considered as the sign of the exploration power of methods becoming competitive for solving real-world problems.&#x20;
* **Solving the general-soldier control problem with Multi-Agent RL**: This is a project for controlling multiple agents in a competitive and cooperative environment. We will introduce multi-agent system concepts and basic methods. And we introduce useful designs for constructing a framework to solve the general-soldier control problem.
* **Building a Deep RL dialogue model with ACER**: In this project, we will build a Deep RL dialogue model. We will introduce the development of dialogue generation, especially using RL techniques and give an implementation of the agent that can learn from interactions and generate reasonable results. Finally, we discuss in-depth algorithms like ACER that can generate more interactive responses and generate more natural conversations in dialogue simulation.
* **Building an agent for playing RTS game Starcraft II**: Starcraft II is a classic RTS game. Players need to have better control both in the strategic layer and tactic layers. Recently there is much progress in RTS game AI. We use this project to test the RL control algorithms in our book. So that readers can practice with the methods to see the potential of RL in game playing.&#x20;
* **Building a self-driving agent with deep RL**: Self-driving problems can be solved by Deep RL methods. DeepDrive is an interesting self-drive environment. We use this project to test the RL control algorithms in our book. So that readers can practice with the methods to see the potential of RL in the self-driving cars.

Through each project, you will get the corresponding deep reinforcement learning methods used for tackling the problems related to that project. The difficulty of each project grows step by step, including more and more advanced and complicated algorithms. The techniques you learned can be utilized to solve real-world problems.&#x20;

## Summary

So far, we have overviewed these fascinating topics on deep reinforcement learning and the arrangement of this book. Now let us jump into the first project to get familiar with how to make your first workable deep reinforcement learning agent with Python and TensorFlow 2.0.&#x20;


# 用因果影响图建模通用人工智能安全框架

**Ramana Kumar, DeepMind**\
**Translated by Xiaohu Zhu, University AI**\
\
**我们写了一篇论文，将用来设计安全通用人工智能（AGI）的各种框架（例如，带有奖励建模的强化学习，合作式逆强化学习 CIRL，辩论 debate 等）表示为因果影响图（CID），以帮助我们比较框架并更好地理解相应的智能体激励机制。**\
\
**我们很乐意收到评论，特别是关于**\
**1. 介绍的框架是否可以被准确表示？**\
**2. CID表示有用吗？**\
**3. 我们没有包含的框架建模成这种模型有用吗？**\
\
**论文的摘要：安全的通用人工智能系统（AGI）的提议通常在框架层面进行，规定了如何训练所提议系统的组件并相互交互。在本文中，我们使用因果影响图来模拟和比较最有希望的 AGI 安全框架。图显示了框架的优化目标和因果假设。统一的表示可以让我们轻松地比较框架及其假设。我们希望这些图可以作为主要 AGI 安全框架的一个易接受和可视化的介绍。**\
\
[**本文对齐论坛地址**](https://www.alignmentforum.org/posts/HE5DL6XeomYxFab74/modeling-agi-safety-frameworks-with-causal-influence-1?fbclid=IwAR3dppEJjDITl-PpAAXdJnlSjrj9xXUB6b0faxXJypnQLI0M3F2lYYCiSNU)<br>


# test


# Measuring and avoiding side effects using relative reachability

Victoria Krakovna, Laurent Orseau, Miljan Martic, Shane Legg

from DeepMind&#x20;

## Abstract&#x20;

How can we design reinforcement learning agents that avoid causing unnecessary disruptions to their environment? We argue that current approaches to penalizing side effects can introduce bad incentives in tasks that require irreversible actions, and in environments that contain sources of change other than the agent. For example, some approaches give the agent an incentive to prevent any irreversible changes in the environment, including the actions of other agents. We introduce a general definition of side effects, based on relative reachability of states compared to a default state, that avoids these undesirable incentives. Using a set of gridworld experiments illustrating relevant scenarios, we empirically compare relative reachability to penalties based on existing definitions and show that it is the only penalty among those tested that produces the desired behavior in all the scenarios.


