<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1 plus MathML 2.0 plus SVG 1.1//EN" "http://www.w3.org/2002/04/xhtml-math-svg/xhtml-math-svg.dtd">
<html xmlns="http://www.w3.org/1999/xhtml">
  <head>
    <meta http-equiv="Content-Type" content="application/xhtml+xml; charset=utf-8"/>
    <title>Project-Team:SEQUEL</title>
    <link rel="stylesheet" href="../static/css/raweb.css" type="text/css"/>
    <meta name="description" content="Research Program - Decision-making Under Uncertainty"/>
    <meta name="dc.title" content="Research Program - Decision-making Under Uncertainty"/>
    <meta name="dc.subject" content=""/>
    <meta name="dc.publisher" content="INRIA"/>
    <meta name="dc.date" content="(SCHEME=ISO8601) 2015-01"/>
    <meta name="dc.type" content="Report"/>
    <meta name="dc.language" content="(SCHEME=ISO639-1) en"/>
    <meta name="projet" content="SEQUEL"/>
    <!-- Piwik -->
    <script type="text/javascript" src="/rapportsactivite/piwik.js"></script>
    <noscript><p><img src="//piwik.inria.fr/piwik.php?idsite=49" style="border:0;" alt="" /></p></noscript>
    <!-- End Piwik Code -->
  </head>
  <body>
    <div class="tdmdiv">
      <div class="logo">
        <a href="http://www.inria.fr">
          <img style="align:bottom; border:none" src="../static/img/icons/logo_INRIA-coul.jpg" alt="Inria"/>
        </a>
      </div>
      <div class="TdmEntry">
        <div class="tdmentete">
          <a href="uid0.html">Project-Team Sequel</a>
        </div>
        <span>
          <a href="uid1.html">Members</a>
        </span>
      </div>
      <div class="TdmEntry">Overall Objectives<ul><li><a href="./uid3.html">Presentation</a></li></ul></div>
      <div class="TdmEntry">Research Program<ul><li><a href="uid15.html&#10;&#9;&#9;  ">In Short</a></li><li class="tdmActPage"><a href="uid18.html&#10;&#9;&#9;  ">Decision-making Under Uncertainty</a></li><li><a href="uid29.html&#10;&#9;&#9;  ">Statistical analysis of time series</a></li></ul></div>
      <div class="TdmEntry">Application Domains<ul><li><a href="uid36.html&#10;&#9;&#9;  ">Recommendation systems in a broad sense</a></li><li><a href="uid44.html&#10;&#9;&#9;  ">Spoken dialog systems</a></li><li><a href="uid45.html&#10;&#9;&#9;  ">Adaptive/learning systems more generally</a></li><li><a href="uid46.html&#10;&#9;&#9;  ">Prediction in general</a></li></ul></div>
      <div class="TdmEntry">
        <a href="./uid48.html">Highlights of the Year</a>
      </div>
      <div class="TdmEntry">New Software and Platforms<ul><li><a href="uid56.html&#10;&#9;&#9;  ">Function optimization</a></li></ul></div>
      <div class="TdmEntry">New Results<ul><li><a href="uid59.html&#10;&#9;&#9;  ">Decision-making Under Uncertainty</a></li><li><a href="uid66.html&#10;&#9;&#9;  ">Statistical analysis of time series</a></li><li><a href="uid68.html&#10;&#9;&#9;  ">Statistical Learning and Bayesian Analysis</a></li><li><a href="uid70.html&#10;&#9;&#9;  ">Applications</a></li></ul></div>
      <div class="TdmEntry">Bilateral Contracts and Grants with Industry<ul><li><a href="uid74.html&#10;&#9;&#9;  ">Bilateral Contracts with Industry</a></li><li><a href="uid76.html&#10;&#9;&#9;  ">Bilateral Grants with Industry</a></li></ul></div>
      <div class="TdmEntry">Partnerships and Cooperations<ul><li><a href="uid80.html&#10;&#9;&#9;  ">Regional Initiatives</a></li><li><a href="uid86.html&#10;&#9;&#9;  ">National Initiatives</a></li><li><a href="uid121.html&#10;&#9;&#9;  ">European Initiatives</a></li><li><a href="uid131.html&#10;&#9;&#9;  ">International Initiatives</a></li><li><a href="uid159.html&#10;&#9;&#9;  ">International Research Visitors</a></li></ul></div>
      <div class="TdmEntry">Dissemination<ul><li><a href="uid166.html&#10;&#9;&#9;  ">Promoting Scientific Activities</a></li><li><a href="uid252.html&#10;&#9;&#9;  ">Teaching - Supervision - Juries</a></li><li><a href="uid317.html&#10;&#9;&#9;  ">Popularization</a></li></ul></div>
      <div class="TdmEntry">
        <div>Bibliography</div>
      </div>
      <div class="TdmEntry">
        <ul>
          <li>
            <a id="tdmbibentyear" href="bibliography.html">Publications of the year</a>
          </li>
          <li>
            <a id="tdmbibentfoot" href="bibliography.html#References">References in notes</a>
          </li>
        </ul>
      </div>
    </div>
    <div id="main">
      <div class="mainentete">
        <div id="head_agauche">
          <small><a href="http://www.inria.fr">
	    
	    Inria
	  </a> | <a href="../index.html">
	    
	    Raweb 
	    2015</a> | <a href="http://www.inria.fr/en/teams/sequel">Presentation of the Project-Team SEQUEL</a> | <a href="http://sequel.lille.inria.fr/">SEQUEL Web Site
	  </a></small>
        </div>
        <div id="head_adroite">
          <table class="qrcode">
            <tr>
              <td>
                <a href="sequel.xml">
                  <img style="align:bottom; border:none" alt="XML" src="../static/img/icons/xml_motif.png"/>
                </a>
              </td>
              <td>
                <a href="sequel.pdf">
                  <img style="align:bottom; border:none" alt="PDF" src="IMG/qrcode-sequel-pdf.png"/>
                </a>
              </td>
              <td>
                <a href="../sequel/sequel.epub">
                  <img style="align:bottom; border:none" alt="e-pub" src="IMG/qrcode-sequel-epub.png"/>
                </a>
              </td>
            </tr>
            <tr>
              <td/>
              <td>PDF
</td>
              <td>e-Pub
</td>
            </tr>
          </table>
        </div>
      </div>
      <!--FIN du corps du module-->
      <br/>
      <div class="bottomNavigation">
        <div class="tail_aucentre">
          <a href="./uid15.html" accesskey="P"><img style="align:bottom; border:none" alt="previous" src="../static/img/icons/previous_motif.jpg"/> Previous | </a>
          <a href="./uid0.html" accesskey="U"><img style="align:bottom; border:none" alt="up" src="../static/img/icons/up_motif.jpg"/>  Home</a>
          <a href="./uid29.html" accesskey="N"> | Next <img style="align:bottom; border:none" alt="next" src="../static/img/icons/next_motif.jpg"/></a>
        </div>
        <br/>
      </div>
      <div id="textepage">
        <!--DEBUT2 du corps du module-->
        <h2>Section: 
      Research Program</h2>
        <h3 class="titre3">Decision-making Under Uncertainty</h3>
        <p>The phrase “Decision under uncertainty” refers to the problem of taking decisions when we do not have a full knowledge neither of the situation, nor of the consequences of the decisions, as well as when the consequences of decision are non deterministic.</p>
        <p>We introduce two specific sub-domains, namely the Markov decision processes which models sequential decision problems, and bandit problems.</p>
        <a name="uid19"/>
        <h4 class="titre4">Reinforcement Learning</h4>
        <p>Sequential decision processes occupy the heart of the <span class="smallcap">SequeL </span> project; a detailed presentation of this problem may be found in Puterman's book <a href="./bibliography.html#sequel-2015-bid1">[46]</a> .</p>
        <p>A Markov Decision Process (MDP) is defined as the tuple <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>(</mo><mi>𝒳</mi><mo>,</mo><mi>𝒜</mi><mo>,</mo><mi>P</mi><mo>,</mo><mi>r</mi><mo>)</mo></mrow></math></span> where <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝒳</mi></math></span> is the state space, <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝒜</mi></math></span> is the action space, <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi></math></span> is the probabilistic transition kernel, and <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>r</mi><mo>:</mo><mi>𝒳</mi><mo>×</mo><mi>𝒜</mi><mo>×</mo><mi>𝒳</mi><mo>→</mo><mi>I</mi><mspace width="-0.166667em"/><mspace width="-0.166667em"/><mi>R</mi></mrow></math></span> is the reward function. For the sake of simplicity, we assume in this introduction that the state and action spaces are finite. If the current state (at time <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math></span>) is <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span> and the chosen action is <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>a</mi><mo>∈</mo><mi>𝒜</mi></mrow></math></span>, then the Markov assumption means that the transition probability to a new state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>x</mi><mo>'</mo></msup><mo>∈</mo><mi>𝒳</mi></mrow></math></span> (at time <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></math></span>) only depends on <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo>(</mo><mi>x</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow></math></span>. We write <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>p</mi><mo>(</mo><msup><mi>x</mi><mo>'</mo></msup><mo>|</mo><mi>x</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow></math></span> the corresponding transition probability. During a transition <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mrow><mo>(</mo><mi>x</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow><mo>→</mo><msup><mi>x</mi><mo>'</mo></msup></mrow></math></span>, a reward <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>r</mi><mo>(</mo><mi>x</mi><mo>,</mo><mi>a</mi><mo>,</mo><msup><mi>x</mi><mo>'</mo></msup><mo>)</mo></mrow></math></span> is incurred.</p>
        <p>In the MDP (<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>𝒳</mi><mo>,</mo><mi>𝒜</mi><mo>,</mo><mi>P</mi><mo>,</mo><mi>r</mi><mo>)</mo></mrow></math></span>, each initial state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>x</mi><mn>0</mn></msub></math></span> and action sequence <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>a</mi><mn>0</mn></msub><mo>,</mo><msub><mi>a</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo></mrow></math></span> gives rise to a sequence of states <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>x</mi><mn>1</mn></msub><mo>,</mo><msub><mi>x</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo></mrow></math></span>, satisfying <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>ℙ</mi><mfenced separators="" open="(" close=")"><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>=</mo><msup><mi>x</mi><mo>'</mo></msup><mrow><mo>|</mo></mrow><msub><mi>x</mi><mi>t</mi></msub><mo>=</mo><mi>x</mi><mo>,</mo><msub><mi>a</mi><mi>t</mi></msub><mo>=</mo><mi>a</mi></mfenced><mo>=</mo><mi>p</mi><mrow><mo>(</mo><msup><mi>x</mi><mo>'</mo></msup><mo>|</mo><mi>x</mi><mo>,</mo><mi>a</mi><mo>)</mo></mrow><mo>,</mo></mrow></math></span> and rewards (Note that for simplicity, we considered the case of a deterministic reward function, but in many applications, the reward <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>r</mi><mi>t</mi></msub></math></span> itself is a random variable.) <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>r</mi><mn>1</mn></msub><mo>,</mo><msub><mi>r</mi><mn>2</mn></msub><mo>,</mo><mo>...</mo></mrow></math></span> defined by <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>r</mi><mi>t</mi></msub><mo>=</mo><mi>r</mi><mrow><mo>(</mo><msub><mi>x</mi><mi>t</mi></msub><mo>,</mo><msub><mi>a</mi><mi>t</mi></msub><mo>,</mo><msub><mi>x</mi><mrow><mi>t</mi><mo>+</mo><mn>1</mn></mrow></msub><mo>)</mo></mrow></mrow></math></span>.</p>
        <p>The history of the process up to time <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math></span> is defined to be <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>H</mi><mi>t</mi></msub><mo>=</mo><mrow><mo>(</mo><msub><mi>x</mi><mn>0</mn></msub><mo>,</mo><msub><mi>a</mi><mn>0</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mrow><mi>t</mi><mo>-</mo><mn>1</mn></mrow></msub><mo>,</mo><msub><mi>a</mi><mrow><mi>t</mi><mo>-</mo><mn>1</mn></mrow></msub><mo>,</mo><msub><mi>x</mi><mi>t</mi></msub><mo>)</mo></mrow></mrow></math></span>. A policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> is a sequence of functions <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>π</mi><mn>0</mn></msub><mo>,</mo><msub><mi>π</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo></mrow></math></span>, where <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>t</mi></msub></math></span> maps the space of possible histories at time <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math></span> to the space of probability distributions over the space of actions <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝒜</mi></math></span>. To follow a policy means that, in each time step, we assume that the process history up to time <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>t</mi></math></span> is <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>x</mi><mn>0</mn></msub><mo>,</mo><msub><mi>a</mi><mn>0</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mi>t</mi></msub></mrow></math></span> and the probability of selecting an action <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>a</mi></math></span> is equal to <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>π</mi><mi>t</mi></msub><mrow><mo>(</mo><msub><mi>x</mi><mn>0</mn></msub><mo>,</mo><msub><mi>a</mi><mn>0</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mi>t</mi></msub><mo>)</mo></mrow><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></math></span>. A policy is called stationary (or Markovian) if <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>π</mi><mi>t</mi></msub></math></span> depends only on the last visited state. In other words, a policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>π</mi><mo>=</mo><mo>(</mo><msub><mi>π</mi><mn>0</mn></msub><mo>,</mo><msub><mi>π</mi><mn>1</mn></msub><mo>,</mo><mo>...</mo><mo>)</mo></mrow></math></span> is called stationary if <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mi>π</mi><mi>t</mi></msub><mrow><mo>(</mo><msub><mi>x</mi><mn>0</mn></msub><mo>,</mo><msub><mi>a</mi><mn>0</mn></msub><mo>,</mo><mo>...</mo><mo>,</mo><msub><mi>x</mi><mi>t</mi></msub><mo>)</mo></mrow><mo>=</mo><msub><mi>π</mi><mn>0</mn></msub><mrow><mo>(</mo><msub><mi>x</mi><mi>t</mi></msub><mo>)</mo></mrow></mrow></math></span> holds for all <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>t</mi><mo>≥</mo><mn>0</mn></mrow></math></span>. A policy is called deterministic if the probability distribution prescribed by the policy for any history is concentrated on a single action. Otherwise it is called a stochastic policy.</p>
        <p>We move from an MD process to an MD problem by formulating the goal of the agent, that is what the sought policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> has to optimize? It is very often formulated as maximizing (or minimizing), in expectation, some functional of the sequence of future rewards. For example, an usual functional is the infinite-time horizon sum of discounted rewards. For a given (stationary) policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span>, we define the value function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>V</mi><mi>π</mi></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow></math></span> of that policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> at a state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span> as the expected sum of discounted future rewards given that we state from the initial state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math></span> and follow the policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span>:</p>
        <div align="center" class="mathdisplay">
          <a name="uid21"/>
          <table width="100%">
            <tr valign="middle">
              <td align="center">
                <math xmlns="http://www.w3.org/1998/Math/MathML">
                  <mrow>
                    <msup>
                      <mi>V</mi>
                      <mi>π</mi>
                    </msup>
                    <mrow>
                      <mo>(</mo>
                      <mi>x</mi>
                      <mo>)</mo>
                    </mrow>
                    <mo>=</mo>
                    <mi>𝔼</mi>
                    <mfenced separators="" open="[" close="]">
                      <munderover>
                        <mo>∑</mo>
                        <mrow>
                          <mi>t</mi>
                          <mo>=</mo>
                          <mn>0</mn>
                        </mrow>
                        <mi>∞</mi>
                      </munderover>
                      <msup>
                        <mi>γ</mi>
                        <mi>t</mi>
                      </msup>
                      <msub>
                        <mi>r</mi>
                        <mi>t</mi>
                      </msub>
                      <mo>|</mo>
                      <msub>
                        <mi>x</mi>
                        <mn>0</mn>
                      </msub>
                      <mo>=</mo>
                      <mi>x</mi>
                      <mo>,</mo>
                      <mi>π</mi>
                    </mfenced>
                    <mo>,</mo>
                  </mrow>
                </math>
              </td>
              <td class="eqno" width="10" align="right">(1)</td>
            </tr>
          </table>
        </div>
        <p>where <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝔼</mi></math></span> is the expectation operator and <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>γ</mi><mo>∈</mo><mo>(</mo><mn>0</mn><mo>,</mo><mn>1</mn><mo>)</mo></mrow></math></span> is the discount factor. This value function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mi>π</mi></msup></math></span> gives an evaluation of the performance of a given policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span>. Other functionals of the sequence of future rewards may be considered, such as the undiscounted reward (see the stochastic shortest path problems <a href="./bibliography.html#sequel-2015-bid2">[45]</a> ) and average reward settings. Note also that, here, we considered the problem of maximizing a reward functional, but a formulation in terms of minimizing some cost or risk functional would be equivalent.</p>
        <p>In order to maximize a given functional in a sequential framework, one usually applies Dynamic Programming (DP)  <a href="./bibliography.html#sequel-2015-bid3">[43]</a> , which introduces the optimal value function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>V</mi><mo>*</mo></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow></math></span>, defined as the optimal expected sum of rewards when the agent starts from a state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math></span>. We have <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>V</mi><mo>*</mo></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow><mo>=</mo><msub><mo movablelimits="true" form="prefix">sup</mo><mi>π</mi></msub><msup><mi>V</mi><mi>π</mi></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow></math></span>. Now, let us give two definitions about policies:</p>
        <ul>
          <li>
            <p class="notaparagraph"><a name="uid22"> </a>We say that a policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> is optimal, if it attains the optimal values <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>V</mi><mo>*</mo></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow></math></span> for any state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span>, <i>i.e.</i>, if <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msup><mi>V</mi><mi>π</mi></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow><mo>=</mo><msup><mi>V</mi><mo>*</mo></msup><mrow><mo>(</mo><mi>x</mi><mo>)</mo></mrow></mrow></math></span> for all <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span>. Under mild conditions, deterministic stationary optimal policies exist <a href="./bibliography.html#sequel-2015-bid4">[44]</a> . Such an optimal policy is written <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>π</mi><mo>*</mo></msup></math></span>.</p>
          </li>
          <li>
            <p class="notaparagraph"><a name="uid23"> </a>We say that a (deterministic stationary) policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> is greedy with respect to (w.r.t.) some function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>V</mi></math></span> (defined on <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝒳</mi></math></span>) if, for all <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span>,</p>
            <div align="center" class="mathdisplay">
              <math xmlns="http://www.w3.org/1998/Math/MathML">
                <mrow>
                  <mi>π</mi>
                  <mrow>
                    <mo>(</mo>
                    <mi>x</mi>
                    <mo>)</mo>
                  </mrow>
                  <mo>∈</mo>
                  <mo form="prefix">arg</mo>
                  <munder>
                    <mo movablelimits="true" form="prefix">max</mo>
                    <mrow>
                      <mi>a</mi>
                      <mo>∈</mo>
                      <mi>𝒜</mi>
                    </mrow>
                  </munder>
                  <munder>
                    <mo>∑</mo>
                    <mrow>
                      <msup>
                        <mi>x</mi>
                        <mo>'</mo>
                      </msup>
                      <mo>∈</mo>
                      <mi>𝒳</mi>
                    </mrow>
                  </munder>
                  <mi>p</mi>
                  <mrow>
                    <mo>(</mo>
                    <msup>
                      <mi>x</mi>
                      <mo>'</mo>
                    </msup>
                    <mo>|</mo>
                    <mi>x</mi>
                    <mo>,</mo>
                    <mi>a</mi>
                    <mo>)</mo>
                  </mrow>
                  <mfenced separators="" open="[" close="]">
                    <mi>r</mi>
                    <mrow>
                      <mo>(</mo>
                      <mi>x</mi>
                      <mo>,</mo>
                      <mi>a</mi>
                      <mo>,</mo>
                      <msup>
                        <mi>x</mi>
                        <mo>'</mo>
                      </msup>
                      <mo>)</mo>
                    </mrow>
                    <mo>+</mo>
                    <mi>γ</mi>
                    <mi>V</mi>
                    <mrow>
                      <mo>(</mo>
                      <msup>
                        <mi>x</mi>
                        <mo>'</mo>
                      </msup>
                      <mo>)</mo>
                    </mrow>
                  </mfenced>
                  <mo>.</mo>
                </mrow>
              </math>
            </div>
            <p><a name="uid23"> </a> </p>
            <p class="notaparagraph"><a name="uid23"> </a>where <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mo form="prefix">arg</mo><msub><mo movablelimits="true" form="prefix">max</mo><mrow><mi>a</mi><mo>∈</mo><mi>𝒜</mi></mrow></msub><mi>f</mi><mrow><mo>(</mo><mi>a</mi><mo>)</mo></mrow></mrow></math></span> is the set of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>a</mi><mo>∈</mo><mi>𝒜</mi></mrow></math></span> that maximizes <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>f</mi><mo>(</mo><mi>a</mi><mo>)</mo></mrow></math></span>. For any function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>V</mi></math></span>, such a greedy policy always exists because <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>𝒜</mi></math></span> is finite.</p>
          </li>
        </ul>
        <p>The goal of Reinforcement Learning (RL), as well as that of dynamic programming, is to design an optimal policy (or a good approximation of it).</p>
        <p>The well-known Dynamic Programming equation (also called the Bellman equation) provides a relation between the optimal value function at a state <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math></span> and the optimal value function at the successors states <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>x</mi><mo>'</mo></msup></math></span> when choosing an optimal action: for all <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>x</mi><mo>∈</mo><mi>𝒳</mi></mrow></math></span>,</p>
        <div align="center" class="mathdisplay">
          <a name="uid24"/>
          <table width="100%">
            <tr valign="middle">
              <td align="center">
                <math xmlns="http://www.w3.org/1998/Math/MathML">
                  <mrow>
                    <msup>
                      <mi>V</mi>
                      <mo>*</mo>
                    </msup>
                    <mrow>
                      <mo>(</mo>
                      <mi>x</mi>
                      <mo>)</mo>
                    </mrow>
                    <mo>=</mo>
                    <munder>
                      <mo movablelimits="true" form="prefix">max</mo>
                      <mrow>
                        <mi>a</mi>
                        <mo>∈</mo>
                        <mi>𝒜</mi>
                      </mrow>
                    </munder>
                    <munder>
                      <mo>∑</mo>
                      <mrow>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>∈</mo>
                        <mi>𝒳</mi>
                      </mrow>
                    </munder>
                    <mi>p</mi>
                    <mrow>
                      <mo>(</mo>
                      <msup>
                        <mi>x</mi>
                        <mo>'</mo>
                      </msup>
                      <mo>|</mo>
                      <mi>x</mi>
                      <mo>,</mo>
                      <mi>a</mi>
                      <mo>)</mo>
                    </mrow>
                    <mfenced separators="" open="[" close="]">
                      <mi>r</mi>
                      <mrow>
                        <mo>(</mo>
                        <mi>x</mi>
                        <mo>,</mo>
                        <mi>a</mi>
                        <mo>,</mo>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>)</mo>
                      </mrow>
                      <mo>+</mo>
                      <mi>γ</mi>
                      <msup>
                        <mi>V</mi>
                        <mo>*</mo>
                      </msup>
                      <mrow>
                        <mo>(</mo>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>)</mo>
                      </mrow>
                    </mfenced>
                    <mo>.</mo>
                  </mrow>
                </math>
              </td>
              <td class="eqno" width="10" align="right">(2)</td>
            </tr>
          </table>
        </div>
        <p>The benefit of introducing this concept of optimal value function relies on the property that, from the optimal value function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mo>*</mo></msup></math></span>, it is easy to derive an optimal behavior by choosing the actions according to a policy greedy w.r.t. <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mo>*</mo></msup></math></span>. Indeed, we have the property that a policy greedy w.r.t. the optimal value function is an optimal policy:</p>
        <div align="center" class="mathdisplay">
          <a name="uid25"/>
          <table width="100%">
            <tr valign="middle">
              <td align="center">
                <math xmlns="http://www.w3.org/1998/Math/MathML">
                  <mrow>
                    <msup>
                      <mi>π</mi>
                      <mo>*</mo>
                    </msup>
                    <mrow>
                      <mo>(</mo>
                      <mi>x</mi>
                      <mo>)</mo>
                    </mrow>
                    <mo>∈</mo>
                    <mo form="prefix">arg</mo>
                    <munder>
                      <mo movablelimits="true" form="prefix">max</mo>
                      <mrow>
                        <mi>a</mi>
                        <mo>∈</mo>
                        <mi>𝒜</mi>
                      </mrow>
                    </munder>
                    <munder>
                      <mo>∑</mo>
                      <mrow>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>∈</mo>
                        <mi>𝒳</mi>
                      </mrow>
                    </munder>
                    <mi>p</mi>
                    <mrow>
                      <mo>(</mo>
                      <msup>
                        <mi>x</mi>
                        <mo>'</mo>
                      </msup>
                      <mo>|</mo>
                      <mi>x</mi>
                      <mo>,</mo>
                      <mi>a</mi>
                      <mo>)</mo>
                    </mrow>
                    <mfenced separators="" open="[" close="]">
                      <mi>r</mi>
                      <mrow>
                        <mo>(</mo>
                        <mi>x</mi>
                        <mo>,</mo>
                        <mi>a</mi>
                        <mo>,</mo>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>)</mo>
                      </mrow>
                      <mo>+</mo>
                      <mi>γ</mi>
                      <msup>
                        <mi>V</mi>
                        <mo>*</mo>
                      </msup>
                      <mrow>
                        <mo>(</mo>
                        <msup>
                          <mi>x</mi>
                          <mo>'</mo>
                        </msup>
                        <mo>)</mo>
                      </mrow>
                    </mfenced>
                    <mo>.</mo>
                  </mrow>
                </math>
              </td>
              <td class="eqno" width="10" align="right">(3)</td>
            </tr>
          </table>
        </div>
        <p>In short, we would like to mention that most of the reinforcement learning methods developed so far are built on one (or both) of the two following approaches ( <a href="./bibliography.html#sequel-2015-bid5">[49]</a> ):</p>
        <ul>
          <li>
            <p class="notaparagraph"><a name="uid26"> </a>Bellman's dynamic programming approach, based on the introduction of the value function. It consists in learning a “good” approximation of the optimal value function, and then using it to derive a greedy policy w.r.t. this approximation. The hope (well justified in several cases) is that the performance <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mi>π</mi></msup></math></span> of the policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span> greedy w.r.t. an approximation <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>V</mi></math></span> of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mo>*</mo></msup></math></span> will be close to optimality. This approximation issue of the optimal value function is one of the major challenges inherent to the reinforcement learning problem. <b>Approximate dynamic programming</b> addresses the problem of estimating performance bounds (<i>e.g.</i> the loss in performance <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mrow><mo>|</mo><mo>|</mo></mrow><msup><mi>V</mi><mo>*</mo></msup><mo>-</mo><msup><mi>V</mi><mi>π</mi></msup><mrow><mo>|</mo><mo>|</mo></mrow></mrow></math></span> resulting from using a policy <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>π</mi></math></span>-greedy w.r.t. some approximation <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>V</mi></math></span>- instead of an optimal policy) in terms of the approximation error <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mrow><mo>|</mo><mo>|</mo></mrow><msup><mi>V</mi><mo>*</mo></msup><mo>-</mo><mi>V</mi><mrow><mo>|</mo><mo>|</mo></mrow></mrow></math></span> of the optimal value function <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>V</mi><mo>*</mo></msup></math></span> by <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>V</mi></math></span>. Approximation theory and Statistical Learning theory provide us with bounds in terms of the number of sample data used to represent the
functions, and the capacity and approximation power of the considered function spaces.</p>
          </li>
          <li>
            <p class="notaparagraph"><a name="uid27"> </a>Pontryagin's maximum principle approach, based on sensitivity analysis of the performance measure w.r.t. some control parameters. This approach, also called <b>direct policy search</b> in the Reinforcement Learning community aims at directly finding a good feedback control law in a parameterized policy space without trying to approximate the value function. The method consists in estimating the so-called <b>policy gradient</b>, <i>i.e.</i> the sensitivity of the performance measure (the value function) w.r.t. some parameters of the current policy. The idea being that an optimal control problem is replaced by a parametric optimization problem in the space of parameterized policies. As such, deriving a policy gradient estimate would lead to performing a stochastic gradient method in order to search for a local optimal parametric policy.</p>
          </li>
        </ul>
        <p>Finally, many extensions of the Markov decision processes exist, among which the Partially Observable MDPs (POMDPs) is the case where the current state does not contain all the necessary information required to decide for sure of the best action.</p>
        <a name="uid28"/>
        <h4 class="titre4">Multi-arm Bandit Theory</h4>
        <p>Bandit problems illustrate the fundamental difficulty of decision making in the face of uncertainty: A decision maker must choose between what seems to be the best choice (“exploit”), or to test (“explore”) some alternative, hoping to discover a choice that beats the current best choice.</p>
        <p>The classical example of a bandit problem is deciding what treatment to give each patient in a clinical trial when the effectiveness of the treatments are initially unknown and the patients arrive sequentially. These bandit problems became popular with the seminal paper <a href="./bibliography.html#sequel-2015-bid6">[47]</a> , after which they have found applications in diverse fields, such as control, economics, statistics, or learning theory.</p>
        <p>Formally, a K-armed bandit problem (<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>K</mi><mo>≥</mo><mn>2</mn></mrow></math></span>) is specified by K real-valued distributions. In each time step a decision maker can select one of the distributions to obtain a sample from it. The samples obtained are considered as rewards. The distributions are initially unknown to the decision maker, whose goal is to maximize the sum of the rewards received, or equivalently, to minimize the regret which is defined as the loss compared to the total payoff that can be achieved given full knowledge of the problem, <i>i.e.</i>, when the arm giving the highest expected reward is pulled all the time.</p>
        <p>The name “bandit” comes from imagining a gambler playing with K slot machines. The gambler can pull the arm of any of the machines, which produces a random payoff as a result: When arm k is pulled, the random payoff is drawn from the distribution associated to k. Since the payoff distributions are initially unknown, the gambler must use exploratory actions to learn the utility of the individual arms. However, exploration has to be carefully controlled since excessive exploration may lead to unnecessary losses. Hence, to play well, the gambler must carefully balance exploration and exploitation. Auer <i>et al.</i> <a href="./bibliography.html#sequel-2015-bid7">[42]</a>  introduced the algorithm UCB (Upper Confidence Bounds) that follows what is now called the “optimism in the face of uncertainty principle”. Their algorithm works by computing upper confidence bounds for all the arms and then choosing the arm with the highest such bound. They proved that the expected regret of their algorithm increases at most at a logarithmic rate
with the number of trials, and that the algorithm achieves the smallest possible regret up to some sub-logarithmic factor (for the considered family of distributions).</p>
      </div>
      <!--FIN du corps du module-->
      <br/>
      <div class="bottomNavigation">
        <div class="tail_aucentre">
          <a href="./uid15.html" accesskey="P"><img style="align:bottom; border:none" alt="previous" src="../static/img/icons/previous_motif.jpg"/> Previous | </a>
          <a href="./uid0.html" accesskey="U"><img style="align:bottom; border:none" alt="up" src="../static/img/icons/up_motif.jpg"/>  Home</a>
          <a href="./uid29.html" accesskey="N"> | Next <img style="align:bottom; border:none" alt="next" src="../static/img/icons/next_motif.jpg"/></a>
        </div>
        <br/>
      </div>
    </div>
  </body>
</html>
