<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1 plus MathML 2.0 plus SVG 1.1//EN" "http://www.w3.org/2002/04/xhtml-math-svg/xhtml-math-svg.dtd">
<html xmlns="http://www.w3.org/1999/xhtml">
  <head>
    <meta http-equiv="Content-Type" content="application/xhtml+xml; charset=utf-8"/>
    <title>Project-Team:DATASHAPE</title>
    <link rel="stylesheet" href="../static/css/raweb.css" type="text/css"/>
    <meta name="description" content="New Results - Statistical aspects of topological and geometric data analysis"/>
    <meta name="dc.title" content="New Results - Statistical aspects of topological and geometric data analysis"/>
    <meta name="dc.creator" content="Claire Brécheteau"/>
    <meta name="dc.creator" content="Clément Levrard"/>
    <meta name="dc.creator" content="Mathieu Carrière"/>
    <meta name="dc.creator" content="Bertrand Michel"/>
    <meta name="dc.creator" content="Steve Oudot"/>
    <meta name="dc.creator" content="Thomas Bonis"/>
    <meta name="dc.creator" content="Steve Oudot"/>
    <meta name="dc.creator" content="Théo Lacombe"/>
    <meta name="dc.creator" content="Steve Oudot"/>
    <meta name="dc.creator" content="Claire Brécheteau"/>
    <meta name="dc.creator" content="Clément Levrard"/>
    <meta name="dc.creator" content="Frédéric Chazal"/>
    <meta name="dc.creator" content="Vincent Divol"/>
    <meta name="dc.creator" content="Vincent Divol"/>
    <meta name="dc.creator" content="Frédéric Chazal"/>
    <meta name="dc.creator" content="Bertrand Michel"/>
    <meta name="dc.creator" content="Frédéric Chazal"/>
    <meta name="dc.creator" content="Bertrand Michel"/>
    <meta name="dc.subject" content=""/>
    <meta name="dc.publisher" content="INRIA"/>
    <meta name="dc.date" content="(SCHEME=ISO8601) 2018-01"/>
    <meta name="dc.type" content="Report"/>
    <meta name="dc.language" content="(SCHEME=ISO639-1) en"/>
    <meta name="projet" content="DATASHAPE"/>
    <script type="text/javascript" src="https://cdn.mathjax.org/mathjax/latest/MathJax.js?config=TeX-MML-AM_CHTML">
      <!--MathJax-->
    </script>
  </head>
  <body>
    <div class="tdmdiv">
      <div class="logo">
        <a href="http://www.inria.fr">
          <img style="align:bottom; border:none" src="../static/img/icons/logo_INRIA-coul.jpg" alt="Inria"/>
        </a>
      </div>
      <div class="TdmEntry">
        <div class="tdmentete">
          <a href="uid0.html">Project-Team Datashape</a>
        </div>
        <span>
          <a href="uid1.html">Team, Visitors, External Collaborators</a>
        </span>
      </div>
      <div class="TdmEntry">
        <a href="./uid3.html">Overall Objectives</a>
      </div>
      <div class="TdmEntry">Research Program<ul><li><a href="uid5.html&#10;&#9;&#9;  ">Algorithmic aspects of topological and geometric data analysis</a></li><li><a href="uid6.html&#10;&#9;&#9;  ">Statistical aspects of topological and geometric data analysis</a></li><li><a href="uid7.html&#10;&#9;&#9;  ">Topological approach for multimodal data processing</a></li><li><a href="uid8.html&#10;&#9;&#9;  ">Experimental research and software development</a></li></ul></div>
      <div class="TdmEntry">Application Domains<ul><li><a href="uid11.html&#10;&#9;&#9;  ">Main application domains</a></li></ul></div>
      <div class="TdmEntry">
        <a href="./uid13.html">Highlights of the Year</a>
      </div>
      <div class="TdmEntry">New Software and Platforms<ul><li><a href="uid19.html&#10;&#9;&#9;  ">GUDHI</a></li></ul></div>
      <div class="TdmEntry">New Results<ul><li><a href="uid24.html&#10;&#9;&#9;  ">Algorithmic aspects of topological and geometric data analysis</a></li><li class="tdmActPage"><a href="uid38.html&#10;&#9;&#9;  ">Statistical aspects of topological and geometric data analysis</a></li><li><a href="uid48.html&#10;&#9;&#9;  ">Topological approach for multimodal data processing</a></li><li><a href="uid51.html&#10;&#9;&#9;  ">Experimental research and software development</a></li><li><a href="uid55.html&#10;&#9;&#9;  ">Miscellaneous</a></li></ul></div>
      <div class="TdmEntry">Bilateral Contracts and Grants with Industry<ul><li><a href="uid58.html&#10;&#9;&#9;  ">Bilateral Contracts with Industry</a></li><li><a href="uid61.html&#10;&#9;&#9;  ">Bilateral Grants with Industry</a></li></ul></div>
      <div class="TdmEntry">Partnerships and Cooperations<ul><li><a href="uid64.html&#10;&#9;&#9;  ">National Initiatives</a></li><li><a href="uid67.html&#10;&#9;&#9;  ">European Initiatives</a></li><li><a href="uid77.html&#10;&#9;&#9;  ">International Research Visitors</a></li></ul></div>
      <div class="TdmEntry">Dissemination<ul><li><a href="uid85.html&#10;&#9;&#9;  ">Promoting Scientific Activities</a></li><li><a href="uid126.html&#10;&#9;&#9;  ">Teaching - Supervision - Juries</a></li><li><a href="uid155.html&#10;&#9;&#9;  ">Popularization</a></li></ul></div>
      <div class="TdmEntry">
        <div>Bibliography</div>
      </div>
      <div class="TdmEntry">
        <ul>
          <li>
            <a id="tdmbibentmajor" href="bibliography.html">Major publications</a>
          </li>
          <li>
            <a id="tdmbibentyear" href="bibliography.html#year">Publications of the year</a>
          </li>
        </ul>
      </div>
    </div>
    <div id="main">
      <div class="mainentete">
        <div id="head_agauche">
          <small><a href="http://www.inria.fr">
	    
	    Inria
	  </a> | <a href="../index.html">
	    
	    Raweb 
	    2018</a> | <a href="http://www.inria.fr/en/teams/datashape">Presentation of the Project-Team DATASHAPE</a> | <a href="https://team.inria.fr/datashape/">DATASHAPE Web Site
	  </a></small>
        </div>
        <div id="head_adroite">
          <table class="qrcode">
            <tr>
              <td>
                <a href="datashape.xml">
                  <img style="align:bottom; border:none" alt="XML" src="../static/img/icons/xml_motif.png"/>
                </a>
              </td>
              <td>
                <a href="datashape.pdf">
                  <img style="align:bottom; border:none" alt="PDF" src="IMG/qrcode-datashape-pdf.png"/>
                </a>
              </td>
              <td>
                <a href="../datashape/datashape.epub">
                  <img style="align:bottom; border:none" alt="e-pub" src="IMG/qrcode-datashape-epub.png"/>
                </a>
              </td>
            </tr>
            <tr>
              <td/>
              <td>PDF
</td>
              <td>e-Pub
</td>
            </tr>
          </table>
        </div>
      </div>
      <!--FIN du corps du module-->
      <br/>
      <div class="bottomNavigation">
        <div class="tail_aucentre">
          <a href="./uid24.html" accesskey="P"><img style="align:bottom; border:none" alt="previous" src="../static/img/icons/previous_motif.jpg"/> Previous | </a>
          <a href="./uid0.html" accesskey="U"><img style="align:bottom; border:none" alt="up" src="../static/img/icons/up_motif.jpg"/>  Home</a>
          <a href="./uid48.html" accesskey="N"> | Next <img style="align:bottom; border:none" alt="next" src="../static/img/icons/next_motif.jpg"/></a>
        </div>
        <br/>
      </div>
      <div id="textepage">
        <!--DEBUT2 du corps du module-->
        <h2>Section: 
      New Results</h2>
        <h3 class="titre3">Statistical aspects of topological and geometric data analysis</h3>
        <a name="uid39"/>
        <h4 class="titre4">Robust Bregman Clustering</h4>
        <p class="participants"><span class="part">Participants</span> :
	Claire Brécheteau, Clément Levrard.</p>
        <p class="bold">
          <p>In collaboration with Aurélie Fischer (Université Paris-Diderot).</p>
        </p>
        <p>Using a trimming approach, in <a href="./bibliography.html#datashape-2018-bid13">[38]</a>, we investigate a k-means type method based on Bregman divergences for clustering data possibly corrupted with clutter noise. The main interest of Bregman divergences is that the standard Lloyd algorithm adapts to these distortion measures, and they are well-suited for clustering data sampled according to mixture models from exponential families. We prove that there exists an optimal codebook, and that an empirically optimal codebook converges a.s. to an optimal codebook in the distortion sense. Moreover, we obtain the sub-Gaussian rate of convergence for k-means 1 √ n under mild tail assumptions. Also, we derive a Lloyd-type algorithm with a trimming parameter that can be selected from data according to some heuristic, and present some experimental results.</p>
        <a name="uid40"/>
        <h4 class="titre4">Statistical analysis and parameter selection for Mapper</h4>
        <p class="participants"><span class="part">Participants</span> :
	Mathieu Carrière, Bertrand Michel, Steve Oudot.</p>
        <p>In <a href="./bibliography.html#datashape-2018-bid14">[15]</a> we study the question of the
statistical convergence of the 1-dimensional Mapper to its continuous
analogue, the Reeb graph. We show that the Mapper is an optimal
estimator of the Reeb graph, which gives, as a byproduct, a method to
automatically tune its parameters and compute confidence regions on
its topological features, such as its loops and flares. This allows to
circumvent the issue of testing a large grid of parameters and keeping
the most stable ones in the brute-force setting, which is widely used
in visualization, clustering and feature selection with the Mapper.</p>
        <a name="uid41"/>
        <h4 class="titre4">A Fuzzy Clustering Algorithm for the Mode-Seeking Framework</h4>
        <p class="participants"><span class="part">Participants</span> :
	Thomas Bonis, Steve Oudot.</p>
        <p>In <a href="./bibliography.html#datashape-2018-bid15">[13]</a> we propose a new soft clustering
algorithm based on the mode-seeking framework. Given a point cloud
in <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mi>d</mi></msup></math></span>, we define regions of high density that we call cluster
cores, then we implement a random walk on a neighborhood graph built
on top of the data points. This random walk is designed in such a
way that it is attracted by high-density regions, the intensity of
the attraction being controlled by a temperature parameter <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>β</mi><mo>&gt;</mo><mn>0</mn></mrow></math></span>. The membership of a point to a given cluster is then the
probability for the random walk starting at this point to hit the
corresponding cluster core before any other. While many properties
of random walks (such as hitting times, commute distances, etc) are
known to eventually encode purely local information when the number
of data points grows to infinity, the
regularization introduced by the use of cluster cores allows the
output of our algorithm to converge to quantities involving the
global structure of the underlying density function. Empirically,
we show how the choice of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>β</mi></math></span> influences the behavior of our
algorithm: for small values of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>β</mi></math></span> the result is really close to
hard mode-seeking, while for values of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>β</mi></math></span> close to 1 the
result is similar to the output of the (soft) spectral clustering.
We also demonstrate the scalability of our approach experimentally.</p>
        <a name="uid42"/>
        <h4 class="titre4">Large Scale computation of Means and Clusters for Persistence Diagrams using Optimal Transport</h4>
        <p class="participants"><span class="part">Participants</span> :
	Théo Lacombe, Steve Oudot.</p>
        <p class="bold">
          <p>In collaboration with Marco Cuturi (ENSAE).</p>
        </p>
        <p>Persistence diagrams (PDs) are at the core of topological data analysis. They provide succinct descriptors encoding the underlying topology of sophisticated data. PDs are backed-up by strong theoretical results regarding their stability and have been used in various learning contexts. However, they do not live in a space naturally endowed with a Hilbert structure where natural metrics are not even differentiable, thus not suited to optimization process. Therefore, basic statistical notions such as the barycenter of a finite sample of PDs are not properly defined. In <a href="./bibliography.html#datashape-2018-bid16">[30]</a> we provide a theoretically good and computationally tractable framework to estimate the barycenter of a set of persistence diagrams. This construction is based on the theory of Optimal Transport (OT) and endows the space of PDs with a metric inspired from regularized Wasserstein distances.</p>
        <a name="uid43"/>
        <h4 class="titre4">The k-PDTM : a coreset for robust geometric inference</h4>
        <p class="participants"><span class="part">Participants</span> :
	Claire Brécheteau, Clément Levrard.</p>
        <p>Analyzing the sub-level sets of the distance to a compact sub-manifold of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mi>d</mi></msup></math></span> is a common method in TDA to understand its topology. The distance to measure (DTM) was introduced by Chazal, Cohen-Steiner and Mérigot to face the non-robustness of the distance to a compact set to noise and outliers. This function makes possible the inference of the topology of a compact subset of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mi>d</mi></msup></math></span> from a noisy cloud of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi></math></span> points lying nearby in the Wasserstein sense. In practice, these sub-level sets may be computed using approximations of the DTM such as the q-witnessed distance or other power distance. These approaches lead eventually to compute the homology of unions of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi></math></span> growing balls, that might become intractable whenever <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi></math></span> is large. To simultaneously face the two problems of large number of points and noise, we introduce in <a href="./bibliography.html#datashape-2018-bid17">[39]</a> the <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span>-power distance to measure (<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span>-PDTM). This new approximation of the distance to measure may be thought of as a <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span>-coreset based approximation of the DTM. Its sublevel sets consist in union of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span>-balls, <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>k</mi><mo>&lt;</mo><mo>&lt;</mo><mi>n</mi></mrow></math></span>, and this distance is also proved robust to noise. We assess the quality of this approximation for <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span> possibly dramatically smaller than <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>n</mi></math></span>, for instance <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>k</mi><mo>=</mo><mi>n</mi><mn>13</mn></mrow></math></span> is proved to be optimal for 2-dimensional shapes. We also provide an algorithm to compute this <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>k</mi></math></span>-PDTM.</p>
        <a name="uid44"/>
        <h4 class="titre4">The density of expected persistence diagrams and its kernel based estimation</h4>
        <p class="participants"><span class="part">Participants</span> :
	Frédéric Chazal, Vincent Divol.</p>
        <p>Persistence diagrams play a fundamental role in Topological Data Analysis where they are used
as topological descriptors of filtrations built on top of data. They consist in discrete multisets
of points in the plane <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mn>2</mn></msup></math></span>
that can equivalently be seen as discrete measures in <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mn>2</mn></msup></math></span>. When the
data come as a random point cloud, these discrete measures become random measures whose
expectation is studied in this paper. In <a href="./bibliography.html#datashape-2018-bid18">[28]</a> we first show that for a wide class of filtrations, including
the Čech and Rips-Vietoris filtrations, the expected persistence diagram, that is a deterministic
measure on <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>ℝ</mi><mn>2</mn></msup></math></span>, has a density with respect to the Lebesgue measure. Second, building on the
previous result we show that the persistence surface recently introduced by Adams et al can be seen as a
kernel estimator of this density. We propose a cross-validation scheme for selecting an optimal
bandwidth, which is proven to be a consistent procedure to estimate the density.</p>
        <a name="uid45"/>
        <h4 class="titre4">On the choice of weight functions for linear representations of persistence diagrams</h4>
        <p class="participants"><span class="part">Participant</span> :
	Vincent Divol.</p>
        <p class="bold">
          <p>In collaboration with Wolfgang Polonik (UC Davis)</p>
        </p>
        <p>Persistence diagrams are efficient descriptors of the topology of a point cloud. As they do not naturally belong to a Hilbert space, standard statistical methods cannot be directly applied to them. Instead, feature maps (or representations) are commonly used for the analysis. A large class of feature maps, which we call linear, depends on some weight functions, the choice of which is a critical issue. An important criterion to choose a weight function is to ensure stability of the feature maps with respect to Wasserstein distances on diagrams. In <a href="./bibliography.html#datashape-2018-bid19">[42]</a>, we improve known results on the stability of such maps, and extend it to general weight functions. We also address the choice of the weight function by considering an asymptotic setting; assume that <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>X</mi><mi>n</mi></msub></math></span> is an i.i.d. sample from a density on <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mrow><mo>[</mo><mn>0</mn><mo>,</mo><mn>1</mn><mo>]</mo></mrow><mi>d</mi></msup></math></span>. For the Cech and Rips filtrations, we characterize the weight functions for which the corresponding feature maps converge as n approaches infinity, and by doing so, we prove laws of large numbers for the total persistence of such diagrams. Both approaches lead to the same simple heuristic for tuning weight functions: if the data lies near a <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>d</mi></math></span>-dimensional manifold, then a sensible choice of weight function is the persistence to the power <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>α</mi></math></span> with <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>α</mi><mo>≥</mo><mi>d</mi></mrow></math></span>.</p>
        <a name="uid46"/>
        <h4 class="titre4">Estimating the Reach of a Manifold</h4>
        <p class="participants"><span class="part">Participants</span> :
	Frédéric Chazal, Bertrand Michel.</p>
        <p class="bold">
          <p>In collaboration with E. Aamari (CNRS Paris 7), J.Kim, A. Rinaldo and L. Wasserman (Carnegie Mellon University).</p>
        </p>
        <p>Various problems in manifold estimation make use of a quantity called the reach, denoted by
<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mi>M</mi></msub></math></span>, which is a measure of the regularity of the
manifold. <a href="./bibliography.html#datashape-2018-bid20">[32]</a> is the first investigation into the problem of how to
estimate the reach. First, we study the geometry of the reach through an
approximation perspective. We derive new geometric results on the reach
for submanifolds without boundary. An estimator <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mover accent="true"><mi>τ</mi><mo>^</mo></mover></math></span> of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>τ</mi><mi>M</mi></msub></math></span> is proposed
in a framework where tangent spaces are known, and bounds assessing
its efficiency are derived. In the case of i.i.d. random point cloud
<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi>𝕏</mi><mi>n</mi></msub></math></span>, <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mrow><mi>τ</mi><mo>(</mo><msub><mi>𝕏</mi><mi>n</mi></msub><mo>)</mo></mrow></math></span> is showed to achieve uniform expected loss bounds over a
<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><msup><mi>𝒞</mi><mn>3</mn></msup></math></span>-like model. Finally, we obtain upper and lower bounds on the minimax rate
for estimating the reach.</p>
        <a name="uid47"/>
        <h4 class="titre4">Robust Topological Inference: Distance To a Measure and Kernel Distance</h4>
        <p class="participants"><span class="part">Participants</span> :
	Frédéric Chazal, Bertrand Michel.</p>
        <p class="bold">
          <p>In collaboration with B. Fasy (Univ. Montana) and F. Lecci, A. Rinaldo and L. Wasserman (Carnegie Mellon University).</p>
        </p>
        <p>Let <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi></math></span> be a distribution with support <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>S</mi></math></span>. The salient features of <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>S</mi></math></span> can be quantified with
persistent homology, which summarizes topological features of the sublevel sets of the distance function (the distance of any point
<span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>x</mi></math></span> to <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>S</mi></math></span>). Given a sample from <span class="math"><math xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi></math></span> we can infer
the persistent homology using an empirical version of the distance function. However, the
empirical distance function is highly non-robust to noise and outliers. Even one outlier is
deadly. The distance-to-a-measure (DTM), introduced by Chazal et al. (2011), and the kernel distance, introduced by Phillips et al. (2014), are smooth functions that provide useful
topological information but are robust to noise and outliers. Chazal et al. (2015) derived
concentration bounds for DTM. Building on these results, in <a href="./bibliography.html#datashape-2018-bid21">[16]</a>, we derive limiting distributions
and confidence sets, and we propose a method for choosing tuning parameters.</p>
      </div>
      <!--FIN du corps du module-->
      <br/>
      <div class="bottomNavigation">
        <div class="tail_aucentre">
          <a href="./uid24.html" accesskey="P"><img style="align:bottom; border:none" alt="previous" src="../static/img/icons/previous_motif.jpg"/> Previous | </a>
          <a href="./uid0.html" accesskey="U"><img style="align:bottom; border:none" alt="up" src="../static/img/icons/up_motif.jpg"/>  Home</a>
          <a href="./uid48.html" accesskey="N"> | Next <img style="align:bottom; border:none" alt="next" src="../static/img/icons/next_motif.jpg"/></a>
        </div>
        <br/>
      </div>
    </div>
  </body>
</html>
