<?xml version="1.0" encoding="utf-8"?>
<raweb xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="en" year="2013">
  <identification id="compsys" isproject="true">
    <shortname>COMPSYS</shortname>
    <projectName>Compilation and Embedded Computing Systems</projectName>
    <theme-de-recherche>Architecture, Languages and Compilation</theme-de-recherche>
    <domaine-de-recherche>Algorithmics, Programming, Software and Architecture</domaine-de-recherche>
    <urlTeam>http://www.ens-lyon.fr/LIP/COMPSYS/index.html.en</urlTeam>
    <datecreation>2004 January 01</datecreation>
    <structure_exterieure type="Labs">
      <libelle>Laboratoire de l'Informatique du Parallélisme (LIP)</libelle>
    </structure_exterieure>
    <structure_exterieure type="Organism">
      <libelle>CNRS</libelle>
    </structure_exterieure>
    <structure_exterieure type="Organism">
      <libelle>Université Claude Bernard (Lyon 1)</libelle>
    </structure_exterieure>
    <structure_exterieure type="Organism">
      <libelle>Ecole normale supérieure de Lyon</libelle>
    </structure_exterieure>
    <UR name="Grenoble"/>
    <keywords>
      <term>Compilation</term>
      <term>Combinatorial Optimization</term>
      <term>Hardware Accelerators</term>
      <term>High-level Synthesis</term>
      <term>High Performance Computing</term>
    </keywords>
    <moreinfo>
      <p>Compsys is an EPC (équipe-projet commune), i.e., a research project-team that
is common to several institutions: Inria, Ecole normale supérieure de Lyon
(ENS-Lyon), CNRS, and Université Claude Bernard of Lyon (UCB-Lyon). It is located
at Ecole normale supérieure de Lyon and exists since January 2002 as part of
the computer science laboratory (Laboratoire de l'Informatique du
Parallélisme, Lip, UMR CNRS ENS-Lyon UCB-Lyon Inria 5668) and as an
Inria pre-project. It became a full Inria project in January 2004. It
has been evaluated by Inria in Spring 2007 and extended 4 more years. It
has been evaluated by AERES in December 2010 and received the mark A+. It
has been evaluated positively again by Inria in Spring 2012 and extended 4
more years. Thus, it should end around 2016.</p>
      <p>The goal of Compsys is to develop compilation techniques, more precisely
code analysis and code optimization techniques, for programming or designing
embedded computing systems. So far, Compsys focused on both low-level
(back-end) optimizations for embedded processors and high-level (front-end,
mainly source-to-source) transformations, in particular for high-level
synthesis of hardware accelerators. Recent activities also include a shift
towards dynamic compilation, compilation for GPUs and multicores, and the
analysis of parallel languages. The main characteristic of Compsys is its
focus on combinatorial optimization problems (graph algorithms, linear
programming, polyhedral optimizations) coming from code optimization problems
(register allocation, memory optimization, scheduling, automatic generation
of interfaces, etc.) and the validation of these techniques in the
development of compilation tools.</p>
    </moreinfo>
  </identification>
  <team id="uid1">
    <person key="compsys-2006-id18253">
      <firstname>Christophe</firstname>
      <lastname>Alias</lastname>
      <categoryPro>Chercheur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria, Researcher</moreinfo>
    </person>
    <person key="compsys-2005-id18414">
      <firstname>Florent</firstname>
      <lastname>Bouchez</lastname>
      <categoryPro>Technique</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria, ManycoreLabs funding, Caisse des
Dépôts et Consignations, from May 2013
to Aug. 2013</moreinfo>
    </person>
    <person key="compsys-2005-id18078">
      <firstname>Alain</firstname>
      <lastname>Darte</lastname>
      <categoryPro>Chercheur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Team leader, CNRS, Senior
Researcher</moreinfo>
      <hdr>oui</hdr>
    </person>
    <person key="compsys-2005-id18170">
      <firstname>Paul</firstname>
      <lastname>Feautrier</lastname>
      <categoryPro>Enseignant</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>ENS-Lyon, Professor,
Emeritus</moreinfo>
      <hdr>oui</hdr>
    </person>
    <person key="compsys-2008-id18330">
      <firstname>Laure</firstname>
      <lastname>Gonnord</lastname>
      <categoryPro>Enseignant</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Univ. Lyon I, Associate Professor, since
Sep. 2013 (formerly in Lille University)</moreinfo>
    </person>
    <person key="compsys-2005-id18199">
      <firstname>Fabrice</firstname>
      <lastname>Rastello</lastname>
      <categoryPro>Chercheur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria, Researcher</moreinfo>
      <hdr>oui</hdr>
    </person>
    <person key="compsys-2013-idp140327620765808">
      <firstname>Lukasz</firstname>
      <lastname>Domagala</lastname>
      <categoryPro>PostDoc</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria, ManycoreLabs funding, Caisse des
Dépôts et Consignations, from Jan. 2013</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620768288">
      <firstname>Alexandros</firstname>
      <lastname>Lamprineas</lastname>
      <categoryPro>Technique</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria, ManycoreLabs funding, Caisse
des Dépôts et Consignations, from Sep. 2013</moreinfo>
    </person>
    <person key="graal-2009-id60684">
      <firstname>Adrian</firstname>
      <lastname>Muresan</lastname>
      <categoryPro>CollaborateurExterieur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>ENS-Lyon, Zettice funding from
Jan. 2013 to Mar. 2013</moreinfo>
    </person>
    <person key="compsys-2006-id18441">
      <firstname>Alexandru</firstname>
      <lastname>Plesco</lastname>
      <categoryPro>CollaborateurExterieur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Zettice, Inria funding until
Apr. 2013</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620775440">
      <firstname>François</firstname>
      <lastname>Gindraud</lastname>
      <categoryPro>PhD</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>ENS-Lyon grant, from Jan. 2013</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620777744">
      <firstname>Guillaume</firstname>
      <lastname>Iooss</lastname>
      <categoryPro>PhD</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>ENS-Lyon grant, from Sep. 2011</moreinfo>
    </person>
    <person key="compsys-2012-idp140679680158976">
      <firstname>Alexandre</firstname>
      <lastname>Isoard</lastname>
      <categoryPro>PhD</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>ENS-Lyon grant, from Sep. 2012</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620782352">
      <firstname>Diogo</firstname>
      <lastname>Nunes Sampaio</lastname>
      <categoryPro>PhD</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Brazilian CAPES grant, from Oct. 2013</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620784656">
      <firstname>Duco</firstname>
      <lastname>Van Amstel</lastname>
      <categoryPro>PhD</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Kalray Cifre grant, from Jan. 2013</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620786960">
      <firstname>Laetitia</firstname>
      <lastname>Lecot-Gauthé</lastname>
      <categoryPro>Assistant</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620789264">
      <firstname>Maria Immaculada</firstname>
      <lastname>Presseguer</lastname>
      <categoryPro>Assistant</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Inria</moreinfo>
    </person>
    <person key="compsys-2013-idp140327620791568">
      <firstname>Raphael Ernani</firstname>
      <lastname>Rodrigues</lastname>
      <categoryPro>Visiteur</categoryPro>
      <research-centre>Grenoble</research-centre>
      <moreinfo>Univ. Lyon I, Internship, from May 2013 to
Jul. 2013</moreinfo>
    </person>
  </team>
  <presentation id="uid2">
    <bodyTitle>Overall Objectives</bodyTitle>
    <subsection id="uid3" level="1">
      <bodyTitle>Introduction</bodyTitle>
      <descriptionlist>
        <label>Keywords:</label>
        <li id="uid4">
          <p noindent="true">Compilation, code analysis, code optimization, memory
optimization, combinatorial optimization, algorithmics, polyhedral
optimization, hardware accelerators, high-level synthesis, high-performance
computing.</p>
        </li>
      </descriptionlist>
      <p>The objective of Compsys is to adapt and to extend code analysis and code
optimization techniques primarily designed in compilers/parallelizers for
high performance computing to the special case of <i>embedded computing
systems</i>. In particular, Compsys works on back-end optimizations for
specialized processors and on high-level program transformations, in
particular for the compilation towards or the synthesis of hardware
accelerators. The main characteristic of Compsys is its focus on
combinatorial problems (graph algorithms, linear programming, polyhedra)
coming from code optimizations (register allocation, cache and memory
optimizations, scheduling, optimizations for power, automatic generation of
software/hardware interfaces, etc.) and the validation of techniques
developed in compilation tools.</p>
      <p>Compsys started as an Inria project in 2004, after 2 years of maturation.
This first period of Compsys, Compsys I, was positively evaluated in
Spring 2007 after its first 4 years period (2004-2007). It was again
evaluated by AERES in 2009, as part of the general evaluation of Lip, and got
the best possible mark, A+. The second period (2007-2012), Compsys II, was
again evaluated positively by Inria in Spring 2012 and formally prolongated
into Compsys III at the very end of 2012. The geographical move in 2013 of
Fabrice Rastello to Grenoble was first to expand the activities of Compsys in
the context of Giant, a R&amp;D technology center with several industrial and
academic actors. In 2014, this geographical move is a departure from
Compsys, Fabrice Rastello will now work on his own. The research directions of
Compsys III are nevertheless not modified drastically and are in line with
the research directions presented in the synthesis report provided for the
2012 evaluation  <footnote id="uid5" id-text="1">See <ref xlink:href="http://www.ens-lyon.fr/LIP/COMPSYS/wordpress/wp-content/uploads/2013/09/ficheSynthese.pdf" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>LIP/<allowbreak/>COMPSYS/<allowbreak/>wordpress/<allowbreak/>wp-content/<allowbreak/>uploads/<allowbreak/>2013/<allowbreak/>09/<allowbreak/>ficheSynthese.<allowbreak/>pdf</ref></footnote>. The shift towards dynamic compilation, underlined in this
report, will be pursued by Fabrice Rastello only, while the shift towards the
compilation of streaming programming, the analysis and optimizations of
parallel languages, with an even stronger focus on polyhedral optimizations
are the heart of Compsys III, as well as the development of the Zettice
start-up in which Christophe Alias is involved. The hiring of Laure Gonnord also adds
new forces on the code analysis research aspects.</p>
      <p>Section <ref xlink:href="#uid6" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> defines the
general context of the team's activities.
Section <ref xlink:href="#uid11" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> presents the research
objectives and main achievements in Compsys I, i.e., until 2007, and how
its research directions were modified for Compsys II.
Section <ref xlink:href="#uid17" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> briefly presents the
main achievements of Compsys II, referring to the annual reports from 2008
to 2012 for details. Finally,
Section <ref xlink:href="#uid20" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> highlights the main
novelties of the past year, i.e., 2013.</p>
    </subsection>
    <subsection id="uid6" level="1">
      <bodyTitle>General Presentation</bodyTitle>
      <p>Classically, an embedded computer is a digital system that is part of a
larger system and that is not directly accessible to the user. Examples are
appliances like phones, TV sets, washing machines, game platforms, or even
larger systems like radars and sonars. In particular, this computer is not
programmable in the usual way. Its program, if it exists, is supplied as part
of the manufacturing process and is seldom (or never) modified thereafter.
As the embedded systems market grows and evolves, this view of embedded
systems is becoming obsolete and tends to be too restrictive. Many aspects of
general-purpose computers apply to modern embedded platforms. Nevertheless,
embedded systems remain characterized by a set of specialized application
domains, rigid constraints (cost, power, efficiency,
heterogeneity), and its market structure. The term <i>embedded system</i> has
been used for naming a wide variety of objects. More precisely, there are two
categories of so-called <i>embedded systems</i>: a) control-oriented and hard
real-time embedded systems (automotive, plant control, airplanes, etc.); b)
compute-intensive embedded systems (signal processing, multi-media, stream
processing) processing large data sets with parallel and/or pipelined
execution. Compsys is primarily concerned with this second type of
embedded systems, now referred to as <i>embedded computing systems</i>.</p>
      <p>Today, the industry sells many more embedded processors than general-purpose
processors; the field of embedded systems is one of the few segments of the
computer market where the European industry still has a substantial share,
hence the importance of embedded system research in the European research
initiatives. Our priority towards embedded software is motivated by the
following observations: a) the embedded system market is expanding, among
many factors, one can quote pervasive digitalization, low-cost products,
appliances, etc.; b) research on software for embedded systems is poorly
developed in France, especially if one considers the importance of actors
like Alcatel, STMicroelectronics, Matra, Thales, etc.; c) since embedded systems
increase in complexity, new problems are emerging: computer-aided design,
shorter time-to-market, better reliability, modular design, and component
reuse.</p>
      <p>A specific aspect of embedded computing systems is the use of various kinds
of processors, with many particularities (instruction sets, registers, data
and instruction caches, now multiple cores) and constraints (code size,
performance, storage). The development of <i>compilers</i> is crucial for
this industry, as selling a platform without its programming environment and
compiler would not be acceptable. To cope with such a range of different
processors, the development of robust, generic (retargetable), though
efficient compilers is mandatory. Unlike standard compilers for
general-purpose processors, compilers for embedded processors can be more
aggressive (i.e., take more time to optimize) for optimizing some important
parts of applications. This opens a new range of optimizations. Another
interesting aspect is the introduction of platform-independent intermediate
languages, such as Java bytecode, that is compiled dynamically at runtime
(aka just-in-time). Extreme lightweight compilation mechanisms that run
faster and consume less memory have to be developed. One of the objectives of
Compsys was to revisit existing compilation techniques in the context of
embedded computing systems, to deconstruct these techniques, to improve them,
and to develop new techniques taking constraints of embedded processors into
account.</p>
      <p>As for <i>high-level synthesis</i> (HLS), several compilers/systems have
appeared, after some first unsuccessful industrial attempts in the past.
These tools are mostly based on C or C++ as for example SystemC,
VCC, CatapultC, Altera C2H, Pico-Express.
Academic projects also exist such as Flex
and Raw
at MIT, Piperench
at Carnegie-Mellon University, Compaan
at the University of Leiden, Ugh/Disydent at LIP6 (Paris), Gaut at Lester
(Bretagne), MMAlpha (Insa-Lyon), and others. In general, the support for
parallelism in HLS tools is minimal, especially in industrial tools. Also,
the basic problem that these projects have to face is that the definition of
performance is more complex than in classical systems. In fact, it is a
multi-criteria optimization problem and one has to take into account the
execution time, the size of the program, the size of the data structures, the
power consumption, the manufacturing cost, etc. The impact of the compiler
on these costs is difficult to assess and control. Success will be the
consequence of a detailed knowledge of all steps of the design process, from
a high-level specification to the chip layout. A strong cooperation of the
compilation and chip design communities is needed. The main expertise in
Compsys for this aspect is in the <i>parallelization</i> and optimization
of <i>regular computations</i>. Hence, we will target applications with a
large potential parallelism, but we will attempt to integrate our solutions
into the big picture of CAD environments.</p>
      <p>More generally, the aims of Compsys are to develop new compilation and
optimization techniques for the field of embedded computing system design.
This field is large, and Compsys does not intend to cover it in its
entirety. As previously mentioned, we are mostly interested in the automatic
design of accelerators, for example designing a VLSI or FPGA circuit
for a digital filter, and in the development of new back-end compilation strategies for
embedded processors. We study code transformations that optimize features
such as execution time, power consumption, code and die size, memory
constraints, and compiler reliability. These features are related to embedded
systems but some are not specific to them. The code transformations we
develop are both at source level and at assembly level. A specificity of
Compsys is to mix a solid theoretical basis for all code optimizations we
introduce with algorithmic/software developments. Within Inria, our
project is related to the “architecture and compilation” theme, more
precisely code optimization, as some of the research conducted in Alchemy
(now Parkas), Alf (previously known as Caps), Camus, and to
high-level architectural synthesis, as some of the research in Cairn.</p>
      <p>Most french researchers working on high-performance computing (automatic
parallelization, languages, operating systems, networks) moved to grid
computing at the end of the 90s. We thought that applications, industrial
needs, and research problems were more interesting in the design of embedded
platforms. Furthermore, we were convinced that our expertise on high-level
code transformations could be more useful in this field. This is the reason
why Tanguy Risset came to Lyon in 2002 to create the Compsys team with
Anne Mignotte and Alain Darte, before Paul Feautrier, Antoine Fraboulet,
Fabrice Rastello, and finally Christophe Alias joined the group. Then, Tanguy
Risset left Compsys to become a professor at INSA Lyon, and Antoine Fraboulet
and Anne Mignotte moved to other fields of research. As for Laure Gonnord,
after a post-doc in Compsys, she obtained an assistant professor position
in Lille but remained external collaborator of the team for the
period 2009-2013 and finally obtained an assistant professor
position in Lyon.</p>
      <p>All present and past members of Compsys have a background in automatic
parallelization and high-level program analyses and transformations. Paul Feautrier was
the initiator of the polytope model for program transformations around 1990
and, before coming to Lyon, started to be more interested in programming
models and optimizations for embedded applications, in particular through
collaborations with Philips. Alain Darte worked on mathematical tools and
algorithmic issues for parallelism extraction in programs. He became
interested in the automatic generation of hardware accelerators, thanks to
his stay at HP Labs in the Pico project in Spring 2001. Antoine Fraboulet did
a PhD with Anne Mignotte – who was working on high-level synthesis (HLS) –
on code and memory optimizations for embedded applications. Fabrice Rastello
did a PhD on tiling transformations for parallel machines, then was hired by
STMicroelectronics where he worked on assembly code optimizations for embedded
processors. Tanguy Risset worked for a long time on the synthesis of systolic
arrays, being the main architect of the HLS tool MMAlpha. Christophe Alias
did a PhD on algorithm recognition for program optimizations and
parallelization. He first spent a year in Compsys working on array
contraction, where he started to develop his tool Bee, then a year at Ohio
State University with Prof. P. Sadayappan on memory optimizations. He
finally joined Compsys as an Inria researcher. Laure Gonnord
did a PhD on invariant generation and program analysis and became
interested on application on compilation and code generation since
her postdoc in the team.</p>
      <p>It may be worth to quote Bob Rau and his colleagues (IEEE Computer, sept.
2002):</p>
      <p>
        <i>"Engineering disciplines tend to go through fairly predictable phases:
ad hoc, formal and rigorous, and automation. When the discipline is in its
infancy and designers do not yet fully understand its potential problems
and solutions, a rich diversity of poorly understood design techniques
tends to flourish. As understanding grows, designers sacrifice the
flexibility of wild and woolly design for more stylized and restrictive
methodologies that have underpinnings in formalism and rigorous theory.
Once the formalism and theory mature, the designers can automate the design
process. This life cycle has played itself out in disciplines as diverse as
PC board and chip layout and routing, machine language parsing, and logic
synthesis.</i>
      </p>
      <p>
        <i>We believe that the computer architecture discipline is ready to enter the
automation phase. Although the gratification of inventing brave new
architectures will always tempt us, for the most part the focus will shift
to the automatic and speedy design of highly customized computer systems
using well-understood architecture and compiler technologies.”</i>
      </p>
      <p>We share this view of the future of architecture and compilation. Without
targeting too ambitious objectives, we were convinced of two complementary
facts: a) the mathematical tools developed in the past for manipulating
programs in automatic parallelization were lacking in high-level synthesis
and embedded computing optimizations and, even more, they started to be
rediscovered frequently in less mature forms, b) before being able to really
use these techniques in HLS and embedded program optimizations, we needed to
learn a lot from the application side, from the electrical engineering side,
and from the embedded architecture side. Our primary goal was thus twofold:
to increase our knowledge of embedded computing systems and to adapt/extend
code optimization techniques, primarily designed for high performance
computing, to the special case of embedded computing systems. In the initial
Compsys proposal, we proposed four research directions, centered on
compilation methods for embedded applications, both for software and
accelerators design:</p>
      <simplelist>
        <li id="uid7">
          <p noindent="true">Code optimization for specific processors (mainly DSP and VLIW
processors);</p>
        </li>
        <li id="uid8">
          <p noindent="true">Platform-independent loop transformations (including memory
optimization);</p>
        </li>
        <li id="uid9">
          <p noindent="true">Silicon compilation and hardware/software codesign;</p>
        </li>
        <li id="uid10">
          <p noindent="true">Development of polyhedral (but not only) optimization tools.</p>
        </li>
      </simplelist>
      <p>These research activities were primarily supported by a marked investment in
polyhedra manipulation tools and, more generally, solid mathematical and
algorithmic studies, with the aim of constructing operational software tools,
not just theoretical results. Hence the fourth research theme was centered on
the development of these tools.</p>
    </subsection>
    <subsection id="uid11" level="1">
      <bodyTitle>Summary of Compsys I Achievements</bodyTitle>
      <p>The Compsys team has been evaluated by Inria for the first time in April
2007. The evaluation, conducted by Erik Hagersted (Uppsala University), Vinod
Kathail (Synfora, inc), J. (Ram) Ramanujam (Baton Rouge University) was
positive. Compsys I thus continued into Compsys II for 4-5 years but in
a new configuration as Tanguy Risset and Antoine Fraboulet left the project to
follow research directions closer to their host laboratory at Insa-Lyon. The main
achievements of Compsys I, for this period, were the following:</p>
      <simplelist>
        <li id="uid12">
          <p noindent="true">The development of a strong collaboration with the compilation group at
STMicroelectronics, with important results in aggressive optimizations for
instruction cache and register allocation.</p>
        </li>
        <li id="uid13">
          <p noindent="true">New results on the foundation of high-level program
transformations, including scheduling techniques for process networks
and a general technique for array contraction (memory reuse) based on the
theory of lattices.</p>
        </li>
        <li id="uid14">
          <p noindent="true">Many original contributions with partners closer to hardware constraints,
including CEA, related to SoC simulation, hardware/software interfaces, power
models, and simulators.</p>
        </li>
      </simplelist>
      <p>Due to Compsys size reduction (from 5 permanent researchers to 3 in 2008,
then 4 again in 2009), the team then focused, in Compsys II, on two research
directions only:</p>
      <simplelist>
        <li id="uid15">
          <p noindent="true">Code generation for embedded processors, on the two opposite, though
connected, aspects: aggressive compilation and just-in-time compilation.</p>
        </li>
        <li id="uid16">
          <p noindent="true">High-level program analysis and transformations for high-level synthesis
tools.</p>
        </li>
      </simplelist>
    </subsection>
    <subsection id="uid17" level="1">
      <bodyTitle>Quick view of Compsys II
Achievements and directions for Compsys III</bodyTitle>
      <p>The main achievements of Compsys II were:</p>
      <simplelist>
        <li id="uid18">
          <p noindent="true">the great success of the collaboration with STMicroelectronics with many deep
results on SSA (Static Single Assignment), register allocation, and
intermediate program representations;</p>
        </li>
        <li id="uid19">
          <p noindent="true">the design of high-level program analysis, optimizations, and tools,
mainly related to high-level synthesis, some leading to the development of
the Zettice start-up.</p>
        </li>
      </simplelist>
      <p>For more details on the past years of Compsys II, see the previous annual
reports from 2008 to 2012. Compsys II was positively evaluated in Spring
2012 by Inria. The evaluation committee members were Walid Najjar
(University of California Riverside), Paolo Faraboschi (HP Labs), Scott Mahlke
(University of Michigan), Pedro Diniz (University of Southern California),
Peter Marwedel (TU Dortmund), and Pierre Paulin (STMicroelectronics, Canada),
the last three assigned specifically to Compsys.</p>
      <p>For Compsys III, the changes in the permanent members (departure of
Fabrice Rastello and arrival of Laure Gonnord (while she was only external collaborator of
Compsys until Sep. 2013) reduces the forces on back-end code optimizations,
and in particular dynamic compilation, but increases the forces on program
analysis. In this context, Compsys III will continue to develop fundamental
concepts or techniques whose applicability should go beyond a particular
architectural or language trend, as well as stand-alone tools (either as proofs
of concepts or to be used as basic blocks in larger tools/compilers developed
by others) and our own experimental prototypes. One of the main objectives of
Compsys III is to try to push the polyhedral model beyond its present limits
both in terms of analysis techniques (possibly integrating approximation and
runtime support) and of applicability (e.g., analysis of parallel or streaming
languages, program verification, compilation towards accelerators such as GPU
or multicores).
</p>
    </subsection>
    <subsection id="uid20" level="1">
      <bodyTitle>Highlights of the Year</bodyTitle>
      <p>For 2013, from the point of view of organization, funding, collaborations,
the main points to highlight are:</p>
      <simplelist>
        <li id="uid21">
          <p noindent="true">The Zettice startup project, initiated by Alexandru Plesco and
Christophe Alias, won the <i>concours OSEO 2013</i> grant (Banque Publique
d'Investissement, 40 Keuros) and the <i>“most promising start-up
award”</i> at SAME 2013. See more details in
Section <ref xlink:href="#uid122" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
        </li>
        <li id="uid22">
          <p noindent="true">Laure Gonnord was hired as assistant professor at ENS-Lyon, she is now a
permanent member of Compsys. Fabrice Rastello has left Compsys and will
continue his research in Grenoble.</p>
        </li>
        <li id="uid23">
          <p noindent="true">The collaborations with Colorado State University (S. Rajopadhye) and
Ohio State University (Sadayappan) were very successful. New topics of
collaboration with the Inria Parkas and Camus teams have started.</p>
        </li>
        <li id="uid24">
          <p noindent="true">From April 2013 to July 2013, Compsys organized 4 scientific events
on compilation, regrouped in a larger and coherent <i>thematic quarter on
compilation</i>  <footnote id="uid25" id-text="2"><ref xlink:href="http://labexcompilation.ens-lyon.fr" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>labexcompilation.<allowbreak/>ens-lyon.<allowbreak/>fr</ref></footnote>, with
international audience and visibility. It was mainly funded by the Labex
MILYON, see details in Section <ref xlink:href="#uid143" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
        </li>
      </simplelist>
      <p>From a scientific point of view, the shift, in Compsys III, towards the
analysis of parallel programs, the extensions of the polyhedral model, both
in terms of techniques and applications, and the code optimizations based on
trace analysis has been already fruitful, see the section “New Results”, in
particular:</p>
      <simplelist>
        <li id="uid26">
          <p noindent="true">Innovative contributions on parametric
tiling <ref xlink:href="#compsys-2013-bid0" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, <ref xlink:href="#compsys-2013-bid1" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> as extensions of the
polyhedral model.</p>
        </li>
        <li id="uid27">
          <p noindent="true">A groundbreaking introduction of polyhedral techniques for the analysis of
parallel programs, in particular X10 <ref xlink:href="#compsys-2013-bid2" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, <ref xlink:href="#compsys-2013-bid3" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
        </li>
        <li id="uid28">
          <p noindent="true">Several important contributions (e.g., <ref xlink:href="#compsys-2013-bid4" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) that
demonstrate the interest of mixing trace analysis and static analysis for
code (in particular locality) improvements.</p>
        </li>
      </simplelist>
    </subsection>
  </presentation>
  <fondements id="uid29">
    <bodyTitle>Research Program</bodyTitle>
    <subsection id="uid30" level="1">
      <bodyTitle>Generalities</bodyTitle>
      <p>The embedded system design community is facing two challenges:</p>
      <simplelist>
        <li id="uid31">
          <p noindent="true">The complexity of embedded applications is increasing at a rapid rate.</p>
        </li>
        <li id="uid32">
          <p noindent="true">The needed increase in processing power is no longer obtained by
increases in the clock frequency, but by increased parallelism.</p>
        </li>
      </simplelist>
      <p>While, in the past, each type of embedded application was implemented in a
separate appliance, the present tendency is toward a universal hand-held
object, which must serve as a cell-phone, as a personal digital assistant, as a
game console, as a camera, as a Web access point, and much more. One may say
that embedded applications are of the same level of complexity as those running
on a PC, but they must use a more constrained platform in terms of processing
power, memory size, and energy consumption. Furthermore, most of them depend
on international standards (e.g., in the field of radio digital communication),
which are evolving rapidly. Lastly, since ease of use is at a premium for
portable devices, these applications must be integrated seamlessly to a degree
that is unheard of in standard computers.</p>
      <p>All of this dictates that modern embedded systems retain some form of
programmability. For increased designer productivity and reduced
time-to-market, programming must be done in some high-level language, with
appropriate tools for compilation, run-time support, and debugging. This does
not mean that all embedded systems (or all of an embedded system) must be
processor based. Another solution is the use of field programmable gate arrays
(FPGA), which may be programmed at a much finer grain than a processor,
although the process of FPGA “programming” is less well understood than
software generation. Processors are better than application-specific circuits
at handling complicated control and unexpected events. On the other hand,
FPGAs may be tailored to just meet the needs of their application, resulting in
better energy and silicon area usage. It is expected that most embedded systems
will use a combination of general-purpose processors, specific processors like
DSPs, and FPGA accelerators.
Such a combination is already present in recent versions of the Atom
Intel processor.</p>
      <p>As a consequence, parallel programming, which has long been confined to the
high-performance community, must become the common place rather than the
exception. In the same way that sequential programming moved from assembly code
to high-level languages at the price of a slight loss in performance, parallel
programming must move from low-level tools, like OpenMP or even MPI, to
higher-level programming environments. While fully-automatic parallelization
is a Holy Grail that will probably never be reached in our lifetimes, it will
remain as a component in a comprehensive environment, including general-purpose
parallel programming languages, domain-specific parallelizers, parallel
libraries and run-time systems, back-end compilation, dynamic parallelization.
The landscape of embedded systems is indeed very diverse and many design flows
and code optimization techniques must be considered. For example, embedded
processors (micro-controllers, DSP, VLIW) require powerful back-end
optimizations that can take into account hardware specificities, such as
special instructions and particular organizations of registers and memories.
FPGA and hardware accelerators, to be used as small components in a larger
embedded platform, require “hardware compilation”, i.e., design flows and
code generation mechanisms to generate non-programmable circuits. For the
design of a complete system-on-chip platform, architecture models, simulators,
debuggers are required. The same is true for multi-cores of any kind, GPGPU
(“general-purpose” graphical processing units), CGRA (coarse-grain
reconfigurable architectures), which require specific methodologies and
optimizations, although all these techniques converge or have connections. In
other words, embedded systems need all usual aspects of the process that
transforms some specification down to an executable, software or hardware. In
this wide range of topics, Compsys concentrates on the code optimizations
aspects in this transformation chain, restricting to compilation (transforming
a program to a program) for embedded processors and to high-level synthesis
(transforming a program into a circuit description) for FPGAs.</p>
      <p>Actually, it is not a surprise to see compilation and high-level synthesis
getting closer. Now that high-level synthesis has grown up sufficiently to be
able to rely on place-and-route tools, or even to synthesize C-like languages,
standard techniques for back-end code generation (register allocation,
instruction selection, instruction scheduling, software pipelining) are used in
HLS tools. At the higher level, programming languages for programmable parallel
platforms share many aspects with high-level specification languages for HLS,
for example, the description and manipulations of nested loops, or the model of
computation/communication (e.g., Kahn process networks). In all aspects, the
frontier between software and hardware is vanishing. For example, in terms of
architecture, customized processors (with processor extension as proposed by
Tensilica) share features with both general-purpose processors and hardware
accelerators. FPGAs are both hardware and software as they are fed with
“programs” representing their hardware configurations. In other words, this
convergence in code optimizations explains why Compsys studies both program
compilation and high-level synthesis. Besides, Compsys has a tradition of
building free software tools for linear programming and optimization in
general, and will continue it, as needed for our current research.</p>
      <p>
        <b>The next two sections give an overview of the main directions that were
explored by Compsys II and partially extended in 2013 in Compsys III:
back-end code optimizations for embedded processors (including aggressive and
just-in-time compilation) and high-level program analysis and
transformations, primarily for high-level synthesis. For Compsys III, the
shifts towards dynamic compilation on one hand and more advanced polyhedral
techniques for program analysis and optimization on the other hand are not
detailed here but appear clearly in the section “New Results”. Indeed, the
first axis (dynamic compilation and trace analysis) will not be pursued in
2014, due to the departure of Fabrice Rastello, it will thus be described in his
activity report for 2014. The second axis (polyhedral extensions and
high-level program analysis) will be detailed more deeply in 2014. But, it is
already not limited to high-level synthesis as can be seen from the different
contributions on X10, OpenStream, parametric tiling, etc.</b>
      </p>
    </subsection>
    <subsection id="uid33" level="1">
      <bodyTitle>Back-End Code Optimizations for Embedded Processors</bodyTitle>
      <p>Compilation is an old activity, in particular back-end code optimizations. We
first give some elements that explain why the development of embedded systems
makes compilation come back as a research topic. We then detail the code
optimizations that we are interested in, both for aggressive and just-in-time
compilation.</p>
      <subsection id="uid34" level="2">
        <bodyTitle>Embedded Systems and the Revival of Compilation &amp; Code
Optimizations</bodyTitle>
        <p>Applications for embedded computing systems generate complex programs and need
more and more processing power. This evolution is driven, among others, by the
increasing impact of digital television, the first instances of UMTS
networks, and the increasing size of digital supports, like recordable DVD,
and even Internet applications. Furthermore, standards are evolving very
rapidly (see for instance the successive versions of MPEG). As a consequence,
the industry has rediscovered the interest of programmable structures, whose
flexibility more than compensates for their larger size and power consumption.
The appliance provider has a choice between hard-wired structures (Asic),
special-purpose processors (Asip), or (quasi) general-purpose processors
(DSP for multimedia applications). Our cooperation with STMicroelectronics led us to
investigate the last solution, as implemented in the ST100 (DSP processor)
and the ST200 (VLIW DSP processor) family for example. Compilation and,
in particular, back-end code optimizations find a second life in the context of
such embedded computing systems.</p>
        <p>At the heart of this progress is the concept of <i>virtualization</i>, which is
the key for more portability, more simplicity, more reliability, and of course
more security. This concept, implemented through binary translation,
just-in-time compilation, etc., consists in hiding the architecture-dependent
features as far as possible during the compilation process. It has been used
for quite a long time for servers such as HotSpot, a bit more recently for
workstations, and it is quite recent for embedded computing for reasons we now
explain.</p>
        <p>As previously mentioned, the definition of “embedded systems” is rather
imprecise. However, one can at least agree on the following features:</p>
        <simplelist>
          <li id="uid35">
            <p noindent="true">Even for processors that are programmable (as opposed to hardware
accelerators), processors have some architectural specificities, and are very
diverse;</p>
          </li>
          <li id="uid36">
            <p noindent="true">Many processors (but not all of them) have limited resources, in
particular in terms of memory;</p>
          </li>
          <li id="uid37">
            <p noindent="true">For some processors, power consumption is an issue;</p>
          </li>
          <li id="uid38">
            <p noindent="true">In some cases, aggressive compilation (through cross-compilation) is
possible, and even highly desirable for important functions.</p>
          </li>
        </simplelist>
        <p>This diversity is one of the reason why virtualization, which starts to be more
mature, is becoming more and more common in programmable embedded systems, in
particular through CIL (a standardization of MSIL). This implies a late
compilation of programs, through just-in-time (JIT), including dynamic
compilation. Some people even think that dynamic compilation, which can have
more information because performed at run-time, can outperform the performances
of “ahead-of-time” compilation.</p>
        <p>Performing code generation (and some higher-level optimizations) in a late
phase is potentially advantageous, as it can exploit architectural
specificities and run-time program information such as constants and aliasing,
but it is more constrained in terms of time and available resources. Indeed,
the processor that performs the late compilation phase is, <i>a priori</i>, less
powerful (in terms of memory for example) than a processor used for
cross-compilation. The challenge is thus to spread the compilation process in
time by deferring some optimizations (“deferred compilation”) and by
propagating some information for those whose computation is expensive (“split
compilation”). Classically, a compiler has to deal with different intermediate
representations (IR) where high-level information (i.e., more
target-independent) co-exist with low-level information. The split compilation
has to solve a similar problem where, this time, the compactness of the
information representation, and thus its pertinence, is also an important
criterion. Indeed, the IR is evolving not only from a target-independent
description to a target-dependent one, but also from a situation where the
compilation time is almost unlimited (cross-compilation) to one where any type
of resource is limited. This is also a reason why static single assignment
(SSA) is becoming specific to embedded compilation, even if it was first used
for workstations. Indeed, SSA is a sparse (i.e., compact) representation of
liveness information. In other words, if time constraints are common to all JIT
compilers (not only for embedded computing), the benefit of using SSA is also
in terms of its good ratio pertinence/storage of information. It also enables
to simplify algorithms, which is also important for increasing the reliability
of the compiler.</p>
      </subsection>
      <subsection id="uid39" level="2">
        <bodyTitle>Aggressive and Just-in-Time Optimizations of Assembly-Level Code</bodyTitle>
        <p>Compilation for embedded processors is difficult because the architecture and
the operations are specially tailored to the task at hand, and because the
amount of resources is strictly limited. For instance, the potential for
instruction level parallelism (SIMD, MMX), the limited number of registers
and the small size of the memory, the use of direct-mapped instruction caches,
of predication, but also the special form of applications  <ref xlink:href="#compsys-2013-bid5" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>
generate many open problems. Our goal is to contribute to their understanding
and their solutions.</p>
        <p>As previously explained, compilation for embedded processors include both
aggressive and just in time (JIT) optimizations. Aggressive compilation
consists in allowing more time to implement costly solutions (so, looking for
complete, even expensive, studies is mandatory): the compiled program is loaded
in permanent memory (ROM, flash, etc.) and its compilation time is not
significant; also, for embedded systems, code size and energy consumption
usually have a critical impact on the cost and the quality of the final
product. Hence, the application is cross-compiled, in other words, compiled on
a powerful platform distinct from the target processor. Just-in-time
compilation corresponds to compiling applets on demand on the target processor.
For compatibility and compactness, the source languages are CIL or Java. The
code can be uploaded or sold separately on a flash memory. Compilation is
performed at load time and even dynamically during execution. Used heuristics,
constrained by time and limited resources, are far from being aggressive. They
must be fast but smart enough.</p>
        <p>Our aim is, in particular, to develop exact or heuristic solutions to <i>combinatorial</i> problems that arise in compilation for VLIW and DSP
processors, and to integrate these methods into industrial compilers for DSP
processors (mainly ST100, ST200, Strong ARM). Such combinatorial problems can
be found for example in register allocation, in opcode selection, or in code
placement for optimization of the instruction cache. Another example is the
problem of removing the multiplexer functions (known as <formula type="inline"><math xmlns="http://www.w3.org/1998/Math/MathML" overflow="scroll"><mi>φ</mi></math></formula> functions) that
are inserted when converting into SSA form. These optimizations are usually
done in the last phases of the compiler, using an assembly-level intermediate
representation. In industrial compilers, they are handled in independent phases
using heuristics, in order to limit the compilation time. Our initial goal was
to develop a more global understanding of these optimization problems to derive
both aggressive heuristics and JIT techniques, the main tool being the SSA
representation.</p>
        <p>In particular, we investigated the interaction of register allocation,
coalescing, and spilling, with the different code representations, such as
SSA. One of the challenging features of today's processors is
predication  <ref xlink:href="#compsys-2013-bid6" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, which interferes with all optimization phases, as
the SSA form does. Many classical algorithms become inefficient for
predicated code. This is especially surprising, since, beside giving a better
trade-off between the number of conditional branches and the length of the
critical path, converting control dependences into data dependences increases
the size of basic blocks and hence creates new opportunities for local
optimization algorithms. One has to adapt classical algorithms to predicated
code  <ref xlink:href="#compsys-2013-bid7" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and also to study the impact of predicated code on the
whole compilation process.</p>
        <p>As mentioned in Section <ref xlink:href="#uid11" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, a lot
of progress has already been done in this direction in our past collaborations
with STMicroelectronics. In particular, the goal of the Sceptre project was to revisit,
in the light of SSA, some code optimizations in an aggressive context, i.e.,
by looking for the best performances without limiting, <i>a priori</i>, the
compilation time and the memory usage. One of the major results of this
collaboration was to propose to exploit SSA so as to design a register
allocator in two phases, with one spilling phase relatively target-independent,
then the allocator itself, which takes into account architectural constraints
and optimizes other aspects (in particular, coalescing). This new way of
considering register allocation has shown its interest for aggressive static
compilation. But it offered three other perspectives:</p>
        <simplelist>
          <li id="uid40">
            <p noindent="true">A simplification of the allocator, which again goes toward a more
reliable compiler design, based on static single assignment.</p>
          </li>
          <li id="uid41">
            <p noindent="true">The possibility to handle the hardest part, the spilling phase, as a
preliminary phase, thus a good candidate for split compilation.</p>
          </li>
          <li id="uid42">
            <p noindent="true">The possibility of a fast allocator, with a much higher quality than
usual JIT approaches such as “linear scan”, thus suitable for
virtualization and JIT compilation.</p>
          </li>
        </simplelist>
        <p>These additional possibilities have been the heart of our research on back-end
optimizations in Compsys II. The objective of the Mediacom project with
STMicroelectronics was to address them. More generally, in Compsys II, our goal was
to continue to develop our activity on code optimizations, exploiting SSA
properties, following our two-phases strategy:</p>
        <simplelist>
          <li id="uid43">
            <p noindent="true">First, revisit code optimizations in an aggressive context to develop
better strategies, without eliminating too quickly solutions that may have
been considered as too expensive in the past.</p>
          </li>
          <li id="uid44">
            <p noindent="true">Then, exploit the new concepts introduced in the aggressive context to
design better algorithms in a JIT context, focusing on the speed of
algorithms and their memory footprint, without compromising too much on the
quality of the generated code.</p>
          </li>
        </simplelist>
        <p>An important challenge was also to consider more code optimizations and more
architectural features, such as registers with aliasing, predication, and,
possibly in a longer term, vectorization/parallelization.
</p>
      </subsection>
    </subsection>
    <subsection id="uid45" level="1">
      <bodyTitle>High-Level Program Analysis and
Transformations</bodyTitle>
      <subsection id="uid46" level="2">
        <bodyTitle>High-Level Synthesis Context</bodyTitle>
        <p>High-level synthesis has become a necessity, mainly because the exponential
increase in the number of gates per chip far outstrips the productivity of
human designers. Besides, applications that need hardware accelerators usually
belong to domains, like telecommunications and game platforms, where fast
turn-around and time-to-market minimization are paramount. We believe that our
expertise in compilation and automatic parallelization can contribute to the
development of the needed tools.</p>
        <p>Today, synthesis tools for FPGAs or ASICs come in many shapes. At the lowest
level, there are proprietary Boolean, layout, and place-and-route tools, whose
input is a VHDL or Verilog specification at the structural or register-transfer
level (RTL). Direct use of these tools is difficult, for several reasons:</p>
        <simplelist>
          <li id="uid47">
            <p noindent="true">A structural description is completely different from an usual
algorithmic language description, as it is written in term of interconnected
basic operators. One may say that it has a spatial orientation, in place of
the familiar temporal orientation of algorithmic languages.</p>
          </li>
          <li id="uid48">
            <p noindent="true">The basic operators are extracted from a library, which poses problems of
selection, similar to the instruction selection problem in ordinary
compilation.</p>
          </li>
          <li id="uid49">
            <p noindent="true">Since there is no accepted standard for VHDL synthesis, each tool has its
own idiosyncrasies and reports its results in a different format. This makes
it difficult to build portable HLS tools.</p>
          </li>
          <li id="uid50">
            <p noindent="true">HLS tools have trouble handling loops. This is particularly true for
logic synthesis systems, where loops are systematically unrolled (or
considered as sequential) before synthesis. An efficient treatment of loops
needs the polyhedral model. This is where past results from the automatic
parallelization community are useful.</p>
          </li>
          <li id="uid51">
            <p noindent="true">More generally, a VHDL specification is too low level to allow the
designer to perform, easily, higher-level code optimizations, especially on
multi-dimensional loops and arrays, which are of paramount importance to
exploit parallelism, pipelining, and perform communication and memory
optimizations.</p>
          </li>
        </simplelist>
        <p>Some intermediate tools exist that generate VHDL from a specification in
restricted C, both in academia (such as SPARK,
Gaut,
UGH,
CloogVHDL),
and in industry (such as C2H),
CatapultC,
Pico-Express.
All these tools use only the most elementary form of parallelization,
equivalent to instruction-level parallelism in ordinary compilers, with some
limited form of block pipelining. Targeting one of these tools for low-level
code generation, while we concentrate on exploiting loop parallelism, might be
a more fruitful approach than directly generating VHDL. However, it may be
that the restrictions they impose preclude efficient use of the underlying
hardware.</p>
        <p>Our first experiments with these HLS tools reveal two important issues. First,
they are, of course, limited to certain types of input programs so as to make
their design flows successful. It is a painful and tricky task for the user to
transform the program so that it fits these constraints and to tune it to get
good results. Automatic or semi-automatic program transformations can help the
user achieve this task. Second, users, even expert users, have only a very
limited understanding of what back-end compilers do and why they do not lead to
the expected results. An effort must be done to analyze the different design
flows of HLS tools, to explain what to expect from them, and how to use them to
get a good quality of results. Our first goal is thus to develop high-level
techniques that, used in front of existing HLS tools, improve their
utilization. This should also give us directions on how to modify them.</p>
        <p>More generally, we want to consider HLS as a more global parallelization
process. So far, no HLS tools is capable of generating designs with
communicating <i>parallel</i> accelerators, even if, in theory, at least for the
scheduling part, a tool such as Pico-Express could have such capabilities.
The reason is that it is, for example, very hard to automatically design
parallel memories and to decide the distribution of array elements in memory
banks to get the desired performances with parallel accesses. Also, how to
express communicating processes at the language level? How to express
constraints, pipeline behavior, communication media, etc.? To better exploit
parallelism, a first solution is to extend the source language with parallel
constructs, as in all derivations of the Kahn process networks model, including
communicating regular processes (CRP, see later). The other solution is a form
of automatic parallelization. However, classical methods, which are mostly
based on scheduling, are not directly applicable, firstly because they pay poor
attention to locality, which is of paramount importance in hardware. Besides,
their aim is to extract all the parallelism in the source code; they rely on
the runtime system to tailor the parallelism degree to the available resources.
Obviously, there is no runtime system in hardware. The real challenge is thus
to invent new scheduling algorithms that take both resource and locality into
account, and then to infer the necessary hardware from the schedule. This is
probably possible only for programs that fit into the polyhedral model.</p>
        <p>In summary, as for our activity on back-end code optimizations, which is
decomposed into two complementary activities, aggressive and just-in-time
compilation, we focus our activity on high-level synthesis on two aspects:</p>
        <simplelist>
          <li id="uid52">
            <p noindent="true">Developing high-level transformations, especially for loops and
memory/communication optimizations, that can be used in front of HLS tools so
as to improve their use.</p>
          </li>
          <li id="uid53">
            <p noindent="true">Developing concepts and techniques in a more global view of high-level
synthesis, starting from specification languages down to hardware
implementation.</p>
          </li>
        </simplelist>
        <p>We now give more details on the program optimizations and transformations we
want to consider and on our methodology.</p>
      </subsection>
      <subsection id="uid54" level="2">
        <bodyTitle>Specifications, Transformations, Code Generation for High-Level Synthesis</bodyTitle>
        <p>Before contributing to high-level synthesis, one has to decide which execution
model is targeted and where to intervene in the design flow. Then one has to
solve scheduling, placement, and memory management problems. These three
aspects should be handled as a whole, but present state of the art dictates
that they be treated separately. One of our aims will be to find more
comprehensive solutions. The last task is code generation, both for the
processing elements and the interfaces between FPGAs and the host processor.</p>
        <p>There are basically two execution models for embedded systems: one is the
classical accelerator model, in which data is deposited in the memory of the
accelerator, which then does its job, and returns the results. In the streaming
model, computations are done on the fly, as data flow from an input channel to
the output. Here, data is never stored in (addressable) memory. Other models
are special cases, or sometimes compositions of the basic models. For instance,
a systolic array follows the streaming model, and sometimes extends it to
higher dimensions. Software radio modems follow the streaming model in the
large, and the accelerator model in detail. The use of first-in first-out
queues (FIFO) in hardware design is an application of the streaming model.
Experience shows that designs based on the streaming model are more efficient
that those based on memory. One of the point to be investigated is whether it
is general enough to handle arbitrary (regular) programs. The answer is
probably negative. One possible implementation of the streaming model is as a
network of communicating processes either as Kahn process networks (FIFO based)
or as our more recent model of communicating regular processes (CRP, memory
based). It is an interesting fact that several researchers have investigated
translation from process networks  <ref xlink:href="#compsys-2013-bid8" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and to process
networks  <ref xlink:href="#compsys-2013-bid9" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, <ref xlink:href="#compsys-2013-bid10" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
        <p>Kahn process networks (KPN) were introduced 30 years ago as a notation for
representing parallel programs. Such a network is built from processes that
communicate via perfect FIFO channels. Because the channel histories are
deterministic, one can define a semantics and talk meaningfully about the
equivalence of two implementations. As a bonus, the dataflow diagrams used by
signal processing specialists can be translated on-the-fly into process
networks. The problem with KPNs is that they rely on an asynchronous execution
model, while VLIW processors and FPGAs are synchronous or partially
synchronous. Thus, there is a need for a tool for synchronizing KPNs. This is
best done by computing a schedule that has to satisfy data dependences within
each process, a causality condition for each channel (a message cannot be
received before it is sent), and real-time constraints. However, there is a
difficulty in writing the channel constraints because one has to count messages
in order to establish the send/receive correspondence and, in multi-dimensional
loop nests, the counting functions may not be affine. In order to bypass this
difficulty, one can define another model, <i>communicating regular
processes</i> (CRP), in which channels are represented as write-once/read-many
arrays. One can then dispense with counting functions. One can prove that the
determinacy property still holds  <ref xlink:href="#compsys-2013-bid11" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. As an added benefit, a
communication system in which the receive operation is not destructive is
closer to the expectations of system designers.</p>
        <p>The main difficulty with this approach is that ordinary programs are usually
not constructed as process networks. One needs automatic or semi-automatic
tools for converting sequential programs into process networks. One
possibility is to start from array dataflow analysis  <ref xlink:href="#compsys-2013-bid12" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. Each
statement (or group of statements) may be considered a process, and the source
computation indicates where to implement communication channels. Another
approach attempts to construct threads, i.e., pieces of sequential code with
the smallest possible interactions. In favorable cases, one may even find
outermost parallelism, i.e., threads with no interactions whatsoever. Here,
communications are associated to so-called uncut dependences, i.e., dependences
which cross thread boundaries. In both approaches, the main question is
whether the communications can be implemented as FIFOs, or need a reordering
memory. One of our research directions will be to try to take advantage of the
reordering allowed by dependences to force a FIFO implementation.</p>
        <p>Whatever the chosen solution (FIFO or addressable memory) for communicating
between two accelerators or between the host processor and an accelerator, the
problems of optimizing communication between processes and of optimizing
buffers have to be addressed. Many local memory optimization problems have
already been solved theoretically. Some examples are loop fusion and loop
alignment for array contraction and for minimizing the length of the reuse
vector  <ref xlink:href="#compsys-2013-bid13" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, techniques for data allocation in scratch-pad memory,
or techniques for folding multi-dimensional arrays  <ref xlink:href="#compsys-2013-bid14" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
Nevertheless, the problem is still largely open. Some questions are: how to
schedule a loop sequence (or even a process network) for minimal scratch-pad
memory size? How is the problem modified when one introduces unlimited and/or
bounded parallelism? How does one take into account latency or throughput
constraints, or bandwidth constraints for input and output channels? All loop
transformations are useful in this context, in particular loop tiling, and may
be applied either as source-to-source transformations (when used in front of
HLS tools) or as transformations to generate directly VHDL codes. One should
keep in mind that theory will not be sufficient to solve these problems.
Experiments are required to check the relevance of the various models
(computation model, memory model, power consumption model) and to select the
most important factors according to the architecture. Besides, optimizations do
interact: for instance, reducing memory size and increasing parallelism are
often antagonistic. Experiments will be needed to find a global compromise
between local optimizations.</p>
        <p>Finally, there remains the problem of code generation for accelerators. It is a
well-known fact that modern methods for program optimization and
parallelization do not generate a new program, but just deliver blueprints for
program generation, in the form, e.g., of schedules, placement functions, or
new array subscripting functions. A separate code generation phase must be
crafted with care, as a too naïve implementation may destroy the benefits
of high-level optimization. There are two possibilities here as suggested
before; one may target another high-level synthesis tool, or one may target
directly VHDL. Each approach has its advantages and drawbacks. However, in both
situations, all such tools
require that the input program respects some strong constraints on the code
shape, array accesses, memory accesses, communication protocols, etc.
Furthermore, to get the tool to do what the user wants requires a lot of
program tuning, i.e., of program rewriting. What can be automated in this
rewriting process? Semi-automated?</p>
      </subsection>
    </subsection>
  </fondements>
  <domaine id="uid55">
    <bodyTitle>Application Domains</bodyTitle>
    <subsection id="uid56" level="1">
      <bodyTitle>Compilers for Embedded Computing Systems</bodyTitle>
      <p>The previous sections described our main activities in terms of research
directions, but also places Compsys within the embedded computing systems
domain, especially in Europe. We will therefore not come back here to the
importance, for industry, of compilation and embedded computing systems
design.</p>
      <p>In terms of application domain, the embedded computing systems we consider
are mostly used for multimedia: phones, TV sets, game platforms, etc. But,
more than the final applications developed as programs, our main application
is <em style="UNDERLINE">the computer itself</em>: how the system is organized
(architecture) and designed, how it is programmed (software), how programs
are mapped to it (compilation and high-level synthesis).</p>
      <p>The industry that can be impacted by our research is thus all the companies
that develop embedded systems and processors, and those (the same plus other)
that need software tools to map applications to these platforms, i.e., that
need to use or even develop programming languages, program optimization
techniques, compilers, operating systems. Compsys do not focus on all
these critical parts, but our activities are connected to them.
</p>
    </subsection>
  </domaine>
  <logiciels id="uid57">
    <bodyTitle>Software and Platforms</bodyTitle>
    <subsection id="uid58" level="1">
      <bodyTitle>Introduction</bodyTitle>
      <p>This section lists and briefly describes the software developments conducted
within Compsys. Most are tools that we extend and maintain over the years.
They mainly concern three activities: a) the development of research tools,
in general available on demand, linked to polyhedra and loop/array
transformations, b) the development of tools linked to the start-up Zettice,
in general not available, c) the development of algorithms within the
back-end compilers of STMicroelectronics and/or Kalray.</p>
      <p>Many tools based on the polyhedral representation of codes with nested loops
are now available. They have been developed and maintained over the years by
different teams, after the introduction of Paul Feautrier's Pip, a tool for
parametric integer linear programming. This “polytope model” view of codes
is now widely accepted: it used by Inria projects-teams Cairn and
Alchemy/Parkas, PIPS at École des Mines de Paris, Suif from Stanford
University, Compaan at Berkeley and Leiden, PiCo from the HP-Labs
(continued as PicoExpress by Synfora and now Synopsis), the DTSE
methodology at Imec, Sadayappan's group at Ohio State University,
Rajopadhye's group at Colorado State's University, etc. More recently,
several compiler groups have shown their interest in polyhedral methods,
e.g., the Gcc group, IBM, and Reservoir Labs, a company that develops a
compiler fully based on the polytope model and on the techniques that we (the
french community) introduced for loop and array transformations.
Polyhedra are also used in test and certification projects (Verimag, Lande,
Vertecs). Now that these techniques are well-established and disseminated in
and by other groups, we prefer to focus on the development of new techniques
and tools, which are described here. Some of these tools can be used through
a web interface on the Compsys tool demonstrator web page
<ref xlink:href="http://compsys-tools.ens-lyon.fr/" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>compsys-tools.<allowbreak/>ens-lyon.<allowbreak/>fr/</ref>.</p>
      <p>The other activity concerns the developments within the compilers of
industrial partners such as STMicroelectronics and Kalray. These are not
stand-alone tools, which could be used externally, but algorithms and data
structures implemented inside the LAO back-end compiler or other compiler
branches, year after year, with the help of STMicroelectronics or Kalray
colleagues. They are also completed by important efforts for integration and
evaluation within the complete compiler toolchains. They concern exact
(ILP-based) methods, algorithms for aggressive optimizations, techniques for
just-in-time compilation, code representations, and for improving the design
of the compiler.</p>
      <p>More recently, an important development activity has been started in the
context of the Zettice start-up project (see
Section <ref xlink:href="#uid122" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). An important effort of
applied research and software development has been achieved since, which
results, in particular, in two major software developments: Dcc (DPN C
Compiler) and IceGEN. These tools are outlined in
Sections <ref xlink:href="#uid75" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>
and <ref xlink:href="#uid76" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid59" level="1">
      <bodyTitle>Pip</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Cédric</firstname>
          <lastname>Bastoul</lastname>
          <moreinfo>professor, Strasbourg University and Inria/CAMUS</moreinfo>
        </person>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
      </participants>
      <p>Paul Feautrier is the main developer of Pip (Parametric Integer
Programming) since its inception in 1988. Basically, Pip is an “all
integer” implementation of the Simplex, augmented for solving integer
programming problems (the Gomory cuts method), which also accepts parameters in
the non-homogeneous term. Pip is freely available under the GPL at
<ref xlink:href="http://www.piplib.org" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>piplib.<allowbreak/>org</ref>. It is widely used in the automatic
parallelization community for testing dependences, scheduling, several kind of
optimizations, code generation, and others. Beside being used in several
parallelizing compilers, Pip has found applications in some unconnected
domains, as for instance in the search for optimal polynomial approximations of
elementary functions (see the Inria project Arénaire).
</p>
    </subsection>
    <subsection id="uid60" level="1">
      <bodyTitle>Syntol</bodyTitle>
      <participants>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
      </participants>
      <p>Syntol is a modular process network scheduler. The source
language is C augmented with specific constructs for representing communicating
regular process (CRP) systems. The present version features a syntax analyzer,
a semantic analyzer to identify DO loops in C code, a dependence computer, a
modular scheduler, and interfaces for CLooG (loop generator developed by
C. Bastoul) and Cl@k (see Sections <ref xlink:href="#uid63" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and <ref xlink:href="#uid70" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). The dependence
computer now handles casts, records (<tt>structures</tt>), and the modulo operator
in subscripts and conditional expressions. The latest developments are,
firstly, a new code generator, and secondly, several experimental tools for the
construction of bounded parallelism programs.</p>
      <simplelist>
        <li id="uid61">
          <p noindent="true">The new code generator, based on the ideas of Boulet and
Feautrier  <ref xlink:href="#compsys-2013-bid15" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, generates a counter automaton that can be
presented as a C program, as a rudimentary VHDL program at the RTL level, as
an automaton in the Aspic input format, or as a drawing specification for the
DOT tool.</p>
        </li>
        <li id="uid62">
          <p noindent="true">Hardware synthesis can only be applied to bounded parallelism programs.
Our present aim is to construct threads with the objective of minimizing
communications and simplifying synchronization. The distribution of
operations among threads is specified using a placement function, which is
found using techniques of linear algebra and combinatorial optimization.</p>
        </li>
      </simplelist>
    </subsection>
    <subsection id="uid63" level="1">
      <bodyTitle>Cl@k</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Fabrice</firstname>
          <lastname>Baray</lastname>
          <moreinfo>Mentor, Former post-doc in
Compsys</moreinfo>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
      </participants>
      <p>Cl@k (Critical LAttice Kernel) is a stand-alone optimization
tool useful for the automatic derivation of array mappings that enable memory
reuse, based on the notions of admissible lattice and of modular allocation
(linear mapping plus modulo operations). It has been developed in 2005-2006 by
Fabrice Baray, former post-doc Inria under Alain Darte's supervision. It computes
or approximates the critical lattice for a given 0-symmetric polytope. (An
admissible lattice is a lattice whose intersection with the polytope is reduced
to 0; a critical lattice is an admissible lattice with minimal determinant.)</p>
      <p>Its application to array contraction has been implemented by Christophe Alias in a
tool called Bee (see Section <ref xlink:href="#uid70" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). Bee uses Rose as a parser,
analyzes the lifetimes of the elements of the arrays to be compressed, and
builds the necessary input for Cl@k, i.e., the 0-symmetric polytope of
conflicting differences. Then, Bee computes the array contraction mapping
from the lattice provided by Cl@k and generates the final program with
contracted arrays. More details on the underlying theory are available in
previous reports. Cl@k can be viewed as a complement to the Polylib
suite, enabling yet another kind of optimizations on polyhedra. Initially,
Bee was the complement of Cl@k in terms of its application to memory
reuse. Now, Bee is a stand-alone tool that contains more and more features
for program analysis and loop transformations.
</p>
    </subsection>
    <subsection id="uid64" level="1">
      <bodyTitle>PoCo</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
      </participants>
      <p>PoCo is a polyhedral compilation framework providing many
features to quickly prototype program analysis and optimizations in the
polyhedral model. Essentially, PoCo provides:</p>
      <simplelist>
        <li id="uid65">
          <p noindent="true">A C front-end extracting the polyhedral representation of the
input program. The parser itself is based on EDG (<i>via</i>
Rose), an industrial C/C++ parser from Edison group used in
Intel compilers.</p>
        </li>
        <li id="uid66">
          <p noindent="true">An extended language of pragmas to feed the source code with
compilation directives (a schedule, for example).</p>
        </li>
        <li id="uid67">
          <p noindent="true">A symbolic layer on polyhedral libraries Polylib (set operations on
polyhedra) and Piplib (parameterized ILP, see
Section <ref xlink:href="#uid59" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). This feature simplifies
drastically the developer task.</p>
        </li>
        <li id="uid68">
          <p noindent="true">Some dependence analysis (polyhedral dependence graph, array dataflow
analysis), array region analysis, array liveness analysis.</p>
        </li>
        <li id="uid69">
          <p noindent="true">A C and VHDL code generation based on the ideas of P. Boulet and
P. Feautrier  <ref xlink:href="#compsys-2013-bid15" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
        </li>
      </simplelist>
      <p>The array dataflow analysis (ADA) of PoCo has been extended to a FADA (Fuzzy
ADA) by M. Belaoucha, former PhD student at Université de Versailles. FADALib
is available at <ref xlink:href="https://bitbucket.org/mbelaoucha/fadalib" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">https://<allowbreak/>bitbucket.<allowbreak/>org/<allowbreak/>mbelaoucha/<allowbreak/>fadalib</ref>. PoCo has
been developed by Christophe Alias. It represents more than 19000 lines of C++
code. The tools Bee, Chuba, and RanK presented thereafter make an
extensive use of PoCo abstractions.
</p>
    </subsection>
    <subsection id="uid70" level="1">
      <bodyTitle>Bee</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
      </participants>
      <p>Bee is a source-to-source optimizer that contracts the temporary arrays of a
program under scheduling constraints. Bee bridges the gap between the
mathematical optimization framework described in  <ref xlink:href="#compsys-2013-bid14" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and
implemented in Cl@k (Section <ref xlink:href="#uid63" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), and effective source-to-source
array contraction. Bee applies a precise lifetime analysis for arrays to
build the mathematical input of Cl@k. Then, Bee derives the array
allocations from the basis found by Cl@k and generates the C code
accordingly. Bee is – to our knowledge – the only complete array
contraction tool.</p>
      <p>Bee is sensitive to the program schedule. This latter feature enlarges the
application field of array contraction to parallel programs. For instance, it
is possible to mark a loop to be software-pipelined (with an affine schedule)
and to let Bee find an optimized array contraction. But the most important
application is the ability to optimize communicating regular processes (CRP).
Given a schedule for every process, Bee can compute an optimized size for the
channels, together with their access functions (the corresponding allocations).
We currently use this feature in source-to-source transformations for
high-level synthesis (see Section <ref xlink:href="#uid45" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).</p>
      <simplelist>
        <li id="uid71">
          <p noindent="true">Bee was made available to STMicroelectronics as a binary.</p>
        </li>
        <li id="uid72">
          <p noindent="true">Bee has been transferred to the (incubated) start-up Zettice,
initiated by Alexandru Plesco.</p>
        </li>
        <li id="uid73">
          <p noindent="true">Bee has been used as an external tool by the compiler Gecos
developed in the Cairn team at Irisa.</p>
        </li>
      </simplelist>
      <p>Bee has been implemented by Christophe Alias, using the compiler infrastructure
PoCo (see Section <ref xlink:href="#uid64" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). It represents more
than 2400 lines of C++ code.</p>
    </subsection>
    <subsection id="uid74" level="1">
      <bodyTitle>Chuba</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="compsys-2006-id18441">
          <firstname>Alexandru</firstname>
          <lastname>Plesco</lastname>
          <moreinfo>Compsys/Zettice</moreinfo>
        </person>
      </participants>
      <p>Chuba is a source-level optimizer that improves a C program in the context
of the high-level synthesis (HLS) of hardware. Chuba is an implementation of
the work described in the PhD thesis of Alexandru Plesco. The optimized program
specifies a system of multiple communicating accelerators, which optimize the
data transfers with the external DDR memory. The program is divided into
blocks of computations obtained thanks to tiling techniques, and, in each
block, data are fetched by block to reduce the penalty due to line changes in
the DDR accesses. Four accelerators achieve data transfers in a
macro-pipeline fashion so that data transfers and computations (performed by a
fifth accelerator) are overlapped.</p>
      <p>So far, the back-end of Chuba is specific to the HLS tool C2H but the
analysis is quite general and adapting Chuba to other HLS tools should be
possible. Besides, it is interesting to mention that the program analysis and
optimizations implemented in Chuba address a problem that is also very
relevant in the context of GPGPUs. The underlying theory and corresponding
experiments are described in <ref xlink:href="#compsys-2013-bid16" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
      <p>Chuba has been implemented by Christophe Alias, using the compiler
infrastructure PoCo (see Section <ref xlink:href="#uid64" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). It
represents more than 900 lines of C++. The reduced size of Chuba is mainly
due to the high-level abstractions provided by PoCo.</p>
    </subsection>
    <subsection id="uid75" level="1">
      <bodyTitle>Dcc</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2006-id18441">
          <firstname>Alexandru</firstname>
          <lastname>Plesco</lastname>
          <moreinfo>Compsys/Zettice</moreinfo>
        </person>
      </participants>
      <p>Dcc (DPN C Compiler) is the <i>front-end</i> of the HLS tool transferred to
the start-up Zettice (see Section <ref xlink:href="#uid122" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).
Dcc takes as input a C program annotated with pragmas and produces an
optimized data-aware process network (DPN). A DPN is a regular process network
that makes explicit the I/O transfers and the synchronizations. Dcc features
throughput optimization, communication vectorization, and automatic
parallelization. Furthermore, Dcc applies analysis to build the DPN
circuitry: multiplexing, channels sizing and allocation, FSM generation. To do
so, Dcc uses extensively the analysis implemented in PoCo
(Section <ref xlink:href="#uid64" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), in particular dataflow analysis
and control generation, and Bee (Section <ref xlink:href="#uid70" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>
for buffer sizing. The DPN specific analysis of Dcc is currently under
patent deposit.</p>
      <p>Dcc represents more than 3000 lines of C++ code.
</p>
    </subsection>
    <subsection id="uid76" level="1">
      <bodyTitle>IceGEN</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2006-id18441">
          <firstname>Alexandru</firstname>
          <lastname>Plesco</lastname>
          <moreinfo>Compsys/Zettice</moreinfo>
        </person>
      </participants>
      <p>IceGEN (Integrated Circuit Generator) is the <i>back-end</i> of the HLS tool
transferred to the start-up Zettice (see
Section <ref xlink:href="#uid122" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). IceGEN takes as input the
DPN produced by Dcc (see
Section <ref xlink:href="#uid75" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and generates:</p>
      <simplelist>
        <li id="uid77">
          <p noindent="true">a SystemC description relevant for fast and accurate circuit
simulation.</p>
        </li>
        <li id="uid78">
          <p noindent="true">a VHDL description of the circuit, which can be mapped
efficiently to an FPGA.</p>
        </li>
      </simplelist>
      <p>IceGEN makes an extensive use of the pipelined arithmetic operators
of the tool FloPoCo  <ref xlink:href="#compsys-2013-bid17" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> developed by Florent De Dinechin, formerly from
Inria ARIC team.</p>
      <p>IceGEN represents more than 6000 lines of C++ code.
</p>
    </subsection>
    <subsection id="uid79" level="1">
      <bodyTitle>C2fsm</bodyTitle>
      <participants>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
      </participants>
      <p>C2fsm is a general tool that converts an arbitrary C program into a counter
automaton. This tool reuses the parser and pre-processor of Syntol (see
Section <ref xlink:href="#uid60" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), which has been greatly
extended to handle <tt>while</tt> and <tt>do while</tt> loops, <tt>goto</tt>, <tt>break</tt>, and <tt>continue</tt> statements. C2fsm reuses also part of the code
generator of Syntol and has several output formats, including FAST (the
input format of Aspic, see Section <ref xlink:href="#uid80" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), a
rudimentary VHDL generator, and a DOT generator which draws the output
automaton. C2fsm is also able to do elementary transformations on the
automaton, such as eliminating useless states, transitions and variables,
simplifying guards, or selecting cut-points, i.e., program points on loops that
can be used by RanK (see Section <ref xlink:href="#uid81" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) to prove program termination.
</p>
    </subsection>
    <subsection id="uid80" level="1">
      <bodyTitle>Aspic</bodyTitle>
      <participants>
        <person key="compsys-2008-id18330">
          <firstname>Laure</firstname>
          <lastname>Gonnord</lastname>
        </person>
      </participants>
      <p>Aspic is an invariant generator for general counter automata. Used with
C2fsm (see Section <ref xlink:href="#uid79" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), it can be used
to derivate invariant for numerical C programs, and also prove safety. It is
also part of the WTC toolsuite (see
<ref xlink:href="http://compsys-tools.ens-lyon.fr/wtc/index.html" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>compsys-tools.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>wtc/<allowbreak/>index.<allowbreak/>html</ref>), a set of examples to
demonstrate the capability of the RanK tool (see
Section <ref xlink:href="#uid81" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) for evaluating worse-case time
complexity (number of transitions when executing an automaton).</p>
      <p>Aspic implements the theoretical results of Laure Gonnord's PhD
thesis on acceleration techniques and has been maintained since 2007.</p>
    </subsection>
    <subsection id="uid81" level="1">
      <bodyTitle>RanK</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
        <person key="compsys-2008-id18330">
          <firstname>Laure</firstname>
          <lastname>Gonnord</lastname>
          <moreinfo>Compsys</moreinfo>
        </person>
      </participants>
      <p>RanK is a software tool that can prove the termination of a program (in some
cases) by computing a <i>ranking function</i>, i.e., a mapping from the
operations of the program to a well-founded set that <i>decreases</i> as the
computation advances. In case of success, RanK can also provide an upper
bound of the worst-case time complexity of the program as a symbolic affine
expression involving the input variables of the program (parameters), when it
exists. In case of failure, RanK tries to prove the non-termination of the
program and then to exhibit a counter-example input. This last feature is of
great help for program understanding and debugging, and has already been
experimented. The theory underlying RanK was presented at
SAS'10  <ref xlink:href="#compsys-2013-bid18" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
      <p>The input of RanK is an integer automaton, computed by C2fsm (see
Section <ref xlink:href="#uid79" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), representing the control
structure of the program to be analyzed. RanK uses the Aspic tool (see
Section <ref xlink:href="#uid80" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), developed by Laure Gonnord during
her PhD thesis, to compute automaton invariants. RanK has been used to
discover successfully the worst-case time complexity of many benchmarks
programs of the community (see the WTC benchmark suite
<ref xlink:href="http://compsys-tools.ens-lyon.fr/wtc/index.html" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>compsys-tools.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>wtc/<allowbreak/>index.<allowbreak/>html</ref>). It uses the libraries
Piplib (Section <ref xlink:href="#uid59" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and Polylib.</p>
      <p>RanK has been implemented by Christophe Alias, using the compiler
infrastructure PoCo (Section <ref xlink:href="#uid64" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). It
represents more than 3000 lines of C++. The tool has been presented at the
CSTVA'13 workshop <ref xlink:href="#compsys-2013-bid19" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid82" level="1">
      <bodyTitle>SToP</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Guillaume</firstname>
          <lastname>Andrieu</lastname>
          <moreinfo>University of Lille</moreinfo>
        </person>
        <person key="compsys-2008-id18330">
          <firstname>Laure</firstname>
          <lastname>Gonnord</lastname>
          <moreinfo>Compsys</moreinfo>
        </person>
      </participants>
      <p>SToP (Scalable Termination of Programs) is the implementation of the
modular termination technique presented at the TAPAS'12
workshop  <ref xlink:href="#compsys-2013-bid20" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. It takes as input a large irregular
C program and conservatively checks its termination. To do so, SToP
generates a set of small programs whose termination implies the termination of
the whole input program. Then, the termination of each small program is checked
thanks to RanK (see Section <ref xlink:href="#uid81" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). In case
of success, SToP infers a ranking (schedule) for the whole program. This
schedule can be used in a subsequent analysis to optimize the program.</p>
      <p>SToP represents more than 2000 lines of C++.</p>
    </subsection>
    <subsection id="uid83" level="1">
      <bodyTitle>Simplifiers</bodyTitle>
      <participants>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
      </participants>
      <p>The aim of the <tt>simple</tt> library is to simplify Boolean formulas on affine
inequalities. It works by detecting redundant inequalities in the
representation of the subject formula as an ordered binary decision diagram
(OBDD), see details in  <ref xlink:href="#compsys-2013-bid21" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. It uses PIP (see
Section <ref xlink:href="#uid59" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) for testing the feasibility –
or unfeasibility – of a conjunction of affine inequalities.</p>
      <p>The library is written in Java and is presented as a collection of class files.
For experimentation, several front-ends have been written. They differ mainly
in their input syntax, among which are a C like syntax, the Mathematica and
SMTLib syntaxes, and an ad hoc Quast (quasi-affine syntax tree) syntax.
</p>
    </subsection>
    <subsection id="uid84" level="1">
      <bodyTitle>LAO Developments in Aggressive Compilation</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Benoit</firstname>
          <lastname>Boissinot</lastname>
        </person>
        <person key="compsys-2005-id18414">
          <firstname>Florent</firstname>
          <lastname>Bouchez</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Florian</firstname>
          <lastname>Brandner</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Quentin</firstname>
          <lastname>Colombet</lastname>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Benoît</firstname>
          <lastname>Dupont de Dinechin</lastname>
          <moreinfo>Kalray</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Christophe</firstname>
          <lastname>Guillon</lastname>
          <moreinfo>STMicroelectronics</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Sebastian</firstname>
          <lastname>Hack</lastname>
          <moreinfo>Former
post-doc in Compsys</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Cédric</firstname>
          <lastname>Vincent</lastname>
          <moreinfo>Former student in Compsys</moreinfo>
        </person>
      </participants>
      <p>Our past aggressive optimization techniques are all implemented in
stand-alone experimental tools (as for example for register coalescing
algorithms) or within LAO, the back-end compiler of STMicroelectronics, or both. They
concern SSA construction and destruction, instruction-cache optimizations,
register allocation. Here, we report only our activities related to register
allocation.</p>
      <p>Our developments on register allocation within the STMicroelectronics compiler
started when Cédric Vincent (bachelor degree, under Alain Darte supervision)
developed a complete register allocator in LAO, the assembly-code optimizer
of STMicroelectronics. This was the first time a complete implementation was done
with success, outside the MCDT (now CEC) team, in their optimizer. This
continued with developments made during the master internships and PhD theses
of Florent Bouchez, Benoit Boissinot, and Quentin Colombet, and post-doctoral works of
Sebastian Hack and Florian Brandner. In 2009, Quentin Colombet started to develop and
integrate into the main trunk of LAO a full implementation of a two-phases
register allocation. This implementation now includes two different decoupled
spilling phases, the first one as described in Sebastian Hack's PhD thesis and
a second ILP-based solution. It also includes an up-to-date graph-based
register coalescing. Finally, since all these optimizations take place under
SSA form, it includes also a mechanism for going out of colored-SSA
(register-allocated SSA) form that can handle critical edges and does further
optimizations. See details in the “new results” presented in previous Compsys
activity reports.
</p>
    </subsection>
    <subsection id="uid85" level="1">
      <bodyTitle>LAO Developments in JIT Compilation</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Benoit</firstname>
          <lastname>Boissinot</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Florian</firstname>
          <lastname>Brandner</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Quentin</firstname>
          <lastname>Colombet</lastname>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Benoît</firstname>
          <lastname>Dupont de Dinechin</lastname>
          <moreinfo>Kalray</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Christophe</firstname>
          <lastname>Guillon</lastname>
          <moreinfo>STMicroelectronics</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
      </participants>
      <p>The other side of our work in the STMicroelectronics compiler LAO has been to adapt
the compiler to make it more suitable for JIT compilation. This means
lowering the time and space complexity of several algorithms. In particular
we implemented our fast out-of-SSA translation method, and we programmed and
tested various ways to compute the liveness information. Recent efforts also
focused on developing a tree-scan register allocator for the JIT part of the
compiler, in particular a JIT conservative coalescing. The technique is to
bias the tree-scan coalescing, taking into account register constraints, with
the result of a JIT aggressive coalescing. See details in the “new results”
presented in previous Compsys activity reports.
</p>
    </subsection>
    <subsection id="uid86" level="1">
      <bodyTitle>Low-Level Exchange Format (TireX) and
Minimalist Intermediate Representation (MinIR)</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Christophe</firstname>
          <lastname>Guillon</lastname>
          <moreinfo>STMicroelectronics</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Benoît</firstname>
          <lastname>Dupont de Dinechin</lastname>
          <moreinfo>Kalray</moreinfo>
        </person>
      </participants>
      <p>Most compilers define their own intermediate representation (IR) to be able
to work on a program. Sometimes, they even use a different representation for
each representation level, from source code parsing to the final object code
generation. MinIR (Minimalist Intermediate Representation) is a new
intermediate representation, designed to ease the interconnection of
compilers, static analyzers, code generators, and other tools. In addition
to the specification of MinIR, generic core tools have been developed to
offer a basic toolkit and to help the connection of client tools. MinIR
generators exist for several compilers, and different analyzers are developed
as a testbed to rapidly prototype different static analyses over SSA code.
This new common format enables the comparison of the code generator of
several production compilers, and simplifies the connection of external tools
to existing compilers.</p>
      <p>MinIR has been extended into TireX, a Textual Intermediate Representation
for EXchanging target-level information between compiler optimizers and whole
or parts of code generators (a.k.a., compiler back-end). The first motivation for
this intermediate representation is to factor target-specific compiler
optimizations into a single component, in case several compilers need to be
maintained for a particular target (e.g., operating system compiler and
application code compiler). Another motivation is to reduce the run-time cost
of JIT compilation and of mixed mode execution, since the program to compile
is already in a representation lowered to the level of the target processor.
Beside the lowering at the target level, the extensions of MinIR include the
program data stream and loop scoped information. TireX is currently
produced by the Open64/Path64 and the LLVM compilers, with a GCC producer
under work. It is used by the LAO code generator.</p>
      <p>Detailed information, generic core tools, and LLVM IR based generator for
MinIR are available at <ref xlink:href="http://www.assembla.com/spaces/minir-dev/wiki" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>assembla.<allowbreak/>com/<allowbreak/>spaces/<allowbreak/>minir-dev/<allowbreak/>wiki</ref>.
MinIR was presented at
WIR'11  <ref xlink:href="#compsys-2013-bid22" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
  </logiciels>
  <resultats id="uid87">
    <bodyTitle>New Results</bodyTitle>
    <subsection id="uid88" level="1">
      <bodyTitle>Parameterized Construction of Program Representations for Sparse Dataflow Analysiss</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>André</firstname>
          <lastname>Tavares</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Benoit</firstname>
          <lastname>Boissinot</lastname>
          <moreinfo>Ex-Compsys, Google Zurich</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Fernando</firstname>
          <lastname>Magno Quintão Pereira</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
      </participants>
      <p>Data-flow analysis usually associates information with control flow regions.
Informally, if these regions are too small like a point between two consecutive
statements, we call the analysis dense. On the other hand, if these regions
include many such points, then we call it sparse. This work presents a
systematic method to build program representations that support sparse
analyses. To pave the way to this framework, we clarify the literature about
well-known intermediate program representations. We show that our approach,
subsumes, up to parameter choices, many of these representations, such as the
SSA, SSI, and e-SSA forms. In particular, our algorithms are faster, simpler
and more frugal than the previous techniques used to construct SSI (static
single information) form programs. We produce intermediate representations
isomorphic to Choi <i>et al.</i>'s sparse evaluation graphs (SEG) for the family
of data-flow problems that can be partitioned by variables. However, contrary
to SEGs, we can handle - sparsely - problems that are not in this family. We
have tested our ideas in the LLVM compiler, comparing different program
representations in terms of size and construction time.</p>
      <p>This work is part of the collaboration with UFMG (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and has
been accepted for presentation and publication at CC'14 (Compiler Construction
Conference) <ref xlink:href="#compsys-2013-bid23" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid89" level="1">
      <bodyTitle>A Framework for Enhancing Data Reuse via Associative Reordering</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Kevin</firstname>
          <lastname>Stock</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Louis-Noël</firstname>
          <lastname>Pouchet</lastname>
          <moreinfo>UCLA, Los Angeles, USA</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>J.</firstname>
          <lastname>Ramanujam</lastname>
          <moreinfo>LSU, Houston, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>P.</firstname>
          <lastname>Sadayappan</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
      </participants>
      <p>The freedom to reorder computations involving associative operators
has been widely recognized and exploited in designing parallel
algorithms and to a more limited extent in optimizing compilers.
However, the use of associative reordering for enhancing data
locality has not been previously explored to our knowledge.</p>
      <p>In this work, we develop a novel framework for utilizing associativity
of operations in regular loop computations to enhance register reuse.
Stencils represent a particular class of important computations where
our optimization framework can be applied to enhance performance.
We use a multi-dimensional retiming formalism to characterize the
space of valid transformations and to generate the transformed code.
Experimental results demonstrate the effectiveness of the framework.</p>
      <p>This work has been submitted to PLDI'14 and is part of the
collaboration with P. Sadayappan from the University of Columbus (OSU) (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).
</p>
    </subsection>
    <subsection id="uid90" level="1">
      <bodyTitle>Function Cloning Revisited</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Matheus</firstname>
          <lastname>Vilela</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Guilherme</firstname>
          <lastname>Balena</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Guilherme</firstname>
          <lastname>Marques</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Fernando</firstname>
          <lastname>Magno Quintão Pereira</lastname>
          <moreinfo>UFMG, Belo Horizonte, Brazil</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
      </participants>
      <p>Compilers rely on two main techniques to implement optimizations that depend on
the calling context of functions: inlining and cloning. Historically, function
inlining has seen more widespread use, as it tends to be more effective in
practice. Yet, function cloning provides benefits that inline leaves behind. In
particular, cloning gives the program developer a way to fight performance
bugs, because it generates reusable code. Furthermore, it deals with recursion
more naturally. Finally, it might lead to less code expansion, the inlining's
nemesis.</p>
      <p>In this work, we revisited function cloning under the light of these benefits.
We discuss four independent code specialization techniques based on function
cloning, which, although simple, find wide applicability, even in highly
optimized benchmarks, such as SPEC CPU 2006. We claim that our optimizations
are easy to implement and to deploy. We use Wu and Larus's well-known static
profiling heuristic to measure the profitability of a clone. This metric gives
us a concrete way to point out to program developers potential performance
bugs, and gives us a metric to decide if we should keep a clone or not. By
implementing our ideas in LLVM, we have been able to speed up some of the SPEC
benchmarks by up to 6% on top of the -O2 optimization level.</p>
      <p>This work is part of the collaboration with UFMG (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and was
also done in the context of the collaboration with Kalray and the ManycoreLabs
project (see Section <ref xlink:href="#uid121" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).
</p>
    </subsection>
    <subsection id="uid91" level="1">
      <bodyTitle>Register Allocation and Promotion through Combined Instruction Scheduling, Loop Splitting and Unrolling</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>P.</firstname>
          <lastname>Sadayappan</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Lukasz</firstname>
          <lastname>Domanaga</lastname>
        </person>
      </participants>
      <p>Register allocation is a much studied problem. A particularly
important context for optimizing register allocation is within loops,
since a significant fraction of the execution time of programs is
often inside loop code. A variety of algorithms have been proposed in
the past for register allocation, but the complexity of the problem
has resulted in a decoupling of several important aspects, including
loop unrolling, loop fission, register promotion, and instruction
reordering.</p>
      <p>In this work, we develop an approach to register allocation and
promotion in a unified optimization framework that simultaneously
considers the impact of loop unrolling, loop splitting, and
instruction scheduling. This is done via a novel instruction tiling
approach where instructions within a loop are represented along one
dimension and innermost loop iterations along the other dimension. By
exploiting the regularity along the loop dimension, and a constrained
intra-tile execution order, the problem of optimizing register
pressure is cast in a constraint programming formalism. Experimental
results are provided from thousands of innermost loops extracted from
the SPEC benchmarks, demonstrating improvements over the current
state of the art.</p>
      <p>This work is part of the collaboration with OSU (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and was
also done in the context of the collaboration with Kalray and the ManycoreLabs
project (see Section <ref xlink:href="#uid121" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). It
contributes to the developments of the Tirex toolbox
(see <ref xlink:href="#uid86" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). It has also been submitted to
PLDI'14.
</p>
    </subsection>
    <subsection id="uid92" level="1">
      <bodyTitle>Beyond Reuse Distance Analysis: Dynamic Analysis for Characterization of Data Locality Potential</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Naznin</firstname>
          <lastname>Fauzia</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Venmugil</firstname>
          <lastname>Elango</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Mahesh</firstname>
          <lastname>Ravishankar</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>J. (ram)</firstname>
          <lastname>Ramanujam</lastname>
          <moreinfo>LSU, Houston, USA</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Atanas</firstname>
          <lastname>Rountev</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Louis-Noël</firstname>
          <lastname>Pouchet</lastname>
          <moreinfo>UCLA, Los Angeles, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>P.</firstname>
          <lastname>Sadayappan</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
      </participants>
      <p>Emerging computer architectures will feature drastically decreased
flops/byte (ratio of peak processing rate to memory bandwidth) as
highlighted by recent studies on Exascale architectural
trends. Further, flops are getting cheaper while the energy cost of
data movement is increasingly dominant. The understanding and
characterization of data locality properties of computations is
critical in order to guide efforts to enhance data locality.</p>
      <p>Reuse distance analysis of memory address traces is a valuable tool to
perform data locality characterization of programs. A single reuse
distance analysis can be used to estimate the number of cache misses
in a fully associative LRU cache of any size, thereby providing
estimates on the minimum bandwidth requirements at different levels of
the memory hierarchy to avoid being bandwidth bound. However, such an
analysis only holds for the particular execution order that produced
the trace. It cannot estimate potential improvement in data locality
through dependence preserving transformations that change the
execution schedule of the operations in the computation.</p>
      <p>In this work, we develop a novel dynamic analysis approach to
characterize the inherent locality properties of a computation and
thereby assess the potential for data locality enhancement via
dependence preserving transformations.
The execution trace of a code is analyzed to extract a computational
directed acyclic graph (CDAG) of the data dependences. The CDAG is
then partitioned into convex subsets, and the convex partitioning is
used to reorder the operations in the execution trace to enhance data
locality. The approach enables us to go beyond reuse distance analysis
of a single specific order of execution of the operations of a
computation in characterization of its data locality properties. It
can serve a valuable role in identifying promising code regions for
manual transformation, as well as assessing the effectiveness of
compiler transformations for data locality enhancement. We demonstrate
the effectiveness of the approach using a number of benchmarks,
including case studies where the potential shown by the analysis is
exploited to achieve lower data movement costs and better performance.</p>
      <p>This work is part of the collaboration with OSU (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and has
been accepted for publication at ACM TACO <ref xlink:href="#compsys-2013-bid4" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid93" level="1">
      <bodyTitle>Characterizing the Inherent Data Movement Complexity of Computations via Lower Bounds</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>P.</firstname>
          <lastname>Sadayappan</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Venmugil</firstname>
          <lastname>Elango</lastname>
          <moreinfo>OSU, Columbus, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>J. (ram)</firstname>
          <lastname>Ramanujam</lastname>
          <moreinfo>LSU, Houston, USA</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Louis-Noël</firstname>
          <lastname>Pouchet</lastname>
          <moreinfo>UCLA, Los Angeles, USA</moreinfo>
        </person>
        <person key="compsys-2005-id18199">
          <firstname>Fabrice</firstname>
          <lastname>Rastello</lastname>
        </person>
      </participants>
      <p>Technology trends will cause data movement to account for the majority of
energy expenditure and execution time on emerging computers. Therefore,
computational complexity will no longer be a sufficient metric for comparing
algorithms, and a fundamental characterization of data access complexity will
be increasingly important. Although the problem of characterizing data access
complexity has been modeled previously using the formalism of Hong &amp; Kung's
red/blue pebble game  <ref xlink:href="#compsys-2013-bid24" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, applicability of previously-developed approaches has
been extremely limited. We improve on prior work in several ways: 1) we
develop an approach to composing lower bounds from arbitrary decompositions of
computational directed acyclic graphs, thereby eliminating a significant
limitation of previous approaches that required homogeneity of analyzed
computations, 2) we develop a complementary graph min-cut based strategy to
Hong &amp; Kung's S-partitioning approach, and 3) we develop an automated
approach to generate concrete I/O lower bounds of arbitrary, possibly
irregular computational directed acyclic graphs. We provide experimental
results demonstrating the utility of the developed approach.</p>
      <p>This work has been submitted to PLDI'14 and is part of an informal
collaboration with P. Sadayappan from the University of Columbus (OSU) (see
Section <ref xlink:href="#uid131" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).
</p>
    </subsection>
    <subsection id="uid94" level="1">
      <bodyTitle>Enhancing the Compilation of Synchronous Dataflow Programs</bodyTitle>
      <participants>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Abdoulaye</firstname>
          <lastname>Gamatié</lastname>
          <moreinfo>LIRMM, Montpellier</moreinfo>
        </person>
        <person key="compsys-2008-id18330">
          <firstname>Laure</firstname>
          <lastname>Gonnord</lastname>
        </person>
      </participants>
      <p>In this work <ref xlink:href="#compsys-2013-bid25" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, which is an extension
of  <ref xlink:href="#compsys-2013-bid26" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, we propose an enhancement of the
compilation of synchronous programs with a combined numerical-Boolean
abstraction. While our approach applies to synchronous dataflow languages in
general, here, we consider the SIGNAL language for illustration. In the new
abstraction, every signal in a program is associated with a pair of the form
(<tt>clock</tt>, <tt>value</tt>), where <tt>clock</tt> is a Boolean function
and <tt>value</tt> is a Boolean or numeric function. Given the performance
level reached by recent progress in satisfiability modulo theory (SMT), we use
an SMT solver to reason on this abstraction. Through sample examples, we show
how our solution is used to determine absence of reaction captured by empty
clocks; mutual exclusion captured by two or more clocks whose associated
signals never occur at the same time; or hierarchical control of component
activations via clock inclusion. We also show that the analysis improves the
quality of the code generated automatically by a compiler, e.g., a code with
smaller footprint, or a code executed more efficiently thanks to optimizations
enabled by the new abstraction. The implementation of the whole approach
includes a translator of synchronous programs towards the standard input
format of SMT solvers, and an ad hoc SMT solver that integrates advanced
functionalities to cope with the issues of interest in this work. These
results have been published in 2013 (but considered as published in 2012) in
the CSI Journal of Computing  <ref xlink:href="#compsys-2013-bid27" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid95" level="1">
      <bodyTitle>Synthesis of Ranking Functions
using Extremal Counter-Examples</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>David</firstname>
          <lastname>Monniaux</lastname>
          <moreinfo>Verimag, Grenoble</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Lucas</firstname>
          <lastname>Séguinot</lastname>
          <moreinfo>Student at ENS Cachan Bretagne</moreinfo>
        </person>
        <person key="compsys-2008-id18330">
          <firstname>Laure</firstname>
          <lastname>Gonnord</lastname>
        </person>
      </participants>
      <p>In  <ref xlink:href="#compsys-2013-bid18" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>, we presented a new algorithm adapted from scheduling
techniques to synthesize (multi-dimensional) affine functions from general
flowcharts programs. But, as for other methods, our algorithm tried to solve
linear constraints on each control point and each transition, which can lead
to quasi-untractable linear programming instances.</p>
      <p>In contrast to these approaches, we proposed a new algorithm based on the
following observations:</p>
      <simplelist>
        <li id="uid96">
          <p noindent="true">Searching for ranking functions for loop headers is sufficient to prove
termination.</p>
        </li>
        <li id="uid97">
          <p noindent="true">Furthermore, there exist loops such that there is a linear lexicographic
ranking function that decreases along each path inside the loop, from one
loop iteration to the next, but such that there is no lexicographic linear
ranking function that decreases at each step along these paths. For these
reasons, it is tempting to treat each path inside a loop as a single
transition.</p>
        </li>
      </simplelist>
      <p>Unfortunately the number of paths may be exponential in the size of the
program, thus the constraint system may become very large, even though it
features fewer variables. To face this theoretical complexity, even though the
number of paths may be large, we argue that, in practice, few of them actually
matter in the constraint system (we formalize this concept by giving a
characterization as geometric extremal points). Our algorithm therefore builds
the constraint system lazily, taking paths into account <i>on demand</i>.</p>
      <p>We are currently testing our preliminary implementation and submitting
a paper on these new results.</p>
    </subsection>
    <subsection id="uid98" level="1">
      <bodyTitle>Data-Aware Process Networks</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="compsys-2006-id18441">
          <firstname>Alexandru</firstname>
          <lastname>Plesco</lastname>
        </person>
      </participants>
      <p>The following results concern the applied research activities directly linked
to the Zettice start-up (see
Section <ref xlink:href="#uid122" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), which aims at applying
polyhedral techniques to high-level circuit synthesis (HLS). Following the
guidelines of Inria DTI, as this research aims to be transferred, these results
are not published before being “protected” or exploited. An Inria patent
deposit is currently processed.</p>
      <simplelist>
        <li id="uid99">
          <p noindent="true"><b>Data-aware process networks (DPN).</b> This is the
intermediate representation of the HLS flow. DPN is a parallel
execution model fitting the hardware constraints of circuit
synthesis, in which the data transfer and the synchronizations are
made explicit. We formally described the DPN model and a translation
scheme from C programs, and we showed the consistency in the meaning
where any terminating sequential program is translated to an
equivalent DPN, guaranteed to be deadlock free.</p>
        </li>
        <li id="uid100">
          <p noindent="true"><b>Front-end analysis.</b> We designed many program analyses to
produce a quality DPN from a C program:</p>
          <simplelist>
            <li id="uid101">
              <p noindent="true"><i>Throughput optimization.</i> A I/O scheme has been designed, with
the corresponding compiler analysis, to minimize the I/O traffic
with the external memory. This allows us to balance efficiently the
spilling of temporary value to the memory, and the local buffer
size. This scheme impacts the DPN structure itself.</p>
            </li>
            <li id="uid102">
              <p noindent="true"><i>Communication vectorization.</i> The matrix structure of the memory
allows us to load data by chunks. A polyhedral analysis has been
designed to solve this issue.</p>
            </li>
            <li id="uid103">
              <p noindent="true"><i>Synchronization scheme.</i> As parallel units need to
communicate intermediate results, synchronizations must be
ensured.Unlike KPN, DPN do not use FIFO, but buffers, which
required an efficient synchronization mechanism.</p>
            </li>
          </simplelist>
        </li>
        <li id="uid104">
          <p noindent="true"><b>Back-end analysis.</b> Once generated, a DPN must be mapped to an FPGA. This
raises many interesting issues:</p>
          <simplelist>
            <li id="uid105">
              <p noindent="true"><i>Pipeline completion.</i> Data paths make an extensive use of
pipelined operators, which delays the signal. An algorithm has
been designed to enforce the time coherence of signals.</p>
            </li>
            <li id="uid106">
              <p noindent="true"><i>Polyhedral units.</i> DPNs make an extensive use of
piece-wise affine functions, which must be mapped properly to
ensure the efficiency of the whole system. A preliminary algorithm
has been designed to reach a correct trade-off between critical
path size and LUT usage.</p>
            </li>
          </simplelist>
        </li>
      </simplelist>
      <p>All these analyses have been fully implemented. The tool Dcc (DPN
C Compiler) implements all the front-end analyses. The tool IceGEN
implements the back-end analysis.
</p>
    </subsection>
    <subsection id="uid107" level="1">
      <bodyTitle>Program Equivalence Modulo A/C (Associativity/Commutativity)</bodyTitle>
      <participants>
        <person key="compsys-2013-idp140327620777744">
          <firstname>Guillaume</firstname>
          <lastname>Iooss</lastname>
          <moreinfo>PhD student</moreinfo>
        </person>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Sanjay</firstname>
          <lastname>Rajopadhye</lastname>
          <moreinfo>Colorado State University</moreinfo>
        </person>
      </participants>
      <p>Program equivalence is a well-known problem with a wide range of applications,
such as algorithm recognition, program verification, and program optimization.
This problem is also known to be undecidable if the class of programs is rich
enough, in which case semi-algorithms are commonly used. We focus on programs
represented as a system of affine recurrence equations (SARE), defined over
parametric polyhedral domains, a well-known formalism for the <i>polyhedral
model</i>, which includes as a proper subset, the class of affine control loop
programs. Several semi-algorithms for program equivalence have already been
proposed for this class. A few of them take into account algebraic properties
such as associativity and commutativity. However, to the best of our
knowledge, none of them is able to manage reductions, i.e., accumulations of a
parametric number of sub-expressions using an associative and commutative
operator.</p>
      <p>Our contributions are:</p>
      <simplelist>
        <li id="uid108">
          <p noindent="true">An equivalence checking algorithm able to manage associativity and
commutativity properties. Our method subsumes the previous approaches and
is, to the best of our knowledge, the first one able to manage these
properties over a parametric number of expressions.</p>
        </li>
        <li id="uid109">
          <p noindent="true">A semi-algorithm to construct a perfect matching problem on a
parametric bipartite graph. We partially solve this problem through a
heuristic based on the augmenting path algorithm. This heuristic is able
to find a set of non-interfering augmenting paths to improve a proposed
maximum matching, as long as these augmenting paths do not have a
parametric length.</p>
        </li>
      </simplelist>
      <p>A preliminary implementation is under development. This work has
been submitted to ESOP'14.
</p>
    </subsection>
    <subsection id="uid110" level="1">
      <bodyTitle>Constant Aspect-Ratio Parametric Tiling</bodyTitle>
      <participants>
        <person key="compsys-2013-idp140327620777744">
          <firstname>Guillaume</firstname>
          <lastname>Iooss</lastname>
          <moreinfo>PhD student</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Sanjay</firstname>
          <lastname>Rajopadhye</lastname>
          <moreinfo>Colorado State University</moreinfo>
        </person>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Yun</firstname>
          <lastname>Zou</lastname>
          <moreinfo>PhD student, Colorado State University</moreinfo>
        </person>
      </participants>
      <p>Parametric tiling is a well-known transformation that is widely used to
improve locality, parallelism, and granularity. However, parametric tiling
is also a non-linear transformation and this prevents polyhedral analysis or
further polyhedral transformation after parametric tiling. It is therefore
generally applied during the code generation phase.</p>
      <p>This result consists on a method to stay polyhedral in a special
case of parametric tiling, where all the dimensions are tiled and
all the tile sizes are constant multiples of a single tile size
parameter. We call this <i>Constant Aspect Ratio Tiling</i>. We
show how to mathematically transform a polyhedron and an affine
function into their tiled counterpart and show how to obtain good
generated code.</p>
      <p>This work has been accepted for publication at IMPACT'14
<ref xlink:href="#compsys-2013-bid0" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid111" level="1">
      <bodyTitle>Parametric Tiling with Inter-Tile
Data Reuse</bodyTitle>
      <participants>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="compsys-2012-idp140679680158976">
          <firstname>Alexandre</firstname>
          <lastname>Isoard</lastname>
        </person>
      </participants>
      <p>Loop tiling is a loop transformation widely used to improve spatial and
temporal data locality, increase computation granularity, and enable blocking
algorithms, which are particularly useful when offloading kernels on platforms
with small memories. When hardware caches are not available, data transfers
must be software-managed: they can be reduced by exploiting data reuse between
tiles and, this way, avoid some useless external communications. An important
parameter of loop tiling is the sizes of the tiles, which impact the size of
the necessary local memory. However, for most analyzes that involve several
tiles, which is the case for inter-tile data reuse, the tile sizes induce
non-linear constraints, unless they are numerical constants. This complicates
or prevents a parametric analysis. In this work, we showed that, actually,
parametric tiling with inter-tile data reuse is nevertheless possible.</p>
      <p>Our solution is the first parametric solution for generating the memory
transfers needed when a kernel is offloaded to a distant accelerator, tile by
tile after loop tiling, and when all intermediate results are stored locally on
the accelerator. For such computations, there is a complete decoupling between
loads and stores, and when a value has been defined in a previous tile, it has
to be loaded from the local memory and not from the distant memory as this
memory is not yet up-to-date. In other words, inter-tile reuse is mandatory.
This also saves external communications. Our solution is parametric in the
sense that we derive the set of loads and stores from and to the distant memory
with the tile sizes as parameters. Although the direct formulation is
quadratic, we can still solve it in an affine way by developing techniques that
consider, in the analysis, all (unaligned) possible tiles obtained by
translation and not just those that belong to a tiling (partitioning) of the
iteration space. We were able to use a similar technique to also parameterize
the computations of local memory sizes, thanks to parametric lifetime analysis
and folding with modulos, even for pipeline schedules similar to double
buffering. Our method is currently implemented with the <tt>iscc</tt>
calculator of <tt>ISL</tt>, a library for the manipulation of integer sets
defined with Presburger arithmetic.</p>
      <p>Also, the whole analysis can handle approximations thanks to the introduction
of the concept of pointwise functions, well suited to deal with unaligned
tiles. We believe that this technique can be used for other applications linked
to the extension of the polyhedral model as it turns out to be fairly powerful.
Our future work will be to derive efficient approximation techniques, either
because the program cannot be fully analyzable, or because approximations can
speed-up or simplify the results of the analysis without losing much in terms
of memory transfers and/or memory sizes.</p>
      <p>This work has been accepted for publication at
IMPACT'14 <ref xlink:href="#compsys-2013-bid1" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid112" level="1">
      <bodyTitle>Data Races in the Parallel Language X10</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Tomofumi</firstname>
          <lastname>Yuki</lastname>
          <moreinfo>Colorado State University and Inria/IRISA</moreinfo>
        </person>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Sanjay</firstname>
          <lastname>Rajopadhye</lastname>
          <moreinfo>Colorado State University</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Vijay</firstname>
          <lastname>Saraswat</lastname>
          <moreinfo>IBM Research</moreinfo>
        </person>
      </participants>
      <p>Parallel programmers are now required to efficiently utilize the massive
amount of parallelism provided by multi-core and many-core systems. Parallel
programming is difficult, and the existing tools are mostly low-level
extensions to sequential languages or libraries. As an effort to improve this
situation, several groups have initiated the design of parallel programming
languages, mostly based on the partitioned global address space (PGAS)
paradigm. One of these languages is X10, which is developed at IBM Research by
a team led by Vijay Saraswat.</p>
      <p>While such languages hide the low-level details of parallel programming, they
cannot guarantee that the object code will be correct by construction.
Parallelism introduces two new types of bugs: non-determinism and deadlocks,
and experience shows that it is possible to guarantee the absence of one type
but not both. X10 programs are guaranteed deadlock-free but may have
non-determinism. Non-determinism can be detected at runtime, but this approach
cannot give absolute guarantees. However, it is possible, at least for a
restricted class of X10 programs, to check for non-determinism at compile
time.</p>
      <p>The first step in this direction is to define the <i>polyhedral fragment</i>
of X10, in which the only control constructs are <tt>for</tt> loops with
affine bounds, and the only data structures are arrays with affine subscripts.
X10 has many parallel constructs: as a first effort, we focused on
<tt>async</tt>, which creates an activity (lightweight thread) and
<tt>finish</tt>, which waits for termination of all impending activities. The
execution order (or <i>happens-before relation</i>) of such a program is an
incomplete lexicographic order, in which terms relating operations in
different activities are removed. The dataflow analysis method
of  <ref xlink:href="#compsys-2013-bid12" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> has to be adapted to a partial execution order, which
may have many extrema instead of a unique maximum. Multiple extrema denote
data races, thus non-determinism. A detector along these lines has been
implemented and presented at PPoPP'13 (Symposium on Principles and Practice of
Parallel Programming) <ref xlink:href="#compsys-2013-bid2" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
      <p>X10 other parallel programming primitive directives are <i>clocks</i> and
<tt>atomic</tt>. The <tt>at</tt> construct allows downloading a computation
to another <i>place</i>. Clocks are a dynamic version of barriers. Their
analysis involves counting their instances. For polyhedral programs, this can
be done using the Ehrhart and Barvinok theories; the results are polynomials.
Checking whether clocks remove non-determinism involves finding integer roots
and hence is undecidable. However, modern SMT solvers are able to solve most
of these problems. The resulting paper <ref xlink:href="#compsys-2013-bid28" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> has been
submitted to the ECOOP conference.
</p>
    </subsection>
    <subsection id="uid113" level="1">
      <bodyTitle>Clock Removal in X10</bodyTitle>
      <participants>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
        <person key="PASUSERID">
          <firstname>Eric</firstname>
          <lastname>Violard</lastname>
          <moreinfo>Inria/Camus</moreinfo>
        </person>
        <person key="PASUSERID">
          <firstname>Alain</firstname>
          <lastname>Ketterlin</lastname>
          <moreinfo>Inria/Camus</moreinfo>
        </person>
      </participants>
      <p>In the light of the previous work on the determinism of X10, a natural
question is: are the parallel programming directives of X10 redundant? The
answer is yes, at least for static control programs, i.e., programs in which
the set of operations and their execution order do not depend on the input
data. The basic idea is that the synchronization which occurs when several
activities execute an advance is similar to the synchronization at the
end of a finish. If one is able to count advances, one may construct a front
by gathering all operations with the same advance count. Each front is
executed inside one finish, and fronts are executed sequentially in order
of increasing counts. For polyhedral programs, advance counting can be done
at compile time. If the counts are affine functions, the restructuring can
be done by classical polyhedral code generators like CLooG, and no overhead
is incurred. For polynomial counts, one overall enclosing loop must be added,
but the resulting program can usually be optimized by simple loop
transformations, e.g., pushing guards into enclosing loop bounds.
For arbitrary programs, the counts have to be computed
dynamically; this is possible only if the program has static control.</p>
      <p>This result does not contradict the previous undecidability proof
(Section <ref xlink:href="#uid112" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>), as the translation of a
polyhedral program is usually not polyhedral. Application of the method to a
set of simple kernels has shown significant speedups. The interpretation of
this result is that, at least in the present state of the X10 runtime, the
implementation of the <tt>async</tt> primitive is more mature than the
implementation of clocks. A paper on this topic has been accepted at CC'14
(Compiler Construction Conference) <ref xlink:href="#compsys-2013-bid3" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.
</p>
    </subsection>
    <subsection id="uid114" level="1">
      <bodyTitle>Static Analysis of OpenStream Programs</bodyTitle>
      <participants>
        <person key="PASUSERID">
          <firstname>Albert</firstname>
          <lastname>Cohen</lastname>
          <moreinfo>Inria, Parkas</moreinfo>
        </person>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="compsys-2005-id18170">
          <firstname>Paul</firstname>
          <lastname>Feautrier</lastname>
        </person>
      </participants>
      <p>The objective of the collaboration between the Compsys and Parkas teams in the
ManycoreLabs project (Section <ref xlink:href="#uid121" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) is
to evaluate the possibility of applying polyhedral techniques to the parallel
language OpenStream, which is developed by Inria Parkas. When applicable,
these techniques are invaluable for compile-time debugging and for improving
the target code for a better adaptation to the target architecture.</p>
      <p>OpenStream is a two-level language, in which a sequential control code directs
the initialization of parallel task instances that communicate through
<i>streams</i>. OpenStream programs are deterministic by construction, but may
have deadlocks. If the control code is polyhedral, one may statically compute,
for each task instance, its read and write indices for each stream. These
indices may be polynomials of arbitrary degree. When linear, the full power of
the polyhedral model may be brought to bear for dependence and dataflow
analysis, scheduling and deadlock detection, and program transformations.</p>
      <p>In the general case, one can think of two approaches: the first one consists
in over-approximating dependences until problems become linear. In the second
approach, one first leverages modern developments in SMT solvers, which allow
them to solve polynomial problems, albeit with no guarantee of success.
Furthermore, the task index functions have special properties that may be used
to construct original analysis algorithms. Three preliminary results in this
direction:</p>
      <simplelist>
        <li id="uid115">
          <p noindent="true">the proof that deadlock detection is undecidable in general, thanks to an
adaptation of the proof designed for
X10 (Section <ref xlink:href="#uid112" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>),</p>
        </li>
        <li id="uid116">
          <p noindent="true">a characterization of deadlocks in terms of dependence graphs, which
implies that streams can be safely bounded as soon as a schedule exists with
such sizes,</p>
        </li>
        <li id="uid117">
          <p noindent="true">a preliminary analysis of some solvable cases.</p>
        </li>
      </simplelist>
      <p>A document is available as Deliverable 2.5.3 for the ManycoreLabs project.
</p>
    </subsection>
    <subsection id="uid118" level="1">
      <bodyTitle>Array Contraction in
Parallel Programs</bodyTitle>
      <participants>
        <person key="compsys-2005-id18078">
          <firstname>Alain</firstname>
          <lastname>Darte</lastname>
        </person>
        <person key="compsys-2012-idp140679680158976">
          <firstname>Alexandre</firstname>
          <lastname>Isoard</lastname>
        </person>
      </participants>
      <p>Array contraction is a technique to reuse array elements when they are dead, in
a form of array folding. A standard technique for array contraction is to use
affine remappings with modulos. When the modulo is equal to 1, this
corresponds to the removal of the corresponding array dimension. Array
contraction is well-known for sequential programs, after element-wise array
liveness analysis. It has also been customized for parallel codes obtained
through affine schedules by Lefebvre and Feautrier, and Quilleré-Rajopadhye,
both frameworks being generalized by the lattice-based memory allocation
framework of Darte, Schreiber, and Villard  <ref xlink:href="#compsys-2013-bid14" location="biblio" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and the
construction of the set of conflicting array indices. We showed how the same
framework can be used for a larger range of parallel programs, including
programs with outer parallel loops, programs exhibiting pipelining, a subset of
X10, etc. The optimality of the construction can be shown, despite a related
(but actually non-contradictory here) NP-completeness result for
worst-case of register pressure in the context of register allocation.
A research report on this topic is in preparation.
</p>
    </subsection>
  </resultats>
  <contrats id="uid119">
    <bodyTitle>Bilateral Contracts and Grants with Industry</bodyTitle>
    <subsection id="uid120" level="1">
      <bodyTitle>Tirex Contract with Kalray</bodyTitle>
      <p>Compsys
has a contract with Kalray called Tirex. The goal of this project is to
prototype within the TireX toolbox (see
Section <ref xlink:href="#uid86" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) some new profiling/analysis
techniques necessary to enable cloning. Because of the current financial
problems encountered by Kalray, the efforts related to this project have been
frozen until further notice.
</p>
    </subsection>
    <subsection id="uid121" level="1">
      <bodyTitle>ManycoreLabs Project with Kalray</bodyTitle>
      <p>Compsys is part of a bilateral grant with Kalray called ManycoreLabs,
funded by “Investissements d'avenir pour le développement de l'économie
numérique”. The goal of this project is to allow the company Kalray, based
on a collaboration with several partners, to become the European leader of
the market of many-core chips for embedded systems. Industrial partners of
this project include Bull, CAPS Entreprise, Digigram, Thales, Renault.
Academic partners are CEA, Inria (Parkas and Compsys), VERIMAG.</p>
      <p>The cloning/specialization work summarized in
Section <ref xlink:href="#uid90" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> and the generalized register
tiling work summarized in Section <ref xlink:href="#uid91" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>
have been done in the context of this grant and correspond to WP 3.3.3. The
research on OpenStream described in
Section <ref xlink:href="#uid114" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/> corresponds to WP 2.5.3.
</p>
    </subsection>
    <subsection id="uid122" level="1">
      <bodyTitle>Technological Transfer Towards Zettice Start-Up</bodyTitle>
      <participants>
        <person key="compsys-2006-id18253">
          <firstname>Christophe</firstname>
          <lastname>Alias</lastname>
        </person>
        <person key="graal-2009-id60684">
          <firstname>Adrian</firstname>
          <lastname>Muresan</lastname>
          <moreinfo>Zettice</moreinfo>
        </person>
        <person key="compsys-2006-id18441">
          <firstname>Alexandru</firstname>
          <lastname>Plesco</lastname>
          <moreinfo>Zettice</moreinfo>
        </person>
      </participants>
      <p>The Zettice start-up project has been initiated by Alexandru Plesco and
Christophe Alias in March 2011, with the idea of transferring some of the
research concepts emerging from the polyhedral model to the context of
high-level circuit synthesis. Since, an important amount of applied research
has been achieved to propose an effective technology ready for industrial
transfer. From an academic perspective, Zettice is a unique opportunity to
cover all the aspects of high-level synthesis from the front-end aspects
(polyhedral code analysis and optimization) to the back-end aspects
(pipelining, retiming, FPGA mapping) providing a global knowledge of relevant
industrial issues.</p>
      <p>Zettice received in 2012 the <i>“lean start-up award”</i> of the startup
weekend labs 2012, the <i>“most exciting start-up mention”</i> at SAME 2012,
and the <i>concours Crealys Excel&amp;Rate 2012</i> grant (30 Keuros). In 2013,
Zettice won the <i>concours OSEO 2013</i> grant (Banque Publique
d'Investissement, 40 Keuros) and the <i>“most promising start-up award”</i> at
SAME 2013.</p>
      <p>A patent is under deposit. The research results related to Zettice are
presented in Section <ref xlink:href="#uid98" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. The software tools
developed in the context of Zettice are Dcc (see
Section <ref xlink:href="#uid75" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) and IceGEN (see
Section <ref xlink:href="#uid76" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>).
</p>
    </subsection>
  </contrats>
  <partenariat id="uid123">
    <bodyTitle>Partnerships and Cooperations</bodyTitle>
    <subsection id="uid124" level="1">
      <bodyTitle>Regional Initiatives</bodyTitle>
      <p>Compsys has increased its relationship with the CITI laboratory
(Insa-Lyon) and, in particular, the team of Tanguy Risset (Socrate Inria
project <ref xlink:href="http://www.citi-lab.fr/team/socrate/" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>citi-lab.<allowbreak/>fr/<allowbreak/>team/<allowbreak/>socrate/</ref>). Compsys and Socrate
made several common working groups in 2012 and 2013, and are mutually invited
to seminars organized by the other team. Streaming languages are a common
topic of interest. In this context, Socrate, with the help of Compsys,
will organize a thematic day (April 14, 2014) on the “compilation and
execution of streaming programs”, in Domaine des Hautannes, St Germain au
Mont d'Or. Lionel Morel and Laure Gonnord have also common topics of interest.</p>
      <p>Compsys has stronger connections with the Grame music/computer laboratory
(<ref xlink:href="http://www.grame.fr" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>grame.<allowbreak/>fr</ref>) in Lyon and, in particular, Yann Orlarey, also
due to common interests on streaming languages, in particular the language
Faust developed by Grame. Yann Orlarey was one of the invited speaker of the
keynotes on parallel languages (see the description the thematic quarter on
compilation in Section <ref xlink:href="#uid155" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>). Alexandre Isoard's Master 1
training period was on Faust, co-advised by Alain Darte and Yann Orlarey. For
2014, Laure Gonnord and Yann Orlarey proposed a Master research topic on the
generation of invariants for the Faust language.</p>
      <p>Compsys is also involved in the Labex MILYON (Mathématiques et
Informatique Fondamentale de Lyon), which regroups Institut Camille Jordan,
and the mathematics and computer science labs of ENS-Lyon. The aim of MILYON
is “to strengthen our international relationships, in particular by
organizing thematic quarters which will allow world experts of a subject to
gather in Lyon and work together in a stimulating environment.” In this
context, Compsys organized a thematic quarter on compilation from April
2013 to July 2013, see details in Section <ref xlink:href="#uid155" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. Compsys
also follows or participates to the activities of LyonCalcul
(<ref xlink:href="http://lyoncalcul.univ-lyon1.fr/" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>lyoncalcul.<allowbreak/>univ-lyon1.<allowbreak/>fr/</ref>), a network to federate activities on
computing in Lyon.
</p>
    </subsection>
    <subsection id="uid125" level="1">
      <bodyTitle>National Initiatives</bodyTitle>
      <subsection id="uid126" level="2">
        <bodyTitle>CNRS PEPS</bodyTitle>
        <p>Christophe Alias and Laure Gonnord initiated with the DART/Emeraude team at LIFL
Laboratory (University of Lille) a CNRS PEPS (“Projets Exploratoire Premier
Soutien”) called “HLS and real time” (8 Keuros/year, during two years in
2011-2013). The goal of this project is to investigate how to introduce
real-time constraints in the high-level synthesis workflow.</p>
      </subsection>
      <subsection id="uid127" level="2">
        <bodyTitle>Inria AEN MULTICORE</bodyTitle>
        <p>Fabrice Rastello is part of an Inria Large Scale Initiative (AEN: action d'envergure
nationale) called MULTICORE, which regroups researchers from seven teams:
Camus, Regal, Alf, Runtime, Algorille, Dali, and thus Compsys on “Large
scale multicore virtualization for performance scaling and portability”. One
of the goals of this project is to enable loop transformations by combining
dynamic and static analysis/compilation techniques.</p>
      </subsection>
      <subsection id="uid128" level="2">
        <bodyTitle>French Compiler Community</bodyTitle>
        <p>The french compiler community is now well identified and is visible through its
web-page <ref xlink:href="http://compilation.gforge.inria.fr/" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>compilation.<allowbreak/>gforge.<allowbreak/>inria.<allowbreak/>fr/</ref>. The “journées françaises
de la compilation” were initiated in 2010 and are still animated by Fabrice Rastello
and Laure Gonnord as a biannual event. Their local organization is handled
alternately by the different research teams: Lyon (by Compsys) in Summer
2010, Aussois in Winter 2010, Dinard in Spring 2011, St Hippolyte in Autumn
2011, Rennes in Summer 2012, Annecy (by Compsys again) in Spring 2013,
Dammarie-les-lys in December 2013.</p>
      </subsection>
    </subsection>
    <subsection id="uid129" level="1">
      <bodyTitle>European Initiatives</bodyTitle>
      <subsection id="uid130" level="2">
        <bodyTitle>Collaborations with Major European Organizations</bodyTitle>
        <p>Alain Darte, Paul Feautrier, and Fabrice Rastello are members or affiliate members of
the European Network of Excellence on High Performance and
Embedded Architecture and Compilation (HiPEAC). Fabrice Rastelloattended the
computing system week in may 2013 (Paris), and the computing system week in
October 2013 (Tallinn). He participated to the organization of two thematic
sessions in Paris: Thread Level Speculation (as chair) and Intermediate
Representation (as co-organizer). The thematic quarter on compilation (see
Section <ref xlink:href="#uid155" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>) was presented in HIPEAC info 35 (July 2013),
the HIPEAC quarterly newsletter
(<ref xlink:href="http://www.hipeac.net/content/hipeacinfo-35-july-2013" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>www.<allowbreak/>hipeac.<allowbreak/>net/<allowbreak/>content/<allowbreak/>hipeacinfo-35-july-2013</ref>) and the keynotes
on HPC languages (third event) recognized as an HIPEAC event.</p>
      </subsection>
    </subsection>
    <subsection id="uid131" level="1">
      <bodyTitle>International Initiatives</bodyTitle>
      <subsection id="uid132" level="2">
        <bodyTitle>Inria International Partners</bodyTitle>
        <subsection id="uid133" level="3">
          <bodyTitle>Declared Inria International Partners</bodyTitle>
          <simplelist>
            <li id="uid134">
              <p noindent="true">Compsys and, in particular Fabrice Rastello, has a regular collaboration with
P. Sadayappan from Ohio State University (USA). This year, this collaboration
led to several results, see Sections <ref xlink:href="#uid89" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>,
<ref xlink:href="#uid91" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>,
<ref xlink:href="#uid92" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>,
and <ref xlink:href="#uid93" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
            </li>
            <li id="uid135">
              <p noindent="true">Fabrice Rastello and Laure Gonnord have a regular collaboration with Fernando Magno
Quintao Pereira from the University of Mineas Gerais (Brazil). This year,
this collaboration led to several results, see
Sections <ref xlink:href="#uid88" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>
and <ref xlink:href="#uid90" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>. Compsys also hosted Raphael
Ernani Rodrigues, from the group of F. Pereira, who made part of his master
in Lyon supervised by Laure Gonnord and Christophe Alias.</p>
            </li>
            <li id="uid136">
              <p noindent="true">Compsys and, in particular Christophe Alias, has a regular collaboration
with S. Rajopadhye from Colorado State University (CSU). Guillaume Iooss is
preparing a PhD through a PhD convention between Ecole normale supérieure de
Lyon and Colorado State University, co-advised by Christophe Alias and Sanjay
Rajopadhye. In 2013, Guillaume Iooss spent part of the summer at CSU, joined by
Christophe Alias for a week. Paul Feautrier and Fabrice Rastello also made regular visits at
Colorado State University in the previous years. This year, this
collaboration led to several results, see
Sections <ref xlink:href="#uid107" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>,
<ref xlink:href="#uid110" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>,
and <ref xlink:href="#uid112" location="intern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest"/>.</p>
            </li>
          </simplelist>
        </subsection>
      </subsection>
    </subsection>
    <subsection id="uid137" level="1">
      <bodyTitle>International Research Visitors</bodyTitle>
      <subsection id="uid138" level="2">
        <bodyTitle>Visits of International Scientists</bodyTitle>
        <subsection id="uid139" level="3">
          <bodyTitle>Invited Researchers</bodyTitle>
          <p>Fernando Magno Quintão Pereira is visiting Fabrice Rastello for 1.5 month in early
2014. The goal of his visit is to work on dynamic analysis and cloning for loop
transformations (so called hybrid compilation).</p>
        </subsection>
        <subsection id="uid140" level="3">
          <bodyTitle>Internships</bodyTitle>
          <p>Raphael Ernani Rodrigues made part of his master Internship in Lyon in
June/July 2013 under the supervision of Laure Gonnord and
Christophe Alias. He worked on synthesizing preconditions that (may)
ensure termination. We are currently pursuing the collaboration with
him and his supervisor in Brazil, Fernando Magno Quintao Pereira (Univ.
Mineas Gerais).</p>
        </subsection>
      </subsection>
      <subsection id="uid141" level="2">
        <bodyTitle>Visits to International Teams</bodyTitle>
        <p>Fabrice Rastello visited the group of P. Sadayappan (OSU) during two months, in
June-July 2013, in addition to shorter stays. He worked on dynamic analysis
and generalized tiling.</p>
        <p>Alexandre Isoard did an internship at Xilinx, during 2.5 months, from June to
September 2013, under the supervision of Stephen Neuendorffer, working on
exploring polyhedral tools for Xilinx HLS tool.</p>
      </subsection>
    </subsection>
  </partenariat>
  <diffusion id="uid142">
    <bodyTitle>Dissemination</bodyTitle>
    <subsection id="uid143" level="1">
      <bodyTitle>Scientific Animation</bodyTitle>
      <subsection id="uid144" level="2">
        <bodyTitle>Program Committees, Editorial Boards, and Reviewing Activities</bodyTitle>
        <simplelist>
          <li id="uid145">
            <p noindent="true">Christophe Alias was a member of the steering committee of IMPACT
2013 (International Workshop on Polyhedral Compilation Techniques,
Berlin, Germany).</p>
          </li>
          <li id="uid146">
            <p noindent="true">Christophe Alias, Alain Darte, and Paul Feautrier were members of the program
committees of IMPACT 2013 and IMPACT 2014 (Vienna, Austria).</p>
          </li>
          <li id="uid147">
            <p noindent="true">Christophe Alias was member of the program committee of ODES 2013 (i.e.,
ODES-10, 10th Workshop on Optimizations for DSP and Embedded Systems,
Shenzen, China).</p>
          </li>
          <li id="uid148">
            <p noindent="true">Fabrice Rastello was member of the program committees of CGO 2014
(International Symposium on Code Generation and Optimization, Orlando,
Florida) and CRI 2013 (Conférence de Recherche en Informatique, Yaoundé,
Cameroun).</p>
          </li>
          <li id="uid149">
            <p noindent="true">Alain Darte was member of the program committees of DATE 2013 (Design,
Automation, and Test in Europe, Grenoble, France) and DATE 2014 (Dresden,
Germany), IPDPS 2013 (International Parallel and Distributed Processing
Symposium, Boston, Massachusetts) and IPDPS 2014 (Phoenix, Arizona).</p>
          </li>
          <li id="uid150">
            <p noindent="true">Alain Darte was member of the editorial board of IEEE TECS (Transactions on
Embedded Computing Systems) until end of 2013.</p>
          </li>
          <li id="uid151">
            <p noindent="true">Christophe Alias was a reviewer for the journals JPDC (Journal of Parallel
and Distributed Computing), MICPRO (Microprocessors and Microsystems), PPL
(Parallel Processing Letters), ACM TRETS (Transactions on Reconfigurable
Technology and Systems), TSI (Technique et Science Informatique), CDT (IET
Computers and Digital Techniques), IPL (Information Processing Letters).</p>
          </li>
          <li id="uid152">
            <p noindent="true">Paul Feautrier was a reviewer for ACM TECS, IJPP, IEEE TPDS, ACM TOPLAS,
DATE14, IMPACT 2014.</p>
          </li>
          <li id="uid153">
            <p noindent="true">Alain Darte was a reviewer for DATE'14, IPDPS'14, IMPACT'14, Parallel
Computing, ACM TACO, and ACM TECS.</p>
          </li>
          <li id="uid154">
            <p noindent="true">Laure Gonnord was a reviewer for MSR'13, DAC'13 and AMT'13.</p>
          </li>
        </simplelist>
      </subsection>
      <subsection id="uid155" level="2">
        <bodyTitle>Thematic Quarter on Compilation</bodyTitle>
        <p>Compsys is part of the Labex MILYON, which regroups Institut Camille Jordan,
and the mathematics and computer science labs of ENS-Lyon. One of its goal is
“to strengthen our international relationships, in particular by organizing
thematic quarters which will allow world experts of a subject to gather in Lyon
and work together in a stimulating environment.” In this context, Alain Darte,
helped by Alexandre Isoard and Laetitia Lecot, organized, from April to July 2013,
a thematic quarter on compilation techniques
(<ref xlink:href="http://labexcompilation.ens-lyon.fr" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>labexcompilation.<allowbreak/>ens-lyon.<allowbreak/>fr</ref>), with a special focus on the
interactions with languages and architectures for high performance computing.
This thematic quarter (with a total budget of 100 Keuros), consisted, in addition to
the “french compilation days” organized separately in Annecy by Laure Gonnord and
Fabrice Rastello (April 4-7, 2013), in three international scientific events
organized in Lyon or the vicinity.</p>
        <simplelist>
          <li id="uid156">
            <p noindent="true">A <b>spring school on polyhedral code analysis and optimizations</b>
(<ref xlink:href="http://labexcompilation.ens-lyon.fr/polyhedral-school" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>labexcompilation.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>polyhedral-school</ref>), May 13-17,
2013, in Domaine des Hautannes in St Germain au Mont d'Or, the first
international school on the polyhedral model and related optimizations. The
school covered scheduling theory, algorithms and modeling with integer sets
and relations, abstract interpretation, compilation for distributed
platforms, array region analysis, vectorization and SIMD optimizations,
through courses given by S. Rajopadhye (Colorado State Univ.), P. Feautrier
(Compsys, ENS-Lyon), L.-N. Pouchet (UCLA), S. Verdoolaege (ENS Paris), A.
Miné (ENS Paris), U. Bondhugula (IIS Bangalore), A. Darte (Compsys,
CNRS), B. Creusillet (Silkan), P. Sadayappan (Ohio State Univ.), N.
Vasilache (Reservoir Labs, New York). The school attracted 56 participants,
half from France, but also from Germany, the USA, England, Belgium, Spain,
China, India, Ireland, and Italy and, interestingly, also from groups that
are not familiar with polyhedral optimizations. Roughly half of the
participants were PhD students.</p>
          </li>
          <li id="uid157">
            <p noindent="true">A <b>dive in languages for high-performance computing</b>
(<ref xlink:href="http://labexcompilation.ens-lyon.fr/hpc-languages" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>labexcompilation.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>hpc-languages</ref>), June 29-July 2,
2013 in Résidence Villemanzy in Lyon, organized as a set of long keynotes on
CAF (Coarray Fortran), UPC (Unified Parallel C), X10, Chapel, OpenACC &amp;
OpenHMPP, Liquid Metal, OmpSs, OpenStream, and some DSL approaches. The
keynotes were given by a panel of international experts on compilation for
high-performance computing: J. Mellor-Crummey and V. Sarkar (Rice), K.
Yelick (Berkeley), R. Schreiber (HP Labs), B. Chamberlain (Cray), D. Grove
and R. Rabbah (IBM Watson), A. Cohen (Inria, ENS Paris), R. Badia (UPC
Barcelona), F. Bodin (Univ. Rennes, previously Caps Entreprise), Y. Orlarey
(Grame), K. Knobe (Intel, Massachusets), P. Sadayappan (Ohio State Univ.).
This event regrouped 71 participants, including speakers, and, as we hoped,
also attracted people from industry, and not only computer industry.</p>
          </li>
          <li id="uid158">
            <p noindent="true"><b>CPC'13, the 17th international workshop on compilers for parallel
computing</b> (<ref xlink:href="http://labexcompilation.ens-lyon.fr/cpc2013" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>labexcompilation.<allowbreak/>ens-lyon.<allowbreak/>fr/<allowbreak/>cpc2013</ref>), July 3-5,
2013, in Musée Gadagne, in (old) Lyon, a venue that is held every 18 months
in Europe since 1989 and that encompasses all areas of parallelism and
optimization linked to compilers. The program consisted in 29 talks, from
the international community on compilers for HPC (from Japan &amp; Taiwan to the
USA, and of course Europe), with 47 participants.</p>
          </li>
        </simplelist>
        <p>During this compilation thematic quarter, Paul Feautrier and Alain Darte gave
the following talks:</p>
        <simplelist>
          <li id="uid159">
            <p noindent="true">“Array Dataflow Analysis for Polyhedral X10 Programs” (Paul Feautrier) and
“Modèles et algorithmes: comprendre de quoi on parle” (Alain Darte) at the
French Compiler Community meeting (April 2-4, 2013),</p>
          </li>
          <li id="uid160">
            <p noindent="true">“The Care and Feeding of Polyhedra” (Paul Feautrier)
and “Array Contraction with Lattice−Based Memory Allocation” (Alain Darte) at
the Spring School on Polyhedral Code Analysis and Optimizations (May 13-17,
2013),</p>
          </li>
          <li id="uid161">
            <p noindent="true">“Determinacy Analysis of Polyhedral X10 Programs” (Paul Feautrier), paper with
Alain Ketterlin and Eric Violard, at the CPC Workshop (July 3-5, 2013),</p>
          </li>
        </simplelist>
      </subsection>
    </subsection>
    <subsection id="uid162" level="1">
      <bodyTitle>Teaching - Supervision - Juries</bodyTitle>
      <subsection id="uid163" level="2">
        <bodyTitle>Teaching</bodyTitle>
        <sanspuceslist>
          <li id="uid164">
            <p noindent="true">Licence:</p>
            <simplelist>
              <li id="uid165">
                <p noindent="true">Laure Gonnord, Algorithmique et programmation C (60h), L3,
Université de Lille 1, Polytech'Lille.</p>
              </li>
              <li id="uid166">
                <p noindent="true">Laure Gonnord, Architecture des
ordinateurs (25h), L3, Université de Lille 1, Polytech'Lille.</p>
              </li>
              <li id="uid167">
                <p noindent="true">Laure Gonnord, Algorithmique et programmation fonctionnelle et récursive (42h), L1,
Université Lyon 1 Claude Bernard.</p>
              </li>
              <li id="uid168">
                <p noindent="true">Guillaume Iooss, LIF3: algorithmique et programmation fonctionnelle et récursive (28h), L1, Université Lyon 1 Claude Bernard.</p>
              </li>
              <li id="uid169">
                <p noindent="true">Guillaume Iooss, Programmation 1 (12h), L3, ENS-Lyon.</p>
              </li>
              <li id="uid170">
                <p noindent="true">Christophe Alias, Architecture des ordinateurs (21h TP), Université Lyon 1.</p>
              </li>
              <li id="uid171">
                <p noindent="true">Christophe Alias, Compilation (6h CM, 6h TP), ENSI Bourges.</p>
              </li>
              <li id="uid172">
                <p noindent="true">Christophe Alias,
Correction de copies, concours E3A, épreuve informatique MPSI.</p>
              </li>
              <li id="uid173">
                <p noindent="true">Alexandre Isoard, LIF12: Système et Réseau (32h TP), L3, Université Lyon 1 Claude Bernard.</p>
              </li>
            </simplelist>
          </li>
          <li id="uid174">
            <p noindent="true">Master:</p>
            <simplelist>
              <li id="uid175">
                <p noindent="true">Laure Gonnord, Compilation (24h), M1, Université Lyon 1 Claude
Bernard.</p>
              </li>
              <li id="uid176">
                <p noindent="true">Laure Gonnord, Introduction aux systèmes et réseaux (52h), M2 Pro, Université
Lyon 1.</p>
              </li>
              <li id="uid177">
                <p noindent="true">Christophe Alias, Compilation (24h CM), M1, ENS-Lyon.</p>
              </li>
              <li id="uid178">
                <p noindent="true">Christophe Alias,
Compilation avancée (8h CM), M2, ENS-Lyon.</p>
              </li>
              <li id="uid179">
                <p noindent="true">Fabrice Rastello, Compilation avancée
(6h CM), M2, ENS-Lyon.</p>
              </li>
              <li id="uid180">
                <p noindent="true">Fabrice Rastello, SSA-based compiler design (2 days), CRI
Cameroun.</p>
              </li>
              <li id="uid181">
                <p noindent="true">Alexandre Isoard, Compilation (24h TP), M1, ENS-Lyon.</p>
              </li>
              <li id="uid182">
                <p noindent="true">Guillaume Iooss, Image (24h TP), M1, ENS-Lyon.</p>
              </li>
              <li id="uid183">
                <p noindent="true">Laure Gonnord also organized, for the Computer Science Department of ENS Lyon,
a <i>research school</i> for Master students, on synchronous programming. The
program can be found at the url:
<ref xlink:href="http://laure.gonnord.org/pro/research/sync_research_school.html" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>laure.<allowbreak/>gonnord.<allowbreak/>org/<allowbreak/>pro/<allowbreak/>research/<allowbreak/>sync_research_school.<allowbreak/>html</ref>.</p>
              </li>
            </simplelist>
          </li>
        </sanspuceslist>
      </subsection>
      <subsection id="uid184" level="2">
        <bodyTitle>Supervision</bodyTitle>
        <simplelist>
          <li id="uid185">
            <p noindent="true">PhD in progress: Guillaume Iooss, “Semantic Tiling”, started on
September 2011, advisors: Christophe Alias and Sanjay Rajopadhye (Associate
Professor, Colorado State University).</p>
          </li>
          <li id="uid186">
            <p noindent="true">PhD in progress: François Gindraud, started on January 2013, advisors Fabrice Rastello, Albert Cohen (Parkas Inria team)</p>
          </li>
          <li id="uid187">
            <p noindent="true">PhD in progress: Duco Van Amstel, started on January 2013, advisors Fabrice Rastello, Benoit Dupont-de-Dinechin (Kalray)</p>
          </li>
          <li id="uid188">
            <p noindent="true">PhD in progress: Diogo Nunes Sampaio, started on October 2013, advisor Fabrice Rastello</p>
          </li>
          <li id="uid189">
            <p noindent="true">PhD in progress: Alexandre Isoard, started in September 2012, advisor Alain Darte</p>
          </li>
        </simplelist>
      </subsection>
      <subsection id="uid190" level="2">
        <bodyTitle>Juries</bodyTitle>
        <simplelist>
          <li id="uid191">
            <p noindent="true">Laure Gonnord participated to the Jury of Clement Guy's PhD defense (in
Rennes) entitled “Facilités de typage pour l'ingénierie des Langages”. This
PhD was supervised by J.M. Jézéquiel (Professor, Rennes University), and B.
Combemale (Assistant Professor, Rennes University) and S. Derrien (Professor,
Rennes university).</p>
          </li>
          <li id="uid192">
            <p noindent="true">Christophe Alias participated to the Jury of Antoine Morvan's PhD defense
(in Rennes) entitled “Utilisation du modèle polyédrique pour la synthèse
d'architectures pipelinées”. This PhD thesis was supervised by S. Derrien
(Professor, Rennes University), and P. Quinton (Professor, Rennes
University).</p>
          </li>
          <li id="uid193">
            <p noindent="true">Fabrice Rastello participated to the jury of Alexandre Carbon's PhD defense,
entitled “Accélération matérielle de la compilation à la volée pour les
systèmes embarqués”.</p>
          </li>
          <li id="uid194">
            <p noindent="true">Paul Feautrier was a reviewer for the HDR of Stephane Mancini (Grenoble) and for
the PhD of Amira Mensi (Paris).</p>
          </li>
          <li id="uid195">
            <p noindent="true">Alain Darte was a reviewer for the PhD thesis of Cupertino Miranda (Paris 11),
entitled “Erbium: Reconciling languages, runtimes, compilation and
optimizations for streaming applications” and supervised by Albert Cohen (DR Inria, Parkas team).</p>
          </li>
        </simplelist>
      </subsection>
    </subsection>
  </diffusion>
  <biblio id="bibliography" html="bibliography" numero="10" titre="Bibliography">
    
    <biblStruct id="compsys-2013-bid29" type="article" rend="year" n="cite:brandner:hal-00768781">
      <identifiant type="doi" value="10.1016/j.cl.2012.09.001"/>
      <identifiant type="hal" value="hal-00768781"/>
      <analytic>
        <title level="a">Elimination of parallel copies using code motion on data dependence graphs</title>
        <author>
          <persName key="compsys-2009-id59559">
            <foreName>Florian</foreName>
            <surname>Brandner</surname>
            <initial>F.</initial>
          </persName>
          <persName key="compsys-2007-id18566">
            <foreName>Quentin</foreName>
            <surname>Colombet</surname>
            <initial>Q.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-editorial-board="yes" x-international-audience="yes" id="rid00438">
        <idno type="issn">1477-8424</idno>
        <title level="j">Computer Languages, Systems and Structures</title>
        <imprint>
          <biblScope type="volume">39</biblScope>
          <biblScope type="number">1</biblScope>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <biblScope type="pages">25 - 47</biblScope>
          <ref xlink:href="http://hal.inria.fr/hal-00768781" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00768781</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid4" type="article" rend="year" n="cite:fauzia:hal-00920031">
      <identifiant type="hal" value="hal-00920031"/>
      <analytic>
        <title level="a">Beyond Reuse Distance Analysis: Dynamic Analysis for Characterization of Data Locality Potential</title>
        <author>
          <persName>
            <foreName>Naznin</foreName>
            <surname>Fauzia</surname>
            <initial>N.</initial>
          </persName>
          <persName>
            <foreName>Venmugil</foreName>
            <surname>Elango</surname>
            <initial>V.</initial>
          </persName>
          <persName>
            <foreName>Mahesh</foreName>
            <surname>Ravishankar</surname>
            <initial>M.</initial>
          </persName>
          <persName>
            <foreName>J.</foreName>
            <surname>Ramanujam</surname>
            <initial>J.</initial>
          </persName>
          <persName key="compsys-2005-id18199">
            <foreName>Fabrice</foreName>
            <surname>Rastello</surname>
            <initial>F.</initial>
          </persName>
          <persName>
            <foreName>Atanas</foreName>
            <surname>Rountev</surname>
            <initial>A.</initial>
          </persName>
          <persName key="alchemy-2006-id18739">
            <foreName>Louis-Noël</foreName>
            <surname>Pouchet</surname>
            <initial>L.-N.</initial>
          </persName>
          <persName>
            <foreName>P.</foreName>
            <surname>Sadayappan</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-editorial-board="yes" x-international-audience="yes" id="rid00017">
        <idno type="issn">1544-3566</idno>
        <title level="j">Transaction on Architecture and Code Optimization</title>
        <imprint>
          <biblScope type="volume">10</biblScope>
          <biblScope type="number">4</biblScope>
          <dateStruct>
            <month>December</month>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00920031" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00920031</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid30" type="article" rend="year" n="cite:gonnord:hal-00876627">
      <identifiant type="doi" value="10.1016/j.scico.2013.09.016"/>
      <identifiant type="hal" value="hal-00876627"/>
      <analytic>
        <title level="a">Abstract Acceleration in Linear Relation Analysis</title>
        <author>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
          <persName key="pop_art-2009-id59830">
            <foreName>Peter</foreName>
            <surname>Schrammel</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-editorial-board="yes" x-international-audience="yes" id="rid01845">
        <idno type="issn">0167-6423</idno>
        <title level="j">Science of Computer Programming</title>
        <imprint>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00876627" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00876627</ref>
          <ref xlink:href="http://hal.inria.fr/hal-00787212/en" type="hal" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00787212/<allowbreak/>en</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid19" type="inproceedings" rend="year" n="cite:alias:hal-00801571">
      <identifiant type="hal" value="hal-00801571"/>
      <analytic>
        <title level="a">Rank: a tool to check program termination and computational complexity</title>
        <author>
          <persName key="compsys-2006-id18253">
            <foreName>Christophe</foreName>
            <surname>Alias</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2005-id18078">
            <foreName>Alain</foreName>
            <surname>Darte</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="no">
        <title level="m">Constraints in Software Testing Verification and Analysis</title>
        <loc>Luxembourg</loc>
        <imprint>
          <dateStruct>
            <month>March</month>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00801571" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00801571</ref>
        </imprint>
        <meeting id="cid393923">
          <title>Workshop on Constraints in Software Testing, Verification and Analysis</title>
          <num>5</num>
          <abbr type="sigle">CSTVA</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid16" type="inproceedings" rend="year" n="cite:alias:hal-00761533">
      <identifiant type="hal" value="hal-00761533"/>
      <analytic>
        <title level="a">Optimizing Remote Accesses for Offloaded Kernels: Application to High-Level Synthesis for FPGA</title>
        <author>
          <persName key="compsys-2006-id18253">
            <foreName>Christophe</foreName>
            <surname>Alias</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2005-id18078">
            <foreName>Alain</foreName>
            <surname>Darte</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2006-id18441">
            <foreName>Alexandru</foreName>
            <surname>Plesco</surname>
            <initial>A.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <title level="m">Design, Automation, and Test in Europe (DATE'13)</title>
        <loc>Grenoble, France</loc>
        <imprint>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00761533" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00761533</ref>
        </imprint>
        <meeting id="cid58552">
          <title>Design, Automation, and Test in Europe</title>
          <num>15</num>
          <abbr type="sigle">DATE</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid1" subtype="nonparu-n" type="inproceedings" rend="year" n="cite:darte:hal-00915831">
      <identifiant type="hal" value="hal-00915831"/>
      <analytic>
        <title level="a">Parametric Tiling with Inter-Tile Data Reuse</title>
        <author>
          <persName key="compsys-2005-id18078">
            <foreName>Alain</foreName>
            <surname>Darte</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2012-idp140679680158976">
            <foreName>Alexandre</foreName>
            <surname>Isoard</surname>
            <initial>A.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <editor role="editor">
          <persName>
            <foreName>Sanjay</foreName>
            <surname>Rajopadhye</surname>
            <initial>S.</initial>
          </persName>
          <persName key="alchemy-2010-id59677">
            <foreName>Sven</foreName>
            <surname>Verdoolaege</surname>
            <initial>S.</initial>
          </persName>
        </editor>
        <title level="m">4th International Workshop on Polyhedral Compilation Techniques (IMPACT'14)</title>
        <loc>Vienna, Austria</loc>
        <imprint>
          <dateStruct>
            <year>2014</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00915831" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00915831</ref>
        </imprint>
        <meeting id="cid596150">
          <title>International Workshop on Polyhedral Compilation Techniques</title>
          <num>4</num>
          <abbr type="sigle">IMPACT</abbr>
        </meeting>
      </monogr>
      <note type="bnote">To be published</note>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid31" type="inproceedings" rend="year" n="cite:diouf:hal-00911887">
      <identifiant type="doi" value="10.1109/CGO.2013.6495005"/>
      <identifiant type="hal" value="hal-00911887"/>
      <analytic>
        <title level="a">A Polynomial Spilling Heuristic: Layered Allocation</title>
        <author>
          <persName key="alchemy-2007-id18971">
            <foreName>Boubacar</foreName>
            <surname>Diouf</surname>
            <initial>B.</initial>
          </persName>
          <persName key="alchemy-2005-id18146">
            <foreName>Albert</foreName>
            <surname>Cohen</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2005-id18199">
            <foreName>Fabrice</foreName>
            <surname>Rastello</surname>
            <initial>F.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <title level="m">CGO 2013 - International Symposium on Code Generation and Optimization</title>
        <loc>Shenzhen, China</loc>
        <imprint>
          <publisher>
            <orgName>IEEE</orgName>
          </publisher>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00911887" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00911887</ref>
        </imprint>
        <meeting id="cid312438">
          <title>International Symposium on Code Generation and Optimization</title>
          <num>11</num>
          <abbr type="sigle">CGO</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid3" type="inproceedings" rend="year" n="cite:feautrier:hal-00924206">
      <identifiant type="hal" value="hal-00924206"/>
      <analytic>
        <title level="a">Improving X10 Program Performances by Clock Removal</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName key="calvi-2005-id18214">
            <foreName>Eric</foreName>
            <surname>Violard</surname>
            <initial>E.</initial>
          </persName>
          <persName key="camus-2010-id59483">
            <foreName>Alain</foreName>
            <surname>Ketterlin</surname>
            <initial>A.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <title level="m">Compiler Construction 2014</title>
        <loc>Grenoble, France</loc>
        <imprint>
          <dateStruct>
            <month>January</month>
            <year>2014</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00924206" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00924206</ref>
        </imprint>
        <meeting id="cid114893">
          <title>International Conference on Compiler Construction</title>
          <num>23</num>
          <abbr type="sigle">CC</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid0" type="inproceedings" rend="year" n="cite:iooss:hal-00915827">
      <identifiant type="hal" value="hal-00915827"/>
      <analytic>
        <title level="a">CART: Constant Aspect Ratio Tiling</title>
        <author>
          <persName>
            <foreName>Guillaume</foreName>
            <surname>Iooss</surname>
            <initial>G.</initial>
          </persName>
          <persName>
            <foreName>Sanjay</foreName>
            <surname>Rajopadhye</surname>
            <initial>S.</initial>
          </persName>
          <persName key="compsys-2006-id18253">
            <foreName>Christophe</foreName>
            <surname>Alias</surname>
            <initial>C.</initial>
          </persName>
          <persName>
            <foreName>Yun</foreName>
            <surname>Zou</surname>
            <initial>Y.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <title level="m">IMPACT 2014</title>
        <loc>Vienna, Austria</loc>
        <imprint>
          <dateStruct>
            <month>January</month>
            <year>2014</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00915827" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00915827</ref>
        </imprint>
        <meeting id="cid596150">
          <title>International Workshop on Polyhedral Compilation Techniques</title>
          <num>4</num>
          <abbr type="sigle">IMPACT</abbr>
        </meeting>
      </monogr>
      <note type="bnote">Not yet published</note>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid23" type="inproceedings" rend="year" n="cite:tavares:hal-00921461">
      <identifiant type="hal" value="hal-00921461"/>
      <analytic>
        <title level="a">Parameterized Construction of Program Representations for Sparse Dataflow Analyses</title>
        <author>
          <persName>
            <foreName>André</foreName>
            <surname>Tavares</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2005-id18199">
            <foreName>Fabrice</foreName>
            <surname>Rastello</surname>
            <initial>F.</initial>
          </persName>
          <persName key="compsys-2005-id18400">
            <foreName>Benoit</foreName>
            <surname>Boissinot</surname>
            <initial>B.</initial>
          </persName>
          <persName>
            <foreName>Fernando</foreName>
            <surname>Pereira</surname>
            <initial>F.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <editor role="editor">
          <persName key="alchemy-2005-id18146">
            <foreName>Albert</foreName>
            <surname>Cohen</surname>
            <initial>A.</initial>
          </persName>
        </editor>
        <title level="m">CC 2014 - 23rd International Conference on Compiler Construction</title>
        <loc>Grenoble, France</loc>
        <imprint>
          <publisher>
            <orgName>Springer</orgName>
          </publisher>
          <dateStruct>
            <year>2014</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00921461" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00921461</ref>
        </imprint>
        <meeting id="cid114893">
          <title>International Conference on Compiler Construction</title>
          <num>23</num>
          <abbr type="sigle">CC</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid2" type="inproceedings" rend="year" n="cite:yuki:hal-00761537">
      <identifiant type="hal" value="hal-00761537"/>
      <analytic>
        <title level="a">Array Dataflow Analysis for Polyhedral X10 Programs</title>
        <author>
          <persName key="cairn-2012-idp140691522818672">
            <foreName>Tomofumi</foreName>
            <surname>Yuki</surname>
            <initial>T.</initial>
          </persName>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName>
            <foreName>Sanjay</foreName>
            <surname>Rajopadhye</surname>
            <initial>S.</initial>
          </persName>
          <persName>
            <foreName>Vijay</foreName>
            <surname>Saraswat</surname>
            <initial>V.</initial>
          </persName>
        </author>
      </analytic>
      <monogr x-international-audience="yes" x-proceedings="yes">
        <title level="m">18th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP'13)</title>
        <loc>Shenzhen, China</loc>
        <imprint>
          <publisher>
            <orgName>ACM</orgName>
          </publisher>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00761537" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00761537</ref>
        </imprint>
        <meeting id="cid22707">
          <title>ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming</title>
          <num>18</num>
          <abbr type="sigle">PPOPP</abbr>
        </meeting>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid25" type="techreport" rend="year" n="cite:feautrier:hal-00780521">
      <identifiant type="hal" value="hal-00780521"/>
      <monogr>
        <title level="m">Enhancing the Compilation of Synchronous Dataflow Programs with a Combined Numerical-Boolean Abstraction</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName key="dart-2005-id18257">
            <foreName>Abdoulaye</foreName>
            <surname>Gamatié</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
        <imprint>
          <dateStruct>
            <month>July</month>
            <year>2013</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00780521" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00780521</ref>
        </imprint>
      </monogr>
      <note type="typdoc">Report</note>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid28" type="unpublished" rend="year" n="cite:yuki:hal-00907723">
      <identifiant type="hal" value="hal-00907723"/>
      <monogr>
        <title level="m">Checking Race Freedom of Clocked X10 Programs</title>
        <author>
          <persName key="cairn-2012-idp140691522818672">
            <foreName>Tomofumi</foreName>
            <surname>Yuki</surname>
            <initial>T.</initial>
          </persName>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName>
            <foreName>Sanjay</foreName>
            <surname>Rajopadhye</surname>
            <initial>S.</initial>
          </persName>
          <persName>
            <foreName>Vijay</foreName>
            <surname>Saraswat</surname>
            <initial>V.</initial>
          </persName>
        </author>
        <imprint>
          <dateStruct>
            <year>2013</year>
          </dateStruct>
          <biblScope type="pages">11</biblScope>
          <ref xlink:href="http://hal.inria.fr/hal-00907723" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00907723</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid18" type="inproceedings" rend="foot" n="footcite:AliasDFG10">
      <analytic>
        <title level="a">Multi-dimensional Rankings, Program Termination, and Complexity Bounds of Flowchart Programs</title>
        <author>
          <persName key="compsys-2006-id18253">
            <foreName>Christophe</foreName>
            <surname>Alias</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2005-id18078">
            <foreName>Alain</foreName>
            <surname>Darte</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">17th International Static Analysis Symposium (SAS'10)</title>
        <loc>Perpignan, France</loc>
        <imprint>
          <publisher>
            <orgName>ACM press</orgName>
          </publisher>
          <dateStruct>
            <month>September</month>
            <year>2010</year>
          </dateStruct>
          <biblScope type="pages">117-133</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid20" type="inproceedings" rend="foot" n="footcite:andrieu:hal-00760926">
      <identifiant type="hal" value="hal-00760926"/>
      <analytic>
        <title level="a">SToP: Scalable Termination Analysis of (C) Programs (Tool Presentation)</title>
        <author>
          <persName>
            <foreName>Guillaume</foreName>
            <surname>Andrieu</surname>
            <initial>G.</initial>
          </persName>
          <persName key="compsys-2006-id18253">
            <foreName>Christophe</foreName>
            <surname>Alias</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Workshop on Tools for Automatic Program Analysis (TAPAS'12)</title>
        <loc>Deauville, France</loc>
        <imprint>
          <dateStruct>
            <month>September</month>
            <year>2012</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/hal-00760926" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>hal-00760926</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid15" type="inproceedings" rend="foot" n="footcite:Boul:98">
      <analytic>
        <title level="a">Scanning Polyhedra without DO loops</title>
        <author>
          <persName key="dart-2005-id18117">
            <foreName>Pierre</foreName>
            <surname>Boulet</surname>
            <initial>P.</initial>
          </persName>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Conference on Parallel Architecture and Compilation Techniques (PACT'98)</title>
        <loc>Paris, France</loc>
        <imprint>
          <publisher>
            <orgName>IEEE Computer Society</orgName>
          </publisher>
          <dateStruct>
            <month>October</month>
            <year>1998</year>
          </dateStruct>
          <biblScope type="pages">4-11</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid14" type="article" rend="foot" n="footcite:DarteSV05">
      <analytic>
        <title level="a">Lattice-Based Memory Allocation</title>
        <author>
          <persName key="compsys-2005-id18078">
            <foreName>Alain</foreName>
            <surname>Darte</surname>
            <initial>A.</initial>
          </persName>
          <persName>
            <foreName>Robert</foreName>
            <surname>Schreiber</surname>
            <initial>R.</initial>
          </persName>
          <persName key="arenaire-2005-id18078">
            <foreName>Gilles</foreName>
            <surname>Villard</surname>
            <initial>G.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="j">IEEE Transactions on Computers</title>
        <imprint>
          <biblScope type="volume">54</biblScope>
          <biblScope type="number">10</biblScope>
          <dateStruct>
            <month>October</month>
            <year>2005</year>
          </dateStruct>
          <biblScope type="pages">1242-1257</biblScope>
        </imprint>
      </monogr>
      <note type="bnote">Special Issue: Tribute to B. Ramakrishna (Bob) Rau</note>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid8" type="article" rend="foot" n="footcite:Feau:2006b">
      <analytic>
        <title level="a">Scalable and Structured Scheduling</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="j">International Journal of Parallel Programming</title>
        <imprint>
          <biblScope type="volume">34</biblScope>
          <biblScope type="number">5</biblScope>
          <dateStruct>
            <month>October</month>
            <year>2006</year>
          </dateStruct>
          <biblScope type="pages">459–487</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid11" type="incollection" rend="foot" n="footcite:Feau:2011b">
      <analytic>
        <title level="a">Bernstein's Conditions</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <editor role="editor">
          <persName>
            <foreName>David</foreName>
            <surname>Padua</surname>
            <initial>D.</initial>
          </persName>
        </editor>
        <title level="m">Encyclopedia of Parallel Programming</title>
        <imprint>
          <publisher>
            <orgName>Springer</orgName>
          </publisher>
          <dateStruct>
            <year>2011</year>
          </dateStruct>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid21" type="techreport" rend="foot" n="footcite:Feau:2011a">
      <identifiant type="hal" value="inria-00609519"/>
      <monogr>
        <title level="m">Simplification of Boolean Affine Formulas</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
        </author>
        <imprint>
          <biblScope type="number">RR-7689</biblScope>
          <publisher>
            <orgName type="institution">Inria</orgName>
          </publisher>
          <dateStruct>
            <month>July</month>
            <year>2011</year>
          </dateStruct>
          <ref xlink:href="http://hal.inria.fr/inria-00609519/PDF/RR-7689.pdf" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>hal.<allowbreak/>inria.<allowbreak/>fr/<allowbreak/>inria-00609519/<allowbreak/>PDF/<allowbreak/>RR-7689.<allowbreak/>pdf</ref>
        </imprint>
      </monogr>
      <note type="typdoc">Technical report</note>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid12" type="article" rend="foot" n="footcite:Feau:91">
      <analytic>
        <title level="a">Dataflow Analysis of Scalar and Array References</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="j">International Journal of Parallel Programming</title>
        <imprint>
          <biblScope type="volume">20</biblScope>
          <biblScope type="number">1</biblScope>
          <dateStruct>
            <month>February</month>
            <year>1991</year>
          </dateStruct>
          <biblScope type="pages">23–53</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid27" type="article" rend="foot" n="footcite:FeaGG12">
      <analytic>
        <title level="a">Enhancing the Compilation of Synchronous Dataflow Programs with a Combined Numerical-Boolean Abstraction</title>
        <author>
          <persName key="compsys-2005-id18170">
            <foreName>Paul</foreName>
            <surname>Feautrier</surname>
            <initial>P.</initial>
          </persName>
          <persName key="dart-2005-id18257">
            <foreName>Abdoulaye</foreName>
            <surname>Gamatié</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="j">CSI Journal of Computing</title>
        <imprint>
          <biblScope type="volume">1</biblScope>
          <biblScope type="number">4</biblScope>
          <dateStruct>
            <year>2012</year>
          </dateStruct>
          <biblScope type="pages">8:86</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid13" type="inproceedings" rend="foot" n="footcite:isss01">
      <analytic>
        <title level="a">Loop Fusion for Memory Space Optimization</title>
        <author>
          <persName key="compsys-2005-id18213">
            <foreName>Antoine</foreName>
            <surname>Fraboulet</surname>
            <initial>A.</initial>
          </persName>
          <persName>
            <foreName>Karen</foreName>
            <surname>Godary</surname>
            <initial>K.</initial>
          </persName>
          <persName>
            <foreName>Anne</foreName>
            <surname>Mignotte</surname>
            <initial>A.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Symposium on System Synthesis (ISSS'01)</title>
        <loc>Montréal, Canada</loc>
        <imprint>
          <publisher>
            <orgName>IEEE Press</orgName>
          </publisher>
          <dateStruct>
            <month>October</month>
            <year>2001</year>
          </dateStruct>
          <biblScope type="pages">95–100</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid26" type="inproceedings" rend="foot" n="footcite:GamatieGonnord:lctes11">
      <analytic>
        <title level="a">Static Analysis of Synchronous Programs in Signal for Efficient Design of Multi-Clocked Embedded Systems</title>
        <author>
          <persName key="dart-2005-id18257">
            <foreName>Abdoulaye</foreName>
            <surname>Gamatié</surname>
            <initial>A.</initial>
          </persName>
          <persName key="compsys-2008-id18330">
            <foreName>Laure</foreName>
            <surname>Gonnord</surname>
            <initial>L.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Conference on Languages, Compilers, and Tools for Embedded Systems (LCTES'11)</title>
        <loc>Chicago, USA</loc>
        <imprint>
          <dateStruct>
            <month>April</month>
            <year>2011</year>
          </dateStruct>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid24" type="inproceedings" rend="foot" n="footcite:hong.81.stoc">
      <analytic>
        <title level="a">I/O Complexity: The Red-Blue Pebble Game</title>
        <author>
          <persName>
            <foreName>Jia-Wei</foreName>
            <surname>Hong</surname>
            <initial>J.-W.</initial>
          </persName>
          <persName>
            <foreName>H. T.</foreName>
            <surname>Kung</surname>
            <initial>H. T.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">13th Annual ACM Symposium on Theory of Computing (STOC'81)</title>
        <imprint>
          <publisher>
            <orgName>ACM</orgName>
          </publisher>
          <dateStruct>
            <year>1981</year>
          </dateStruct>
          <biblScope type="pages">326–333</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid6" type="inproceedings" rend="foot" n="footcite:PPG">
      <analytic>
        <title level="a">Analysis Techniques for Predicated Code</title>
        <author>
          <persName>
            <foreName>R.</foreName>
            <surname>Johnson</surname>
            <initial>R.</initial>
          </persName>
          <persName>
            <foreName>M.</foreName>
            <surname>Schlansker</surname>
            <initial>M.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">29th Annual ACM/IEEE International Symposium on Microarchitecture (MICRO-29)</title>
        <loc>Paris, France</loc>
        <imprint>
          <publisher>
            <orgName>IEEE Computer Society</orgName>
          </publisher>
          <dateStruct>
            <year>1996</year>
          </dateStruct>
          <biblScope type="pages">100–113</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid7" type="inproceedings" rend="foot" n="footcite:psi-ssa">
      <analytic>
        <title level="a">Efficient Static Single Assignment Form for Predication</title>
        <author>
          <persName>
            <foreName>Arthur</foreName>
            <surname>Stoutchinin</surname>
            <initial>A.</initial>
          </persName>
          <persName>
            <foreName>François</foreName>
            <surname>De Ferrière</surname>
            <initial>F.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">34th Annual ACM/IEEE International Symposium on Microarchitecture (MICRO-34)</title>
        <loc>Austin, Texas</loc>
        <imprint>
          <publisher>
            <orgName>IEEE Computer Society</orgName>
          </publisher>
          <dateStruct>
            <year>2001</year>
          </dateStruct>
          <biblScope type="pages">172–181</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid9" type="inproceedings" rend="foot" n="footcite:turjan04">
      <analytic>
        <title level="a">Translating Affine Nested-Loop Programs to Process Networks</title>
        <author>
          <persName>
            <foreName>Alexandru</foreName>
            <surname>Turjan</surname>
            <initial>A.</initial>
          </persName>
          <persName>
            <foreName>Bart</foreName>
            <surname>Kienhuis</surname>
            <initial>B.</initial>
          </persName>
          <persName>
            <foreName>Ed</foreName>
            <surname>Deprettere</surname>
            <initial>E.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Conference on Compilers, Architecture, and Synthesis for Embedded Systems (CASES'04)</title>
        <loc>New York, NY, USA</loc>
        <imprint>
          <publisher>
            <orgName>ACM</orgName>
          </publisher>
          <dateStruct>
            <year>2004</year>
          </dateStruct>
          <biblScope type="pages">220–229</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid10" type="inproceedings" rend="foot" n="footcite:skimo06">
      <analytic>
        <title level="a">Improved Derivation of Process Networks</title>
        <author>
          <persName key="alchemy-2010-id59677">
            <foreName>Sven</foreName>
            <surname>Verdoolaege</surname>
            <initial>S.</initial>
          </persName>
          <persName>
            <foreName>Hristo</foreName>
            <surname>Nikolov</surname>
            <initial>H.</initial>
          </persName>
          <persName>
            <foreName>Nikolov</foreName>
            <surname>Todor</surname>
            <initial>N.</initial>
          </persName>
          <persName>
            <foreName>Plamenov</foreName>
            <surname>Stefanov</surname>
            <initial>P.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">International Workshop on Optimization for DSP and Embedded Systems (ODES'06)</title>
        <imprint>
          <dateStruct>
            <year>2006</year>
          </dateStruct>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid17" type="inproceedings" rend="foot" n="footcite:flopoco">
      <identifiant type="hal" value="ensl-00379154"/>
      <analytic>
        <title level="a">Generating High-Performance Custom Floating-Point Pipelines</title>
        <author>
          <persName>
            <foreName>Florent</foreName>
            <surname>De Dinechin</surname>
            <initial>F.</initial>
          </persName>
          <persName key="graal-2009-id60562">
            <foreName>Cristian</foreName>
            <surname>Klein</surname>
            <initial>C.</initial>
          </persName>
          <persName key="arenaire-2008-id18482">
            <foreName>Bogdan</foreName>
            <surname>Pasca</surname>
            <initial>B.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">Field Programmable Logic and Applications</title>
        <imprint>
          <publisher>
            <orgName>IEEE</orgName>
          </publisher>
          <dateStruct>
            <month>August</month>
            <year>2009</year>
          </dateStruct>
          <ref xlink:href="http://prunel.ccsd.cnrs.fr/ensl-00379154/" location="extern" xlink:type="simple" xlink:show="replace" xlink:actuate="onRequest">http://<allowbreak/>prunel.<allowbreak/>ccsd.<allowbreak/>cnrs.<allowbreak/>fr/<allowbreak/>ensl-00379154/</ref>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid5" type="inproceedings" rend="foot" n="footcite:cfab14">
      <analytic>
        <title level="a">Parallel Execution of Saturated Reductions</title>
        <author>
          <persName>
            <foreName>Benoît</foreName>
            <surname>Dupont de Dinechin</surname>
            <initial>B.</initial>
          </persName>
          <persName>
            <foreName>Christophe</foreName>
            <surname>Monat</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2005-id18199">
            <foreName>Fabrice</foreName>
            <surname>Rastello</surname>
            <initial>F.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <title level="m">Workshop on Signal Processing Systems (SIPS'01)</title>
        <imprint>
          <publisher>
            <orgName>IEEE Computer Society Press</orgName>
          </publisher>
          <dateStruct>
            <year>2001</year>
          </dateStruct>
          <biblScope type="pages">373-384</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
    
    <biblStruct id="compsys-2013-bid22" type="inproceedings" rend="foot" n="footcite:LeGuenGR11">
      <analytic>
        <title level="a">MinIR, a Minimalistic Intermediate Representation</title>
        <author>
          <persName>
            <foreName>Julien</foreName>
            <surname>Le Guen</surname>
            <initial>J.</initial>
          </persName>
          <persName>
            <foreName>Christophe</foreName>
            <surname>Guillon</surname>
            <initial>C.</initial>
          </persName>
          <persName key="compsys-2005-id18199">
            <foreName>Fabrice</foreName>
            <surname>Rastello</surname>
            <initial>F.</initial>
          </persName>
        </author>
      </analytic>
      <monogr>
        <editor role="editor">
          <persName key="compsys-2005-id18414">
            <foreName>Florent</foreName>
            <surname>Bouchez</surname>
            <initial>F.</initial>
          </persName>
          <persName key="compsys-2006-id18274">
            <foreName>Sebastian</foreName>
            <surname>Hack</surname>
            <initial>S.</initial>
          </persName>
          <persName>
            <foreName>Eelco</foreName>
            <surname>Visser</surname>
            <initial>E.</initial>
          </persName>
        </editor>
        <title level="m">Workshop on Intermediate Representations (WIR'11), held with CGO'11</title>
        <loc>Chamonix</loc>
        <imprint>
          <dateStruct>
            <month>April</month>
            <year>2011</year>
          </dateStruct>
          <biblScope type="pages">5-12</biblScope>
        </imprint>
      </monogr>
    </biblStruct>
  </biblio>
</raweb>
