<?xml version="1.0" encoding="UTF-8"?><rss version="2.0">
  <channel>
    <title>Hell Oh Entropy!</title>
    <link>https://gotplt.org/</link>
    <description>Life, Code and everything in between</description>
    <managingEditor>siddhesh.poyarekar@gmail.com (Siddhesh)</managingEditor>
    <pubDate>28 Sep 25 23:16 UTC</pubDate>
    <item>
      <title>Finding reason, finding belonging</title>
      <link>https://gotplt.org/posts/finding-reason-finding-belonging.html</link>
      <description>&lt;!--&#xA;.. title: Finding reason, finding belonging&#xA;.. slug: finding-reason-finding-belonging&#xA;.. date: 2025-09-28T15:49:48-07:00&#xA;.. tags: cauldron, glibc, gcc, binutils, gnutools, gnu, toolchain&#xA;.. link:&#xA;.. description: An emotional retrospective of the 2025 GNU Tools Cauldron&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;I&amp;rsquo;ve been around in the Free and Open Source world for over 20 years now,&#xA;initially as a user but largely as a hacker.  Since I started making money&#xA;using (and growing) my technical skills, it never once occurred to me that I&amp;rsquo;d&#xA;be doing anything other than writing FOS software.  Which is why over the years&#xA;as I grew up the Engineer Value Stack* in my career, I started shedding some&#xA;things that once used to give my joy.  With that train of thought, I almost did&#xA;not go to the 2025 GNU Tools Cauldron that just concluded in Porto today.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Dear reader, if you&amp;rsquo;re expecting a review of all of the technically awesome&#xA;things that happened at Cauldron this weekend, stop reading and wait a couple&#xA;of days.  Jonathan Corbett was there so I assume he will have something&#xA;interesting to write in that space.  Go watch the &lt;a href=&#34;https://lwn.net/Archives&#34;&gt;LWN&#xA;feed&lt;/a&gt; and maybe even buy a subscription if you&#xA;haven&amp;rsquo;t already.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;So yeah, I almost decided to not go to Porto for Cauldron, because for the past&#xA;year or so, I didn&amp;rsquo;t feel like I did anything of consequence in the GNU&#xA;toolchain community.  Sitting alone in my basement in Waterloo, I had already&#xA;concluded to myself that nobody would miss that I wasn&amp;rsquo;t there.  Things would&#xA;go on as usual.  I had already forgotten whatever work I had done over the last&#xA;years; they didn&amp;rsquo;t feel valuable enough.  I had concluded that I was mostly a&#xA;glorified Jira wrangler (the modern equivalent of the &amp;ldquo;paper pusher&amp;rdquo; slur one&#xA;could use to denigrate anybody who doesn&amp;rsquo;t do Real Work&amp;trade;) and I wasn&amp;rsquo;t&#xA;needed.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I did fly in the end, and arrived at a hot (OK, 24C, but I&amp;rsquo;m practically&#xA;Canadian now so forgive me) Porto, still unsure why I was there.  To try and&#xA;brush off the fatigue, I walked with Carlos (there&amp;rsquo;s nothing like a loud,&#xA;always driven Argentinian man to lift your spirits) to FEUP, ending our day&#xA;downtown meeting many of the attendees, many friends.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;That seemed like a great day, and a great evening.  But still, was it worth&#xA;flying the 7 hours?  I wasn&amp;rsquo;t sure.  Anyway, that&amp;rsquo;s one down, 3 more days to&#xA;go.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;&lt;strong&gt;Cut to Sunday night, I&amp;rsquo;m now wondering where those 3 days went in a blur, and&#xA;wondering what washed away those doubts I had coming in.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it was the wonderful Belgian beer we had on Friday night.  Or was it&#xA;because I was with old friends and new, talking about everything under the sun?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it was the chilled Londrina house beer that stayed ice cold to the last&#xA;drop.  Or was it the joy of finding a shared love for working with wood with a&#xA;colleague who I had only occasionally chatted over the intertubes all these&#xA;years?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it was the &lt;a href=&#34;https://en.wikipedia.org/wiki/Francesinha&#34;&gt;Francesinha&lt;/a&gt;, a&#xA;sinfully tasty but likely just as unhealthy Portuguese dish.  Or was it the&#xA;colleague who suggested the dish, who also had welcomed me earlier, genuinely&#xA;and warmly, saying &amp;ldquo;here comes the great Sid!&amp;rdquo; instantly making me feel like I&#xA;matter?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it was the wonderful port wine and fancy 3 course meal we had at&#xA;Taylor&amp;rsquo;s.  Or was it the intimate conversations I had with some new friends and&#xA;old about our failures and insecurities, and how they shaped us?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it is the wonderful catering Cupertino and folks arranged for our lunches&#xA;at Cauldron.  Or was it the new connections I made with people I looked at with&#xA;admiration across the hall all these years but never had the courage to walk&#xA;across and introduce myself?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Was it the fact that all of the most amazing leaders in the GNU tools ecosystem&#xA;were there?  Or was it the relief at seeing an old friend and mentor (I don&amp;rsquo;t&#xA;know if he knows how much my interactions with him meant to me) safe and doing&#xA;well?  Or, in fact, was it the realization of how much I owe it to pretty much&#xA;every person who has been coming to Cauldron regularly, probably with their own&#xA;personal reasons, but leaving their own, indelible impression on me as a&#xA;person?  Or, of course, the annual JL (if you know you know) therapy session?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Maybe it was the fantastic surprise musical performance by a group of school&#xA;kids, which reminded me of my kiddo back home.  OK maybe that one actually had&#xA;me longing to return home soon.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Anyway, the material experiences tend to get washed away days after I return&#xA;from these experiences, but the personal and emotional ones are not permanent&#xA;either.  They&amp;rsquo;ve shaped me and made me the person I am, but months later, I&#xA;know I will have forgotten why I loved being here, with my friends, people with&#xA;whom I share my commitment to Free and Open Source Software.  I&amp;rsquo;m writing this&#xA;with the hope that I&amp;rsquo;ll come back here to remind myself of why it matters, to&#xA;remind myself that I belong, to be grateful to all of those people who made me&#xA;feel like I belong.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;If you were here the first time, note that I&amp;rsquo;ve been here for over 13 years&#xA;now, and you belong, just as I do.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;* Engineer Value Stack: A hierarchy that companies tend to have in their&#xA;engineering organizations based on an engineer&amp;rsquo;s ability to effectively&#xA;communicate their ideas and work with their peers, as opposed to the&#xA;superiority of their technical skills, which in itself is also a nebulous&#xA;concept that only serves to promote a deep impostor syndrome among most&#xA;competent developers.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>28 Sep 25 22:49 UTC</pubDate>
    </item>
    <item>
      <title>_FORTIFY_SOURCE=3 performance</title>
      <link>https://gotplt.org/posts/fortify-source-3-performance.html</link>
      <description>&lt;!--&#xA;.. title: _FORTIFY_SOURCE=3 performance&#xA;.. slug: fortify-source-3-performance&#xA;.. date: 2023-01-05T16:28:10-05:00&#xA;.. tags: security, c, fortify, c++, gcc, glibc, clang, fedora&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;So early last year I finished implemented everything needed for a fully working &lt;code&gt;_FORTIFY_SOURCE=3&lt;/code&gt; so that disrtributions can use it out of the box.  OpenSUSE adopted it almost immediately and Gentoo started the work of adding it to their hardened profile. I proposed to make it the default for Fedora 38 after some tests but people quoted to me &lt;a href=&#34;https://developers.redhat.com/articles/2022/09/17/gccs-new-fortification-level#&#34;&gt;this blog post that some guy wrote&lt;/a&gt;, telling me that there&amp;rsquo;s a performance issue.  Since my explanations and clarifications in the Fedora wiki or on the Fedora devel list is not sufficient (the feature was approved but the &amp;ldquo;&lt;code&gt;_FORTIFY_SOURCE=3&lt;/code&gt; has performance overhead&amp;rdquo; claims don&amp;rsquo;t seem to stop), here&amp;rsquo;s a blog post for a blog post, stating conclusively that the performance issue is theoretical and overstated, the guy didn&amp;rsquo;t know what he was talking about when he wrote it.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;That guy is working on a clarification blog post of his own, describing in some more detail why the concern is overblown but he has to jump through editorial hoops of a multi-billion dollar corporation that pays his salary, so his apology to me is going to take a while.  Whenever he gets to publish his work, I&amp;rsquo;ll link it here so that it&amp;rsquo;s two blog posts against one.  Take that!&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>05 Jan 23 21:28 UTC</pubDate>
    </item>
    <item>
      <title>The beauty of equations in Physics</title>
      <link>https://gotplt.org/posts/the-beauty-of-equations-in-physics.html</link>
      <description>&lt;!--&#xA;.. title: The beauty of equations in Physics&#xA;.. slug: the-beauty-of-equations-in-physics&#xA;.. date: 2022-06-23T06:48:35+05:30&#xA;.. tags: physics, equations, maths&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;I have been stereotyped by many I know (and understandably so) as a quintessential computer geek, probably someone who dabbled in computers since his childhood and is his first love. That is however far from the truth because I first programmed a computer at the age of 20 and started my career soon after as a lowly back office outsourced engineer that a lot of the geek community looks down upon. What I did grow up with was something related but still different - mathematics and physics.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Around the age of 15 my mother enrolled me for IIT coaching classes (a huge financial struggle and a social shock for me, but that&amp;rsquo;s another story) and I found teachers that ignited a love for these subjects. I was always technically inclined thanks to my father who was an engineer.  After his early death everyone around me wanted me to be his &amp;lsquo;successor&amp;rsquo; as an engineer and I happily (almost proudly then, what does a 9 year old know!) obliged.  Whatever technical inclination I had was due to watching my father as a 6-9 year old (another interesting story) and it was enough to make me an above average math and science student.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;In the IIT coaching classes, I learned in Physics, a way to look at the world and in mathematics, a way to express that view.  Despite the ragging due to my social status (backward caste and poor, in a predominantly rich, upper caste, South Bombay clique), I was floating in the air those two years, making up problems to trick friends and solving equations for fun.  I don&amp;rsquo;t think I&amp;rsquo;ve done maths just for the heck of it since.  The love affair with physics and its mathematics did not live too long though as I flunked my IIT examinations and had to rethink everything, including my view of myself; it woouldn&amp;rsquo;t be the last time either.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I went into what I now recognize as depression for over a year.  As I recovered, I was ankle deep in Linux and FOSS and it was the beginning of the next chapter of my life, a life that has more extensive documentation on the intertubes.  The physics and maths got beaten out of me in the process though.  One thing however seems to have stuck, my obsession for beauty and rhythm whenever I encounter a mathematical equation. If a result looked large and wieldy, I would get very uncomfortable and keep working at it, refactoring it till it looked beautiful and reading it out sounded like I was reciting a poem. It was an obsession that my teachers loved and hated depending on how they were feeling on that day.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I rediscovered that love for symmetry and rhythm when I spent some time working with multiple precision math nearly a decade ago. I discovered it once again some years ago with just a few minutes of hacking away at a physics problem at reserved-bit where I came up with equations for a little maze that the kids at the makerspace wanted to build. The immense satisfaction of seeing the equation being easy on the eyes and almost musical to read is a feeling I cannot express in words. Then there is the beauty of discovering little facts by reading the equation (like location of an object at any point in the maze being independent of the acceleration due to gravity for the maze) that adds to the wonderful experience.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There is a parallel to such beauty in programming in the form of APIs or algorithms, but it doesn&amp;rsquo;t quite feel the same to me. I guess I enjoy programming quite a lot but no, I don&amp;rsquo;t love it like I did physics and maths.  I don&amp;rsquo;t seem to have the mental space to go back to it though.  I guess it&amp;rsquo;s a first love that I can only look back fondly at for now.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>23 Jun 22 01:18 UTC</pubDate>
    </item>
    <item>
      <title>That is not a number, that is a freed object</title>
      <link>https://gotplt.org/posts/that-is-not-a-number-that-is-a-freed-object.html</link>
      <description>&lt;!--&#xA;.. title: That is not a number, that is a freed object&#xA;.. slug: that-is-not-a-number-that-is-a-freed-object&#xA;.. date: 2022-04-20T12:01:49+05:30&#xA;.. tags: malloc, security, c, standards, use-after-free&#xA;.. link:&#xA;.. description: Using even the value of pointers after they have been freed is dangerous and should be avoided&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;How many of you have written this kind of code in the past:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;o = xmalloc (old_size);&#xA;...&#xA;n = xrealloc (o, new_size);&#xA;&#xA;if (n != o)&#xA;  {&#xA;    o = n;&#xA;    /* Update other pointers that referred to o or offsets from it.  */&#xA;  }&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;Not uncommon right?  We&amp;rsquo;re not dereferencing the freed &lt;code&gt;o&lt;/code&gt; and the pointer is after all, a number and hence should be perfectly safe to check, right?  And more optimal too since we&amp;rsquo;re not updating pointers if it&amp;rsquo;s not necessary.  Well&amp;hellip;&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Better Fortification&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;TLDR; I broke this &amp;lsquo;safety&amp;rsquo; in my implementation for &lt;code&gt;__builtin_dynamic_object_size&lt;/code&gt; in gcc but I&amp;rsquo;m not wrong, you are!  See the last section for why.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Now for those of you interested in the &lt;em&gt;story&lt;/em&gt;, it all began with the implementation for &lt;code&gt;__builtin_dynamic_object_size&lt;/code&gt;.  This builtin was implemented first in clang and promised to be a better &lt;code&gt;__builtin_object_size&lt;/code&gt;, which was severely limited by its necessity to emit a constant.  That restriction meant that (1) there were many cases where it just couldn&amp;rsquo;t arrive at a constant size and (2) where it did, it would come up with an upper or lower estimate and not necessarily a precise size.  Given that the builtin is primarily used to implement &lt;code&gt;_FORTIFY_SOURCE&lt;/code&gt; (there&amp;rsquo;s a &lt;a href=&#34;https://www.redhat.com/en/blog/security-technologies-fortifysource&#34;&gt;more detailed blog post&lt;/a&gt; describing its mechanism out there), this directly reduces the scope of this security protection.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;&lt;code&gt;__builtin_dynamic_object_size&lt;/code&gt; had deeper implications than just being a dynamic version of &lt;code&gt;__builtin_object_size&lt;/code&gt; however, which had led to &lt;a href=&#34;https://www.mail-archive.com/gcc@gcc.gnu.org/msg87232.html&#34;&gt;initial pushback&lt;/a&gt; in the gcc community.  Now that the implementation is due to come out in gcc 12.1 and is being tested with distribution rebuilds, new and interesting implications are being discovered.  One of these (and so far the most fascinating to me) was its impact on using (not dereferencing, mind you) a freed pointer.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;How gcc deduces object sizes&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The object size computation is largely (there are some caveats here but not important for the purposes of this post) done in a &lt;a href=&#34;https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/tree-object-size.cc;h=fc062b94d7628d139fea8f9abce4a1e7517d779d;hb=HEAD&#34;&gt;separate pass&lt;/a&gt;. The pass runs twice, once &lt;a href=&#34;https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/passes.def;h=375d3d62d519329d62b4ba0ccd994503ac14710b;hb=HEAD#l79&#34;&gt;very early in the pass chain&lt;/a&gt; and finally, &lt;a href=&#34;https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/passes.def;h=375d3d62d519329d62b4ba0ccd994503ac14710b;hb=HEAD#l203&#34;&gt;near the end&lt;/a&gt; of the tree passes.  The early run is a hack that tries to record subobject size estimates before subsequent passes simplify subobject references to references to their parent object, thus returning a more precise subobject size.  The late run is where the actual fun happens.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The object sizes pass, &lt;a href=&#34;https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/tree-object-size.cc;h=fc062b94d7628d139fea8f9abce4a1e7517d779d;hb=HEAD#l1017&#34;&gt;at a high level&lt;/a&gt;, tracks the pointer passed to either &lt;code&gt;__builtin_object_size&lt;/code&gt; or &lt;code&gt;__builtin_dynamic_object_size&lt;/code&gt; to &lt;a href=&#34;https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/tree-object-size.cc;h=fc062b94d7628d139fea8f9abce4a1e7517d779d;hb=HEAD#l1591&#34;&gt;all possible objects&lt;/a&gt; it may point to and subsequently, to the site of their assignment, to derive the size.  In the static case (i.e. &lt;code&gt;__builtin_object_size&lt;/code&gt;), it tries to come up with either the maximum or the minimum estimate while in the dynamic case it builds a fancy expression that would evaluate to the precise size at that point. Of course, &amp;lsquo;precise&amp;rsquo; shouldn&amp;rsquo;t be taken for granted because there could be future changes that make the expressions imprecise in the interest of broadening coverage.  If the pass is unable to deduce a size of any of the target objects of the pointer for any reason (passed through a call, non-constant in the static case, etc.), the call is replaced with &lt;code&gt;(size_t) -1&lt;/code&gt; or &lt;code&gt;(size_t) 0&lt;/code&gt; as appropriate.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;I can&amp;rsquo;t judge what I can&amp;rsquo;t see&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;As the pass tracks origins of the pointer in question, it unfortunately does not take into account any uses between the allocation and the reference in the builtin that may alter the nature of the pointer.  This means that if the pointer was reallocated between its first allocation and the builtin call, the pass won&amp;rsquo;t notice unless the pointer was explicitly updated.  This is a benign limitation in the static case because for the above example, it would simply compute the maximum of &lt;code&gt;new_size&lt;/code&gt; and &lt;code&gt;old_size&lt;/code&gt; and return the result.  In fact in most real world cases since the reallocation is bound to be dynamic, it would simply bail out, resulting in a missed fortification.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;With dynamic sizes though, one will now get the new size for &lt;code&gt;n != o&lt;/code&gt; but not for the &lt;code&gt;n == o&lt;/code&gt; case.  As a result, any fortified function call based on this information will see the old size and abort fearing a buffer overflow even though there technically wasn&amp;rsquo;t any.  This was seen in &lt;a href=&#34;https://sourceforge.net/p/autogen/bugs/212/&#34;&gt;autogen&lt;/a&gt;, which had this precise pattern and hence stumbled when it was built with &lt;code&gt;_FORTIFY_SOURCE=3&lt;/code&gt;.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;It&amp;rsquo;s a bug, it&amp;rsquo;s not a bug&amp;hellip;&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;After a bit of back and forth, Martin Liška very helpfully came up with a &lt;a href=&#34;https://gcc.gnu.org/bugzilla/show_bug.cgi?id=105217&#34;&gt;contained reproducer&lt;/a&gt; that allowed us to see what had actually happened.  I had broken a pretty common idiom, which meant that those applications would have false positive &lt;em&gt;aborts&lt;/em&gt;, something that hadn&amp;rsquo;t happened with &lt;code&gt;_FORTIFY_SOURCE&lt;/code&gt; before.  That is until I found an excuse that I could use to point the finger back at you (which includes past me, who is clearly a different person, no?), the developer!&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Object Lifetimes&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;clang 13 also broke with the test case Martin shared after I altered it a bit to fortify &lt;code&gt;fread&lt;/code&gt;.  That gave me first relief because clearly whatever I did wrong, the smart folks in the clang community did wrong too.  So I wasn&amp;rsquo;t &lt;em&gt;that&lt;/em&gt; stupid.  Then of course, there was this, which put our collective &amp;lsquo;stupidity&amp;rsquo; into perspective, kinda letting us off the hook:&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Section 6.2.4 of the ISO C standard (I&amp;rsquo;m referring to an April 2011 draft because who even in their right minds pays for their copy?!) has this in point number 2:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;The lifetime of an object is the portion of program execution during which storage is&#xA;guaranteed to be reserved for it. An object exists, has a constant address, and retains&#xA;its last-stored value throughout its lifetime. If an object is referred to outside of its&#xA;lifetime, the behavior is undefined. The value of a pointer becomes indeterminate when&#xA;the object it points to (or just past) reaches the end of its lifetime.&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;It clearly states that even the value of the pointer pointing to the object is not reliable after it has been freed, so not only should one avoid dereferencing the pointer after it is freed, they should refrain from using it altogether.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Essentially, the comparison with the old pointer results in undefined behaviour.  I don&amp;rsquo;t think the standards committee intended to invalidate this specific idiom with that language, but it does allow compilers the freedom to make assumptions about pointer validity and this idiom ends up trouncing on it.  It is possible for the compiler to look for a dominating realloc and update its expectations for size in very specific cases, but it still remains largely unsupported.  It won&amp;rsquo;t, for example, work in cases where a reallocation has been wrapped in a function without any &lt;code&gt;malloc&lt;/code&gt; attribute annotations.  In fact, gcc 12 has a new &lt;code&gt;-Wuse-after-free&lt;/code&gt; option that warns users of this that I, admittedly, once thought was too harsh.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;EDIT 2022-04-21: This spawned a conversation in the rust community and &lt;a href=&#34;https://zulip-archive.rust-lang.org/stream/136281-t-lang/wg-unsafe-code-guidelines/topic/pointer.20comparisons.20after.20free.2Frealloc.html#279604865&#34;&gt;Ralf Jung&lt;/a&gt; pointed out a way to think about this in pointer provenance terms and does not rely on the above C standard indeterminate pointer clause.  This is very relevant because what the object size pass does in this context is essentially pointer provenance (albeit limited and somewhat incomplete), which makes it natural for it to trip on this implicit assumption of &lt;code&gt;o == n&lt;/code&gt;.  Continuing to use &lt;code&gt;o&lt;/code&gt; (and any pointers derived from it) in this context is incorrect.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Getting better together&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;I&amp;rsquo;m going to try and support some of these simple cases in gcc during the gcc 13 cycle but in general, this is undefined behaviour.  If your code uses this idiom, you should start weaning away from it if it&amp;rsquo;s not performance sensitive and unconditionally update pointers once their lifetime ends.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Deploying &lt;code&gt;_FORTIFY_SOURCE=3&lt;/code&gt; more widely has been a learning experience (all owing to Martin Liška&amp;rsquo;s efforts since he was the one building thousands of packages and reporting bugs!) in the deeper implications that &lt;code&gt;__builtin_dynamic_object_size&lt;/code&gt; would have when replacing &lt;code&gt;__builtin_object_size&lt;/code&gt;.  Another interesting implication was the misuse of &lt;code&gt;malloc_usable_size&lt;/code&gt; and equivalent interfaces that we discovered with &lt;a href=&#34;https://github.com/systemd/systemd/issues/22801&#34;&gt;systemd&lt;/a&gt; and &lt;a href=&#34;https://github.com/jemalloc/jemalloc/issues/2238&#34;&gt;jemalloc&lt;/a&gt; that open up deeper design questions for malloc interfaces.  More on that in a separate post either here or on one of the Red Hat blogs.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;A simple change of more precise object sizes and wider coverage ended up not only weeding out actual overflows, but also some interesting corner cases and &amp;ldquo;adventurous&amp;rdquo; programming practices.  I&amp;rsquo;m going to start rolling some of this out into Fedora near the end of the year and we&amp;rsquo;ll hopefully have better mitigations in Linux distributions very soon.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>20 Apr 22 06:31 UTC</pubDate>
    </item>
    <item>
      <title>Relocations: fantastic symbols, but where to find them?</title>
      <link>https://gotplt.org/posts/relocations-fantastic-symbols-but-where-to-find-them.html</link>
      <description>&lt;!--&#xA;.. title: Relocations: fantastic symbols, but where to find them?&#xA;.. slug: relocations-fantastic-symbols-but-where-to-find-them&#xA;.. date: 2020-04-10T12:36:45-07:00&#xA;.. tags: gnutools, glibc, gcc, gas, ld, linker, compiler, aarch64, arm, toolchain&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;&lt;em&gt;Update: There&amp;rsquo;s a Links section at the end that should give you a list of all the reference you&amp;rsquo;ll ever need to undrstand Linkers and Relocations!&lt;/em&gt;&lt;/p&gt;&#xA;&#xA;&lt;p&gt;When I started out in compilers years ago, I found relocations especially hard to wrap my head around.  They&amp;rsquo;re just simple math in the end but they combine elements from different places that make it complicated enough that many (as did I back then) assume it to be some kind of black art.  The &lt;a href=&#34;https://docs.oracle.com/cd/E26502_01/html/E26507/chapter3-29.html&#34;&gt;Oracle documentation&lt;/a&gt; on relocation processing and &lt;a href=&#34;https://docs.oracle.com/cd/E23824_01/html/819-0690/chapter6-54839.html&#34;&gt;relocation sections&lt;/a&gt; is pretty much the best thing I&amp;rsquo;ve found on the internet that explains relocations and they&amp;rsquo;re a great start if you already know what you&amp;rsquo;re doing.  This is why I figured I ought to try writing something more accessible that puts some of the bits into practice.  The fact that I&amp;rsquo;m currently working on relocation processing makes it that much simpler for me to just bash at the keyboard and commit the stuff in my memory to a more persistent medium before it gets swapped out to make space for kitten photos.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;While this tutorial is meant to be beginner friendly, it does assume though that the reader has some awareness of the ELF format, at least to the extent of knowing about different ELF sections.  It also assumes that the reader has some undersanding of assembly language programming, bonus if you know aarch64 assembly since that&amp;rsquo;s the flavour of the examples.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Finding each other in this crowded world&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;The idea of relocation is quite simple: when compiling programs, we need flexibility to build components of programs independently and then have them link together.  This could be in the same source file where we don&amp;rsquo;t know where parts of the source would end up, multiple source files built into different object files or sets of object files built into different libraries that reference objects in each other.  This flexibility is achieved using relocations.  Here&amp;rsquo;s a very simple example using aarch64 assembly:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;.data&#xA;.globl somedata&#xA;somedata:&#xA;&#x9;.8byte 0x42&#xA;&#xA;.text&#xA;.globl start&#xA;_start:&#xA;&#x9;nop&#xA;&#x9;ldr&#x9;x2, somedata&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;This is a simple program that loads &lt;code&gt;somedata&lt;/code&gt; into register &lt;code&gt;x2&lt;/code&gt;.  It doesn&amp;rsquo;t do much and if you try to run the program it will crash, but it is a useful example that shows an assembly source where parts of it end up in different ELF sections.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The interesting bit here is a load instruction in the text section that is reading a variable &lt;code&gt;somedata&lt;/code&gt; from the data section.  The load instruction encodes within itself, the offset of &lt;code&gt;somedata&lt;/code&gt; from itself, aka the PC-relative offset.  The assembler can see both, the variable and the instruction but it cannot say for sure where they will be in memory at this point, because it does not know how far the data section will be from the text section in the final library or executable.  To work around this limitation, it needs to leave the offset field in the ldr instruction blank so that the linker can finally fill it in.  It also needs to provide instructions to the linker to tell it how to fill in this field.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;This is where relocations come into play.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;In this specific case, one may assemble the example and disassemble it using &lt;code&gt;objdump -Dr&lt;/code&gt; to find this disassembly:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Disassembly of section .text:&#xA;&#xA;0000000000000000 &amp;lt;_start&amp;gt;:&#xA;   0:&#x9;d503201f &#x9;nop&#xA;   4:&#x9;58000002 &#x9;ldr&#x9;x2, 0 &amp;lt;_start&amp;gt;&#xA;&#x9;&#x9;&#x9;4: R_AARCH64_LD_PREL_LO19&#x9;somedata&#xA;&#xA;Disassembly of section .data:&#xA;&#xA;0000000000000000 &amp;lt;somedata&amp;gt;:&#xA;   0:&#x9;00000042 &#x9;.inst&#x9;0x00000042 ; undefined&#xA;   4:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;The &lt;code&gt;R_AARCH64_LD_PREL_LO19&lt;/code&gt; in that output is the relocation.  How did it land in there in between the instructions you ask?  Well, it didn&amp;rsquo;t.  The relocations are actually in a separate section of their own, as is evident with &lt;code&gt;objdump -r&lt;/code&gt;:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;RELOCATION RECORDS FOR [.text]:&#xA;OFFSET           TYPE              VALUE &#xA;0000000000000004 R_AARCH64_LD_PREL_LO19  somedata&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;even better with &lt;code&gt;readelf -r&lt;/code&gt; because it tells you that the relocations are essentially just a table of entries in the &lt;code&gt;.rela.text&lt;/code&gt; section:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Relocation section &#39;.rela.text&#39; at offset 0x110 contains 1 entry:&#xA;  Offset          Info           Type           Sym. Value    Sym. Name + Addend&#xA;000000000004  000500000111 R_AARCH64_LD_PREL 0000000000000000 somedata + 0&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;All sections with relocation entries have names with prefix &lt;code&gt;.rela&lt;/code&gt; or &lt;code&gt;.rel&lt;/code&gt; followed by the name of the section for which the relocations need to be applied.  Based on these section names, it&amp;rsquo;s evident that there are two types of relocation entries, REL and RELA.  There are a number of important pieces of information the assembler leaves for the linker here:&lt;/p&gt;&#xA;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The memory address it needs to fix up&lt;/li&gt;&#xA;&lt;li&gt;Which symbol the memory address is referring to&lt;/li&gt;&#xA;&lt;li&gt;What is the offset from the symbol that it should finally consider as the result&lt;/li&gt;&#xA;&lt;li&gt;How should it perform the computation&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;p&gt;All of this can be seen in the above relocation table.  Each entry in the relocation table is basically a C structure of the following form for REL type relocations:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;typedef struct {&#xA;        Elf64_Addr      r_offset;&#xA;        Elf64_Xword     r_info;&#xA;} Elf64_Rel;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;and for RELA:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;typedef struct {&#xA;        Elf64_Addr      r_offset;&#xA;        Elf64_Xword     r_info;&#xA;        Elf64_Sxword    r_addend;&#xA;} Elf64_Rela;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;&lt;code&gt;r_offset&lt;/code&gt; corresponds to the &lt;code&gt;Offset&lt;/code&gt; entry in the readelf output above and is typically the memory address that needs to be fixed up.  The offset from the symbol, aka the addendum is present only in RELA type relocations and it corresponds to the &lt;code&gt;r_addend&lt;/code&gt; element in the structure and the &lt;code&gt;Addend&lt;/code&gt; field in the readelf output.  The symbol and computation related information is encoded in the &lt;code&gt;r_info&lt;/code&gt; field.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The &lt;code&gt;Elf32_Rel&lt;/code&gt; and &lt;code&gt;Elf32_Rela&lt;/code&gt; structures are similar, except for the data sizes of the elements.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Symbol hunting&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;The &amp;lsquo;what to write&amp;rsquo; is where the &lt;code&gt;r_info&lt;/code&gt; field comes in.  That&amp;rsquo;s what the linker needs to figure out before the where and how, which comes later.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;This field is split into two 32-bit parts (16-bit for 32-bit architectures).  The lower part tells the linker how to perform the computation and the upper part tells the linker what the target symbol is.  The upper part is a symbol ID, which basically is an index into the symbol table.  In our example above, the &lt;code&gt;r_info&lt;/code&gt; is &lt;code&gt;0x000500000111&lt;/code&gt;, which means that the symbol id is 0x5.  We can pull out the symbol table using &lt;code&gt;readelf -s&lt;/code&gt;:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Symbol table &#39;.symtab&#39; contains 7 entries:&#xA;   Num:    Value          Size Type    Bind   Vis      Ndx Name&#xA;     0: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT  UND &#xA;     1: 0000000000000000     0 SECTION LOCAL  DEFAULT    1 &#xA;     2: 0000000000000000     0 SECTION LOCAL  DEFAULT    3 &#xA;     3: 0000000000000000     0 SECTION LOCAL  DEFAULT    4 &#xA;     4: 0000000000000000     0 NOTYPE  LOCAL  DEFAULT    1 $x&#xA;     5: 0000000000000000     0 NOTYPE  GLOBAL DEFAULT    3 somedata&#xA;     6: 0000000000000000     0 NOTYPE  GLOBAL DEFAULT    1 _start&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;and we find that the symbol id 0x5 is &lt;code&gt;somedata&lt;/code&gt; and has the &lt;code&gt;Value&lt;/code&gt; (i.e. the address of the symbol relative to its section) of 0x0.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Now we need to figure out how to combine all of this information together using the lower part of &lt;code&gt;r_info&lt;/code&gt;, which is 0x111.  This number corresponds to &lt;code&gt;R_AARCH64_PREL_LO19&lt;/code&gt;, which has a specific meaning.  An architecture defines a number of such relocations with descriptions of how they&amp;rsquo;re supposed to be applied.  In case of &lt;code&gt;R_AARCH64_PREL_LO19&lt;/code&gt; it means &amp;ldquo;Add the symbol address and addend and subtract from it, the location of the memory address that is being fixed up&amp;rdquo;.  Think about that a bit and you&amp;rsquo;ll notice that it is how you would compute the PC-relative offset of &lt;code&gt;somedata&lt;/code&gt; from the instruction, i.e. subtract the location of the fixup (i.e. the LDR instruction) from the address of &lt;code&gt;somedata&lt;/code&gt;.  In short form (and you&amp;rsquo;ll see this and similar notations to describe relocations), it is written as &lt;code&gt;S + A - P&lt;/code&gt;.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Putting things together&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;Now that we know what the target symbol is and how to compute the pc-relative offset, we need to compute the final symbol address, the final target address and then do the actual fixup.  This is done near the end of the linking process in the GNU linker, when all sections have been laid out and we finally are in a position to know the relative addresses.  The linker will then make the computation (i.e. S+A-P) and patch in the result into the LDR instruction before writing the output to the final binary.  Here is what our result looks like:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Disassembly of section .text:&#xA;&#xA;00000000004000b0 &amp;lt;_start&amp;gt;:&#xA;  4000b0:&#x9;d503201f &#x9;nop&#xA;  4000b4:&#x9;58080022 &#x9;ldr&#x9;x2, 4100b8 &amp;lt;somedata&amp;gt;&#xA;&#xA;Disassembly of section .data:&#xA;&#xA;00000000004100b8 &amp;lt;somedata&amp;gt;:&#xA;  4100b8:&#x9;00000042 &#x9;.inst&#x9;0x00000042 ; undefined&#xA;  4100bc:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;Notice that the opcode of the LDR instuction is now different (and as a result the instruction itself is also different) and includes the 0x8002, which is basically the encoded difference (0x10004) from &lt;code&gt;somedata&lt;/code&gt;.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Raising the stakes: Dynamic Relocations&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;This is all great when all our symbols are local and have predictable layouts such that PC-relative relocations such as &lt;code&gt;R_AARCH64_PREL_LO19&lt;/code&gt; are sufficient to describe and resolve in between assembling and linking a program.  What happens however, when these symbol references cross boundaries of sections in ways that we cannot predict at compile or link time?  What happens when your symbol references cross boundaries of your shared object?  These are problems that need to be solved to make Position Independent Code (PIC) possible.  PIC is when your program (and sections within your program) could get mapped anywhere in memory and you need your code to adapt to that fact.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Take this very simple example:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;.data&#xA;somedata:&#xA;        .8byte 0x42&#xA;somedata2:&#xA;&#x9;.8byte somedata&#xA;&#xA;.text&#xA;.globl _start&#xA;_start:&#xA;&#x9;ret&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;There&amp;rsquo;s next to nothing here; just &lt;code&gt;somedata&lt;/code&gt; like in our previous example and a &lt;code&gt;somedata2&lt;/code&gt; which points to &lt;code&gt;somedata&lt;/code&gt;.  However, in this next-to-nothing example lies an interesting complication that needs to be resolved at runtime: the value in &lt;code&gt;somedata2&lt;/code&gt; cannot be computed at compile time; it needs a fixup at runtime!  Let&amp;rsquo;s walk through the compilation to see how we get to the final result.  First, the disassembly to understand what the assembler did for us:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Disassembly of section .text:&#xA;&#xA;0000000000000000 &amp;lt;_start&amp;gt;:&#xA;   0:&#x9;d65f03c0 &#x9;ret&#xA;&#xA;Disassembly of section .data:&#xA;&#xA;0000000000000000 &amp;lt;somedata&amp;gt;:&#xA;   0:&#x9;00000042 &#x9;.inst&#x9;0x00000042 ; undefined&#xA;   4:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&#xA;0000000000000008 &amp;lt;somedata2&amp;gt;:&#xA;&#x9;...&#xA;&#x9;&#x9;&#x9;8: R_AARCH64_ABS64&#x9;.data&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;We see now that the address to be relocated is &lt;code&gt;somedata2&lt;/code&gt; in the &lt;code&gt;.data&lt;/code&gt; section and it is of type &lt;code&gt;R_AARCH64_ABS64&lt;/code&gt;.  This is simple relocation that instructs the linker to compute &lt;code&gt;S + A&lt;/code&gt; to get the result, i.e. get the symbol address of &lt;code&gt;somedata&lt;/code&gt; and add the addendum (0 again in this case).  This in fact would be the final result for a statically linked result (using &lt;code&gt;ld -static&lt;/code&gt;) and we&amp;rsquo;d lose the relocation in favour of the absolute address written into &lt;code&gt;somedata2&lt;/code&gt;:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Disassembly of section .text:&#xA;&#xA;00000000004000b0 &amp;lt;_start&amp;gt;:&#xA;  4000b0:&#x9;d65f03c0 &#x9;ret&#xA;&#xA;Disassembly of section .data:&#xA;&#xA;00000000004100b4 &amp;lt;somedata&amp;gt;:&#xA;  4100b4:&#x9;00000042 &#x9;.inst&#x9;0x00000042 ; undefined&#xA;  4100b8:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&#xA;00000000004100bc &amp;lt;somedata2&amp;gt;:&#xA;  4100bc:&#x9;004100b4 &#x9;.inst&#x9;0x004100b4 ; undefined&#xA;  4100c0:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;When compiling a shared object however (i.e. &lt;code&gt;ld -shared&lt;/code&gt;) we intend to produce a position independent DSO (dynamic shared object) and to achieve that the linker now emits a relocation to describe how to compute the final address to assign to &lt;code&gt;somedata2&lt;/code&gt; and where the memory address can be located.  In this example, it is the &lt;code&gt;R_AARCH64_RELATIVE&lt;/code&gt; dynamic relocation, as seen using &lt;code&gt;objdump -DR&lt;/code&gt; (output snipped to retain only useful bits):&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;Disassembly of section .data:&#xA;&#xA;0000000000011000 &amp;lt;somedata&amp;gt;:&#xA;   11000:&#x9;00000042 &#x9;.inst&#x9;0x00000042 ; undefined&#xA;   11004:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&#xA;0000000000011008 &amp;lt;somedata2&amp;gt;:&#xA;   11008:&#x9;00011000 &#x9;.inst&#x9;0x00011000 ; undefined&#xA;&#x9;&#x9;&#x9;11008: R_AARCH64_RELATIVE&#x9;*ABS*+0x11000&#xA;   1100c:&#x9;00000000 &#x9;.inst&#x9;0x00000000 ; undefined&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;This relocation is interesting not just for the reason that it is dynamic, but also because it is a &lt;code&gt;S+A&lt;/code&gt; type relocation that puts the non-relocated address (i.e. the link time address) of &lt;code&gt;somedata&lt;/code&gt; into its addend.  This relocation also does not reference a symbol; instead it references an &lt;code&gt;*ABS*&lt;/code&gt; value, which is basically the offset at which this DSO would be loaded during execution.  It is the dynamic linker in the C runtime library (ld.so in GNU systems) that reads these relocations from the &lt;code&gt;.rela.dyn&lt;/code&gt; section.  Because this relocation is based on an absolute address computed by the static linker, the dynamic linker does not have to do a symbol lookup to resolve the relocation.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The other difference from static relocations is that when a dynamic relocation references a symbol in its &lt;code&gt;r_info&lt;/code&gt;, it is looked up in the &lt;code&gt;.dynsym&lt;/code&gt; section, i.e. in the dynamic symbol table and not in the regular symbol table.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Final Thoughts&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;There are a number of other cases that the linker needs to cater for when it comes to relocations such as entries in Global Offset Tables, resolution of intermediate functions and Thread-Local Storage.  Thankfully though, the first principles behind all those relocations are the same as the above and you can apply this knowledge to GOT, TLS and IFUNC relocations as well.  GOT relocations for example reference GOT base, which the linker knows where to find (since it sets up the &lt;code&gt;.got&lt;/code&gt; section in the first place) and can use that information to compute the location to fix up.  Other than this special knowledge, everything else remains pretty much the same.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Once you&amp;rsquo;re equipped with these first principles, the next task is to figure out where documentation for specific relocations is for every architecture.  While the &lt;a href=&#34;https://sourceware.org/binutils/docs/as/AArch64_002dRelocations.html&#34;&gt;binutils documentation&lt;/a&gt; makes some effort to document the public facing part of relocations, the detailed documentation is usually distributed by the architecture chip vendors.  The &lt;a href=&#34;https://developer.arm.com/docs/ihi0056/g/elf-for-the-arm-64-bit-architecture-aarch64-abi-2019q4-documentation&#34;&gt;AArch64 ELF documentation&lt;/a&gt; for example is hosted on the Arm website.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Links&lt;/h2&gt;&#xA;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://twitter.com/gnutools&#34;&gt;@gnutools&lt;/a&gt; pointed me to a full blog series by Ian Lance Taylor on Linkers, &lt;a href=&#34;https://lwn.net/Articles/276782/&#34;&gt;indexed by LWN&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://twitter.com/matt_dz&#34;&gt;@matt_dz&lt;/a&gt; maintains a comprehensive set of links on &lt;a href=&#34;https://github.com/MattPD/cpplinks/blob/master/executables.linking_loading.md#readings-relocations&#34;&gt;linking and loading&lt;/a&gt; on GitHub.  This looks like everything you&amp;rsquo;ll ever need to understand ELF and linkers!&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>10 Apr 20 19:36 UTC</pubDate>
    </item>
    <item>
      <title>gcc under the hood</title>
      <link>https://gotplt.org/posts/gcc-under-the-hood.html</link>
      <description>&lt;!--&#xA;.. title: gcc under the hood&#xA;.. slug: gcc-under-the-hood&#xA;.. date: 2019-10-03T14:12:24-07:00&#xA;.. tags: gcc, compiler, internals, cpu, mtune&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;My background in computers is a bit hacky for a compiler engineer.  I never had the theoretical computer science base that the average compiler geek does (yes I have a Masters in Computer Applications, but it&amp;rsquo;s from India and yes I&amp;rsquo;m going to leave that hanging without explanation) and I have pretty much winged it all these years.  What follows is a bunch of thoughts (high five to those who get that reference!) from my winging it for almost a decade with the GNU toolchain.  If you&amp;rsquo;re a visual learner then I would recommend watching my &lt;a href=&#34;https://www.youtube.com/watch?v=brxAm99w8D8&#34;&gt;talk video&lt;/a&gt; at Linaro Connect 2019 instead of reading this.  This is an imprecise transcript of my talk, with less silly quips and in a more neutral accent.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Hell Oh World!&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;It all started as a &lt;a href=&#34;https://siddhesh.in/posts/science-hack-day-belgaum.html&#34;&gt;lightning talk at SHD Belgaum&lt;/a&gt; where I did a 5 minute demonstration of how a program goes from source code to a binary program.  I got many people asking me to talk about this in more detail and it eventually became a &lt;a href=&#34;https://siddhesh.in/posts/hello-fossasia-revisiting-the-event-and-the-first-program-we-write-in-c.html&#34;&gt;full hour workshop at FOSSASIA&lt;/a&gt; in 2017.  Those remained the basis for the first part of my talk at Connect, in fact they&amp;rsquo;re a bit more detailed in their treatment of the subject of purely taking code from source to binary.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Go read those posts, I&amp;rsquo;ll wait.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Under the hood&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;Welcome back! So now you know almost exactly how code goes from source to binary and how it gets executed on your computer.  Congratulations, you now have a headstart into hacking on the dynamic loader bits in glibc!  Now that we&amp;rsquo;ve explored the woods around the swamp, lets step &lt;em&gt;into&lt;/em&gt; the swamp.  Lets take a closer look at what gcc does with our source code.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Babelfish&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The first thing a compiler does is to read the source code and make it into a data structure form that is easy for the computer to manipulate.  Since the computer deals best with binary data, it converts the text form of the source code language into a tree-like structure.  There is plenty of literature available on how that is implemented; in fact most compiler texts end up putting too much focus on this aspect of the compiler and then end up rushing through the real fun stuff that the compiler does with the program you wrote.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The data structure that the compiler translates your source code into is called an &lt;a href=&#34;https://en.wikipedia.org/wiki/Intermediate_representation&#34;&gt;Intermediate Representation (IR)&lt;/a&gt;, and gcc has three of them!&lt;/p&gt;&#xA;&#xA;&lt;p&gt;This part of the compiler that parses the source code into IR is known as the front end.  gcc has many such front ends, one for each language it implements; they are all cluttered into the &lt;code&gt;gcc/&lt;/code&gt; directory.  The C family of languages has a common subset of code in the &lt;code&gt;gcc/c-family&lt;/code&gt; directory and then there are specializations like the C++ front end, implemented in the &lt;code&gt;gcc/cp&lt;/code&gt; directory.  There is a directory for fortran, another for java, yet another for go and so on.  The translation of code from source to IR varies from frontend to frontend because of language differences, but they all have one thing in common; their output is an IR called &lt;code&gt;GENERIC&lt;/code&gt; and it attempts to be language-independent.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Optimisation passes&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;Once gcc translates the source code into its IR form, it runs the IR through a number of passes (about 200 of them, more or less depending on the &lt;code&gt;-O&lt;/code&gt; flag you pass) to try and come up with the most optimal machine code representation of the program.  gcc builds source code a file at a time, which is called a Translation Unit (TU).  A lot of its optimisation passes operate at the function level, i.e. individual functions are seen as independent units and are optimised separately.  Then there are Inter-procedural analysis (IPA) passes that look at interactions of these functions and finally there is Link Time Optimisation(LTO) that attempts to analyse source code across translation units to potentially get even better results.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Optimisation passes can be architecture-independent or architecture-dependent.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Architecture independent passes do not care too much about the underlying machine beyond some basic details like word size, whether the CPU has floating point support, vector support, etc.  These passes have some configuration hooks that allow their behaviour to be modified according to the target CPU architecture but the high level behaviour is architecture-agnostic.  Architecture-independent passes are the holy grail for optimisation because they usually don&amp;rsquo;t get old; optimisations that work today will continue to work regardless of CPU architecture evolution.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Architecture-dependent passes need more information about the architecture and as a result their results may change as architectures evolve.  Register allocation and instruction scheduling for example are very interesting and complex problems that architecture-specific passes handle.  The instruction scheduling problem was a lot more critical back in the day when CPUs could only execute code sequentially.  With out-of-order execution, the scheduling problem has not become less critical, but it has definitely changed in nature.  Similarly the register allocation problem can be complicated by various factors such as number of locigal registers, how they share physical register files, how costly moving between different types of registers is, and so on.  Architecture-dependent passes have to take into consideration all of these factors.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The final pass in architecture-dependent passes does the work of emitting assembly code for the target CPU.  The architecture-independent passes constitute what is known as the middle-end of the compiler and the architecture-dependent passes form the compiler backend.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Each pass is a complex work of art, mathematics and logic and may have one or more highly cited research papers as their basis.  No single gcc engineer would claim to understand all of these passes; there are many who have spent most of their careers on a small subset of passes, such is their complexity.  But then, this is also a testament to how well we can work together as humans to create something that is significantly more complex than what our individual minds can grasp.  What I mean to say with all this is, it&amp;rsquo;s OK to not know all of the passes, let alone know them well; you&amp;rsquo;re definitely not the only one.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Optimisation passes are all listed in &lt;code&gt;gcc/passes.def&lt;/code&gt; and that is the sequence in which they are executed.  A pass is defined as a class with a gate and execute function that determine whether to run and what to do respectively.  Here&amp;rsquo;s what a single pass namespace looks like:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;/* Pass data to help the pass manager classify, prepare and cleanup.  */&#xA;const pass_data pass_data_my_pass =&#xA;{&#xA;  GIMPLE_PASS, /* type */&#xA;  &amp;quot;my_pass&amp;quot;, /* name */&#xA;  OPTGROUP_LOOP, /* optinfo_flags */&#xA;  TV_TREE_LOOP, /* tv_id */&#xA;  PROP_cfg, /* properties_required */&#xA;  0, /* properties_provided */&#xA;  0, /* properties_destroyed */&#xA;  0, /* todo_flags_start */&#xA;  0, /* todo_flags_finish */&#xA;};&#xA;&#xA;/* This is a GIMPLE pass.  I know you don&#39;t know what GIMPLE is yet ;) */&#xA;class pass_my_pass : public gimple_opt_pass&#xA;{&#xA;public:&#xA;  pass_my_pass (gcc::context *ctxt)&#xA;    : gimple_opt_pass (pass_data_my_pass, ctxt)&#xA;  {}&#xA;&#xA;  /* opt_pass methods: */&#xA;  virtual bool gate (function *) { /* Decide whether to run.  */ }&#xA;&#xA;  virtual unsigned int execute (function *fn);&#xA;};&#xA;&#xA;unsigned int&#xA;pass_my_pass::execute (function *)&#xA;{&#xA;  /* Magic!  */&#xA;}&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;We will not go into the anatomy of a pass yet.  That is perhaps a topic for a follow-up post.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;GENERIC&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;GENERIC is the first IR that gets generated from parsing the source code.  It is a tree structure that has been present since the earliest gcc versions.  Its core data structure is a &lt;code&gt;tree_node&lt;/code&gt;, which is a hierarchy of structs, with &lt;code&gt;tree_base&lt;/code&gt; as the Abraham.  OK, if you haven&amp;rsquo;t been following gcc development, this can come as a surprise: a significant portion of gcc is now in c++!&lt;/p&gt;&#xA;&#xA;&lt;p&gt;It&amp;rsquo;s OK, take a minute to mourn/celebrate that.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The &lt;code&gt;tree_node&lt;/code&gt; can mean a lot of things (it is a &lt;code&gt;union&lt;/code&gt;), but the most important one is the &lt;code&gt;tree_typed&lt;/code&gt; node.  There are various types of &lt;code&gt;tree_typed&lt;/code&gt;, like &lt;code&gt;tree_int_cst&lt;/code&gt;, &lt;code&gt;tree_identifier&lt;/code&gt;, &lt;code&gt;tree_string&lt;/code&gt;, etc. that describe the elements of source code that they house.  The base struct, i.e. &lt;code&gt;tree_typed&lt;/code&gt; and even further up, &lt;code&gt;tree_base&lt;/code&gt; have flags that have metadata about the elements that help in optimisation decisions.  You&amp;rsquo;ll almost never use generic in the context of code traversal in optimisation passes, but the nodes are still very important to know about because GIMPLE continues to use them to store operand information.  Take a peek at &lt;code&gt;gcc/tree-core.h&lt;/code&gt; and &lt;code&gt;gcc/tree.def&lt;/code&gt; to take a closer look at all of the types of nodes you could have in GENERIC.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;What is &lt;code&gt;GIMPLE&lt;/code&gt; you ask?  Well, that&amp;rsquo;s our next IR and probably the most important one form the context of optimisations.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;OK so now you have enough background to go look at the guts of the GENERIC tree IR in gcc.  Here&amp;rsquo;s the &lt;a href=&#34;https://gcc.gnu.org/onlinedocs/gccint/GENERIC.html&#34;&gt;gcc internals documentation&lt;/a&gt; that will help you navigate all of the convenience macros to analyze th tree nodes.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;GIMPLE&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The GIMPLE IR is a workhorse of gcc.  Passes that operate on GIMPLE are architecture-independent.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The core structure in GIMPLE is a &lt;code&gt;struct gimple&lt;/code&gt; that holds all of the metadata for a single gimple statement and is also a node in the list of gimple statements.  There are various structures named &lt;code&gt;gimple_statement_with_ops_*&lt;/code&gt; that have the actual operand information based on its type.  Once again like with GENERIC, it is a hierarchy of structs.  Note that the operands are all of the &lt;code&gt;tree&lt;/code&gt; type so we haven&amp;rsquo;t got rid of all of GENERIC.  &lt;code&gt;gcc/gimple.h&lt;/code&gt; is where all of these structures are defined and &lt;code&gt;gcc/gimple.def&lt;/code&gt; is where all of the different types of gimple statements are defined.&lt;/p&gt;&#xA;&#xA;&lt;h4&gt;Where did my control flow go?&lt;/h4&gt;&#xA;&#xA;&lt;p&gt;But how is it that a simple list of gimple statements is sufficient to traverse a program that has conditions and loops?  Here is where things get interesting.  Every sequence of GIMPLE statements is housed in a Basic Block (BB) and the compiler, when translating from GENERIC to GIMPLE (also known as lowering, since GIMPLE is structurally simpler than GENERIC), generates a Control FLow Graph (CFG) that describes the flow of the function from one BB to another.  The only control flow idea one then needs in GIMPLE to traverse code from within the gimple statement context is a jump from one block to another and that is fulfilled by the GIMPLE_GOTO statement type.  The CFG with its basic blocks and edges connecting those blocks, takes care of everything else.  Routines to generate and manipulate the CFG are defined in &lt;code&gt;gcc/cfg.h&lt;/code&gt; and &lt;code&gt;gcc/cfg.c&lt;/code&gt; but beware, you don&amp;rsquo;t modify the CFG directly.  Since CFG is tightly linked with GIMPLE (and RTL, yes that&amp;rsquo;s our third and final IR), it provides hooks to manipulate the graph and update GIMPLE if necessary.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The last interesting detail about CFG is that it has a special construct for loops, because they&amp;rsquo;re typically the most interesting subjects for optimisation: you can splice them, unroll them, distribute them and more to produce some fantastic performance results.  &lt;code&gt;gcc/cfgloop.h&lt;/code&gt; is there you&amp;rsquo;ll find all of the routines you need to traverse and manipulate loops.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The final important detail with regard to GIMPLE is the Single Static Assignment (SSA) form.  Typical source code would have variables that get declared, assigned to, manipulated and then stored back into memory.  Essentially, it is multiple operations on a single variable, sometimes destroying its previous contents as we reuse variables for different things that are logically related in the context of the high level program.  This reuse however makes it complicated for a compiler to understand what&amp;rsquo;s going on and hence it ends up missing a host of optimisation opportunities.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;To make things easier for optimisation passes, these variables are broken up into versions of themselves that have a specific start and end point in their lifecycle.  This is called the Single Static Assignment form of a variable, where each version of the variable as a single starting point, viz. its definition.  So if you have code like this:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    x = 10;&#xA;    x += 20;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;it becomes:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    x_1 = 10;&#xA;    x_2 = x_1 + 20;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;where &lt;code&gt;x_1&lt;/code&gt; and &lt;code&gt;x_2&lt;/code&gt; are versions of &lt;code&gt;x&lt;/code&gt;.  If you have versions of variables in conditions, things get interesting and the compiler deals with it with a mysterious concept called PHI nodes.  So this code:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    if (n &amp;gt; 10)&#xA;      x = 10;&#xA;    else&#xA;      x = 20;&#xA;    return x;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;becomes:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    if (n &amp;gt; 10)&#xA;      x_1 = 10;&#xA;    else&#xA;      x_2 = 20;&#xA;    # x_3 = PHI&amp;lt;x_1, x_2&amp;gt;;&#xA;    return x_3;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;So the PHI node is a conditional selector of the earlier two versions of the variable and depending on the results of the optimisation passes, you could eliminate versions of the variables altogether or use CPU registers more efficiently to store those variable versions.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There you go, now you have everything you need to get started on hacking GIMPLE.  I know this part is a bit heavy but guess what, this is where you can seriously start thinking about hacking on gcc!  When you jump in, you&amp;rsquo;ll need the more detailed information in the gcc internals manual on &lt;a href=&#34;https://gcc.gnu.org/onlinedocs/gccint/GIMPLE.html&#34;&gt;GIMPLE&lt;/a&gt;, &lt;a href=&#34;https://gcc.gnu.org/onlinedocs/gccint/Control-Flow.html#Control-Flow&#34;&gt;CFG&lt;/a&gt; and &lt;a href=&#34;https://gcc.gnu.org/onlinedocs/gccint/Tree-SSA.html#Tree-SSA&#34;&gt;GIMPLE optimisations&lt;/a&gt;.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;RTL&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;We are yet another step closer to generating assembly code for our assembler and linker to build into the final program.  Recall that GIMPLE is largely architecture independent, so it works on high level ideas of statements, expressions and types and their relationships.  RTL is much more primitive in comparison and is designed to mimic sequential architecture instructions.  Its main purpose is to do architecture-specific work, such as register allocation and scheduling instructions in addition to more optimisation (because you can never get enough of that!) passes that make use of architecture information.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Internally, you will encounter two forms of RTL, one being the &lt;code&gt;rtx&lt;/code&gt; struct that is used for most transformations and passes in the compiler.  The other form is in text to map machine instructions to RTL and these are Lisp-like S expressions.  These are used to make machine descriptions, where you can specify machine instructions for all of the common operations such as add, sub, multiply and so on.  For example, this is what the description of a jump looks like for aarch64:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;(define_insn &amp;quot;jump&amp;quot;&#xA;  [(set (pc) (label_ref (match_operand 0 &amp;quot;&amp;quot; &amp;quot;&amp;quot;)))]&#xA;  &amp;quot;&amp;quot;&#xA;  &amp;quot;b\\t%l0&amp;quot;&#xA;  [(set_attr &amp;quot;type&amp;quot; &amp;quot;branch&amp;quot;)]&#xA;)&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;Machine description files are in the &lt;code&gt;gcc/config&lt;/code&gt; directory in their respective architecture subdirectory and have the &lt;code&gt;.md&lt;/code&gt; suffix.  The main aarch64 machine description file for example is &lt;code&gt;gcc/config/aarch64/aarch64.md&lt;/code&gt;.  These files are parsed by a tool in gcc to generate C code with &lt;code&gt;rtx&lt;/code&gt; structures for each of those S-expressions.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;In general, the &lt;code&gt;gcc/config&lt;/code&gt; directory contains source code that handles compilation for that specific architecture.  In addition to the machine descriptions these directories also have C code that enhance the RTL passes to exploit as much of th architecture information as they possibly can.  This is where some of the detailed architecture-specific analysis of the RTL instructions go.  For example, combination of loads into load pairs for aarch64 is an important task and it is done with a combination of machine description and some code to peek into and rearrange neighbouring RTL instructions&lt;/p&gt;&#xA;&#xA;&lt;p&gt;But then you&amp;rsquo;re wondering, why are there multiple description files?  Other than just cleaner layout (put constraint information in a separate file, type information in another, etc.) it is because there are multiple evolutions of an architecture.  The &amp;lsquo;i386&amp;rsquo; architecture is a mess of evolutions that span word sizes and capabilities.  The aarch64 architecture has within it many microarchitectures developed by various Arm licensee vendors like &lt;code&gt;xgene&lt;/code&gt;, &lt;code&gt;thunderxt88&lt;/code&gt;, &lt;code&gt;falkor&lt;/code&gt; and also those developed by Arm such as the &lt;code&gt;cortex-a57&lt;/code&gt;, &lt;code&gt;cortex-a72&lt;/code&gt;, &lt;code&gt;ares&lt;/code&gt;, etc.  All of these have different behaviours and performance characteristics despite sharing an instruction set.  For example on some microarchitecture one may prefer to emit the &lt;code&gt;csel&lt;/code&gt; instruction instead of &lt;code&gt;cmp&lt;/code&gt;, &lt;code&gt;b.cond&lt;/code&gt; and multiple &lt;code&gt;mov&lt;/code&gt; instructions to reduce code size (and hence improve performance) but on some other architecture, the &lt;code&gt;csel&lt;/code&gt; instruction may have been designed really badly and hence it may well be cheaper to execute the 4+ instructions instead of the one &lt;code&gt;csel&lt;/code&gt;.  These are behaviour quirks that you select when you use the &lt;code&gt;-mtune&lt;/code&gt; flag to optimise for a specific CPU.  A lot of this information is also available in the C code of the architecture in the form of structures called cost tables.  These are relative costs of various operations that help the RTL passes and some GIMPLE passes determine the best target code and optimisation behaviour accordingly for the CPU.  Here&amp;rsquo;s an example of register move costs for the Qualcomm Centriq processor:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;static const struct cpu_regmove_cost qdf24xx_regmove_cost =&#xA;{&#xA;  2, /* GP2GP  */&#xA;  /* Avoid the use of int&amp;lt;-&amp;gt;fp moves for spilling.  */&#xA;  6, /* GP2FP  */&#xA;  6, /* FP2GP  */&#xA;  4 /* FP2FP  */&#xA;};&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;This tells us that in general, moving between general purpose registers is cheaper, moving between floating point registers is slightly more expensive and moving between general purpose and floating point registers is most expensive.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The other important detail in such machine description files is the pipeline description.  This is a description of how the pipeline for a specific CPU microarchitecture is designed along with latencies for instructions in that pipeline.  This information is used by the instruction scheduler pass to determine the best schedule of instructions for a CPU.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Now where do I start?&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;This is a lot of information to page in at once and if you&amp;rsquo;re like me, you&amp;rsquo;d want something more concrete to get started with understanding what gcc is doing.  gcc has you covered there because it has flags that allow you to study the IR generated for the compiler at &lt;em&gt;every&lt;/em&gt; stage of compilation.  By every stage, I mean every pass!  The &lt;code&gt;-fdump-*&lt;/code&gt; set of flags allow you to dump the IR to study what gcc did in each pass.  Particularly, &lt;code&gt;-fdump-tree-*&lt;/code&gt; flags dump GIMPLE IR (in a C-like format so that it is not too complicated to read) into a file for each pass and the &lt;code&gt;-fdump-rtl-*&lt;/code&gt; flags do the same for RTL. The GIMPLE IR can be dumped in its raw form as well (e.g. &lt;code&gt;-fdump-tree-all-raw&lt;/code&gt;), which makes it much simpler to correlate with the code that is manipulating the GIMPLE in an optimisation pass.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The easiest way to get into gcc development (and compiler development in general) in my experience is the back end.  Try tweaking the various cost tables to see what effect it has on code generation.  Modify the instructions generated by the RTL descriptions and use that to look closer at one pass that interests you.  Once you&amp;rsquo;re comfortable with making changes to gcc, rebuilding and checking its outputs, you can then try writing a pass, which is a slightly more involved process.  Maybe I&amp;rsquo;ll write about it some day.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;References&lt;/h2&gt;&#xA;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The &lt;a href=&#34;https://gcc.gnu.org/onlinedocs/gccint/index.html#Top&#34;&gt;GCC internals manual&lt;/a&gt; is the canonical place to read (and fix up) documentation for GCC internals.  It is wonderfully detailed and hopelessly outdated at the same time.  Bringing it up to date is a task by itself and the project continues to look for volunteers to do that.&lt;/li&gt;&#xA;&lt;li&gt;David Malcolm has a more functional &lt;a href=&#34;https://dmalcolm.fedorapeople.org/gcc/newbies-guide/&#34;&gt;newbie guide&lt;/a&gt; for fledgling gcc hackers who might struggle with debugging gcc and getting involved in the gcc development process&lt;/li&gt;&#xA;&lt;li&gt;The &lt;a href=&#34;https://www.cse.iitb.ac.in/grc/&#34;&gt;GCC Resource Center&lt;/a&gt; workshop on GCC is where I cut my teeth on gcc internals.  I don&amp;rsquo;t think their workshop is active anymore but they have presentations and other literature there that is still very relevant.&lt;/li&gt;&#xA;&lt;li&gt;I wrote about &lt;a href=&#34;https://siddhesh.in/posts/optimizing-toolchains-for-modern-microprocessors.html&#34;&gt;micro-optimisations&lt;/a&gt; in the past and those ideas are great to try on gcc.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>03 Oct 19 21:12 UTC</pubDate>
    </item>
    <item>
      <title>Wrestling with the register allocator: LuaJIT edition</title>
      <link>https://gotplt.org/posts/wrestling-with-the-register-allocator-luajit-edition.html</link>
      <description>&lt;!--&#xA;.. title: Wrestling with the register allocator: LuaJIT edition&#xA;.. slug: wrestling-with-the-register-allocator-luajit-edition&#xA;.. date: 2019-09-15T17:20:52-07:00&#xA;.. tags: luajit, registers, regalloc, openresty&#xA;.. link:&#xA;.. description: I fixed another bug in luajit and lived to write about it.&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;For some time now, I kept running into one specific piece of code in luajit repeatedly for various reasons and last month I came across a fascinating bug in the register allocator pitfall that I had never encountered before. But then as is the norm after fixing such bugs, I concluded that it&amp;rsquo;s too trivial to write about and I was just stupid to not have found it sooner; all bugs are trivial once they&amp;rsquo;re fixed.&lt;/p&gt;&#xA;&#xA;&lt;div align=&#34;center&#34;&gt;&#xA;&lt;blockquote class=&#34;twitter-tweet&#34;&gt;&lt;p lang=&#34;en&#34; dir=&#34;ltr&#34;&gt;The worst thing about solving tricky problems is deciding to write up about it and then halfway through the writeup, concluding that it was way too simple and I was just stupid to not solve it sooner.&lt;/p&gt;&amp;mdash; Siddhesh Poyarekar (@siddhesh_p) &lt;a href=&#34;https://twitter.com/siddhesh_p/status/1166737530537988096?ref_src=twsrc%5Etfw&#34;&gt;August 28, 2019&lt;/a&gt;&lt;/blockquote&gt; &lt;script async src=&#34;https://platform.twitter.com/widgets.js&#34; charset=&#34;utf-8&#34;&gt;&lt;/script&gt; &#xA;&lt;/div&gt;&#xA;&#xA;&lt;p&gt;After getting over my imposter syndrome (yeah I know, took a while) I finally have the courage to write about it, so like a famous cook+musician says, enough jibberjabber&amp;hellip;&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Looping through a hash table&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;One of the key data structures that luajit uses to implement metatables is a hash table.  Due to its extensive use, the lookup loop into such hash tables is optimised in the JIT using an architecture-specific &lt;code&gt;asm_href&lt;/code&gt; function.  Here is what a snippet of the arm64 version of the function looks like.  A key thing to note is that the assembly code is generated backwards, i.e. the last instruction is emitted first.&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;  /* Key not found in chain: jump to exit (if merged) or load niltv. */&#xA;  l_end = emit_label(as);&#xA;  as-&amp;gt;invmcp = NULL;&#xA;  if (merge == IR_NE)&#xA;    asm_guardcc(as, CC_AL);&#xA;  else if (destused)&#xA;    emit_loada(as, dest, niltvg(J2G(as-&amp;gt;J)));&#xA;&#xA;  /* Follow hash chain until the end. */&#xA;  l_loop = --as-&amp;gt;mcp;&#xA;  emit_n(as, A64I_CMPx^A64I_K12^0, dest);&#xA;  emit_lso(as, A64I_LDRx, dest, dest, offsetof(Node, next));&#xA;  l_next = emit_label(as);&#xA;&#xA;  /* Type and value comparison. */&#xA;  if (merge == IR_EQ)&#xA;    asm_guardcc(as, CC_EQ);&#xA;  else&#xA;    emit_cond_branch(as, CC_EQ, l_end);&#xA;&#xA;  if (irt_isnum(kt)) {&#xA;    if (isk) {&#xA;      /* Assumes -0.0 is already canonicalized to +0.0. */&#xA;      if (k)&#xA;        emit_n(as, A64I_CMPx^k, tmp);&#xA;      else&#xA;        emit_nm(as, A64I_CMPx, key, tmp);&#xA;      emit_lso(as, A64I_LDRx, tmp, dest, offsetof(Node, key.u64));&#xA;    } else {&#xA;      Reg ftmp = ra_scratch(as, rset_exclude(RSET_FPR, key));&#xA;      emit_nm(as, A64I_FCMPd, key, ftmp);&#xA;      emit_dn(as, A64I_FMOV_D_R, (ftmp &amp;amp; 31), (tmp &amp;amp; 31));&#xA;      emit_cond_branch(as, CC_LO, l_next);&#xA;      emit_nm(as, A64I_CMPx | A64F_SH(A64SH_LSR, 32), tisnum, tmp);&#xA;      emit_lso(as, A64I_LDRx, tmp, dest, offsetof(Node, key.n));&#xA;    }&#xA;  } else if (irt_isaddr(kt)) {&#xA;    Reg scr;&#xA;    if (isk) {&#xA;      int64_t kk = ((int64_t)irt_toitype(irkey-&amp;gt;t) &amp;lt;&amp;lt; 47) | irkey[1].tv.u64;&#xA;      scr = ra_allock(as, kk, allow);&#xA;      emit_nm(as, A64I_CMPx, scr, tmp);&#xA;      emit_lso(as, A64I_LDRx, tmp, dest, offsetof(Node, key.u64));&#xA;    } else {&#xA;      scr = ra_scratch(as, allow);&#xA;      emit_nm(as, A64I_CMPx, tmp, scr);&#xA;      emit_lso(as, A64I_LDRx, scr, dest, offsetof(Node, key.u64));&#xA;    }&#xA;    rset_clear(allow, scr);&#xA;  } else {&#xA;    Reg type, scr;&#xA;    lua_assert(irt_ispri(kt) &amp;amp;&amp;amp; !irt_isnil(kt));&#xA;    type = ra_allock(as, ~((int64_t)~irt_toitype(ir-&amp;gt;t) &amp;lt;&amp;lt; 47), allow);&#xA;    scr = ra_scratch(as, rset_clear(allow, type));&#xA;    rset_clear(allow, scr);&#xA;    emit_nm(as, A64I_CMPw, scr, type);&#xA;    emit_lso(as, A64I_LDRx, scr, dest, offsetof(Node, key));&#xA;  }&#xA;&#xA;  *l_loop = A64I_BCC | A64F_S19(as-&amp;gt;mcp - l_loop) | CC_NE;&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;Here, the &lt;code&gt;emit_*&lt;/code&gt; functions emit assembly instructions and the &lt;code&gt;ra_*&lt;/code&gt; functions allocate registers.  In the normal case everything is fine and the table lookup code is concise and effective.  When there is register pressure however, things get interesting.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;As an example, here is what a typical type lookup would look like:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;0x100&#x9;ldr x1, [x16, #52]&#xA;0x104&#x9;cmp x1, x2&#xA;0x108&#x9;beq -&amp;gt; exit&#xA;0x10c&#x9;ldr x16, [x16, #16]&#xA;0x110&#x9;cmp x16, #0&#xA;0x114&#x9;bne 0x100&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;Here, &lt;code&gt;x16&lt;/code&gt; is the table that the loop traverses.  &lt;code&gt;x1&lt;/code&gt; is a key, which if it matches, results in an exit to the interpreter.  Otherwise the loop moves ahead until the end of the table.  The comparison is done with a constant stored in &lt;code&gt;x2&lt;/code&gt;.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The value of &lt;code&gt;x&lt;/code&gt; is loaded later (i.e. earlier in the code, we are emitting code backwards, remember?) whenever that register is needed for reuse through a process called restoration or spilling.  In the restore case, it is loaded into register as a constant or expressed as another constant (look up constant rematerialisation) and in the case of a spill, the register is restored from a slot in the stack.  If there is no register pressure, all of this restoration happens at the head of the trace, which is why if you study a typical trace you will notice a lot of constant loads at the top of the trace.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Like the Spill that ruined your keyboard the other day&amp;hellip;&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;Things get interesting when the allocation of &lt;code&gt;x2&lt;/code&gt; in the loop above results in a restore.  Looking at the code a bit closer:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    type = ra_allock(as, ~((int64_t)~irt_toitype(ir-&amp;gt;t) &amp;lt;&amp;lt; 47), allow);&#xA;    scr = ra_scratch(as, rset_clear(allow, type));&#xA;    rset_clear(allow, scr);&#xA;    emit_nm(as, A64I_CMPw, scr, type);&#xA;    emit_lso(as, A64I_LDRx, scr, dest, offsetof(Node, key));&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;The &lt;code&gt;x2&lt;/code&gt; here is type, which is a constant.  If a register is not available, we have to make one available by either &lt;a href=&#34;https://en.wikipedia.org/wiki/Rematerialization&#34;&gt;rematerializing&lt;/a&gt; or by restoring the register, which would result in something like this:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;0x100   ldr x1, [x16, #52]&#xA;0x104   cmp x1, x2&#xA;0x108&#x9;mov x2, #42&#xA;0x10c   beq -&amp;gt; exit&#xA;0x110   ldr x16, [x16, #16]&#xA;0x114   cmp x16, #0&#xA;0x118   bne 0x100&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;This ends up breaking the loop because the allocator restore/spill logic assumes that the code is linear and the restore will affect only code that follows it, i.e. code that got generated earlier.  To fix this, all of the register allocations should be done before the loop code is generated.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Making things right&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;The result of this analysis was this fix in my &lt;a href=&#34;https://github.com/siddhesh/LuaJIT/commit/d8a7769ef37435f3c81c19779395eb6adb95037a&#34;&gt;LuaJIT fork&lt;/a&gt; that allocates registers for operands that will be used in the loop before generating the body of the loop.  That is, if the registers have to spill, they will do so after the loop (we are generating code in reverse order) and leave the loop compact.  The fix is also in the luajit2 repository in the OpenResty project.  The work was sponsored by OpenResty as this wonderfully vague bug could only be produced by some very complex scripts that are part of the OpenResty product.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>16 Sep 19 00:20 UTC</pubDate>
    </item>
    <item>
      <title>A JIT in Time...</title>
      <link>https://gotplt.org/posts/a-jit-in-time.html</link>
      <description>&lt;!--&#xA;.. title: A JIT in Time...&#xA;.. slug: a-jit-in-time&#xA;.. date: 2019-03-28T13:29:51-07:00&#xA;.. tags: luajit, lua, jit, compiler, toolchain&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;It&amp;rsquo;s been a different 3 months.  For over 6 years I had been working almost exclusively on the GNU toolchain with a focus on glibc and I now had the chance of working on a completely different set of projects, something I had done a lot of during my Red Hat technical support days but not since.  I was to look into Pypy, OpenJDK and LuaJIT, three very different projects with very different development styles, communities and technologies.  The comparison of these projects among themselves and the GNU projects is an interesting point but not the purpose of this post, maybe some other day.  In this post I want to talk about the project I spent the most time on (~1.5 months) and found to be technically the most intriguing: LuaJIT.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;A Just In Time Introduction&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;For those new to the concept, JIT compilation techniques are pretty old and there is a very interesting paper called the &lt;a href=&#34;https://dl.acm.org/citation.cfm?id=857077&#34;&gt;A brief history of just in time&lt;/a&gt; that does what the title states.  The basic concept is quite straightforward - code written in a high level language (in the case of luajit, lua) is interpreted as usual while keeping track of which parts of the code get hit often.  If a part of the code is seen to be executed repeatedly, all or part of that code is compiled into binary and mapped in, with entry and exit branches into the interpreter, also known as exit guards. There are a number of tradeoffs in designing a JIT and the paper I&amp;rsquo;ve linked above gives enough of an introduction to appreciate the complexity of the problem being solved.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The key difference from compilers is that the time required to compile is often as much a performance factor as the quality of the generated code.  Due to this, one needs to be careful about the amount of processing one can do on the code to optimise it.  So while gcc or llvm may end up giving higher quality code, the ~200 passes that are involved in building a TU may well end up eating up all the performance gains compiling just in time would have given.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;LuaJIT: Peeking under the hood&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;The LuaJIT project was started and is mostly written by Mike Pall, that is apparently a pseudonym for a very private and very smart hacker.  I assume that he is male given that Mike is a common male name.  The source code repository is a bit odd.  There is a github repository that is supposed to be official but isn&amp;rsquo;t; it is a mirror created by CloudFlare along with Mike with the aim to broaden the developer community base.  That ride hasn&amp;rsquo;t been the smoothest and I&amp;rsquo;ve talked about it in more detail below.  The latest code with support for other architectures such as arm64 and ppc are in the v2.1 branch, which has only had beta releases come off it, the last one in 2017.  There are tests in a separate repository called &lt;a href=&#34;https://github.com/LuaJIT/LuaJIT-test-cleanup&#34;&gt;LuaJIT-test-cleanup&lt;/a&gt; which has a big fat warning that it is not the official testsuite, although if you look around, it pretty much is the only testsuite worth using for luajit.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Wait, there&amp;rsquo;s also &lt;a href=&#34;https://github.com/vkrasnov/bench_lua&#34;&gt;bench_lua&lt;/a&gt;, which has some benchmarks and a pretty nice driver for the benchmarks, something that the LuaJIT-test-cleanup benchmarks lack.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;LuaJIT uses the concept of &lt;a href=&#34;https://en.wikipedia.org/wiki/Tracing_just-in-time_compilation&#34;&gt;trace compiling&lt;/a&gt; which is pretty simple in concept but has some very interesting side-effects.  The idea of trace compilation, specifically with luajit is quite simple and follows roughly this logic:&lt;/p&gt;&#xA;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Interpret program and profile it while it is running.  Typical candidates for profiling would be loops for the obvious reason that it will likely execute repeatedly.&lt;/li&gt;&#xA;&lt;li&gt;If a loop is hit repeatedly, i.e. it crosses a threshold number of iterations, the JIT compiler is invoked on its next iteration.&lt;/li&gt;&#xA;&lt;li&gt;The JIT compiler first traces execution of the program and generates an IR for the trace of the program.&lt;/li&gt;&#xA;&lt;li&gt;The IR then goes through some optimisation passes and finally code is generated for the desired CPU backend.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;p&gt;This keeps on repeating as the interpreter encounters more hotspots. The interesting bit here is that the only bit that gets compiled is the code that gets executed during the trace. So if you have a branch like so:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&lt;code&gt;    if cond &amp;gt; threshold then&#xA;        i = i + 1&#xA;    else&#xA;        i = i - 1&#xA;    end&#xA;&lt;/code&gt;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;and the else block is executed during the trace, only that bit is compiled and not the if block.  The compiled code then has branches (known as exit guards) to jump back into the interpreter if the condition is true.  This produces an interesting optimisation opportunity that can be done during tracing itself.  If &lt;code&gt;cond &amp;gt; threshold&lt;/code&gt; is found to be always false because they are constants or some other reason, the &lt;code&gt;if&lt;/code&gt; condition can be completely eliminated, which saves compilation time as well as execution time.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Another interesting side effect of tracing that is not seen in typical compilers is that function calls effectively get inlined.  Again, that becomes a very cheap way to achieve something that would otherwise have been done in a separate pass in traditional compilers.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;In addition to very fast tracing and compilation, all of luajit is quite compact.  It&amp;rsquo;s &lt;a href=&#34;http://wiki.luajit.org/SSA-IR-2.0&#34;&gt;IR&lt;/a&gt; is linear array based and is hence allows very fast traversal. It&amp;rsquo;s easy to visualize it using the jit.* debug modules and using the -jdump flag to dump the IR during execution.  The &lt;a href=&#34;https://luajit.org/&#34;&gt;luajit wiki&lt;/a&gt; has some pretty detailed documentation on its internals.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The coding style of the project is a bit too compact to my taste since I personally prefer writing for readability.  There are a lot of constructs throughout the code that need a fair amount of squinting to understand, such as assignments inside the for loop headers and inside conditions.  OK all of you pointing at the macro and makefile soups in glibc and laughing, please be quiet ;)&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There&amp;rsquo;s also the infamous (at least in luajit circles) 47-bit address space limitation for garbage collected objects in luajit because luajit uses the top bits for metadata.  This is known to have correctness issues with Lua userdata objects and also performance issues because luajit repeatedly tries allocations until it finds a suitable address in the 47-bit space.  It doesn&amp;rsquo;t hurt x86 much (because of MAP_32BIT) but arm64 feels it and I imagine so do other architectures.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;My LuaJIT involvement&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;My full time involvement with luajit was brief and will likely end soon (my personal involvement may still continue) so in this short period I wanted to tick off as many short but significant work items as I could.  &lt;a href=&#34;https://github.com/siddhesh/LuaJIT&#34;&gt;My github fork is here&lt;/a&gt;.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/in/sameeradeshpande/&#34;&gt;Sameera Deshpande&lt;/a&gt; started the initial work and then helped me ramp up later on.  We got a couple of CI instances up and running to begin with, one for the &lt;a href=&#34;https://ci.linaro.org/job/luajit-aarch64/&#34;&gt;official repository&lt;/a&gt; and another for my &lt;a href=&#34;https://ci.linaro.org/job/luajit-siddhesh/&#34;&gt;github fork&lt;/a&gt; so that I can review my changes regularly.  If you&amp;rsquo;re interested in adding a node for your architecture to the Ci projects, please feel free to reach out to me, Linaro will happily add the node to the CI matrix.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Register Allocation improvements&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The register allocator in luajit is pretty simple to keep the compilation overhead low.  Registers are allocated sequentially based on their categories (caller saved, callee saved, etc.) and it uses some tricks such as constant rematerialization used to reduce register pressure.  Rematerialization is also very basic in its implementation; whenever constants need to be allocated to registers, it is preferred that they use existing constants, (assuming their live ranges are compatible) either directly or as a constant computation.  This is quite valuable because there is a fair amount of constant usage in the JITted code; exit guard addresses are coded in as constants for example and so are floating point numbers, in addition to the usual integers.  The register modes are not specified during allocation and are defined by the instructions generated in the assembly phase.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There was a bug in the luajit register allocator due to which registers used for constant rematerialization were being clobbered, resulting in corruption.  A &lt;a href=&#34;https://github.com/LuaJIT/LuaJIT/pull/438&#34;&gt;fix was proposed&lt;/a&gt; but the author of the fix was not sure if it was correct.  I posted an &lt;a href=&#34;https://github.com/LuaJIT/LuaJIT/pull/479&#34;&gt;alternative patch&lt;/a&gt; and then realized and explained why my patch is overkill and his approach is optimal.  I added additional cleanups to that to finish it up.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;While working on this problem, I noticed that the arm64 backend was not using XZR often enough and I &lt;a href=&#34;https://github.com/LuaJIT/LuaJIT/pull/482&#34;&gt;posted a patch&lt;/a&gt; to fix that.  I started benchmarking the improvement (the codegen was obviously better, it was saving registers for stores fo zeroes for example) and quickly realized that both bench_lua and the LuaJIT-test-cleanup benchmarks were quite raw and couldn&amp;rsquo;t be relied upon for consistent results.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;So I digressed.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Benchmark improvements and luaJIT-test-cleanup cleanup&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;bench_lua was my more favourite project to hack on benchmarks because it was evident that reviews were very hard to come by in the luajit project.  Also, bench_lua had a benchmark driver that produced repeatable results but it still had some cleanup issues, including the fact that it did not have a license!  The author was very responsive on the license question though and quickly put one in.  I fixed some timing issues in the driver and while doing so, I realized that it might be better if I used this driver on the more extensive set of benchmarks in LuaJIT-test-cleanup.  So that&amp;rsquo;s what I did.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I integrated the bench_lua driver into luajit-test-cleanup and added Makefile targets so that one could easily do make check and make bench to run the tests and benchmarks.  Now I had something I could work with but it was still in a different repo and it was getting quite cumbersome to work with them.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;So I integrated LuaJIT-test-cleanup into LuaJIT.  Now I had a LuaJIT repository that IMO was complete and could handle the standard make/make check workflow.  At the same time, it was modular enough that it could be merged into the upstream LuaJIT with relative ease.  I posted all of these patches as PRs and watched as nothing happened.  The LuaJIT-test-cleanup project had not seen a PR review since about 2016 and the LuaJIT project had seen occassional comments and patches from Mike in the past couple of years, but not much else.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Fusing and combining optimisations&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;Instruction fusion is an architecture dependent feature in luajit and each backend implements its own during the IR to assembly conversion phase, where the IR is traversed from the bottom up and assembly instructions generated sequentially.  Luajit does some trivial reordering in its IR optimisation passes but during assembly, it does not peek ahead to actively look for instruction fusion opportunities; it only tries to fuse neighbouring instructions.  As a result, while there are implementations for instructions like load and store pair in arm64, it is useful in only the most trivial of tests.  Likewise for fmadd/fmsub; a simple intervening load is sufficient to prevent the optimisation.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;In addition to this, it is often seen that optimisations like loop unrolling and vectorisation bring in even more opportunities for combining of loads and stores.  Luajit does some loop peeling but that&amp;rsquo;s about it.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Sameera did some analysis on ways to introduce more aggressive unrolling and possibly some amount of vectorisation but we did not have enough time to implement it.  She did have enough time to implement some instruction fusing and using fnmadd and fnmsub for arm64.  She also looked at load combining opportunities but realized that luajit would need more powerful instruction reordering, similar to the load grouping in the gcc scheduler that makes load pair generation much easier.  So that project was also not small enough for us to complete in the limited time.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Casting floats to unsigned integers&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The C standard defines casting of floating point types to unsigned integer types only for the range (-1.0, UTYPE_MAX), where UTYPE_MAX is the unsigned version of TYPE. Casts to signed types work just fine as long as the number is in the range of that type.  Waters get a bit murky with dynamic types and type narrowing when the default internal representation for all numbers is double.  That was the situation in luajit.  The fix for this was pretty straightforward in theory, which was to add an additional cast from float to signed int and then to unsigned int for floating point values less than zero and sticking to a direct cast to unsigned int for positive numbers.  I have implemented this for the interpreter and for arm64 in my fork.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Project state and the road ahead&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;LuaJIT is a very interesting project that has some very interesting concepts that I learned in the last month or so.  It has a pretty active user community that sings praises of the project and seems to advocate it in a number of areas.  However, the project development itself is in a bit of a crisis.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Around 2015 Mike Pall said he wanted to step back from the project and wanted more people to get involved in the development.  With that intent, Cloudflare created the github organisation and repository to allow for better collaboration.  Based on conversation threads I read, things seemed to go fine when the community stepped in to create the LuaJIT-test-cleanup repository based on some initial tests Mike had written and built it up into a set of 500+ tests.  However in about a year that excitement faded because nobody was made maintainer alongside Mike to carry forward the work and that meant that the LuaJIT project itself would only get sporadic fixes whenever Mike had some free time.  Minor patches were accepted but bigger pieces of code went unreviewed and presumably the developers also lost interest.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Fast forward four years into 2019 and we are still in the same situation, probably worse.  LuaJIT-test-cleanup has not had a patch review since 2016.  LuaJIT has had comments about a couple of times each quarter and bug fixes with similar frequency, but not much else.  The mailing list also has similar traffic - I announced all of the work I did above and did not get any responses.  there are forks of LuaJIT all over the place in projects such as OpenResty and RaptorJIT and the projects seem happy to let things run that way.  Lua language support is in a bit of a limbo with it being mostly 5.1 compliant with some 5.2 bits thrown in.  Overall, it&amp;rsquo;s a great chunk of code that&amp;rsquo;s about to vanish into oblivion.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Then there is the very tricky question of copyright.  The copyright notices all over the code say that Mike Pall has ownership.  However, the code clearly has a number of contributions from others and there is no copyright assignment in place.  While it&amp;rsquo;s likely not an issue from a licensing standpoint (IANAL, etc.), it is definitely something that needs to be addressed if the project is somehow ressurected, at the very least to give more prominent credit to contributors.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I&amp;rsquo;ve posted PRs for my work and tried to engage but I don&amp;rsquo;t have much hope given past history.  I intend to spend at least some of my free time tinkering with this code since it&amp;rsquo;s just a very interesting project and there&amp;rsquo;s a lot that can be done.  I am trawling the PRs and issue lists to look for patches that can be incorporated in my tree so if anyone is interested in contributing patches, you&amp;rsquo;re most welcome.  I will continue to ensure that my tree applies on top of the official repository because I do not want to give up hope of the project coming back to life.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>28 Mar 19 20:29 UTC</pubDate>
    </item>
    <item>
      <title>Optimizing toolchains for modern microprocessors</title>
      <link>https://gotplt.org/posts/optimizing-toolchains-for-modern-microprocessors.html</link>
      <description>&lt;!--&#xA;.. title: Optimizing toolchains for modern microprocessors&#xA;.. slug: optimizing-toolchains-for-modern-microprocessors&#xA;.. date: 2018-02-11T11:37:09-08:00&#xA;.. tags: glibc, cpu, toolchain, strings, gnu, aarch64, arm&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;About 2.5 years ago I left Red Hat to join Linaro in a move that surprised even me for the first few months.  I still work on the GNU toolchain with a glibc focus, but my focus changed considerably.  I am no longer looking at the toolchain in its entirety (although I do that on my own time whenever I can, either as glibc release manager or reviewer); my focus is making glibc routines faster for one specific server microprocessor; no prizes for &lt;a href=&#34;https://www.cygwin.com/ml/binutils/2016-11/msg00031.html&#34;&gt;guessing which processor&lt;/a&gt; that is.  I have read architecture manuals in the past to understand specific behaviours but this is the first time that I have had to pore through the entire manual and optimization guides and try and eek out the last cycle of performance from a chip.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;This post is an attempt to document my learnings and make a high level guide of the various things me and my team looked at to improve performance of the toolchain.  Note that my team is continuing to work on this chip (and I continue to learn new techniques, I may write about it later) so this &amp;lsquo;guide&amp;rsquo; is more of a personal journey.  I may add more follow ups or modify this post to reflect any changes in my understanding of this vast topic.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;All of my examples use ARM64 assembly since that&amp;rsquo;s what I&amp;rsquo;ve been working on and translating the examples to something x86 would have discouraged me enough to not write this at all.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;What am I optimizing for?&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;CPUs today are complicated beasts.  Vulnerabilities like Spectre allude to how complicated CPU behaviour can get but in reality it can get a lot more complicated and there&amp;rsquo;s never really a universal solution to get the best out of them.  Due to this, it is important to figure out what the end goal for the optimization is.  For string functions for example, there are a number of different factors in play and there is no single set of behaviours that trumps over all others.  For compilers in general, the number of such combinations is even higher.  The solution often is to try and ensure that there is a balance and there are no exponentially worse behaviours.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The first line of defence for this is to ensure that the algorithm used for the routine does not exhibit exponential behaviour.  I &lt;a href=&#34;https://developers.redhat.com/blog/2015/01/02/improving-math-performance-in-glibc/&#34;&gt;wrote&lt;/a&gt; about algorithmic changes I did to the multiple precision fallback  implementation in glibc years ago elsewhere so I&amp;rsquo;m not going to repeat that.  I will however state that the first line of attack to improve any function must be algorithmic.  Thankfully barring strcmp, string routines in glibc had a fairly sound algorithmic base.  strcmp fall back to a byte comparison when inputs are not mutually aligned, which is now fixed.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Large strings vs small&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;This is one question that gets asked very often in the context of string functions and different developers have different opinions on it, some differences even leading to flamewars in the past.  One popular approach to &amp;lsquo;solving&amp;rsquo; this is to quote usage of string functions in a popular benchmark and use that as a measuring stick.  For a benchmark like CPU2006 or CPU2017, it means that you optimize for smaller strings because the number of calls to smaller strings is very high in those benchmarks.  There are a few issues to that approach:&lt;/p&gt;&#xA;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;These benchmarks use glibc routines for a very small fraction of time, so you&amp;rsquo;re not going to win a lot of performance in the benchmark by improving small string performance&lt;/li&gt;&#xA;&lt;li&gt;Small string operations have other factors affecting it a lot more, i.e. things like cache locality, branch predictor behaviour, prefether behaviour, etc.  So while it might be fun to tweak behaviour exactly the way a CPU likes it, it may not end up resulting in the kind of gains you&amp;rsquo;re looking for&lt;/li&gt;&#xA;&lt;li&gt;A 10K string (in theory) takes at least 10 times more cycles than a 1K string, often more.  So effectively, there is 10x more incentive to look at improving performance of larger strings than smaller ones.&lt;/li&gt;&#xA;&lt;li&gt;There are CPU features specifically catered for larger sequential string operations and utilizing those microarchitecture quirks will guarantee much better gains&lt;/li&gt;&#xA;&lt;li&gt;There are a significant number of use cases outside of these benchmarks that use glibc far more than the SPEC benchmarks.  There&amp;rsquo;s no established set of benchmarks that represent them though.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;p&gt;I won&amp;rsquo;t conclude with a final answer for this because there is none.  This is also why I had to revisit this question for every single routine I targeted, sometimes even before I decide to target it.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Cached or not?&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;This is another question that comes up for string routines and the answer is actually a spectrum - a string could be cached, not cached or partially cached.  What&amp;rsquo;s the safe assumption then?&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There is a bit more consensus on the answer to this question.  It is generally considered safe to consider that shorter string accesses are cached and then focus on code scheduling and layout for its target code.  If the string is not cached, the cost of getting it into cache far outweighs the savings through scheduling and hence it is pointless looking at that case.  For larger strings, assuming that they&amp;rsquo;re cached does not make sense due to their size.  As a result, the focus for such situations should be on ensuring that cache utilization is optimal.  That is, make sure that the code aids all of the CPU units that populate caches, either through a hardware prefetcher or through judiciously placed software prefetch instructions or by avoiding caching altogether, thus avoiding evicting other hot data.  Code scheduling, alignment, etc. is still important because more often than not you&amp;rsquo;ll have a hot loop that does the loads, compares, stores, etc. and once your stream is primed, you need to ensure that the loop is not suboptimal and runs without stalls.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;My branch is more important than yours&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;Branch predictor units in CPUs are quite complicated and the compiler does not try to model them.  Instead, it tries to do the simpler and more effective thing; make sure that the more probably branch target is accessible through sequential fetching.  This is another aspect of the large strings vs small for string functions and more often than not, smaller sizes are assumed to be more probable for hand-written assembly because it seems to be that way in practice and also the cost of a mispredict hits the smaller size more than it does the larger one.&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Don&amp;rsquo;t waste any part of a &lt;del&gt;pig&lt;/del&gt; CPU&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;CPUs today are complicated beasts.  Yes I know I started the previous section with this exact same line; they&amp;rsquo;re complicated enough to bear repeating that.  However, there is a bit of relief in the fact that the first principles of their design hasn&amp;rsquo;t changed much.  The components of the CPU are all things we heard about in our CS class and the problem then reduces to understanding specific quirks of the processor core.  At a very high level, there are three types of quirks you look for:&lt;/p&gt;&#xA;&#xA;&lt;ol&gt;&#xA;&lt;li&gt;Something the core does exceedingly well&lt;/li&gt;&#xA;&lt;li&gt;Something the core does very badly&lt;/li&gt;&#xA;&lt;li&gt;Something the core does very well or badly under specific conditions&lt;/li&gt;&#xA;&lt;/ol&gt;&#xA;&#xA;&lt;p&gt;Typically this is made easy by CPU vendors when they provide &lt;a href=&#34;http://infocenter.arm.com/help/topic/com.arm.doc.uan0015b/Cortex_A57_Software_Optimization_Guide_external.pdf&#34;&gt;documentation&lt;/a&gt; that specifies a lot of this information.  Then there are cases where you discover these behaviours through profiling.  Oh yes, before I forget:&lt;/p&gt;&#xA;&#xA;&lt;div style=&#34;text-align:center; font-weight:bold;font-size:1.3em;text-transform:uppercase&#34;&gt;Learn how to use perf or similar tool and read its output it will save your life&lt;/div&gt;&#xA;&#xA;&lt;p&gt;For example, the falkor core does something interesting with respect with loads and addressing modes.  Typically, a load instruction would take a specific number of cycles to fetch from L1, more if memory is not cached, but that&amp;rsquo;s not relevant here. If you issue a load instruction with a pre/post-incrementing addressing mode, the microarchitecture issues two micro-instructions; one load and another that updates the base address.  So:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&#xA;   ldr  x1, [x2, 16]!&#xA;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;effectively is:&lt;/p&gt;&#xA;&#xA;&lt;pre&gt;&#xA;  ldr   x1, [x2, 16]&#xA;  add   x2, x2, 16&#xA;&lt;/pre&gt;&#xA;&#xA;&lt;p&gt;and  that increases the net cost of the load.  While it saves us an instruction, this addressing mode isn&amp;rsquo;t always preferred in unrolled loops since you could avoid the base address increment at the end of every instruction and do that at the end.  With falkor however, this operation is very fast and in most cases, this addressing mode is preferred for loads.  The reason for this is the way its hardware prefetcher works.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Hardware Prefetcher&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;A hardware prefetcher is a CPU unit that speculatively loads the memory location after the location requested, in an attempt to speed things up.  This forms a memory stream and larger the string, the more its gains from prefetching.  This however also means that in case of multiple prefetcher units in a core, one must ensure that the same prefetcher unit is hit so that the unit gets trained properly, i.e. knows what&amp;rsquo;s the next block to fetch.  The way a prefetcher typically knows is if sees a consistent stride in memory access, i.e. it sees loads of X, X+16, X+32, etc. in a sequence.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;On falkor the addressing mode plays an important role in determining which hardware prefetcher unit is hit by the load and effectively, a pre/post-incrementing load ensures that the loads hit the same prefetcher.  That combined with a feature called register renaming ensures that it is much quicker to just fetch into the same virtual register and pre/post-increment the base address than to second-guess the CPU and try to outsmart it.  The memcpy and memmove routines use this quirk extensively; comments in the falkor routines even have detailed comments explaining the basis of this behaviour.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Doing something so badly that it is easier to win&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;A colleague once said that the best targets for toolchain optimizations are CPUs that do things badly.  There always is this one behaviour or set of behaviours that CPU designers decided to sacrifice to benefit other behaviours.  On falkor for example, calling the MRS instruction for some registers is painfully slow whereas it is close to single cycle latency for most other processors.  Simply avoiding such slow paths in itself could result in tremendous performance wins; this was evident with the memset function for falkor, which became twice as fast for medium sized strings.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Another example for this is in the compiler and not glibc, where the fact that using a &amp;lsquo;str&amp;rsquo; instruction on 128-bit registers with register addressing mode is very slow on falkor.  Simply avoiding that instruction altogether results in pretty good gains.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;CPU Pipeline&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;Both gcc and llvm allow you to specify a model of the CPU pipeline, i.e.&lt;/p&gt;&#xA;&#xA;&lt;ol&gt;&#xA;&lt;li&gt;The number of each type of unit the CPU has.  That is, the number of load/store units, number of integer math units, number of FP units, etc.&lt;/li&gt;&#xA;&lt;li&gt;The latency for each type of instruction&lt;/li&gt;&#xA;&lt;li&gt;The number of micro-operations each instruction splits into&lt;/li&gt;&#xA;&lt;li&gt;The number of instructions the CPU can fetch/dispatch in a single cycle&lt;/li&gt;&#xA;&lt;/ol&gt;&#xA;&#xA;&lt;p&gt;and so on.  This information is then used to sequence instructions in a function that it optimizes for.  This may also help the compiler choose between instructions based on how long those take.  For example, it may be cheaper to just declare a literal in the code and load from it than to construct a constant using mov/movk.  Similarly, it could be cheaper to use csel to select a value to load to a register than to branch to a different piece of code that loads the register or vice versa.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Optimal instruction sequencing can often result in significant gains.  For example, intespersing load and store instructions with unrelated arithmetic instructions could result in both those instructions executing in parallel, thus saving time.  On the contrary, sequencing multiple load instructions back to back could result in other units being underutilized and all instructions being serialized on to the load unit.  The pipeline model allows the compiler to make an optimal decision in this regard.&lt;/p&gt;&#xA;&#xA;&lt;h3&gt;Vector unit - to use or not to use, that is the question&lt;/h3&gt;&#xA;&#xA;&lt;p&gt;The vector unit is this temptress that promises to double your execution rate, but it doesn&amp;rsquo;t come without cost.  The most important cost is that of moving data between general purpose and vector registers and quite often this may end up eating into your gains.  The cost of the vector instructions themselves may be high, or a CPU might have multiple integer units and just one SIMD unit, because of which code may get a better schedule when executed on the integer units as opposed to via the vector unit.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I had seen an opposite example of this in powerpc years ago when I noticed that much of the integer operations were also implemented in FP in multiple precision math.  This was because the original authors were from IBM and they had noticed a significant performance gain with that on powerpc (possible power7 or earlier given the timelines) because the CPU had 4 FP units!&lt;/p&gt;&#xA;&#xA;&lt;h2&gt;Final Thoughts&lt;/h2&gt;&#xA;&#xA;&lt;p&gt;This is really just the tip of the iceberg when it comes to performance optimization in toolchains and utilizing CPU quirks.  There are more behaviours that could be exploited (such as aliasing behaviour in branch prediction or core topology) but the cost benefit of doing that is questionable.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Despite how much fun it is to hand-write assembly for such routines, the best approach is always to write simple enough code (yes, clever tricks might actually defeat compiler optimization passes!) that the compiler can optimize for you.  If there are missed optimizations, improve compiler support for it.  For glibc and aarch64, there is also the case of impending multiarch explosion.  Due to the presence of multiple vendors, having a perfectly tuned routine for each vendor may pose code maintenance problems and also secondary issues with performance, like code layout in a binary and instruction cache utilization.  There are some random ideas floating about for that already, like making separate text sections for vendor-specific code, but that&amp;rsquo;s something we would like to avoid doing if we can.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>11 Feb 18 19:37 UTC</pubDate>
    </item>
    <item>
      <title>Across the Charles Bridge - GNU Tools Cauldron 2017</title>
      <link>https://gotplt.org/posts/across-the-charles-bridge-gnu-tools-cauldron-2017.html</link>
      <description>&lt;!--&#xA;.. title: Across the Charles Bridge - GNU Tools Cauldron 2017&#xA;.. slug: across-the-charles-bridge-gnu-tools-cauldron-2017&#xA;.. date: 2017-09-11T23:16:28-07:00&#xA;.. tags: cauldron, gnu, toolchain, glibc, tunables&#xA;.. link:&#xA;.. description:&#xA;.. type: text&#xA;--&gt;&#xA;&#xA;&lt;p&gt;Since I joined Linaro back in 2015 around this time, my travel has gone up 3x with 2 Linaro Connects a year added to the one GNU Tools Cauldron.  This year I went to FOSSAsia too, so it&amp;rsquo;s been a busy traveling year.  The special thing about Cauldron though is that it is one of those conferences where I &amp;lsquo;work&amp;rsquo; as well as have a lot of fun.  The fun bit is because I get to meet all of the people that I work with almost every day in person and a lot of them have become great friends over the years.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I still remember the first Cauldron I went to in 2013 at Mountain View where I felt dwarfed by all of the giants I was sitting with.  It was exaggerated because it was the first time I met the likes of Jeff Law, Richard Henderson, etc. in personal meetings since I had joined the Red Hat toolchain team just months before; it was intimidating and exciting all at once.  That was also the first time I met Roland McGrath (I still hadn&amp;rsquo;t met Carlos, he had just had a baby and couldn&amp;rsquo;t come), someone I was terrified of back then because his patch reviews would be quite sharp and incisive.  I had imagined him to be a grim old man hammering out those words from a stern laptop, so it was a surprise to see him use the same kinds of words but with a sarcastic smile, completely changing the context and tone.  That was the first time I truly realized how emails often lack context.  Years later, I still try to visualize people when I read their emails.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Skip to 4 years later and I was at my 5th Cauldron last week and despite my assumptions on how it would go, it was a completely new experience.  A lot of it had to do with my time at Linaro and very little to do with technical growth.  I felt like an equal to Linaro folks all over the world and I seemed to carry that forward here, where I felt like an equal with all of the people present, I felt like I belonged.  I did not feel insecure about my capabilities (I still am intimately aware of my limitations), nor did I feel the need to constantly prove that I belonged.  I was out there seeking toolchain developers (we are hiring btw, email me if you&amp;rsquo;re a fit), comfortable with the idea of leading a team.  The fact that I managed to not screw up the two glibc releases I managed may also have helped :)&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Oh, and one wonderful surprise was that &lt;a href=&#34;http://www.amitshah.net/&#34;&gt;an old friend&lt;/a&gt; decided to drop in an Cauldron and spend a couple of days.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;This year&amp;rsquo;s Cauldron had the most technical talks submitted in recent years.  We had 5 talks in the glibc area, possibly also the highest for us; just as well because we went over time in almost all of them.  I won&amp;rsquo;t say that it&amp;rsquo;s a surprise since that has happened in every single year that I attended.  The first glibc talk was about tunables where I briefly recapped what we have done in tunables so far and talked about the future a bit more at length.  Pedro Alves suggested putting pretty printers for tunables for introspection and maybe also for runtime tuning in the coming future.  There was a significant amount of interest in the idea of auto-tuning, i.e. collecting profiling data about tunable use and coming up with optimal default values and possibly even eliminating such tunables in future if we find that we have a pretty good default.  We also talked about tuning at runtime and the various kinds of support that would be required to make it happen.  Finally there were discussions on tuning profiles and ideas around creating performance-enhanced routines for workloads instead of CPUs.  The video recording of the talk will hopefully be out soon and I&amp;rsquo;ll link the video here when it is available.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;Florian then talked about glibc 3.0, a notional concept (i.e. won&amp;rsquo;t be a soname bump) where we rewrite sections of code that have been rotting due to having to support some legacy platforms.  The most prominent among them is libio, the module in glibc that implements stdio.  When libio was written, it was designed to be compatible with libstdc++ so that FILE streams could be compatible with C++ stdio streams.  The only version of gcc that really supports that is 2.95 since libstdc++ has since moved on.  However because of the way we do things in glibc, we cannot get rid of them even if there is just one user that needs that ABI.  We toyed with the concept of a separate compatibility library that becomes a graveyard for such legacy interfaces so that they don&amp;rsquo;t hold up progress in the library.  It remains to be seen how this pans out, but I would definitely be happy to see this progress; libio was one of my backlog projects for years.  I had to miss Raji&amp;rsquo;s talk on powerpc glibc improvements since I had to be in another meeting, so I&amp;rsquo;ll have to catch it when the video comes out.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;The two BoFs for glibc dealt with a number of administrative and development issues, details of which Carlos will post on the mailing list soon.  The highlights for me were the malloc instrumented benchmarks that Carlos wants to add to benchtests and build and review tools.  Once I clear up my work backlog a bit, I&amp;rsquo;ll attempt to set up something like phabricator or gerrit and see how that works out or the community instead of patchwork.  I am convinced that all of the issues that we want to solve like crediting reviewers, ensuring good git commit logs, running automated builds and tests, etc. can only be effectively solved with a proper review tool in place to review patches.&lt;/p&gt;&#xA;&#xA;&lt;p&gt;There was also a discussion on redoing the makefiles in glibc so that it doesn&amp;rsquo;t spend so much time doing dependecy resolution, but I am going to pretend that it didn&amp;rsquo;t happen because it is an ugly ugly task :/&lt;/p&gt;&#xA;&#xA;&lt;p&gt;I&amp;rsquo;m back home now, recovering from the cold that worsened while I was in Prague before I head out again in a couple of weeks to SFO for Linaro Connect.  I&amp;rsquo;ve booked tickets for whale watching tours there, so hopefully I&amp;rsquo;ll be posting some pictures again after a long break.&lt;/p&gt;&#xA;</description>
      <author>Siddhesh</author>
      <pubDate>12 Sep 17 06:16 UTC</pubDate>
    </item>
  </channel>
</rss>