<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>the Saguaros &#187; c++11</title>
	<atom:link href="https://www.thesaguaros.com/tag/c11/feed" rel="self" type="application/rss+xml" />
	<link>https://www.thesaguaros.com</link>
	<description>Things you will never need from people you&#039;ve never heard about</description>
	<lastBuildDate>Tue, 17 Mar 2026 09:29:46 +0000</lastBuildDate>
	<language>en-EN</language>
		<sy:updatePeriod>hourly</sy:updatePeriod>
		<sy:updateFrequency>1</sy:updateFrequency>
	<generator>https://wordpress.org/?v=3.7.41</generator>
	<item>
		<title>OpenMP-style constructs in C++11, part II</title>
		<link>https://www.thesaguaros.com/openmp-style-constructs-in-c11-part-ii.html</link>
		<comments>https://www.thesaguaros.com/openmp-style-constructs-in-c11-part-ii.html#comments</comments>
		<pubDate>Thu, 15 Sep 2011 09:54:57 +0000</pubDate>
		<dc:creator><![CDATA[Domenico]]></dc:creator>
				<category><![CDATA[General]]></category>
		<category><![CDATA[c++11]]></category>
		<category><![CDATA[OpenMP]]></category>
		<category><![CDATA[parallel programming]]></category>
		<category><![CDATA[threads]]></category>

		<guid isPermaLink="false">http://www.thesaguaros.com/?p=274</guid>
		<description><![CDATA[&#8230;what were we talking about? Last time, we coded a small OpenMP-style parallel construct using some macro directives and a class wrapping a vector of threads. This time we will add a replacement for 2 OpenMP library functions: omp_get_num_threads() and omp_get_thread_num(). These are among the most used (and useful) OpenMP functions. I&#8217;ll show you several [&#8230;]]]></description>
				<content:encoded><![CDATA[<h2>&#8230;what were we talking about? </h2>
<p><a href="http://www.thesaguaros.com/openmp-style-constructs-in-c11.html" title="Last time" target="_blank">Last time</a>, we coded a small OpenMP-style <code>parallel</code> construct using some macro directives and a class wrapping a <code>vector</code> of threads.</p>
<p>This time we will add a replacement for 2 OpenMP library functions: <code>omp_get_num_threads()</code> and <code>omp_get_thread_num()</code>. These are among the most used (and useful) OpenMP functions.</p>
<p>I&#8217;ll show you several implementations, but to start we need a little clean-up in the <code>thread_pool</code> class.</p>
<h2>first try: adding methods</h2>
<p>Following there&#8217;s the source for the thread_pool class, with a few modifications I&#8217;ll explain later. I suppose that you put it in a file called &#8220;thread_pool.h&#8221; in your working directory, if you put it elsewhere, change the #include in the samples to make them work.</p>
<pre class="brush:cpp">
#include &lt;thread&gt;
#include &lt;algorithm&gt;
#include &lt;vector&gt;
#include &lt;iostream&gt;
#include &lt;functional&gt;

using namespace std;

typedef function &lt;void ()&gt; task;
class thread_pool {
  private:
    vector&lt;thread&gt; the_pool;

  public:
    thread_pool(unsigned int num_threads, task tbd) {
      for(int i = 0; i < num_threads; ++i) {
        the_pool.push_back(thread(tbd));
      }
    }

    void join() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.join();});
    }

    void nowait() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.detach();});
    }
    
    int get_num_threads() { return the_pool.size(); }
    
    int get_thread_num() {
      for(int i = 0; i < the_pool.size(); ++i)
        if(the_pool[i].get_id()==this_thread::get_id()) return i;
      return -1;
    }

};

#define parallel_do_(N) thread_pool (N, []()
#define parallel_do parallel_do_(thread::hardware_concurrency())
#define parallel_end ).join();
#define parallel_end_nowait ).nowait();

</pre>
<p>The first important change in the thread_pool class is the definition of task, it's no more a function pointer but &dash; more correctly &dash; a <code>std::function</code> returing void and taking no arguments. This way it works with function pointer and lambda arguments, and allows us to capture variables in lambdas. </p>
<p>The mandatory usage example:</p>
<pre class="brush:cpp">
#include "thread_pool.h"

int main() {

    thread_pool p(4, [&#038;p] () {
        cout << "I'm thread n. " << p.get_thread_num() 
             << " in a pool of " << p.get_num_threads() << endl;
    });

    // You could do other things before joining...
    p.join();

    return 0;
}
</pre>
<p>As you see, there's a thread_pool instance named <code>p</code>, and in the code passed to the constructor (and executed by four threads), the object p itself is used to call the methods <code>get_num_threads()</code> and <code>get_thread_num()</code>. This "magic" is made possible because the variable p has been captured. The square brackets in lamdas are used for this purpose (as usual, I'm not explaining the whole thing, but there's plenty of information on the net).</p>
<p>This solution works, but requires our pools to be named, so we should modify our macro definition to include the pool name as a parameter. We can do better.</p>
<h2>second try: the global map</h2>
<p>I want to say it loud and clear: <strong>I don't like this second solution at all</strong>, it's not elegant and uses a global object. I'm not even going to show you a complete example, just a modified version of the thread_pool class to give you the idea of how it could be done.</p>
<pre class="brush:cpp">
//... includes omitted
int get_thread_num();
int get_num_threads();

typedef function &lt;void ()&gt; task;
class thread_pool {
  private:
    typedef map&lt;thread::id, thread_pool*&gt; thread_map;
    static thread_map allthreads;
    friend int get_thread_num();
    friend int get_num_threads();
    vector&lt;thread&gt; the_pool;

  public:
    thread_pool(unsigned int num_threads, task tbd) {
      for(int i = 0; i < num_threads; ++i) {
        the_pool.push_back(thread(tbd));
        allthreads.insert(
          thread_map::value_type(the_pool[i].get_id(), this));
      }
    }
    
    ~thread_pool() {
       for_each(the_pool.begin(), the_pool.end(), [] (thread&#038; t) { 
         allthreads.erase(t.get_id()); 
        });
    }

    void join() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.join();});
    }

    void nowait() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.detach();});
    }
    
    int get_num_threads() { return the_pool.size(); }
    
    int get_thread_num() {
      for(int i = 0; i < the_pool.size(); ++i)
        if(the_pool[i].get_id()==this_thread::get_id()) return i;
      return -1;
    }

};

int get_thread_num() {
  thread_pool * p = thread_pool::allthreads[this_thread::get_id()];
  return p->get_thread_num();
}

int get_num_threads() {
  thread_pool * p = thread_pool::allthreads[this_thread::get_id()];
  return p->get_num_threads();
}

// This should be in a .cc file!
map&lt;thread::id, thread_pool*&gt; thread_pool::allthreads;

</pre>
<p>How it works:
<ul>
<li>the object <code>allthreads</code> associate every thread in a thread_pool to its pool.</li>
<li>thread_pool's constructor and destructor take care of adding and removing entries to the map</li>
<li>the static friend functions <code>get_num_threads()</code> and <code>get_thread_num()</code> use allthreads to get a pointer to the thread's pool and invoke the homonymous instance methods</li>
</ul>
<p>I'm not going to complete or discuss further this example, because I want to show you a better way.</p>
<h2>a better way: thread local storage</h2>
<p><a href="http://en.wikipedia.org/wiki/Thread-local_storage" title="Thread-local storage" target="_blank">Thread-local storage</a> is a way to let each thread mantain its own version of a global variable or memory region. The idea is to use two thread-local variables to store <code>num_threads</code> and <code>thread_num</code>.</p>
<p>C++11 introduces the storage specifier <code>thread_local</code> to declare thread-local variables. Sadly, many compilers don't support it yet, and GCC is one of them, so I'll use the <code>__thread</code> builtin for this compiler, but the principle is the same.</p>
<p>Here is the resulting thread_pool class.</p>
<pre class="brush:cpp">
#include &lt;thread&gt;
#include &lt;algorithm&gt;
#include &lt;vector&gt;
#include &lt;iostream&gt;
#include &lt;functional&gt;

using namespace std;

#ifdef __GNUG__
static __thread int thread_num;
static __thread int num_threads;
#else
static thread_local int thread_num;
static thread_local int num_threads;
#endif 

typedef function &lt;void ()&gt; task;
class thread_pool {
  private:
    vector&lt;thread&gt; the_pool;

  public:
    thread_pool(unsigned int n_threads, task tbd) {
      for(int i = 0; i < n_threads; ++i) {
        the_pool.push_back(thread([=] () {
          thread_num = i;
          num_threads = n_threads;
          tbd();
        }));
      }
    }
    
    void join() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.join();});
    }

    void nowait() {
      for_each(the_pool.begin(), the_pool.end(), 
        [] (thread&#038; t) {t.detach();});
    }
    
};

#define parallel_(N) thread_pool (N, []()
#define parallel parallel_(thread::hardware_concurrency())
#define parallel_end ).join();
#define parallel_end_nowait ).nowait();
#define single if(thread_num==0) 

</pre>
<p>The local copy of <code>num_threads</code> and <code>thread_num</code> are initialized in thread_pool's constructor. Again, we are using variable capture to access the referenced variables inside the thread code.</p>
<p>Example:</p>
<pre class="brush:cpp">
#include "thread_pool.h"
#include &lt;iostream&gt;

int main() {

    parallel_(4)
    {
      cout << "I'm thread " << thread_num << " of " 
           << num_threads << endl;
      single
      {
        cout << "This region is executed only by thread " 
             << thread_num << endl;
      }
    }
    parallel_end
    
    return 0;
}

</pre>
<p>In the example we have a parallel region executed by four threads, with a nested region executed only by the first thread in the pool. Thanks to the thread-local variables and to the macros, the code is both readable and concise.</p>
<p>That's it. Leave a comment to ask a question, suggest an improvement or share a thought.</p>
]]></content:encoded>
			<wfw:commentRss>https://www.thesaguaros.com/openmp-style-constructs-in-c11-part-ii.html/feed</wfw:commentRss>
		<slash:comments>0</slash:comments>
		</item>
		<item>
		<title>OpenMP-style constructs in C++11</title>
		<link>https://www.thesaguaros.com/openmp-style-constructs-in-c11.html</link>
		<comments>https://www.thesaguaros.com/openmp-style-constructs-in-c11.html#comments</comments>
		<pubDate>Mon, 12 Sep 2011 10:10:52 +0000</pubDate>
		<dc:creator><![CDATA[Domenico]]></dc:creator>
				<category><![CDATA[General]]></category>
		<category><![CDATA[c++11]]></category>
		<category><![CDATA[OpenMP]]></category>
		<category><![CDATA[parallel programming]]></category>
		<category><![CDATA[threads]]></category>

		<guid isPermaLink="false">http://www.thesaguaros.com/?p=220</guid>
		<description><![CDATA[The new C++ standard, called C++11, is finally here. It enriches both the language and its standard library, bringing some features that many users awaited, like lamdas, the &#8220;auto&#8221; type, and so on. But I&#8217;m not going to talk about these, there are a lot of good references on the net. Some days ago, after [&#8230;]]]></description>
				<content:encoded><![CDATA[<p>The new C++ standard, called C++11, is finally <a href="http://www.open-std.org/jtc1/sc22/wg21/" title="here" target="_blank">here</a>.<br />
It enriches both the language and its standard library, bringing some features that many users awaited, like lamdas, the &#8220;auto&#8221; type, and so on. But I&#8217;m not going to talk about these, there are a lot of good references on the net.</p>
<p>Some days ago, after viewing Bartosz Milewski&#8217;s excellent <a href="http://www.corensic.com/Learn/Resources/ConcurrencyTutorialPartOne.aspx" title="tutorial on C++11 concurrency" target="_blank">tutorial on C++11 concurrency</a>, I started playing with the language additions, trying to mimic the behavior of some <a href="http://openmp.org/" title="OpenMP" target="_blank">OpenMP</a> directive. (Again, if you want to learn more about OpenMP, surf the internet. You can start from the official <a href="http://openmp.org/wp/about-openmp/" title="about page" target="_blank">about page</a>).</p>
<p>I&#8217;ll show you some of these experiments. Of course we aren&#8217;t going to fully implement even a single directive, but maybe we can learn something about the new standard. Readers&#8217; comments and suggestion are welcome!</p>
<h2>testing the code</h2>
<p>C++11 is a young standard, and the compilers still don&#8217;t support it fully. I&#8217;ve used GCC 4.7 to compile the code below, but you can try with your favorite C++ compiler. Here&#8217;s a nice table summarizing support for the new features in various popular compilers: <a href="http://wiki.apache.org/stdcxx/C%2B%2B0xCompilerSupport" title="http://wiki.apache.org/stdcxx/C%2B%2B0xCompilerSupport" target="_blank">C++0xCompilerSupport</a>.</p>
<p>To enable support for the new features in g++, add the switch <code> -std=c++0x </code>, to compile OpenMP code add <code>-fopenmp</code> too.</p>
<h2>parallel</h2>
<p>With OpenMP, a programmer can introduce parallelism adding compiler directives and using its library functions. It uses a fork-join execution model. The simplest way to enable parallel execution of a region of code is via the <code>parallel</code> directive. Here&#8217;s a very simple example:</p>
<pre class="brush:cpp">
#include &lt;iostream&gt;
#include &lt;omp.h&gt;

using namespace std;

int main() {

  #pragma omp parallel num_threads(4)
  {
    cout << "I'm thread number " << omp_get_thread_num() << endl;
  }
  cout << "This code is executed by one thread\n";

  return 0;
}
</pre>
<p>The example is self-explanatory: the code block after the OpenMP pragma "parallel" is executed by 4 threads.<br />
At the end of the region there's an implicit barrier, so the last cout is executed only when all the threads have left the parallel region.<br />
Copy this code in a file, compile it, run it and look at the output (eg. <code>g++  -std=c++0x -fopenmp para1.c -o para1; ./para1</code>)</p>
<p>Let's see how to emulate this behavior using C++'s <code>std::thread</code>.</p>
<h2>threads in C++11</h2>
<p>To start a new thread, in C++11 we just need to create a <code>std::thread</code> object. The simpest (and useless!) example I can imagine is this:</p>
<pre class="brush:cpp">
#include &lt;iostream&gt;
#include &lt;thread&gt;

using namespace std;

void hello() {
  cout << "Hello from a thread\n";
}

int main() {
  thread aThread(&#038;hello);
  aThread.join();

  return 0;  
}
</pre>
<p>We can avoid passing a function pointer in line 11 and make the thing nicer using a lambda:<br />
(From now on I'm going to omit some of the includes and other repeated code for brevity. It should be easy to add the missing parts.)</p>
<pre class="brush:cpp">
int main() {
  thread aThread([]() {
    cout << "Hello from a thread\n";
  });
  aThread.join();

  return 0;  
}
</pre>
<h2>let's use the threads</h2>
<p>Ok, we can use threads to execute some work in parallel. Let's write a trivial thread pool class for the purpose.</p>
<pre class="brush:cpp">
using namespace std;

typedef void (*task) ();
class thread_pool {
  private:
    vector<thread> the_pool;

  public:
    thread_pool(unsigned int num_threads, task tbd) {
      for(int i = 0; i < num_threads; ++i) {
        the_pool.push_back(thread(tbd));
      }
    }

    void join() {
      for_each(the_pool.begin(), the_pool.end(), 
                 [] (thread&#038; t) {t.join();});
    }

    void nowait() {
      for_each(the_pool.begin(), the_pool.end(), 
                 [] (thread&#038; t) {t.detach();});
    }
};
</pre>
<p>It's just a wrapper over a <code>vector</code> of threads, with some method we'll find useful later.<br />
We can use it this way:</p>
<pre class="brush:cpp">
thread_pool pool(4, []() {
  cout << "Here I am: " << this_thread::get_id() << endl;
});
cout << "I can do other things before waiting for them to finish!" << endl;
pool.join();
</pre>
<p>Put these lines in a main and run this example. As you can see:</p>
<ul>
<li>Four threads are stared, each of them executes the code in the lambda (says "Here I am")</li>
<li>The main thread can get some other work done before joining them</li>
</ul>
<h2>syntactic sugar</h2>
<p>With a bit of syntactic sugar, we can make the code to resemble the OpenMP version more closely. A couple of macros will help us:</p>
<pre class="brush:cpp">
// class thread_pool omitted...

#define parallel_do_(N) thread_pool (N, []()
#define parallel_end ).join();

int main() {

    parallel_do_(4)
    {
      cout << "Here I am: " << this_thread::get_id() << endl;
    }
    parallel_end

    return 0;
}
</pre>
<h2>a bit more</h2>
<p>As I said you at the beginning, we are not trying to emulate the <code>parallel</code> construct fully, it does a lot more and has a lot of clauses that control its behavior. However, we can easily add support for a couple of nice things:
<ol>
<li>Let the system choose an appropriate number of threads</li>
<li>Avoid the implicit barrier at the end of the parallel region</li>
</ol>
<h3>Omit the number of threads</h3>
<p>When you don't specify the <code>num_threads</code> clause, OpenMP figures out itself the number of threads to start, based on the hardware resources available (and a lot of other things!). We can achieve a similar result using <code>thread::hardware_concurrency()</code> as a default value for num_threads.</p>
<h3>Don't wait at the barrier</h3>
<p>The <code>nowait</code> clause instructs OpenMP to not generate a barrier at the end of the parallel region. We can do this by detaching from the threads in the pool instead of joining them.</p>
<p>The following listing shows the new code and a sample of use.</p>
<pre class="brush:cpp">
// class thread_pool omitted...

#define parallel_do_(N) thread_pool (N, []()
#define parallel_do parallel_do_(thread::hardware_concurrency())
#define parallel_end ).join();
#define parallel_end_nowait ).nowait();

int main() {

    parallel_do_(4)
    {
      cout << "Here I am: " << this_thread::get_id() << endl;
    }
    parallel_end_nowait

    cout << "[MASTER] I can do other things while they complete...\n";

    //With default number of threads
    parallel_do
    {
      cout << "Let's count ourselves. I'm  " 
             << this_thread::get_id() << endl;
    }
    parallel_end

    cout << "[MASTER] Goodbye.\n";
    return 0;
}
</pre>
<p>Today we will stop here. I hope you enjoyed the reading.<br />
Share your thoughts in the comments.</p>
]]></content:encoded>
			<wfw:commentRss>https://www.thesaguaros.com/openmp-style-constructs-in-c11.html/feed</wfw:commentRss>
		<slash:comments>1</slash:comments>
		</item>
	</channel>
</rss>
