<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Wjfisher</id>
	<title>Expertiza_Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Wjfisher"/>
	<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Special:Contributions/Wjfisher"/>
	<updated>2026-10-08T07:44:18Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31152</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31152"/>
		<updated>2010-03-04T03:29:22Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{| align=&amp;quot;center cellpadding=&amp;quot;4&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
!Type of Parallelism&lt;br /&gt;
!Posix Threads&lt;br /&gt;
!Intel&amp;amp;reg; TBB&lt;br /&gt;
!OpenMP 2.0&lt;br /&gt;
!OpenMp 3.0&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
!DOALL&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
!DOACROSS&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
!DOPIPE&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|No&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31151</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31151"/>
		<updated>2010-03-04T03:27:39Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{| align=&amp;quot;center cellpadding=&amp;quot;4&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
|align=&amp;quot;center&amp;quot;|Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;| Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;| Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;| Yes&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31150</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31150"/>
		<updated>2010-03-04T03:26:59Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{| align=&amp;quot;center cellpadding=&amp;quot;4&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
|align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31147</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31147"/>
		<updated>2010-03-04T03:25:07Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{|&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
|align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
|align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31145</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31145"/>
		<updated>2010-03-04T03:24:49Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{|&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| align=&amp;quot;center&amp;quot;&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31144</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31144"/>
		<updated>2010-03-04T03:23:23Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{|&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31143</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31143"/>
		<updated>2010-03-04T03:22:48Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= Summary =&lt;br /&gt;
&lt;br /&gt;
A brief comparison of the code libraries discussed in this article will be made here in the form of a table.  As you can see, none of the libraries contain support for all of the types of parallelism discussed.&lt;br /&gt;
&lt;br /&gt;
{|&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Type of Parallelism&lt;br /&gt;
! Posix Threads&lt;br /&gt;
! Intel&amp;amp;reg; TBB&lt;br /&gt;
! OpenMP 2.0&lt;br /&gt;
! OpenMp 3.0&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOALL&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOACROSS&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! DOPIPE&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Reduction&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
! Functional Parallelism&lt;br /&gt;
| No&lt;br /&gt;
| No&lt;br /&gt;
| Yes&lt;br /&gt;
| Yes&lt;br /&gt;
}&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31136</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31136"/>
		<updated>2010-03-04T03:12:31Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  Generally, parallelism is easily found by investigating the blocks of code that take the longest to execute.  Parallelism can take several different forms but is most often found in loop structures. The focus of this article is to discuss the implementation of  parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  Each of these will be discussed in detail below with references to outside material for more comprehensive information, if desired.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.  Section 3.7 on page 60 of the Solihin text has an overview of OpenMP 2.0 that should be used to supplement the information covered here.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.  See the &amp;quot;Did you know?&amp;quot; box in Chapter 3 of the Solihin text on page 63 for more information about the DOACROSS and DOPIPE parallelism.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31129</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31129"/>
		<updated>2010-03-04T03:03:56Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 //first pipeline operation&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;     // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes pipelined action taken by object    &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 //second pipeline operation&lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);  //execute the actions of this filter in order&lt;br /&gt;
     my_a(a);  // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   //first pipeline action&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
   &lt;br /&gt;
   //second pipeline action&lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
   &lt;br /&gt;
   //execute pipelined operations n-1 times&lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code.  Reduction can be performed for addition and multiplication along with logical operators such as AND and OR.  In TBB, reduction can be specified using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;   // make local copies of constructor arguments&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   // operator function describes pipelined action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   // join performs the final summation of the individual summations&lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)  // set value of local vars to arguments&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   //perform sum by reduction for n-1 iterations&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31127</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31127"/>
		<updated>2010-03-04T02:49:23Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;   // make local copies of constructor arguments&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   // operator function describes parallel action taken by object &lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);   // set value of local vars to arguments&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   // calls SimpleLoop.operator() on range 0 to n-1&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31124</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31124"/>
		<updated>2010-03-04T02:44:52Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 10 and 11.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 25 - 29.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] on pages 18-20.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31093</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31093"/>
		<updated>2010-03-03T23:04:23Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [http://www.threadingbuildingblocks.org/documentation.php Intel&amp;amp;reg;'s open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] found in their documentation.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] found in their documentation.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] found in their documentation.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.  This example is adapted from the [http://software.intel.com/sites/products/documentation/hpc/tbb/tutorial.pdf Intel&amp;amp;reg; TBB Tutorial] found in their documentation.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31089</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31089"/>
		<updated>2010-03-03T22:59:37Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [Intel's open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix Threads&amp;quot;, Human Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31088</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=31088"/>
		<updated>2010-03-03T22:58:29Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#Pthread POSIX thread], also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads pthread_create()] function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads thread] argument is used to provide a unique identifier for the thread you are creating.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads attr] argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads start_routine] argument is the program subroutine that will be executed by the thread being created.  The [https://computing.llnl.gov/tutorials/pthreads/#CreatingThreads arg] argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A [https://computing.llnl.gov/tutorials/pthreads/#MutexOverview mutex variable] is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically [https://computing.llnl.gov/tutorials/pthreads/#MutexCreation create a mutex] variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_lock()] and [https://computing.llnl.gov/tutorials/pthreads/#MutexLocking pthread_mutex_unlock()] functions.  These functions simply lock or unlock the mutex variable specified.&lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
[https://computing.llnl.gov/tutorials/pthreads/#ConVarOverview Conditional variables] allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_wait()] and [https://computing.llnl.gov/tutorials/pthreads/#ConVarSignal pthread_cond_signal()].  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts threads to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.  For more comprehensive information than is covered in this discussion, please see [Intel's open source site].&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is [https://computing.llnl.gov/tutorials/openMP/#Introduction OpenMP]?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: [https://computing.llnl.gov/tutorials/openMP/#ParallelRegion #pragma omp parallel] and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: [https://computing.llnl.gov/tutorials/openMP/#SECTIONS #pragma omp section] (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/6_DD&amp;diff=31064</id>
		<title>CSC/ECE 506 Spring 2010/6 DD</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/6_DD&amp;diff=31064"/>
		<updated>2010-02-27T04:40:45Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* '''Inclusion Property''' */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=='''MULTICORE Architecture'''==&lt;br /&gt;
&lt;br /&gt;
The term '''multicore''' means two or more independent cores. The term '''architecture''' means the design/layout of a particular idea. Hence the combined term multicore architecture simply means the integration of cores and its layout. For e.g. '''dual core''' means each processor contains 2 independent cores.&lt;br /&gt;
&lt;br /&gt;
=='''NUMA architecture'''==&lt;br /&gt;
The term '''NUMA''' is an acronym for Non Uniform Memory Access. It is one way of organizing the '''MIMD''' (Multiple Instruction Multiple Data: independent processors are connected together to form a multiprocessor system.) architecture. NUMA is also known as distributed shared memory. Each processor has an independent cache, memory and they are connected to a common interconnection network like a crossbar, mesh, hypercube, etc.&lt;br /&gt;
&lt;br /&gt;
A '''crossbar''' is a switch which can connect processors to memory in any order. However, it becomes expensive to use, hence as a compromise between crossbar and shared bus sytems, multistage interconnection networks are used (e.g. network on chip interconnects)&lt;br /&gt;
&lt;br /&gt;
=='''CACHE'''== &lt;br /&gt;
'''Cache'' is a storage device having small memory. Why were caches introduced? &amp;lt;sup&amp;gt;[1]&amp;lt;/sup&amp;gt;  &amp;lt;br /&amp;gt;&lt;br /&gt;
The CPU works at clock frequencies of 3 GHZ and most common RAM speeds are about 400 MHz.  Hence, there needs to be a transfer of data to and fro from RAM and CPU. This will increase the latency and CPU will remain idle most times. The main reason of introducing caches was to reduce this latency to access the data from main memory. These small buffers (caches) operate at higher speeds than normal RAM and hence data can be accessed quickly from the cache.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;[[Image: memchart.jpg|120p]]&amp;lt;/center&amp;gt;	 &amp;lt;br /&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;Memory Hierarchy&amp;lt;sup&amp;gt;[2]&amp;lt;/sup&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
The cache which is closest to CPU is called Level 1 Cache or L1 cache. It is also known as primary cache. The size is between 8 KB – 128 KB. The next level is L2 cache (secondary cache) whose size is between 256 KB –1024 KB.  If the data that CPU is looking for is not found in registers, it seeks it in L1 cache, if not found, then in L2 cache,then to the main memory, then finally to an             external storage device. &lt;br /&gt;
&lt;br /&gt;
This leads us to the introduction of two new terms: cache hit and cache miss.&lt;br /&gt;
The data searched for by registers, if found in cache, results in cache hit, and if not found in cache, results in cache miss.&lt;br /&gt;
With the introduction of multiple cache levels, it can be individually stated as L1 cache hit/miss, L2 cache hit/miss and so on.&lt;br /&gt;
&lt;br /&gt;
Most modern computers use the following memory hierarchy.  The figure below shows an uni processor memory hierarchy &amp;amp; a recent dual core memory hierarchy&lt;br /&gt;
&lt;br /&gt;
         &lt;br /&gt;
[[Image:Picture4.png]] 	                                              [[Image: new.png]] &amp;lt;br /&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;Uniprocessor memory hierarchy (L) &amp;lt;sup&amp;gt;[3]&amp;lt;/sup&amp;gt; &amp;amp; Dual core memory hierarchy (R) &amp;lt;sup&amp;gt;[4]&amp;lt;/sup&amp;gt;&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Previously the processors had an unified L1 cache i.e. data and instruction cache were unified in one. If the Instruction fetch unit wants to fetch data from the cache and the memory unit wants to load an address simultaneously, then the memory operation is given priority and the fetch unit is stalled. Hence, L1 cache was split into L1 data cache (L1 – D) and L1 instruction cache (L1 -I) so that it helps in concurrency. One more advantage is fewer number of ports  will be required.&lt;br /&gt;
L2 cache is still unified. As the latency difference between main memory and the fastest cache has become larger, some processors have begun to utilize L3 cache, some don’t. Hence the  above figure shows 2 boxes, one which includes L1 and L2, and the other which combines L3 with them. The L3 cache is roughly about 4 MB – 32 MB in size.&lt;br /&gt;
&lt;br /&gt;
In addition to these 3 levels of cache, some multicore processors use: Inclusion property, Victim cache, Translation LookAside Buffer, Trace Cache, Smart Cache (Intel)&lt;br /&gt;
&lt;br /&gt;
=== '''Inclusion Property''' ===&lt;br /&gt;
Here, the larger lower level cache (L2) includes all the data of upper level (L1) cache. The advantage of this property is that if a value is not found in the lower level cache, it is guaranteed it is not present in the upper level cache, which reduces the amount of checking. The disadvantage is if L2 cache is smaller in size, most of it contains only L1 cache data and hardly any new data, defeating the purpose of another level of cache altogether.&lt;br /&gt;
&lt;br /&gt;
=== '''Victim Cache''' ===&lt;br /&gt;
A '''victim cache''' is a cache used to hold blocks evicted from a CPU cache upon replacement. When main cache evicts a block, the victim cache will take the evicted block. This block is called the victim block. When the main cache misses, it searches the victim cache for recently evicted blocks. So, a hit in victim cache means the main cache doesn’t have to go to the next level of memory.&lt;br /&gt;
&lt;br /&gt;
=== '''Translation LookAside Buffer''' ===&lt;br /&gt;
The '''TLB''' is a small cache holding recently used virtual-to-physical address translations. TLB is a hardware cache that sits alongside the L1 instruction and data caches. &lt;br /&gt;
&lt;br /&gt;
=== '''Trace Cache''' ===&lt;br /&gt;
The '''trace cache''' is located between the decoder unit and the execution unit. Thus the fetch unit grabs data directly from L2 memory cache. &lt;br /&gt;
&lt;br /&gt;
=== '''Smart Cache''' ===&lt;br /&gt;
Intel Advanced '''Smart Cache'''&amp;lt;sup&amp;gt;[5]&amp;lt;/sup&amp;gt; is a multi-core optimized cache that improves performance  by increasing the probability that each execution core of a dual-core processor can access data from a higher-performance, more-efficient cache subsystem. To accomplish this, Intel Core micro architecture shares the Level 2 (L2) cache between the cores. If one of the cores is inactive, the other core will have access to the full cache. It  provides a peak transfer rate of 96 GB/sec (at 3 GHz frequency).&lt;br /&gt;
&lt;br /&gt;
=== '''Parameters of a cache:''' ===&lt;br /&gt;
1. Size:  cache data storage. &amp;lt;br /&amp;gt;&lt;br /&gt;
2. Associativity:  Number of blocks in a set &amp;lt;br /&amp;gt;&lt;br /&gt;
3. Block Size: Total number of bytes in a single block. &amp;lt;br /&amp;gt;&lt;br /&gt;
A set is a fancy term used for the row of a cache.&lt;br /&gt;
&lt;br /&gt;
=='''PLACEMENT POLICY'''==&lt;br /&gt;
The placement policy is the first step in managing a cache. This decides where a block in memory can be placed in the cache. The placement policies are:&lt;br /&gt;
&lt;br /&gt;
=== '''Direct mapped''' ===&lt;br /&gt;
If each block has only one place where it can appear in a cache, it is called direct- mapping. It is also known as one way set associative.&amp;lt;sup&amp;gt;[6]&amp;lt;/sup&amp;gt; &amp;lt;br /&amp;gt;&lt;br /&gt;
The main memory address (32 bits) is broken down as: (Assumed 8 blocks and 8 lines in cache.)&lt;br /&gt;
&lt;br /&gt;
[[Image:table3.jpg]] &amp;lt;br /&amp;gt;&lt;br /&gt;
Number of offset bits = log 2 (Block Size) &amp;lt;br /&amp;gt;&lt;br /&gt;
Number of index bits = log 2 (Number of sets) &amp;lt;br /&amp;gt;&lt;br /&gt;
Number of Tag bits = 32- offset bits- index bits &amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Each line has its own tag associated with it. The index tells which row of cache to look at. To search for a word in the cache, compare the tag bits of the address with the tag of the line. If it matches, the block is in the cache. It is a cache hit. Then select the corresponding word from that line.&lt;br /&gt;
&lt;br /&gt;
=== '''Fully Associative''' ===&lt;br /&gt;
If any line can store the contents of any memory location, it is known as fully associative cache. Hence, no index bits are required. &amp;lt;br /&amp;gt;&lt;br /&gt;
The main memory address (32 bits) is broken down as:&lt;br /&gt;
&lt;br /&gt;
[[Image:table2.jpg]]&lt;br /&gt;
&lt;br /&gt;
It is also known as Content Addressable Memory. To search for a word in the cache, compare the tag bits of the address with the tag of all lines in the cache simultaneously. If any one matches, the word is in the cache. Then select the corresponding word from that line.&lt;br /&gt;
&lt;br /&gt;
=== '''Set Associative''' ===&lt;br /&gt;
&lt;br /&gt;
[[Image:table4.jpg]] &lt;br /&gt;
&lt;br /&gt;
Hence, a compromise between the fully associative cache and the direct mapping technique is made, known as set associative mapping. Almost 90% of the current processors use this mapping. &lt;br /&gt;
Here the cache is divided into ‘s’ sets, where s is a power of 2. The associativity can be found by using the following relation:&lt;br /&gt;
Associativity = Cache size/ (Block Size * Number of sets).  For an associativity of ‘n’, the cache is ‘n- way’ set associative.&lt;br /&gt;
 &lt;br /&gt;
 &amp;lt;center&amp;gt;[[Image: Picture15.png]]&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;2 way set associative mapping &amp;lt;sup&amp;gt;[3]&amp;lt;/sup&amp;gt;&amp;lt;/center&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Here the cache size is the same, but the index field is 1 bit shorter. Based on the index bits, select a particular set. To search for a word in the cache, compare the tag bits of the address with the tag bits of the line in each set (here 2). If it matches with any one of the tags, it is a cache hit. Else, it is a cache miss and the word has to be fetched from main memory.&lt;br /&gt;
&lt;br /&gt;
Conceptually, the direct mapped and fully associative caches are just &amp;quot;special cases&amp;quot;&amp;lt;sup&amp;gt;[6]&amp;lt;/sup&amp;gt; of the N-way set associative cache. If you make n=1, it becomes a &amp;quot;1-way&amp;quot; set associative cache. i.e. there  is only one line per set, which is the same as a direct mapped cache. On the other hand, if you make ‘n’ to be equal to the number of lines in the cache, then you only have one set, containing all of the cache lines, and each and every memory location points to that set. This gives us a fully associative cache.&lt;br /&gt;
&lt;br /&gt;
=='''REPLACEMENT POLICY'''==&lt;br /&gt;
This decides which block needs to be evicted to make space available for a new block entry. The replacement policies used are: LRU, Psuedo LRU, FIFO, OPT, random, etc. &amp;lt;br /&amp;gt;&lt;br /&gt;
The most commonly used is LRU (Least recently used replacement policy). It works as follows: &amp;lt;sup&amp;gt;[3]&amp;lt;/sup&amp;gt; &lt;br /&gt;
&lt;br /&gt;
1. Make small counters per block in a set and the number of bits in each counter = &lt;br /&gt;
log2 (Associativity) &amp;lt;br /&amp;gt;&lt;br /&gt;
2. If there is a cache hit, make the current block’s counter value as 0, which indicates that it is the most recently used one. Increment the counters of the other blocks whose counter values are less than the referenced block’s old counter value &amp;lt;br /&amp;gt;&lt;br /&gt;
3.If there is a cache miss, make the newly allocated block’s counter value as 0 and then increment the counters of all other blocks. &amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
One eg: trace ABCEF. (0-2 indicate counter numbers)&lt;br /&gt;
&amp;lt;center&amp;gt;&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
|  0&lt;br /&gt;
|  1&lt;br /&gt;
|  2&lt;br /&gt;
|-&lt;br /&gt;
|  2&lt;br /&gt;
|  0 &amp;lt;br /&amp;gt; B&lt;br /&gt;
|  1 &amp;lt;br /&amp;gt; A&lt;br /&gt;
|-&lt;br /&gt;
|  0 &amp;lt;br /&amp;gt; C&lt;br /&gt;
|  1 &amp;lt;br /&amp;gt; B&lt;br /&gt;
|  2 &amp;lt;br /&amp;gt; A&lt;br /&gt;
|-&lt;br /&gt;
|  1 &amp;lt;br /&amp;gt; C&lt;br /&gt;
|  2 &amp;lt;br /&amp;gt; B&lt;br /&gt;
|  0 &amp;lt;br /&amp;gt; E&lt;br /&gt;
|-&lt;br /&gt;
|  2 &amp;lt;br /&amp;gt; C&lt;br /&gt;
|  0 &amp;lt;br /&amp;gt; F&lt;br /&gt;
|  1 &amp;lt;br /&amp;gt; E&lt;br /&gt;
|}&amp;lt;/center&amp;gt;&lt;br /&gt;
	&lt;br /&gt;
When E came in, the block with the highest count here 2, block A was evicted since it was least recently used and the value of other counters were incremented accordingly. Similarly, when F came in, block B was evicted. &lt;br /&gt;
&lt;br /&gt;
FIFO policy simply removes the cache block which was brought in earliest in the cache, irrespective of it being least, most recently used.&lt;br /&gt;
&lt;br /&gt;
=='''WRITE POLICY'''== &lt;br /&gt;
&lt;br /&gt;
=== '''Write Through''' ===&lt;br /&gt;
Whenever a value is written in the cache, it is immediately written in the next level of memory hierarchy. The main advantage is that both levels have the recent updated value at all times. No dirty bit is required. The disadvantage is that since it writes every time, it consumes a lot of bandwidth. A bit stored in the cache can flip its value when struck by alpha particles resulting in soft errors&amp;lt;sup&amp;gt;[4]&amp;lt;/sup&amp;gt;. The effect of this is that the value in the cache is lost. Since, the lower level of cache contains the same updated value, it is safe to discard the block, as it can be easily re fetched. So, only error detection is enough.&lt;br /&gt;
&lt;br /&gt;
=== '''Write Back''' ===&lt;br /&gt;
The next level of memory hierarchy is not updated as soon as a value is written in the cache, but is updated only when a block is evicted from the cache. To cope up with the issue of having the most recent value, a dirty bit is associated with the cache block. When there is a write to the cache, the dirty bit is set 1. So, if this block were to be evicted, indication of dirty bit being 1, will make the block to be written back to the main memory. When a new block is written in cache, the dirty bit is made 0. Since, lower level of cache does not contain the same updated value, error correction is necessary. &lt;br /&gt;
&lt;br /&gt;
Thus in most processors, L1 cache uses a WT policy and L2 cache uses a WB policy.&lt;br /&gt;
&lt;br /&gt;
=== '''Write Allocate policy'''===&lt;br /&gt;
On a write miss, the block is brought back into the cache before it is written. A write back policy generally uses write allocate. &lt;br /&gt;
&lt;br /&gt;
=== '''Write No Allocate Policy''':===&lt;br /&gt;
On a write miss, the write is propagated to the lower level of memory hierarchy but it is not brought back into the cache. A write through policy generally uses write no allocate.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=='''RECENT MULTI-CORE ARCHITECTURES'''==&lt;br /&gt;
Cores in multi-core systems may implement architectures such as superscalar, VLIW, vector processing, SIMD, or multithreading. &lt;br /&gt;
&lt;br /&gt;
==='''Super scalar'''===&lt;br /&gt;
'''Superscalar''' architecture implements Instruction level parallelism within a single processor. A superscalar processor executes more than one instruction during a clock cycle. &amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==='''Very long instruction word (VLIW)'''===&lt;br /&gt;
'''VLIW'''&amp;lt;sup&amp;gt;[7]&amp;lt;/sup&amp;gt; refers to a CPU architecture designed to take advantage of instruction level parallelism (ILP). It includes out of order execution. It is a type of MIMD. VLIW CPUs offer significant computational power with less hardware complexity (but greater compiler complexity) than is associated with most superscalar CPUs. One VLIW instruction encodes multiple operations; specifically, one instruction encodes at least one operation for each execution unit of the device. &amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== ''' Simultaneous Multithreading''' ===&lt;br /&gt;
'''SMT''' is a technique for improving the overall efficiency of superscalar CPUs with hardware multithreading. SMT permits many independent threads of execution to make better utilization of the resources provided by modern processor architectures.&amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The following table shows the type of cache structure used by recent processors.&amp;lt;sup&amp;gt;[8]&amp;lt;/sup&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:table31.jpg]]            [[Image:table32.jpg]]  &amp;lt;br /&amp;gt;&lt;br /&gt;
The above tables are not directly cited from anywhere, information has been found and framed.&lt;br /&gt;
	&lt;br /&gt;
== '''RECENT NUMA ARCHITECTURES''' ==&lt;br /&gt;
&lt;br /&gt;
1. 4 '''Opteron 875''' dual core processors: 1 L1- D cache per core (2 per socket), size: 65 KB&lt;br /&gt;
,1 L1- I cache per core, size: 65 KB&lt;br /&gt;
,1 L2 cache per core (2 per socket), size: 1MB&lt;br /&gt;
&lt;br /&gt;
2.AMD '''Hammer''':  L1: 128 KB, 2 way set assoc &amp;amp; L2: 1 MB.&lt;br /&gt;
It also uses a large TLB and a victim cache.	&lt;br /&gt;
&lt;br /&gt;
3. Altix:  The '''Altix UV''' supercomputer architecture combines a development of the NUMAlink interconnect used in the Altix 4000 (NUMAlink 5) with 4/6/8-core &amp;quot;Nehalem-EX&amp;quot; Intel Xeon processors. Altix UV systems run either SuSE Linux Enterprise Server or Red Hat Enterprise Linux, and scale from 32 to 2,048 cores with support for up to 16 TB of shared memory in a single system.  It came in the end of 2009.&lt;br /&gt;
&lt;br /&gt;
Other examples are: IBM p690, SGI Origin. Here are two different structures.&lt;br /&gt;
&lt;br /&gt;
Tilera TILE 64 consists of a mesh network of 64 tiles. It is based on VLIW instruction set. Each of the cores (tile) has its own L1 and L2 cache plus an overall virtual L3 cache, having mutual inclusion property. It was developed in 2007. &amp;lt;br /&amp;gt;&lt;br /&gt;
Tilera, TILE- Gx is a future multi core processor consisting of a mesh network of 100 cores. It is a 64 bit core (3 issue)&lt;br /&gt;
L1- I: 32 KB/ core, L1- D: 32 KB/ core, L2: 256 KB/ core, L3: 26 MB/ chip.&lt;br /&gt;
&lt;br /&gt;
== '''WRITE POLICIES used in recent multi core architectures''' == &lt;br /&gt;
&lt;br /&gt;
1. '''Intel Pentium Processors''': &amp;lt;sup&amp;gt;[10]&amp;lt;/sup&amp;gt; Mainly 2 bits control the cache. CD (cache disable) and NW (not write through)&lt;br /&gt;
&amp;lt;center&amp;gt;CD= 1: cache disabled, CD=0: enables the cache&amp;lt;/center&amp;gt;&lt;br /&gt;
&amp;lt;center&amp;gt;NW= 0: Write Through, NW=1: Write Back&amp;lt;/center&amp;gt;&lt;br /&gt;
A WB &amp;lt;sup&amp;gt;[11]&amp;lt;/sup&amp;gt; line is always fetched into the cache if a cache miss occurs. &amp;lt;br /&amp;gt;&lt;br /&gt;
A WT &amp;lt;sup&amp;gt;[11]&amp;lt;/sup&amp;gt;line is not fetched into the cache on a write miss. For the Pentium Pro processor, a WT hit to the L1 cache updates the L1 cache. A WT hit to L2 cache invalidates the L2 cache. &amp;lt;br /&amp;gt;&lt;br /&gt;
A WP &amp;lt;sup&amp;gt;[11]&amp;lt;/sup&amp;gt; ('''Write Protected''') line is cacheable, but a write to it cannot modify the cache line. A WP line is not fetched into the cache on a write miss. For the Pentium Pro processor, a WP hit to the L2 cache invalidates the line in the L2 cache. &amp;lt;br /&amp;gt;&lt;br /&gt;
An UC &amp;lt;sup&amp;gt;[11]&amp;lt;/sup&amp;gt; ('''Uncacheable''') line is not put into the cache. A UC hit to the L1 or L2 cache invalidates the entry.&lt;br /&gt;
Cache consistency: MESI (Modified, Exclusive, Shared, Invalid) policy is used.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
2.'''Intel IA 32 IA64 architecture''': &amp;lt;sup&amp;gt;[12]&amp;lt;/sup&amp;gt; In IA 32, the number of processor sharing the cache is one plus value represented by CPUID.4.EAX [25:14] &lt;br /&gt;
&lt;br /&gt;
'''Write combining''': Successive writes to the same cache line are combined. &amp;lt;br /&amp;gt;&lt;br /&gt;
'''Write collapsing''': Successive writes to the same bytes result in only the last write being visible. &amp;lt;br /&amp;gt;&lt;br /&gt;
'''Weakly ordered''': No ordering is preserved between WC stores or between other loads or stores. &amp;lt;br /&amp;gt;&lt;br /&gt;
'''Uncacheable &amp;amp; Write No Allocate''': Stored data is written around the cache and it   uses Write No Allocate. &amp;lt;br /&amp;gt;&lt;br /&gt;
'''Non-temporal''': Data which is referenced once and not reused in the immediate future&lt;br /&gt;
&lt;br /&gt;
If the programmer specifies a non-temporal store to UC or WP memory types, then the store behaves like an uncacheable store. There is no conflict if a non temporal store to WC is specified. If the programmer specifies a non-temporal store to WB or WT memory types, and if the data is in the cache, will ensure consistency. However, if it is a miss, the transaction is subject to all WC memory semantics, which will not write allocate.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''3'''.Intel's '''Celeron''' and '''Xeon''' based on '''Netburst microarchitecture''' (i.e. based on Pentium 4) will have the same specs as Pentium 4 but with different L2 cache size, while Celeron and Xeon based on '''Core microarchitecture''' (i.e. based on Core 2 Duo) will have the same specs as Core 2 Duo but with different L2 cache size. &amp;lt;br /&amp;gt;&lt;br /&gt;
Intel Xeon Clowertown: L1: 8 way associative &amp;amp; L2: 16 way associative. &amp;lt;br /&amp;gt;&lt;br /&gt;
Intel Celeron using Intel Core 2 architecture:  L2: 16 way associative &amp;amp; L1- D: 8 way associative &amp;amp; L1- I: 8 way associative.								                It also has 16 entry DTLB. &amp;lt;br /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
4. Intel '''RAID''': These use Write Back policy. &amp;lt;br /&amp;gt;&lt;br /&gt;
5. Intel Core microarchitecture uses FIFO replacement policy. &amp;lt;br /&amp;gt;&lt;br /&gt;
6. AMD uses cache exclusion unlike Intel’s cache inclusion &amp;lt;br /&amp;gt;&lt;br /&gt;
7. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &amp;lt;br /&amp;gt;&lt;br /&gt;
8. In general, WT invalidate protocols gives better performance than WB- MESI protocols. But most modern processors use WB policy.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=='''REFERENCES'''==&lt;br /&gt;
&lt;br /&gt;
[1]:  http://www.karbosguide.com/books/pcarchitecture/chapter10.htm &amp;lt;br /&amp;gt;&lt;br /&gt;
[2]: http://www.real-knowledge.com/memory.htm &amp;lt;br /&amp;gt;&lt;br /&gt;
[3]: Computer Design &amp;amp; Technology- Lectures slides by Prof.Eric Rotenberg &amp;lt;br /&amp;gt;&lt;br /&gt;
[4]: Fundamentals of Parallel Computer Architecture by Prof.Yan Solihin  &amp;lt;br /&amp;gt;&lt;br /&gt;
[5] : http://download.intel.com/technology/architecture/sma.pdf &amp;lt;br /&amp;gt;&lt;br /&gt;
[6]: http://www.pcguide.com/ref/mbsys/cache/funcMapping-c.html &amp;lt;br /&amp;gt;&lt;br /&gt;
[7]:http://en.wikipedia.org/wiki/Very_Long_Instruction_Word &amp;lt;br /&amp;gt;&lt;br /&gt;
[8]: http://en.wikipedia.org/wiki/Multi-core_processor &amp;lt;br /&amp;gt;&lt;br /&gt;
[9]: http://en.wikipedia.org/wiki/Power_Architecture &amp;lt;br /&amp;gt;&lt;br /&gt;
[10]: http://www.intel.com/design/intarch/papers/cache6.pdf  &amp;lt;br /&amp;gt;&lt;br /&gt;
[11]: http://download.intel.com/design/archives/processors/pro/docs/24269001.pdf &amp;lt;br /&amp;gt;&lt;br /&gt;
[12]: http://www.intel.com/Assets/PDF/manual/248966.pdf &amp;lt;br /&amp;gt;&lt;br /&gt;
[13]: http://www.intel.com/pressroom/kits/quickreffam.htm &amp;lt;br /&amp;gt;&lt;br /&gt;
[14]: http://www.hardwaresecrets.com &amp;lt;br /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30934</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30934"/>
		<updated>2010-02-24T15:45:34Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOPIPE Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
   &lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30933</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30933"/>
		<updated>2010-02-24T15:45:05Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOPIPE Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Exploiting DOPIPE parallelism is more complicated than DOALL parallelism in TBB.  The '''pipeline''' and '''filter''' classes are used to implement the pipeline pattern.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop amenable to DOPIPE parallelism&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = a[i-1] + 5.3;&lt;br /&gt;
   b[i] = a[i] + 2.7;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the DOPIPE code below, each section of code that can be separated into a chunk should be defined in a different '''filter''' class instance.  For the sequential code example, there are two separate statements that can be executed in parallel; therefore there are two '''filter''' instances created.  In order to run them in parallel, a '''pipeline''' instance needs to be created and given passed the instances of the created '''filter''' instantiations.  Now, simply run the pipeline over '''n''' iterations and you have pipelined parallelism in TBB.  Notice that the '''operator()''' function is called to actually perform the code execution as for all the previous TBB examples.&lt;br /&gt;
&lt;br /&gt;
 class Filter1: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_a[i] = my_a[i-1] + 5.3;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter1(float *a) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 class Filter2: public filter {&lt;br /&gt;
   float* my_a;&lt;br /&gt;
   float* my_b;&lt;br /&gt;
 public:   &lt;br /&gt;
   void operator() {&lt;br /&gt;
     my_b[i] = my_a[i] + 2.7;&lt;br /&gt;
   }&lt;br /&gt;
   &lt;br /&gt;
   Filter2(float *a, float *b) :&lt;br /&gt;
     filter(serial_in_order);&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 // pipeline creation and execution&lt;br /&gt;
 void main(void) {&lt;br /&gt;
   size_t n;&lt;br /&gt;
   float a[], b[];&lt;br /&gt;
   ...&lt;br /&gt;
   Filter1 filter1(a);&lt;br /&gt;
   pipeline.add_filter(filter1);&lt;br /&gt;
 &lt;br /&gt;
   Filter2 filter2(a,b);&lt;br /&gt;
   pipeline.add_filter(filter2);&lt;br /&gt;
 &lt;br /&gt;
   pipeline.run(n);&lt;br /&gt;
&lt;br /&gt;
   pipeline.clear();&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30932</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30932"/>
		<updated>2010-02-24T15:03:30Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Text&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30931</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30931"/>
		<updated>2010-02-24T04:02:48Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Text&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
== Flexibility ==&lt;br /&gt;
&lt;br /&gt;
TBB also boasts of the ability to combine its package with other threading packages, such as OpenMP and Posix threads.  This allows the programmer more flexibility to mix and match packages to obtain, for example, the functional parallelism benefits of OpenMP 2.0 and still retain the ability to use reduction in TBB.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30930</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30930"/>
		<updated>2010-02-24T03:50:40Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism are not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== DOPIPE Parallelism ==&lt;br /&gt;
&lt;br /&gt;
Text&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30929</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30929"/>
		<updated>2010-02-24T03:46:58Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This function takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30928</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30928"/>
		<updated>2010-02-24T03:46:44Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Reduction */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  This function takes two parameters just like we have seen from '''parallel_for()''' above.  The first is the range of indices of the loop that have an operation to reduce in parallel. The second is a set of operations that can be processed as a unit and are safe to run concurrently.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30927</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30927"/>
		<updated>2010-02-24T03:44:45Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], float b[];&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30926</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30926"/>
		<updated>2010-02-24T03:44:03Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Reduction */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
Reduction is a way of removing conflicts in parallel code and can be specified in TBB using the '''parallel_reduce()''' construct.  The code example described in this section involves calculating a sum of variables in an array.  The sequential code version is shown below.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Instead of using a single '''sum''' variable, in TBB partial sums are calculated in parallel and then summed together to form a total sum at the end.&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(SimpleSum&amp;amp; x, split): my_a(x.my_a), my_sum(0) {}&lt;br /&gt;
 &lt;br /&gt;
   void join(const SimpleSum&amp;amp; y) {&lt;br /&gt;
     my_sum += y.my_sum;&lt;br /&gt;
   }&lt;br /&gt;
 &lt;br /&gt;
   SimpleSum(float a[]) :&lt;br /&gt;
     my_a(a), my_sum(0)&lt;br /&gt;
   {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float a[], sum;&lt;br /&gt;
   int n;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   SimpleSum ss(a);&lt;br /&gt;
   parallel_reduce(blocked_range&amp;lt;size_t&amp;gt;(0,n), ss);&lt;br /&gt;
   sum = ss.my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
In the version of the code using '''parallel_reduce()''', there a few changes to the format used in the '''parallel_for''' example above that are slightly more complex.  A class must be created to perform the parallel operations as before but this time the splitting and combining operations of reduction must be specified.  The splitting is found in the constructor and the joining is found in the '''join()''' function as a member of the class.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30923</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30923"/>
		<updated>2010-02-24T03:20:42Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30922</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30922"/>
		<updated>2010-02-24T03:18:48Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 &amp;lt;source lang=cpp&amp;gt;// parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30921</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30921"/>
		<updated>2010-02-24T03:16:38Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  A simple sequential code loop is showed below followed by the same code being executed using parallel iterations with DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + 5;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism in TBB.  First of all, a class object must be created with a public function called '''operator()'''.  This function should have a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created for any arguments are passed for '''parallel_for''' to operate properly.  This is further explained by the example parallel loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     '''for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;'''&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30920</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30920"/>
		<updated>2010-02-24T03:13:06Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Instead they learn how to use the API for the TBB library with new or existing C++ code.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code with a simple extra command.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS parallelism is not available for explicit usage using the library.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created, if any arguments are passed, for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30919</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30919"/>
		<updated>2010-02-24T03:10:28Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According to the [http://software.intel.com/en-us/intel-tbb/ Intel&amp;amp;reg; Software site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created, if any arguments are passed, for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30918</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30918"/>
		<updated>2010-02-24T03:09:21Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created, if any arguments are passed, for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;POSIX Threads Programming&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/pthreads/ https://computing.llnl.gov/tutorials/pthreads/], January 2010.&lt;br /&gt;
&lt;br /&gt;
* S. Bydiec, &amp;quot;Programming Posix ThreadsHuman Factor website[http://www.humanfactor.com/pthreads/ http://www.humanfactor.com/pthreads/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* OpenMP API website, [http://openmp.org/wp/ http://openmp.org/wp/]&lt;br /&gt;
&lt;br /&gt;
* Blaise Barney, &amp;quot;OpenMP&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/openMP/ https://computing.llnl.gov/tutorials/openMP/], January 2009.&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30917</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30917"/>
		<updated>2010-02-24T03:01:11Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exists, all that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created, if any arguments are passed, for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ...&lt;br /&gt;
  #pragma omp parallel&lt;br /&gt;
  {&lt;br /&gt;
    ...&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        A[i] = A[i] + B[i];&lt;br /&gt;
    #pragma omp section&lt;br /&gt;
    for(i=0;i&amp;lt;n;i++)&lt;br /&gt;
        C[i]=C[i-1] + 5;  &lt;br /&gt;
  } &lt;br /&gt;
&lt;br /&gt;
The code block above shows how to exploit function parallelism.  Each section compiler directive tell the following code to be executed on a single thread.  The code above is broken into 2 sections, which mean the two for loops are handled by 2 threads.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30889</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30889"/>
		<updated>2010-02-23T22:59:00Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code block above shows roughly how use conditional variable in order to exploit DOACROSS parallelism.&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     2&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     double      c[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     int i;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       data-&amp;gt;a[i] = (double) data-&amp;gt;a[i - 1] + (double) data-&amp;gt;b[i];&lt;br /&gt;
       pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
       pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
   &lt;br /&gt;
  void *Bar(void *data)&lt;br /&gt;
  {&lt;br /&gt;
      int i;&lt;br /&gt;
      for(i = 0; i &amp;lt; NUM_ELEMENTS; i++)&lt;br /&gt;
     {&lt;br /&gt;
       pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
       pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
       data-&amp;gt;c[i] = (double) data-&amp;gt;a[i];&lt;br /&gt;
     }&lt;br /&gt;
  }&lt;br /&gt;
  &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
    pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
    pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     &lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[0], NULL, Foo, (void *)data);  // create pthread that runs Foo&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[1], NULL, Bar, (void *)data);  // create pthread that runs Bar&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created and have each thread execute some data set for a single function.  If you you refer back up to the pthread creation example, that code block exploits DOALL parallelism.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
== Overhead ==&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a constructor must also be created, if any arguments are passed, for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
   float *a = my_a;&lt;br /&gt;
   float *b = my_b;&lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *b = my_b;&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt;= r.end(); i++)&lt;br /&gt;
       a[i] = b[i] + 5;&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(float *a, float *b):&lt;br /&gt;
     my_a(a);&lt;br /&gt;
     my_b(b);&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   float *a, float *b;&lt;br /&gt;
   ...&lt;br /&gt;
   ...&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop(a, b));&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Reduction ==&lt;br /&gt;
&lt;br /&gt;
 // sequential loop summation&lt;br /&gt;
 float sum = 0;&lt;br /&gt;
 for(size_t i = 0; i &amp;lt; n; i++)&lt;br /&gt;
   sum += a[i];&lt;br /&gt;
&lt;br /&gt;
Explanation&lt;br /&gt;
&lt;br /&gt;
 // parallel reduction using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleSum {&lt;br /&gt;
   float *my_a;&lt;br /&gt;
 public:&lt;br /&gt;
   float *my_sum;&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     float *a = my_a;&lt;br /&gt;
     float *sum = my_sum;&lt;br /&gt;
     size_t end = r.end();&lt;br /&gt;
     for(size_t i = r.begin(); i &amp;lt; r.end(); i++)&lt;br /&gt;
       sum += a[i];&lt;br /&gt;
     my_sum = sum;&lt;br /&gt;
   }&lt;br /&gt;
 SimpleSum()&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;lt;omp.h&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
 main ()  {&lt;br /&gt;
 &lt;br /&gt;
 int nthreads, tid;&lt;br /&gt;
 &lt;br /&gt;
 /* Fork a team of threads with each thread having a private tid variable */&lt;br /&gt;
 #pragma omp parallel private(tid)&lt;br /&gt;
  {&lt;br /&gt;
 &lt;br /&gt;
  /* Obtain and print thread id */&lt;br /&gt;
   tid = omp_get_thread_num();&lt;br /&gt;
   printf(&amp;quot;Hello World from thread = %d\n&amp;quot;, tid);&lt;br /&gt;
 &lt;br /&gt;
    /* Only master thread does this */&lt;br /&gt;
    if (tid == 0) &lt;br /&gt;
     {&lt;br /&gt;
     nthreads = omp_get_num_threads();&lt;br /&gt;
     printf(&amp;quot;Number of threads = %d\n&amp;quot;, nthreads);&lt;br /&gt;
     }&lt;br /&gt;
  &lt;br /&gt;
   }  /* All threads join master thread and terminate */&lt;br /&gt;
  &lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
  #include&amp;lt;omp.h&amp;gt;&lt;br /&gt;
  ....&lt;br /&gt;
  #pragma omp parallel &lt;br /&gt;
  {&lt;br /&gt;
     ...&lt;br /&gt;
     #pragma omp parallel for default(shared)  // DOALL: all iterations are done in parallel&lt;br /&gt;
     for(i=0; i&amp;lt;n; i++)&lt;br /&gt;
        A[i] = B[i]+C[i];&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The simple code block above shows how OpenMP 2.0 creates a parallel region that exploit DOALL parallelism.  Notice how much this resemble sequential programming.  If the compiler directives were taken out, this code could be run on a sequential machine.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30881</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30881"/>
		<updated>2010-02-23T22:33:17Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
  &lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
   &lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
   &lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
   &lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object must be created with a public function called '''operator()''' with a parameter of the form &amp;quot;const blocked_range&amp;lt;size_t&amp;gt;&amp;amp;&amp;quot; and a simple constructor must also be created in order for '''parallel_for''' to operate properly.  This is further explained by the example DOALL loop below.  This is a parallel version of the sequential loop above.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
 &lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     for(size_t i = r.begin(); i != r.end; i++)&lt;br /&gt;
       a[i] = b[i] + c[i];&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(): {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop());&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30879</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30879"/>
		<updated>2010-02-23T22:26:48Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* DOALL Parallelism */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function, and note their syntax.  For completeness, the example code above shows a simple function being run by every thread.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     50&lt;br /&gt;
  #define NUM_ELEMENTS    50&lt;br /&gt;
   &lt;br /&gt;
  typedef struct &lt;br /&gt;
   {&lt;br /&gt;
     double      a[NUM_ELEMENTS];&lt;br /&gt;
     double      b[NUM_ELEMENTS];&lt;br /&gt;
     long          threadid;&lt;br /&gt;
   } DATA;&lt;br /&gt;
  &lt;br /&gt;
  pthread_mutex_t mutexvar; //mutex variable required for conditional wait and signal functions&lt;br /&gt;
  pthread_cond_t condvar;  //conditional variable required for conditional wait and signal functions&lt;br /&gt;
&lt;br /&gt;
  void *Foo(void *data)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long) data-&amp;gt;threadid;&lt;br /&gt;
     pthread_mutex_lock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_cond_wait(&amp;amp;condvar, &amp;amp;mutexvar);   //wait until safe to continue&lt;br /&gt;
     data-&amp;gt;a[tid] = (double) data-&amp;gt;a[tid - 1] + (double) data-&amp;gt;b[tid];&lt;br /&gt;
     pthread_cond_signal(&amp;amp;condvar);                    //signal that its safe to continue&lt;br /&gt;
     pthread_mutex_unlock(&amp;amp;mutexvar);&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     DATA data;&lt;br /&gt;
&lt;br /&gt;
     //ASSUME data struct is initialized&lt;br /&gt;
&lt;br /&gt;
    //Initialized Mutex and Conditional Variable&lt;br /&gt;
   pthread_mutex_init(&amp;amp;mutexvar, NULL);&lt;br /&gt;
   pthread_cond_init (&amp;amp;condvar, NULL);           &lt;br /&gt;
&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)data);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the '''parallel_for()''' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object of a certain format must be created in order for the '''parallel_for''' function to operate properly.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 class SimpleLoop {&lt;br /&gt;
 &lt;br /&gt;
 public:&lt;br /&gt;
   void operator(const blocked_range&amp;lt;size_t&amp;gt;&amp;amp; r) {&lt;br /&gt;
     for(size_t i = r.begin(); i != r.end; i++)&lt;br /&gt;
       a[i] = b[i] + c[i];&lt;br /&gt;
   }&lt;br /&gt;
   SimpleLoop(): {}&lt;br /&gt;
 };&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), SimpleLoop());&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30877</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30877"/>
		<updated>2010-02-23T22:19:16Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use.  In an attempt to simplify the examples below, this overhead will be omitted.&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
A certain amount of overhead is necessary when creating a loop using DOALL parallelism.  First of all, a class object of a certain format must be created in order for the ''parallel_for'' function to operate properly.&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
 void simple_loop() {&lt;br /&gt;
   for(i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
     a[i] = b[i] + c[i];&lt;br /&gt;
   }&lt;br /&gt;
 }&lt;br /&gt;
 &lt;br /&gt;
 void main(void) {&lt;br /&gt;
   parallel_for(blocked_range&amp;lt;size_t&amp;gt;(0,n), simple_loop());&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30876</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30876"/>
		<updated>2010-02-23T22:02:15Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming or learn an entirely new language.  Also, any compiler that supports ISO C++ can compile Intel&amp;amp;reg; TBB code.  However, as a drawback of its simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage here.&lt;br /&gt;
&lt;br /&gt;
In order to use TBB, you must always insert the following code at the beginning of every file to include the TBB library and make its functions and variables available for use:&lt;br /&gt;
&lt;br /&gt;
 #include &amp;quot;tbb/tbb.h&amp;quot;&lt;br /&gt;
 &lt;br /&gt;
 using namespace tbb;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== DOALL Parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
 // sequential loop&lt;br /&gt;
 for(i = 0; i &amp;lt; n; i++) {&lt;br /&gt;
   a[i] = b[i] + c[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
 // parallel loop using Intel&amp;amp;reg; TBB&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30875</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30875"/>
		<updated>2010-02-23T21:54:59Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.  &lt;br /&gt;
&lt;br /&gt;
A goal of the library is similar to that of OpenMP, where the programmer is not required to have an extensive knowledge of thread programming.  However, as a drawback of simplicity, certain types of parallelism such as DOACROSS and DOPIPE parallelism are not available for explicit usage.&lt;br /&gt;
&lt;br /&gt;
== DOALL parallelism ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel loop can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is a set of operations that can be processed as a unit and are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30874</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30874"/>
		<updated>2010-02-23T21:23:37Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Parallel programming reduces execution time versus sequential programming by taking advantage of code structure.  In practice, there are various C/C++ code libraries that offer parallel programming support without needing to learn a new language or programming model.  These include Posix threads, Intel&amp;amp;reg; Threading Building Blocks, and OpenMP.  The focus of this article is to discuss the implementation of these libraries for parallel programming models related to loop structure, specifically DOALL, DOACROSS and DOPIPE parallelism, reduction, and functional parallelism.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
  #include &amp;lt;pthread.h&amp;gt;&lt;br /&gt;
  #include &amp;lt;stdio.h&amp;gt;&lt;br /&gt;
  #define NUM_THREADS     5&lt;br /&gt;
 &lt;br /&gt;
  void *Foo(void *threadid)&lt;br /&gt;
  {&lt;br /&gt;
     long tid;&lt;br /&gt;
     tid = (long)threadid;&lt;br /&gt;
     ....&lt;br /&gt;
     ....&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
 &lt;br /&gt;
  int main (int argc, char *argv[])&lt;br /&gt;
  {&lt;br /&gt;
     pthread_t threads[NUM_THREADS];&lt;br /&gt;
     int rc;&lt;br /&gt;
     long t;&lt;br /&gt;
     for(t=0; t&amp;lt;NUM_THREADS; t++){&lt;br /&gt;
        printf(&amp;quot;In main: creating thread %ld\n&amp;quot;, t);&lt;br /&gt;
        rc = pthread_create(&amp;amp;threads[t], NULL, Foo, (void *)t);  // create pthread&lt;br /&gt;
        if (rc){&lt;br /&gt;
           printf(&amp;quot;ERROR; return code from pthread_create() is %d\n&amp;quot;, rc);&lt;br /&gt;
           exit(-1);&lt;br /&gt;
        }&lt;br /&gt;
     }&lt;br /&gt;
     pthread_exit(NULL);&lt;br /&gt;
  }&lt;br /&gt;
&lt;br /&gt;
The code above shows a very simple example of how to create threads using pthreads.  Notice the arguments passed to the pthread_create() function&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30867</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30867"/>
		<updated>2010-02-23T21:12:04Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* OpenMP 3.0 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
What is OpenMP?  OpenMP or Open Multi-processing is a multi-platform API used for shared addressed space programming.  There are versions of OpenMP for both C/C++ and Fortran.  The OpenMP libraries provide a list of compiler directives that easily allow for one to write shared memory parallel programs.  Before explaining how to exploit the different types of parallelism supported by OpenMP 2.0 it is necessary to understand how to create threads.&lt;br /&gt;
&lt;br /&gt;
==Parallel Region==&lt;br /&gt;
&lt;br /&gt;
In OpenMP 2.0 in order to create threads, the C/C++ compiler directive used is: #pragma omp parallel and is enclosed by curly brackets (See reference for Fortran directive).  This directive means the start of a parallel region, and at the start of a parallel threads are created.  In order to specify the number of threads that are created in a parallel region the function omp_set_num_threads(int n) is called. &lt;br /&gt;
&lt;br /&gt;
Inside a parallel region different compiler directive can be used to exploit different types of parallelism.  The two types of parallelism discussed below are DOALL and function parallelism.  With OpenMP 2.0 DOACROSS and DOPIPE parallelism cannot be expressed through the use of compiler directives.&lt;br /&gt;
&lt;br /&gt;
==DOALL parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP exploits DOALL parallelism through the use of a simple directive that tells the compiler to execute a section of code on multiple threads.  The C/C++ directive is: #pragma omp parallel for. (See reference for Fortran Directive).  This directive if placed before a normal sequential for loop in C/C++ will execute all iterations of the for loop in parallel.&lt;br /&gt;
&lt;br /&gt;
==Function parallelism==&lt;br /&gt;
&lt;br /&gt;
OpenMP 2.0 can also exploit function parallelism with compiler directives.  The C/C++ directive is: #pragma omp section (See reference for Fortran Directive).  This directive is placed before a section of code that is to be executed by a single thread.  For function parallelism, multiple data independent code blocks can each be placed inside a parallel region with the section compiler directive placed before each block of code.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
In May 2008, OpenMP version 3.0 was released as an upgrade with some features such as tasks, synchronization primitives among others. For the scope of this article, nothing significant was upgraded regarding the execution of DOALL, DOACROSS, and DOPIPE parallelism, reduction, or function parallelism.  Therefore, see sections above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30864</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30864"/>
		<updated>2010-02-23T21:04:56Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* OpenMP 3.0 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
Some features such as tasks, synchronization primitives among a few others were included in the upgrades from OpenMP 2.0 but nothing significant was upgraded that fits the scope of this article.  Therefore, see section above about OpenMP 2.0 since all discussions there are applicable to version 3.0.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30863</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30863"/>
		<updated>2010-02-23T21:01:36Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==Creating and Terminating Pthreads==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
==Mutexes==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==Conditional Variables==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE Parallelism==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL Parallelism==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
Some features such as tasks, synchronization primitives among a few others were included in the upgrades from OpenMP 2.0 but nothing significant was upgraded that fits the scope of this article.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
* Mark Bull, University of Edinburgh, &amp;quot;OpenMP 3.0 Overview&amp;quot;, [http://www.compunity.org/futures/Mark_SC06BOF.pdf http://www.compunity.org/futures/Mark_SC06BOF.pdf]&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30861</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30861"/>
		<updated>2010-02-23T21:00:16Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
==CREATING AND TERMINATING PTHREADS==&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
==MUTEXES==&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
==CONDITIONAL VARIABLES==&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
==DOACROSS==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
==DOPIPE==&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
==DOALL==&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
Some features such as tasks, synchronization primitives among a few others were included in the upgrades from OpenMP 2.0 but nothing significant was upgraded that fits the scope of this article.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
* Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
* http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
* Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
* http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
* https://computing.llnl.gov/tutorials/openMP/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30859</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30859"/>
		<updated>2010-02-23T20:57:15Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* OpenMP 3.0 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
POSIX thread, also referred to as a pthread is used in shared address space architectures for parallel programs.  Through the use of the pthread API, various functions can be used to create and manage pthreads.  In order to fully, understand how pthreads can be used to exploit DOACROSS, DOPIPE, and DOALL parallelism a brief introduction to creating/terminating pthreads, mutexes,  and conditional variables are necessary.&lt;br /&gt;
&lt;br /&gt;
CREATING AND TERMINATING PTHREADS&lt;br /&gt;
&lt;br /&gt;
In order to create a pthread, the API provides the pthread_create() function.  The pthread function accepts 4 arguments: thread, attr, start_routine, and arg.  The thread argument is used to provide a unique identifier for the thread you are creating.  The attr argument is used to specify a threads attribute object, or use default attributes by passing NULL.  For the examples discussed the default attributes will be sufficient; for more information on setting thread attributes please see references.  The start_routine argument is the program subroutine that will be executed by the thread being created.  The arg argument is used to pass an argument to the subroutine that is being executed by the thread being created (value can be set to NULL if no argument is being pass to the subroutine).&lt;br /&gt;
&lt;br /&gt;
In order to terminate pthreads, the API provides the pthread_exit() function.  In order to terminate a pthread, the thread being run simply has to call this function (even the main thread).  NOTE: there are alternate methods for terminating pthreads not discussed here for convenience and simplicity.&lt;br /&gt;
&lt;br /&gt;
MUTEXES&lt;br /&gt;
&lt;br /&gt;
A mutex variable is a variable that must only be accessed by a single thread (mutex is short for mutual exclusion).  The API provides the pthread_mutex_t data type in order statically create a mutex variable, and the pthread_mutex_init() function to create it dynamically.&lt;br /&gt;
&lt;br /&gt;
Mutex variables are used to implement locks, so that multiple pthreads do not access critical data in a program.  The API provides the pthread_mutex_lock() and pthread_mutex_unlock() functions.  These functions simply lock or unlock the mutex variable specified. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
CONDITIONAL VARIABLES&lt;br /&gt;
&lt;br /&gt;
Conditional variables allow for point-to-point synchronization between threads.&lt;br /&gt;
The API provides a few useful functions in order to synchronize threads:  pthread_cond_wait() and pthread_cond_signal().  The pthread_cond_wait() blocks all threads until the specified condition is satisfied.  The pthread_cond_signal() wakes up another thread that is waiting on the condition to be satisfied.  These functions are the pthread specific functions that are analogous to the general wait() and post() function discussed in the Sohlihin text.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
DOACROSS&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOACROSS parallelism using pthreads, conditions are needed in order to synchronize the threads.  Since instructions are executed across iterations, and  data dependencies exist across iterations are assumed (See Sohlin Text), the conditional variables shown above are used to ensure the correct execution of the code. &lt;br /&gt;
&lt;br /&gt;
Lets take a simple example where each thread calculates A[i] = A[i-1] + B[i].  Point-to-point synchronization is necessary in order to make sure A[i-1] is not read before its value is written.  This is where the pthread’s conditional variable comes is useful.  We put pthread_cond_wait() and pthread_cond_signal() around the instruction above, in order to make sure the previous thread has signaled that its completed before the current thread performs its own computation.&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
DOPIPE&lt;br /&gt;
&lt;br /&gt;
In order to exploit DOPIPE parallelism, conditions are also needed in order to synchronize threads.  Instead of instructions being implemented across threads, instructions are implemented on a single thread (i.e. instruction 1 is executed by thread 1, instruction 2 is executed by thread 2, etc).  However, there are loop independent data dependencies which require the conditional variables.&lt;br /&gt;
&lt;br /&gt;
Here we use the pthreads differently from the DOACROSS parallelism.  The DOPIPE parallelism has each thread call a different function.  Each function may have some loop independent dependence with some other function.  Lets say function 2 depends on function 1, so function 1 will call pthread_cond_signal() once it is finished, and function 2 will call pthread_cond_wait().  &lt;br /&gt;
&lt;br /&gt;
The differences between DOPIPE and DOACROSS are that DOPIPE executes different functions on each thread, and it has the signal and wait functions called from different functions.  Whereas DOACROSS executes the same function on each thread (it just uses different data), and it has the signal and wait functions called from the same function.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
DOALL&lt;br /&gt;
&lt;br /&gt;
Since DOALL parallelism just means that all iterations are executed in parallel, and no dependences exits.  All that is necessary is that the threads have to be created.&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
Some features such as tasks, synchronization primitives among a few others were included in the upgrades from OpenMP 2.0 but nothing significant was upgraded that fits the scope of this article.&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
- Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
- http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
- http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/openMP/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_2_aj/Data_Parallel_Programming&amp;diff=30857</id>
		<title>CSC/ECE 506 Spring 2010/ch 2 aj/Data Parallel Programming</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_2_aj/Data_Parallel_Programming&amp;diff=30857"/>
		<updated>2010-02-23T20:54:26Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Example of Data Parallel Programing Model */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Data-Parallel Programming Model &lt;br /&gt;
&lt;br /&gt;
= Introduction =&lt;br /&gt;
&lt;br /&gt;
The data-parallel programming model first appeared in the eighties as a programming model for SIMD (Single Instruction, Multiple Data) parallel machines.  It's defined as multiple processing elements performing an action simultaneously on different parts of a data set and exchanging information globally before processing more code synchronously.  Even though this model is offered in contrast to the shared memory model and the message passing model, data-parallel processing can actually be accomplished by passing messages or sharing variables.  The main distinction between the data-parallel model and the other two models has to do with the outcome of the individual steps instead of the method of communication.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Tasks in this model work in conjunction on the same data structure. A distinct quality of the data parallelism model is that each task performs actions on different partitions of this data structure.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The model is generally associated with applications that involve a data set which is typically organized into a common structure, such as an array or matrix. Data-parallel processing has been found to be effective in situations where computations involve every element of a matrix in a uniform way.  This allows the processing to be divided spatially over memories.  Originally, the need for this model was discovered in scientific calculations and is still used today.&lt;br /&gt;
&lt;br /&gt;
= Data parallelism vs Task parallelism =&lt;br /&gt;
One important feature of data-parallel programming model or data parallelism (SIMD) is the single control flow: there is only one control processor that directs the activities of all the processing elements. In stark contrast to this is task parallelism (MIMD: Multiple Instruction, Multiple Data): characterized by its multiple control flows, it allows the concurrent execution of multiple instruction streams, each manipulates its own data and services separate functions. Below is a contrast between the data parallelism and task parallelism models from wikipedia: [http://en.wikipedia.org/wiki/SIMD SIMD] and [http://en.wikipedia.org/wiki/MIMD MIMD]. In the following subsections we continue to compare and contrast different features of data-parallel model and task-parallel model to help reader understand the unique characteristics of data-parallel programming model.&lt;br /&gt;
[[Image:Smid.png|frame|center|425px|contrast between data parallelism and task parallelism]]&lt;br /&gt;
&lt;br /&gt;
== Synchronous vs Asynchronous ==&lt;br /&gt;
While the [http://en.wikipedia.org/wiki/Lockstep_(computing) lockstep] imposed by data parallelism on all data streams ensures synchronous computation (all PEs perform their tasks at the exact same pace), every processor in task parallelism performs its task at their own pace, which we call asynchronous computation. Thus, at a certain point of a task parallel program's execution, communication and synchronization primitives are needed to allow different instruction streams to coordinate their efforts, and that is where variable-sharing and message-passing come into play.&lt;br /&gt;
&lt;br /&gt;
== Determinism vs. Non-Determinism ==&lt;br /&gt;
Data parallelism's synchronous nature and task parallelism's asynchronism give rise to another pair of features that add to the difference between these two models: determinism versus non-determinism. Data parallelism is deterministic, i.e. computing with the same input will always yield the same result, since its synchronism ensures that issues like relative timing between PEs will not arise. In contrast, task parallelism's asynchronous updates of common data can give rise to non-determinism, i.e, the same input won't always yield the same computation result (the result of a computation will depend also on factors outside the program control, such as scheduling and timing of other PEs). Obviously, non-determinism makes it harder to write and maintain correct programs. This partially explains the advantage of data parallel programming model over data parallelism in terms of development effort (also discussed in section 4.2).&lt;br /&gt;
&lt;br /&gt;
= Example of Data Parallel Programing Model =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
This section shows a simple example adapted from Solihin textbook (pp. 24 - 27) that illustrates  the data-parallel programming model. Each of the codes below are written in pseudo-code style.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Suppose we want to perform the following task on an array &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt;: updating each element of &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt; by the product of itself and its index, and adding together the elements of &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt; into the variable &amp;lt;code&amp;gt;sum&amp;lt;/code&amp;gt;. The corresponding code is shown below.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // simple sequential task&lt;br /&gt;
 sum = 0;&lt;br /&gt;
 '''for''' (i = 0; i &amp;lt; a.length; i++)&lt;br /&gt;
 {&lt;br /&gt;
    a[i] = a[i] * i;&lt;br /&gt;
    sum = sum + a[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
When we implement the task using the data-parallel programming model, the program can be divided into two parts. The first part performs the same operations on separate elements of the array for each processing element (sometimes referred to as PE or pe), and the second part reorganizes data among all processing elements (In our example data reorganization is summing up values across different processing elements). Since data-parallel programming model only defines the overall effects of parallel steps, the second part can be accomplished either through shared memory or message passing. The three code fragments below are examples for the first part of the program, shared-memory version of the second part, and message passing for the second part, respectively.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data parallel programming: let each PE perform the same task on different pieces of distributed data&lt;br /&gt;
 pe_id = getid();&lt;br /&gt;
 my_sum = 0;&lt;br /&gt;
 '''for''' (i = pe_id; i &amp;lt; a.length; i += number_of_pe)         //separate elements of the array are assigned to each PE &lt;br /&gt;
 {&lt;br /&gt;
    a[i] = a[i] * i;&lt;br /&gt;
    my_sum = my_sum + a[i];                               //all PEs accumulate elements assigned to them into local variable my_sum&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the above code, data parallelism is achieved by letting each processing element perform actions on array's separate elements, which are identified using the PE's id. For instance, if three processing elements are used then one processing element would start at i = 0, one would start at i = 1, and the last would start at i = 2. Since there are three processing elements then the index of the array for each will increase by three on each iteration until the task is complete (note that in our example elements assigned to each PE are interleaved instead of continuous). If the length of the array is a multiple of three then each processing element takes the same amount of time to execute its portion of the task.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The picture below illustrates how elements of the array are assigned among different PEs for the specific case: length of the array is 7 and there are 3 PEs available. Elements in the array are marked by their indexes(0 to 6). As shown in the picture, PE0 will work on elements with index 0, 3, 6; PE1 is in charge of elements with index 1, 4; and elements with index 2, 5 are assigned to PE2. In this way, these 3 PEs work collectively on the array, while each PE works on different elements. Thus, data parallelism is achieved.&lt;br /&gt;
&lt;br /&gt;
[[Image:506wiki1.png|frame|center|150px|Illustration of data parallel programming(adapted from [http://computing.llnl.gov/tutorials/parallel_comp/#ModelsData Introduction to Parallel Computing])]]&lt;br /&gt;
&lt;br /&gt;
= Other Parallel Programming Models =&lt;br /&gt;
== Example Continued ==&lt;br /&gt;
As previously mentioned, the shared memory and message passing parallel programming models are commonly used. The following descriptions and code snippets are intended to develop a contrast between the data-parallel programming model and the other two models.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The code below shows how data reorganization is done by the shared memory programming model. A shared variable is declared to store the global sum. Each PE accumulates their local my_sum into this variable. In this example, the lock() and unlock() routines are used to prevent race conditions and barrier ensures that the all the local my_sum variables have been accumulated into the shared variable sum before the code proceeds.  Notice that in this model there is only one copy of the data like in the data-parallel model.  However, in the shared memory model, the data must be protected since any PE could be using it at any point, whereas data is associated with a specific PE in the data-parallel model.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data reorganization via shared memory&lt;br /&gt;
 '''shared''' sum;&lt;br /&gt;
 lock();                                                  //prevent race condition&lt;br /&gt;
 sum = sum + my_sum;                                      //each PE adds up their local my_sum to shared variable sum &lt;br /&gt;
 unlock();&lt;br /&gt;
 barrier;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The final section of code shows how data reorganization is done by the message passing programming model. PE0 acts as the one that collects the local my_sum variable from all the other PEs. This is done by the send_msg() and  the recv_msg() routines: PEs other than PE0 send their my_sum variable as a message to PE0. PE0 receives the messages by specifying the sender in the '''for''' loop. PE ids are used in the message passing model in the same way they are in the data-parallel model.  On the contrary, in the data-parallel model there is only one copy of data that is processed and there is no need to do any final accumulation by a single PE.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data reorganization via message passing&lt;br /&gt;
 '''if''' (pe_id != 0) send_msg (0, my_sum);                    //PE other than PE0 send their local my_sum&lt;br /&gt;
 '''else'''                                                     //PE0 does this&lt;br /&gt;
 {&lt;br /&gt;
    '''for''' (i = 1; i &amp;lt; number_of_pe; i++)                    //for each other PE&lt;br /&gt;
    {&lt;br /&gt;
       recv_msg (i, temp);                                //receive local sum from other PEs&lt;br /&gt;
       my_sum = my_sum + temp;                            //accumulate into total&lt;br /&gt;
    }&lt;br /&gt;
    sum = my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Comparison in Different Aspects ==&lt;br /&gt;
The table below is extended from Table 2.1 on page 22 of the Solihin text.  A column has been added for the data-parallel model.  The table compares some key characteristics of each programming model.  As you can see, the complexity of the data-parallel model comes mainly in the form of dividing up the work among PEs.  Most of this work is done by the programmer and does not necessarily require special hardware, although, specific types of hardware can optimize the benefits of the model.  For instance,  SIMD and SPMD hardware are examples that have efficiently utilized the data-parallel model.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{| align=&amp;quot;center&amp;quot; class=&amp;quot;wikitable&amp;quot; style=&amp;quot;text-align:center;&amp;quot; border=&amp;quot;1&amp;quot; !cellpadding = &amp;quot;2&amp;quot;&lt;br /&gt;
|+ &lt;br /&gt;
! Aspects !! Data-Parallel !! Shared Memory !! Message Passing&lt;br /&gt;
|-&lt;br /&gt;
! Communication&lt;br /&gt;
| implicit - via loads/stores || implicit - via loads/stores || explicit messages&lt;br /&gt;
|-&lt;br /&gt;
! Synchronization&lt;br /&gt;
| none || explicit || implicit - via messages&lt;br /&gt;
|-&lt;br /&gt;
! Hardware support&lt;br /&gt;
| none || typically required || none&lt;br /&gt;
|-&lt;br /&gt;
! Development effort&lt;br /&gt;
| low || low || high&lt;br /&gt;
|- &lt;br /&gt;
! Tuning effort&lt;br /&gt;
| high || high || low&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
&lt;br /&gt;
1) David E. Culler, Jaswinder Pal Singh, and Anoop Gupta, ''Parallel Computer Architecture: A Hardware/Software Approach'', Gulf Professional Publishing, August 1998.&lt;br /&gt;
&lt;br /&gt;
2) Yan Solihin, ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'', Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
3) Philip J. Hatcher, Michael Jay Quinn, ''Data-Parallel Programming on MIMD Computers'', The MIT Press, 1991.&lt;br /&gt;
&lt;br /&gt;
4) Blaise Barney, &amp;quot;Introduction to Parallel Computing: Data Parallel Model&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/parallel_comp/#ModelsData https://computing.llnl.gov/tutorials/parallel_comp/#ModelsData], January 2009.&lt;br /&gt;
&lt;br /&gt;
5) Guy Blelloch, &amp;quot;Is Parallel Programming Hard?&amp;quot;, Carnegie Mellon University, [http://www.cilk.com/multicore-blog/bid/9108/Is-Parallel-Programming-Hard http://www.cilk.com/multicore-blog/bid/9108/Is-Parallel-Programming-Hard], April 2009.&lt;br /&gt;
&lt;br /&gt;
6) Björn Lisper, ''Data parallelism and functional programming'', Lecture Notes in Computer Science, Volume 1132/1996, pp. 220-251, Springer Berlin, 1996.&lt;br /&gt;
&lt;br /&gt;
7) ''SIMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/SIMD http://en.wikipedia.org/wiki/SIMD].&lt;br /&gt;
&lt;br /&gt;
8) ''MIMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/MIMD http://en.wikipedia.org/wiki/MIMD].&lt;br /&gt;
&lt;br /&gt;
9) ''Lockstep'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/Lockstep_(computing) http://en.wikipedia.org/wiki/Lockstep_(computing)].&lt;br /&gt;
&lt;br /&gt;
10) ''SPMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/SPMD http://en.wikipedia.org/wiki/SPMD].&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_2_aj/Data_Parallel_Programming&amp;diff=30856</id>
		<title>CSC/ECE 506 Spring 2010/ch 2 aj/Data Parallel Programming</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_2_aj/Data_Parallel_Programming&amp;diff=30856"/>
		<updated>2010-02-23T20:54:13Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Example of Data Parallel Programing Model */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Data-Parallel Programming Model &lt;br /&gt;
&lt;br /&gt;
= Introduction =&lt;br /&gt;
&lt;br /&gt;
The data-parallel programming model first appeared in the eighties as a programming model for SIMD (Single Instruction, Multiple Data) parallel machines.  It's defined as multiple processing elements performing an action simultaneously on different parts of a data set and exchanging information globally before processing more code synchronously.  Even though this model is offered in contrast to the shared memory model and the message passing model, data-parallel processing can actually be accomplished by passing messages or sharing variables.  The main distinction between the data-parallel model and the other two models has to do with the outcome of the individual steps instead of the method of communication.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Tasks in this model work in conjunction on the same data structure. A distinct quality of the data parallelism model is that each task performs actions on different partitions of this data structure.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The model is generally associated with applications that involve a data set which is typically organized into a common structure, such as an array or matrix. Data-parallel processing has been found to be effective in situations where computations involve every element of a matrix in a uniform way.  This allows the processing to be divided spatially over memories.  Originally, the need for this model was discovered in scientific calculations and is still used today.&lt;br /&gt;
&lt;br /&gt;
= Data parallelism vs Task parallelism =&lt;br /&gt;
One important feature of data-parallel programming model or data parallelism (SIMD) is the single control flow: there is only one control processor that directs the activities of all the processing elements. In stark contrast to this is task parallelism (MIMD: Multiple Instruction, Multiple Data): characterized by its multiple control flows, it allows the concurrent execution of multiple instruction streams, each manipulates its own data and services separate functions. Below is a contrast between the data parallelism and task parallelism models from wikipedia: [http://en.wikipedia.org/wiki/SIMD SIMD] and [http://en.wikipedia.org/wiki/MIMD MIMD]. In the following subsections we continue to compare and contrast different features of data-parallel model and task-parallel model to help reader understand the unique characteristics of data-parallel programming model.&lt;br /&gt;
[[Image:Smid.png|frame|center|425px|contrast between data parallelism and task parallelism]]&lt;br /&gt;
&lt;br /&gt;
== Synchronous vs Asynchronous ==&lt;br /&gt;
While the [http://en.wikipedia.org/wiki/Lockstep_(computing) lockstep] imposed by data parallelism on all data streams ensures synchronous computation (all PEs perform their tasks at the exact same pace), every processor in task parallelism performs its task at their own pace, which we call asynchronous computation. Thus, at a certain point of a task parallel program's execution, communication and synchronization primitives are needed to allow different instruction streams to coordinate their efforts, and that is where variable-sharing and message-passing come into play.&lt;br /&gt;
&lt;br /&gt;
== Determinism vs. Non-Determinism ==&lt;br /&gt;
Data parallelism's synchronous nature and task parallelism's asynchronism give rise to another pair of features that add to the difference between these two models: determinism versus non-determinism. Data parallelism is deterministic, i.e. computing with the same input will always yield the same result, since its synchronism ensures that issues like relative timing between PEs will not arise. In contrast, task parallelism's asynchronous updates of common data can give rise to non-determinism, i.e, the same input won't always yield the same computation result (the result of a computation will depend also on factors outside the program control, such as scheduling and timing of other PEs). Obviously, non-determinism makes it harder to write and maintain correct programs. This partially explains the advantage of data parallel programming model over data parallelism in terms of development effort (also discussed in section 4.2).&lt;br /&gt;
&lt;br /&gt;
= Example of Data Parallel Programing Model =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 This section shows a simple example adapted from Solihin textbook (pp. 24 - 27) that illustrates  the data-parallel programming model. Each of the codes below are written in pseudo-code style.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Suppose we want to perform the following task on an array &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt;: updating each element of &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt; by the product of itself and its index, and adding together the elements of &amp;lt;code&amp;gt;a&amp;lt;/code&amp;gt; into the variable &amp;lt;code&amp;gt;sum&amp;lt;/code&amp;gt;. The corresponding code is shown below.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // simple sequential task&lt;br /&gt;
 sum = 0;&lt;br /&gt;
 '''for''' (i = 0; i &amp;lt; a.length; i++)&lt;br /&gt;
 {&lt;br /&gt;
    a[i] = a[i] * i;&lt;br /&gt;
    sum = sum + a[i];&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
When we implement the task using the data-parallel programming model, the program can be divided into two parts. The first part performs the same operations on separate elements of the array for each processing element (sometimes referred to as PE or pe), and the second part reorganizes data among all processing elements (In our example data reorganization is summing up values across different processing elements). Since data-parallel programming model only defines the overall effects of parallel steps, the second part can be accomplished either through shared memory or message passing. The three code fragments below are examples for the first part of the program, shared-memory version of the second part, and message passing for the second part, respectively.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data parallel programming: let each PE perform the same task on different pieces of distributed data&lt;br /&gt;
 pe_id = getid();&lt;br /&gt;
 my_sum = 0;&lt;br /&gt;
 '''for''' (i = pe_id; i &amp;lt; a.length; i += number_of_pe)         //separate elements of the array are assigned to each PE &lt;br /&gt;
 {&lt;br /&gt;
    a[i] = a[i] * i;&lt;br /&gt;
    my_sum = my_sum + a[i];                               //all PEs accumulate elements assigned to them into local variable my_sum&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
In the above code, data parallelism is achieved by letting each processing element perform actions on array's separate elements, which are identified using the PE's id. For instance, if three processing elements are used then one processing element would start at i = 0, one would start at i = 1, and the last would start at i = 2. Since there are three processing elements then the index of the array for each will increase by three on each iteration until the task is complete (note that in our example elements assigned to each PE are interleaved instead of continuous). If the length of the array is a multiple of three then each processing element takes the same amount of time to execute its portion of the task.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The picture below illustrates how elements of the array are assigned among different PEs for the specific case: length of the array is 7 and there are 3 PEs available. Elements in the array are marked by their indexes(0 to 6). As shown in the picture, PE0 will work on elements with index 0, 3, 6; PE1 is in charge of elements with index 1, 4; and elements with index 2, 5 are assigned to PE2. In this way, these 3 PEs work collectively on the array, while each PE works on different elements. Thus, data parallelism is achieved.&lt;br /&gt;
&lt;br /&gt;
[[Image:506wiki1.png|frame|center|150px|Illustration of data parallel programming(adapted from [http://computing.llnl.gov/tutorials/parallel_comp/#ModelsData Introduction to Parallel Computing])]]&lt;br /&gt;
&lt;br /&gt;
= Other Parallel Programming Models =&lt;br /&gt;
== Example Continued ==&lt;br /&gt;
As previously mentioned, the shared memory and message passing parallel programming models are commonly used. The following descriptions and code snippets are intended to develop a contrast between the data-parallel programming model and the other two models.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The code below shows how data reorganization is done by the shared memory programming model. A shared variable is declared to store the global sum. Each PE accumulates their local my_sum into this variable. In this example, the lock() and unlock() routines are used to prevent race conditions and barrier ensures that the all the local my_sum variables have been accumulated into the shared variable sum before the code proceeds.  Notice that in this model there is only one copy of the data like in the data-parallel model.  However, in the shared memory model, the data must be protected since any PE could be using it at any point, whereas data is associated with a specific PE in the data-parallel model.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data reorganization via shared memory&lt;br /&gt;
 '''shared''' sum;&lt;br /&gt;
 lock();                                                  //prevent race condition&lt;br /&gt;
 sum = sum + my_sum;                                      //each PE adds up their local my_sum to shared variable sum &lt;br /&gt;
 unlock();&lt;br /&gt;
 barrier;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The final section of code shows how data reorganization is done by the message passing programming model. PE0 acts as the one that collects the local my_sum variable from all the other PEs. This is done by the send_msg() and  the recv_msg() routines: PEs other than PE0 send their my_sum variable as a message to PE0. PE0 receives the messages by specifying the sender in the '''for''' loop. PE ids are used in the message passing model in the same way they are in the data-parallel model.  On the contrary, in the data-parallel model there is only one copy of data that is processed and there is no need to do any final accumulation by a single PE.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 // data reorganization via message passing&lt;br /&gt;
 '''if''' (pe_id != 0) send_msg (0, my_sum);                    //PE other than PE0 send their local my_sum&lt;br /&gt;
 '''else'''                                                     //PE0 does this&lt;br /&gt;
 {&lt;br /&gt;
    '''for''' (i = 1; i &amp;lt; number_of_pe; i++)                    //for each other PE&lt;br /&gt;
    {&lt;br /&gt;
       recv_msg (i, temp);                                //receive local sum from other PEs&lt;br /&gt;
       my_sum = my_sum + temp;                            //accumulate into total&lt;br /&gt;
    }&lt;br /&gt;
    sum = my_sum;&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
== Comparison in Different Aspects ==&lt;br /&gt;
The table below is extended from Table 2.1 on page 22 of the Solihin text.  A column has been added for the data-parallel model.  The table compares some key characteristics of each programming model.  As you can see, the complexity of the data-parallel model comes mainly in the form of dividing up the work among PEs.  Most of this work is done by the programmer and does not necessarily require special hardware, although, specific types of hardware can optimize the benefits of the model.  For instance,  SIMD and SPMD hardware are examples that have efficiently utilized the data-parallel model.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{| align=&amp;quot;center&amp;quot; class=&amp;quot;wikitable&amp;quot; style=&amp;quot;text-align:center;&amp;quot; border=&amp;quot;1&amp;quot; !cellpadding = &amp;quot;2&amp;quot;&lt;br /&gt;
|+ &lt;br /&gt;
! Aspects !! Data-Parallel !! Shared Memory !! Message Passing&lt;br /&gt;
|-&lt;br /&gt;
! Communication&lt;br /&gt;
| implicit - via loads/stores || implicit - via loads/stores || explicit messages&lt;br /&gt;
|-&lt;br /&gt;
! Synchronization&lt;br /&gt;
| none || explicit || implicit - via messages&lt;br /&gt;
|-&lt;br /&gt;
! Hardware support&lt;br /&gt;
| none || typically required || none&lt;br /&gt;
|-&lt;br /&gt;
! Development effort&lt;br /&gt;
| low || low || high&lt;br /&gt;
|- &lt;br /&gt;
! Tuning effort&lt;br /&gt;
| high || high || low&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
&lt;br /&gt;
1) David E. Culler, Jaswinder Pal Singh, and Anoop Gupta, ''Parallel Computer Architecture: A Hardware/Software Approach'', Gulf Professional Publishing, August 1998.&lt;br /&gt;
&lt;br /&gt;
2) Yan Solihin, ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'', Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
3) Philip J. Hatcher, Michael Jay Quinn, ''Data-Parallel Programming on MIMD Computers'', The MIT Press, 1991.&lt;br /&gt;
&lt;br /&gt;
4) Blaise Barney, &amp;quot;Introduction to Parallel Computing: Data Parallel Model&amp;quot;, Lawrence Livermore National Laboratory, [https://computing.llnl.gov/tutorials/parallel_comp/#ModelsData https://computing.llnl.gov/tutorials/parallel_comp/#ModelsData], January 2009.&lt;br /&gt;
&lt;br /&gt;
5) Guy Blelloch, &amp;quot;Is Parallel Programming Hard?&amp;quot;, Carnegie Mellon University, [http://www.cilk.com/multicore-blog/bid/9108/Is-Parallel-Programming-Hard http://www.cilk.com/multicore-blog/bid/9108/Is-Parallel-Programming-Hard], April 2009.&lt;br /&gt;
&lt;br /&gt;
6) Björn Lisper, ''Data parallelism and functional programming'', Lecture Notes in Computer Science, Volume 1132/1996, pp. 220-251, Springer Berlin, 1996.&lt;br /&gt;
&lt;br /&gt;
7) ''SIMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/SIMD http://en.wikipedia.org/wiki/SIMD].&lt;br /&gt;
&lt;br /&gt;
8) ''MIMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/MIMD http://en.wikipedia.org/wiki/MIMD].&lt;br /&gt;
&lt;br /&gt;
9) ''Lockstep'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/Lockstep_(computing) http://en.wikipedia.org/wiki/Lockstep_(computing)].&lt;br /&gt;
&lt;br /&gt;
10) ''SPMD'', Wikipedia, the free encyclopedia, [http://en.wikipedia.org/wiki/SPMD http://en.wikipedia.org/wiki/SPMD].&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30834</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30834"/>
		<updated>2010-02-23T03:38:39Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* Intel Threading Building Blocks */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
== Parallel Loops ==&lt;br /&gt;
&lt;br /&gt;
A DOALL parallel construct can be specified using the ''parallel_for()'' construct.  This construct takes two parameters.  The first is the range of indices of the loop that can be run in parallel.  The second is the a solid chunk of operations that can be processed as a units that are safe to run concurrently.  For a DOALL loop, this second parameter should include all possible loop indices.  Also, an optional third parameter can be specified to define the chunk size of the loop and information about cache affinity.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
- Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
- http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
- http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/openMP/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30746</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30746"/>
		<updated>2010-02-22T21:59:42Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: /* References */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
- Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
- http://www.humanfactor.com/pthreads/&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; TBB - Intel&amp;amp;reg; Software Network&amp;quot;, [http://software.intel.com/en-us/intel-tbb/ http://software.intel.com/en-us/intel-tbb/]&lt;br /&gt;
&lt;br /&gt;
- Intel&amp;amp;reg; Corporation, &amp;quot;Intel&amp;amp;reg; Threading Building Blocks 2.2 for Open Source&amp;quot;, [http://www.threadingbuildingblocks.org/ http://www.threadingbuildingblocks.org/]&lt;br /&gt;
&lt;br /&gt;
- http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
- https://computing.llnl.gov/tutorials/openMP/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30745</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30745"/>
		<updated>2010-02-22T21:51:46Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Intel Threading Building Blocks =&lt;br /&gt;
&lt;br /&gt;
According the to Intel&amp;amp;reg; Software [http://software.intel.com/en-us/intel-tbb/ site], Intel&amp;amp;reg; Threading Building Blocks (Intel&amp;amp;reg; TBB) is a C++ template library that abstracts thread to tasks to create reliable, portable, and scalable parallel applications.&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
1) Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
2) https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
3) http://www.threadingbuildingblocks.org/&lt;br /&gt;
&lt;br /&gt;
4) http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
5) https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
6) http://www.humanfactor.com/pthreads/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30744</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30744"/>
		<updated>2010-02-22T19:04:05Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Posix threads =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Intel Thread Building Blocks =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 2.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= OpenMP 3.0 =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
1) Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;br /&gt;
&lt;br /&gt;
2) https://computing.llnl.gov/tutorials/pthreads/&lt;br /&gt;
&lt;br /&gt;
3) http://www.threadingbuildingblocks.org/&lt;br /&gt;
&lt;br /&gt;
4) http://openmp.org/wp/&lt;br /&gt;
&lt;br /&gt;
5) https://computing.llnl.gov/tutorials/openMP/&lt;br /&gt;
&lt;br /&gt;
6) http://www.humanfactor.com/pthreads/&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30740</id>
		<title>CSC/ECE 506 Spring 2010/ch 3 jb/Parallel Programming Model Support</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/ch_3_jb/Parallel_Programming_Model_Support&amp;diff=30740"/>
		<updated>2010-02-22T18:57:28Z</updated>

		<summary type="html">&lt;p&gt;Wjfisher: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Supplement to Chapter 3: Support for parallel-programming models. Discuss how DOACROSS, DOPIPE, DOALL, etc. are implemented in packages such as Posix threads, Intel Thread Building Blocks, OpenMP 2.0 and 3.0.&lt;br /&gt;
&lt;br /&gt;
= Introduction =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Packages =&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Posix threads ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Intel Thread Building Blocks ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== OpenMP 2.0 ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== OpenMP 3.0 ==&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= References =&lt;br /&gt;
1) Yan Solihin, Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems, Solihin Books, August 2009.&lt;/div&gt;</summary>
		<author><name>Wjfisher</name></author>
	</entry>
</feed>