<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Knnavala</id>
	<title>Expertiza_Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Knnavala"/>
	<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Special:Contributions/Knnavala"/>
	<updated>2026-09-11T17:34:22Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40206</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40206"/>
		<updated>2010-11-09T19:53:08Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40205</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40205"/>
		<updated>2010-11-09T19:52:31Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40204</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40204"/>
		<updated>2010-11-09T19:51:53Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40203</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40203"/>
		<updated>2010-11-09T19:50:47Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Template:Pad&amp;diff=40202</id>
		<title>Template:Pad</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Template:Pad&amp;diff=40202"/>
		<updated>2010-11-09T19:49:29Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;&amp;amp;nbsp;&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40201</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40201"/>
		<updated>2010-11-09T19:49:04Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40200</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40200"/>
		<updated>2010-11-09T19:45:36Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40199</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40199"/>
		<updated>2010-11-09T19:44:29Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40198</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40198"/>
		<updated>2010-11-09T19:43:46Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;p&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/p&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40197</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40197"/>
		<updated>2010-11-09T19:43:37Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;p&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/p&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40196</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40196"/>
		<updated>2010-11-09T19:42:58Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:&amp;lt;nowiki&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40195</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40195"/>
		<updated>2010-11-09T19:42:47Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:&lt;br /&gt;
&amp;lt;nowiki&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40194</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40194"/>
		<updated>2010-11-09T19:42:28Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There &amp;lt;nowiki&amp;gt;has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40193</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40193"/>
		<updated>2010-11-09T19:42:11Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There &amp;lt;/nowiki&amp;gt;has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40192</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40192"/>
		<updated>2010-11-09T19:41:20Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic &lt;br /&gt;
&amp;lt;nowiki&amp;gt;environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40191</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40191"/>
		<updated>2010-11-09T19:39:27Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic&amp;lt;br/&amp;gt; environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40190</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40190"/>
		<updated>2010-11-09T19:37:51Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic &lt;br /&gt;
environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40189</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40189"/>
		<updated>2010-11-09T19:37:12Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
::There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40188</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40188"/>
		<updated>2010-11-09T19:34:37Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40187</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40187"/>
		<updated>2010-11-09T19:34:20Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
New paragraph.&amp;lt;/tt&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40186</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40186"/>
		<updated>2010-11-09T19:33:00Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;/tt&amp;gt;There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40185</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40185"/>
		<updated>2010-11-09T19:32:09Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40184</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40184"/>
		<updated>2010-11-09T19:31:44Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
{{pad|4em}}There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40183</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40183"/>
		<updated>2010-11-09T19:30:43Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
:There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40182</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40182"/>
		<updated>2010-11-09T19:30:26Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
::There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40181</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40181"/>
		<updated>2010-11-09T19:22:34Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40180</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40180"/>
		<updated>2010-11-09T19:20:44Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
       There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;       The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
       After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
       Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
       At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
       The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
       For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
       Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
       The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
       The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
       In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40162</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40162"/>
		<updated>2010-11-08T15:26:28Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: Perspectives'''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: Parallel Programming Model'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: Shared Memory Parallel Programming'''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6: Introduction to Memory Hierarchy Organization'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: Introduction to Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8: Bus-Based Coherent Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9: Hardware Support for Synchronization'''&lt;br /&gt;
&lt;br /&gt;
It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: Memory Consistency Models'''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11: Distributed Shared Memory Multiprocessors'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12: Interconnection networks'''&lt;br /&gt;
&lt;br /&gt;
Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40161</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40161"/>
		<updated>2010-11-08T15:19:51Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: '''&lt;br /&gt;
&lt;br /&gt;
The supplement covers an interesting topic of supercomputer evolution. Wiki pages written for this topic includes a lot data from literature. It has interesting topics which are not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the supplement we can see the increase in dominance of Intel’s processors in the consumer market. It also concludes that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) are the earliest style of widely used multiprocessor machine architectures which are replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: '''&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The wiki supplement provides comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concludes that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: '''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE are discussed. These three parallelism techniques are discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provides additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. It also compares DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6:'''&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: The wiki supplement adds additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy is an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic is WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: '''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement is to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency is discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concludes that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization is discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core are also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discusses commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8:'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9:'''&lt;br /&gt;
&lt;br /&gt;
Synchronization: It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier is included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: '''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11:'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12:'''&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40160</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40160"/>
		<updated>2010-11-08T15:10:22Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma Navalakha&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: '''&lt;br /&gt;
&lt;br /&gt;
The supplement covered an interesting topic of supercomputer evolution. Wiki pages written for this topic included a lot data from literature. It has interesting topics which were not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the research we could see the increase in dominance of Intel’s processors in the consumer market. We also conclude that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely used multiprocessor machine architectures which was replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: '''&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The wiki supplement provided comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concluded that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: '''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism techniques were discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provided additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6:'''&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: The wiki supplement added additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic was WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: '''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement was to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency was discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concluded that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization was discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core were also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discussed commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8:'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9:'''&lt;br /&gt;
&lt;br /&gt;
Synchronization: It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier was included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: '''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11:'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12:'''&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40159</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40159"/>
		<updated>2010-11-08T15:09:48Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma NavalakhaA&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: '''&lt;br /&gt;
&lt;br /&gt;
The supplement covered an interesting topic of supercomputer evolution. Wiki pages written for this topic included a lot data from literature. It has interesting topics which were not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From the research we could see the increase in dominance of Intel’s processors in the consumer market. We also conclude that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely used multiprocessor machine architectures which was replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: '''&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The wiki supplement provided comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. It can be noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In the comparisons it concluded that combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: '''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism techniques were discussed with examples in the form of Open MP code as discussed in the text book. Besides it also provided additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally it concludes : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6:'''&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: The wiki supplement added additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was an additional subtopic students threw light on. It also gives definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic was WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: '''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement was to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency was discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. It concluded that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization was discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core were also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. It also discussed commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8:'''&lt;br /&gt;
&lt;br /&gt;
The wiki page discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9:'''&lt;br /&gt;
&lt;br /&gt;
Synchronization: It classifies synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. It also discusses reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier was included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. It shows that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: '''&lt;br /&gt;
&lt;br /&gt;
The wiki supplement discusses the existing bus-based cache coherence in real machines. It goes ahead and classifies the cache coherence protocols based on the year they were introduced and the processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11:'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12:'''&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. It provides in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. It also discusses routing algorithms and deadlock, starvation and livelock associated with it. These topics are covered in in an extremely detailed way. It includes a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40158</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40158"/>
		<updated>2010-11-08T03:30:22Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;center&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;center&amp;gt;Karishma NavalakhaA&amp;lt;/center&amp;gt;&lt;br /&gt;
&lt;br /&gt;
'''Abstract: '''&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
'''Experience with Wiki written text book:'''&lt;br /&gt;
&lt;br /&gt;
&amp;lt;nowiki&amp;gt;The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&amp;lt;/nowiki&amp;gt;&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
'''Chapter wise learning from this independent study:'''&lt;br /&gt;
&lt;br /&gt;
'''Chapter 1: '''&lt;br /&gt;
&lt;br /&gt;
It covered an interesting topic of supercomputer evolution. Wiki pages written for this topic included a lot data from literature. Students came up with interesting topics which were not covered in the text book such as [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers Timeline of supercomputers], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29 First Supercomputer(ENIAC)], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History Cray History], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture Supercomputer Hierarchal Architecture], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System Supercomputer Operating System], [http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer Cooling Supercomputer] and [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family Processor Family]. From their research we could see the increase in dominance of Intel’s processors in the consumer market. We also conclude that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely used multiprocessor machine architectures which was replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/ http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 2: '''&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The students provided comparisons between data parallelism and task parallelism. [http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References Haveraaen (2000)] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. Students noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In their comparisons they concluded combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. W. Daniel Hillis and Guy L. Steele, Jr., [http://portal.acm.org/citation.cfm?id=7903 &amp;quot;Data parallel algorithms,&amp;quot;] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2. Alexander C. Klaiber and Henry M. Levy, [http://portal.acm.org/citation.cfm?id=192020 &amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
'''Chapter3: '''&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism techniques were discussed with examples in the form of Open MP code as discussed in the text book. Besides the students provided additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally they conclude : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf Intel Threading Building Blocks 2.2 for Open Source Reference Manual]&lt;br /&gt;
&lt;br /&gt;
3. [https://computing.llnl.gov/tutorials/pthreads/#Joining POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 6:'''&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: Students added additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was an additional subtopic students threw light on. Students also gave definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic was how students discussed WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://download.intel.com/technology/architecture/sma.pdf http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.intel.com/Assets/PDF/manual/248966.pdf http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.intel.com/design/intarch/papers/cache6.pdf http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
'''Chapter 7: '''&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement was to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency was discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. They concluded that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization was discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core were also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. The students also discussed commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790 http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 8:'''&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf Cache consistency with MESI on Intel processor]&lt;br /&gt;
&lt;br /&gt;
2. [http://techreport.com/articles.x/8236/2 AMD dual core Architecture]&lt;br /&gt;
&lt;br /&gt;
3. [http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913 Silicon Graphics Computer Systems]&lt;br /&gt;
&lt;br /&gt;
4. [http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533 Synapse tightly coupled multiprocessors: a new approach to solve old problems]&lt;br /&gt;
&lt;br /&gt;
5. [http://en.wikipedia.org/wiki/Dragon_protocol Dragon Protocol]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 9:'''&lt;br /&gt;
&lt;br /&gt;
Synchronization: Students classified synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements. &lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. They also discussed reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier was included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. They showed that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1. [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.statemaster.com/encyclopedia/Deadlock http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3. [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 10: '''&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf Shared Memory Consistency Models]&lt;br /&gt;
&lt;br /&gt;
2. [http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273 Designing Memory Consistency Models For Shared-Memory Multiprocessors]&lt;br /&gt;
&lt;br /&gt;
3. [http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html Consistency Models]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Chapter 11:'''&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor. Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
''1. ''Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [http://doi.acm.org/10.1145/325164.325132 &amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;] In ''Proceedings of the 17th Annual International Symposium on Computer Architecture.''&lt;br /&gt;
&lt;br /&gt;
2. David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf &amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;] ''ACM SIGPLAN Notices''.&lt;br /&gt;
&lt;br /&gt;
'''Chapter 12:'''&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. They provided in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. They also discussed routing algorithms and deadlock, starvation and livelock associated with it. These topics were covered in in an extremely detailed way. The students included a diagrammatic representation for every topology. &lt;br /&gt;
&lt;br /&gt;
References: &lt;br /&gt;
&lt;br /&gt;
1. [http://www.top500.org/2007_overview_recent_supercomputers/sci http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2. [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
'''Conclusion:'''&lt;br /&gt;
&lt;br /&gt;
This independent study helped me to increase my knowledge to a great extent in the field of Architecture of parallel computers. There were 4 students working on every chapter and came up with 2 wiki pages per group. We collected a total of 18 wiki supplements. The data collected was enormous. While reviewing their content I kept updating my knowledge base. I also provided the resources from where they can collect data. This helped me to come across latest developments in the field. Interacting with students helped me to increase my communication skills. Constant discussions with Prof. Gehringer helped me to understand key concepts. This idea of writing wiki supplements got selected for KU Village presentation. I got an opportunity to present this paper along with Prof. Gehringer.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40157</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40157"/>
		<updated>2010-11-08T03:17:46Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;html xmlns:o=&amp;quot;urn:schemas-microsoft-com:office:office&amp;quot;&lt;br /&gt;
xmlns:w=&amp;quot;urn:schemas-microsoft-com:office:word&amp;quot;&lt;br /&gt;
xmlns=&amp;quot;http://www.w3.org/TR/REC-html40&amp;quot;&amp;gt;&amp;lt;head&amp;gt;&amp;lt;meta http-equiv=Content-Type content=&amp;quot;text/html; charset=windows-1252&amp;quot;&amp;gt;&amp;lt;meta name=ProgId content=Word.Document&amp;gt;&amp;lt;meta name=Generator content=&amp;quot;Microsoft Word 11&amp;quot;&amp;gt;&amp;lt;meta name=Originator content=&amp;quot;Microsoft Word 11&amp;quot;&amp;gt;&amp;lt;link rel=File-List href=&amp;quot;Independent%20Study_files/filelist.xml&amp;quot;&amp;gt;&amp;lt;title&amp;gt;ECE633 Independent Study: Architecture of Parallel Computers&amp;lt;/title&amp;gt;&amp;lt;!--[if gte mso 9]&amp;gt;&amp;lt;xml&amp;gt;&amp;lt;o:DocumentProperties&amp;gt;&amp;lt;o:Author&amp;gt;karishma navalakha&amp;lt;/o:Author&amp;gt;&amp;lt;o:LastAuthor&amp;gt;Darshan J Pandya&amp;lt;/o:LastAuthor&amp;gt;&amp;lt;o:Revision&amp;gt;2&amp;lt;/o:Revision&amp;gt;&amp;lt;o:TotalTime&amp;gt;1&amp;lt;/o:TotalTime&amp;gt;&amp;lt;o:Created&amp;gt;2010-11-08T02:52:00Z&amp;lt;/o:Created&amp;gt;&amp;lt;o:LastSaved&amp;gt;2010-11-08T02:52:00Z&amp;lt;/o:LastSaved&amp;gt;&amp;lt;o:Pages&amp;gt;1&amp;lt;/o:Pages&amp;gt;&amp;lt;o:Words&amp;gt;4136&amp;lt;/o:Words&amp;gt;&amp;lt;o:Characters&amp;gt;23579&amp;lt;/o:Characters&amp;gt;&amp;lt;o:Lines&amp;gt;196&amp;lt;/o:Lines&amp;gt;&amp;lt;o:Paragraphs&amp;gt;55&amp;lt;/o:Paragraphs&amp;gt;&amp;lt;o:CharactersWithSpaces&amp;gt;27660&amp;lt;/o:CharactersWithSpaces&amp;gt;&amp;lt;o:Version&amp;gt;11.9999&amp;lt;/o:Version&amp;gt;&amp;lt;/o:DocumentProperties&amp;gt;&amp;lt;/xml&amp;gt;&amp;lt;![endif]--&amp;gt;&amp;lt;!--[if gte mso 9]&amp;gt;&amp;lt;xml&amp;gt;&amp;lt;w:WordDocument&amp;gt;&amp;lt;w:PunctuationKerning/&amp;gt;&amp;lt;w:ValidateAgainstSchemas/&amp;gt;&amp;lt;w:SaveIfXMLInvalid&amp;gt;false&amp;lt;/w:SaveIfXMLInvalid&amp;gt;&amp;lt;w:IgnoreMixedContent&amp;gt;false&amp;lt;/w:IgnoreMixedContent&amp;gt;&amp;lt;w:AlwaysShowPlaceholderText&amp;gt;false&amp;lt;/w:AlwaysShowPlaceholderText&amp;gt;&amp;lt;w:Compatibility&amp;gt;&amp;lt;w:BreakWrappedTables/&amp;gt;&amp;lt;w:SnapToGridInCell/&amp;gt;&amp;lt;w:WrapTextWithPunct/&amp;gt;&amp;lt;w:UseAsianBreakRules/&amp;gt;&amp;lt;w:DontGrowAutofit/&amp;gt;&amp;lt;/w:Compatibility&amp;gt;&amp;lt;w:BrowserLevel&amp;gt;MicrosoftInternetExplorer4&amp;lt;/w:BrowserLevel&amp;gt;&amp;lt;/w:WordDocument&amp;gt;&amp;lt;/xml&amp;gt;&amp;lt;![endif]--&amp;gt;&amp;lt;!--[if gte mso 9]&amp;gt;&amp;lt;xml&amp;gt;&amp;lt;w:LatentStyles DefLockedState=&amp;quot;false&amp;quot; LatentStyleCount=&amp;quot;156&amp;quot;&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Normal&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 1&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 2&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 3&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 4&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 5&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 6&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 7&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 8&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;heading 9&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 1&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 2&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 3&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 4&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 5&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 6&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 7&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 8&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;toc 9&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;caption&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Title&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Subtitle&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Strong&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Emphasis&amp;quot;/&amp;gt;&amp;lt;w:LsdException Locked=&amp;quot;true&amp;quot; Name=&amp;quot;Table Grid&amp;quot;/&amp;gt;&amp;lt;/w:LatentStyles&amp;gt;&amp;lt;/xml&amp;gt;&amp;lt;![endif]--&amp;gt;&amp;lt;style&amp;gt;&amp;lt;!--&lt;br /&gt;
/* Font Definitions */&lt;br /&gt;
@font-face&lt;br /&gt;
	{font-family:Wingdings;&lt;br /&gt;
	panose-1:5 0 0 0 0 0 0 0 0 0;&lt;br /&gt;
	mso-font-charset:2;&lt;br /&gt;
	mso-generic-font-family:auto;&lt;br /&gt;
	mso-font-pitch:variable;&lt;br /&gt;
	mso-font-signature:0 268435456 0 0 -2147483648 0;}&lt;br /&gt;
@font-face&lt;br /&gt;
	{font-family:Cambria;&lt;br /&gt;
	panose-1:2 4 5 3 5 4 6 3 2 4;&lt;br /&gt;
	mso-font-charset:0;&lt;br /&gt;
	mso-generic-font-family:roman;&lt;br /&gt;
	mso-font-pitch:variable;&lt;br /&gt;
	mso-font-signature:-1610611985 1073741899 0 0 415 0;}&lt;br /&gt;
@font-face&lt;br /&gt;
	{font-family:Calibri;&lt;br /&gt;
	panose-1:2 15 5 2 2 2 4 3 2 4;&lt;br /&gt;
	mso-font-charset:0;&lt;br /&gt;
	mso-generic-font-family:swiss;&lt;br /&gt;
	mso-font-pitch:variable;&lt;br /&gt;
	mso-font-signature:-520092929 1073786111 9 0 415 0;}&lt;br /&gt;
/* Style Definitions */&lt;br /&gt;
p.MsoNormal, li.MsoNormal, div.MsoNormal&lt;br /&gt;
	{mso-style-parent:&amp;quot;&amp;quot;;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:10.0pt;&lt;br /&gt;
	margin-left:0in;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
h1&lt;br /&gt;
	{mso-style-link:&amp;quot;Heading 1 Char&amp;quot;;&lt;br /&gt;
	mso-style-next:Normal;&lt;br /&gt;
	margin-top:24.0pt;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:0in;&lt;br /&gt;
	margin-left:0in;&lt;br /&gt;
	margin-bottom:.0001pt;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan lines-together;&lt;br /&gt;
	page-break-after:avoid;&lt;br /&gt;
	mso-outline-level:1;&lt;br /&gt;
	font-size:14.0pt;&lt;br /&gt;
	font-family:Cambria;&lt;br /&gt;
	color:#365F91;&lt;br /&gt;
	mso-font-kerning:0pt;}&lt;br /&gt;
h2&lt;br /&gt;
	{mso-style-noshow:yes;&lt;br /&gt;
	mso-style-link:&amp;quot;Heading 2 Char&amp;quot;;&lt;br /&gt;
	mso-style-next:Normal;&lt;br /&gt;
	margin-top:10.0pt;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:0in;&lt;br /&gt;
	margin-left:0in;&lt;br /&gt;
	margin-bottom:.0001pt;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan lines-together;&lt;br /&gt;
	page-break-after:avoid;&lt;br /&gt;
	mso-outline-level:2;&lt;br /&gt;
	font-size:13.0pt;&lt;br /&gt;
	font-family:Cambria;&lt;br /&gt;
	color:#4F81BD;}&lt;br /&gt;
a:link, span.MsoHyperlink&lt;br /&gt;
	{font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:blue;&lt;br /&gt;
	text-decoration:underline;&lt;br /&gt;
	text-underline:single;}&lt;br /&gt;
a:visited, span.MsoHyperlinkFollowed&lt;br /&gt;
	{mso-style-noshow:yes;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:purple;&lt;br /&gt;
	text-decoration:underline;&lt;br /&gt;
	text-underline:single;}&lt;br /&gt;
span.apple-style-span&lt;br /&gt;
	{mso-style-name:apple-style-span;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
span.toctext&lt;br /&gt;
	{mso-style-name:toctext;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.Quote, li.Quote, div.Quote&lt;br /&gt;
	{mso-style-name:Quote;&lt;br /&gt;
	mso-style-link:&amp;quot;Quote Char&amp;quot;;&lt;br /&gt;
	mso-style-next:Normal;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:10.0pt;&lt;br /&gt;
	margin-left:0in;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:black;&lt;br /&gt;
	font-style:italic;}&lt;br /&gt;
span.QuoteChar&lt;br /&gt;
	{mso-style-name:&amp;quot;Quote Char&amp;quot;;&lt;br /&gt;
	mso-style-locked:yes;&lt;br /&gt;
	mso-style-link:Quote;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:black;&lt;br /&gt;
	font-style:italic;}&lt;br /&gt;
span.Heading1Char&lt;br /&gt;
	{mso-style-name:&amp;quot;Heading 1 Char&amp;quot;;&lt;br /&gt;
	mso-style-locked:yes;&lt;br /&gt;
	mso-style-link:&amp;quot;Heading 1&amp;quot;;&lt;br /&gt;
	mso-ansi-font-size:14.0pt;&lt;br /&gt;
	mso-bidi-font-size:14.0pt;&lt;br /&gt;
	font-family:Cambria;&lt;br /&gt;
	mso-ascii-font-family:Cambria;&lt;br /&gt;
	mso-hansi-font-family:Cambria;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:#365F91;&lt;br /&gt;
	font-weight:bold;}&lt;br /&gt;
span.apple-converted-space&lt;br /&gt;
	{mso-style-name:apple-converted-space;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.ListParagraph, li.ListParagraph, div.ListParagraph&lt;br /&gt;
	{mso-style-name:&amp;quot;List Paragraph&amp;quot;;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:10.0pt;&lt;br /&gt;
	margin-left:.5in;&lt;br /&gt;
	mso-add-space:auto;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.ListParagraphCxSpFirst, li.ListParagraphCxSpFirst, div.ListParagraphCxSpFirst&lt;br /&gt;
	{mso-style-name:&amp;quot;List ParagraphCxSpFirst&amp;quot;;&lt;br /&gt;
	mso-style-type:export-only;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:0in;&lt;br /&gt;
	margin-left:.5in;&lt;br /&gt;
	margin-bottom:.0001pt;&lt;br /&gt;
	mso-add-space:auto;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.ListParagraphCxSpMiddle, li.ListParagraphCxSpMiddle, div.ListParagraphCxSpMiddle&lt;br /&gt;
	{mso-style-name:&amp;quot;List ParagraphCxSpMiddle&amp;quot;;&lt;br /&gt;
	mso-style-type:export-only;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:0in;&lt;br /&gt;
	margin-left:.5in;&lt;br /&gt;
	margin-bottom:.0001pt;&lt;br /&gt;
	mso-add-space:auto;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.ListParagraphCxSpLast, li.ListParagraphCxSpLast, div.ListParagraphCxSpLast&lt;br /&gt;
	{mso-style-name:&amp;quot;List ParagraphCxSpLast&amp;quot;;&lt;br /&gt;
	mso-style-type:export-only;&lt;br /&gt;
	margin-top:0in;&lt;br /&gt;
	margin-right:0in;&lt;br /&gt;
	margin-bottom:10.0pt;&lt;br /&gt;
	margin-left:.5in;&lt;br /&gt;
	mso-add-space:auto;&lt;br /&gt;
	line-height:115%;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
p.NoSpacing, li.NoSpacing, div.NoSpacing&lt;br /&gt;
	{mso-style-name:&amp;quot;No Spacing&amp;quot;;&lt;br /&gt;
	mso-style-parent:&amp;quot;&amp;quot;;&lt;br /&gt;
	margin:0in;&lt;br /&gt;
	margin-bottom:.0001pt;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:11.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-fareast-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
span.tocnumber&lt;br /&gt;
	{mso-style-name:tocnumber;&lt;br /&gt;
	font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
span.Heading2Char&lt;br /&gt;
	{mso-style-name:&amp;quot;Heading 2 Char&amp;quot;;&lt;br /&gt;
	mso-style-noshow:yes;&lt;br /&gt;
	mso-style-locked:yes;&lt;br /&gt;
	mso-style-link:&amp;quot;Heading 2&amp;quot;;&lt;br /&gt;
	mso-ansi-font-size:13.0pt;&lt;br /&gt;
	mso-bidi-font-size:13.0pt;&lt;br /&gt;
	font-family:Cambria;&lt;br /&gt;
	mso-ascii-font-family:Cambria;&lt;br /&gt;
	mso-hansi-font-family:Cambria;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
	color:#4F81BD;&lt;br /&gt;
	font-weight:bold;}&lt;br /&gt;
@page Section1&lt;br /&gt;
	{size:8.5in 11.0in;&lt;br /&gt;
	margin:1.0in 1.0in 1.0in 1.0in;&lt;br /&gt;
	mso-header-margin:.5in;&lt;br /&gt;
	mso-footer-margin:.5in;&lt;br /&gt;
	mso-paper-source:0;}&lt;br /&gt;
div.Section1&lt;br /&gt;
	{page:Section1;}&lt;br /&gt;
/* List Definitions */&lt;br /&gt;
@list l0&lt;br /&gt;
	{mso-list-id:220138155;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-935426840 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l0:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l1&lt;br /&gt;
	{mso-list-id:585001264;&lt;br /&gt;
	mso-list-template-ids:1185565082;}&lt;br /&gt;
@list l1:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:.5in;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-ansi-font-size:10.0pt;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
@list l2&lt;br /&gt;
	{mso-list-id:627711828;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-647192220 67698689 67698691 67698693 67698689 67698691 67698693 67698689 67698691 67698693;}&lt;br /&gt;
@list l2:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
@list l3&lt;br /&gt;
	{mso-list-id:631063392;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:288115314 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l3:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l4&lt;br /&gt;
	{mso-list-id:755058200;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-1450139246 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l4:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l5&lt;br /&gt;
	{mso-list-id:766925642;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-598935564 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l5:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l6&lt;br /&gt;
	{mso-list-id:1101224121;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-267988680 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l6:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l7&lt;br /&gt;
	{mso-list-id:1109931714;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-960715140 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l7:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l8&lt;br /&gt;
	{mso-list-id:1402295398;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:763423860 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l8:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l9&lt;br /&gt;
	{mso-list-id:1403603513;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:1497244868 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l9:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l10&lt;br /&gt;
	{mso-list-id:1520926471;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:1496089722 67698689 67698691 67698693 67698689 67698691 67698693 67698689 67698691 67698693;}&lt;br /&gt;
@list l10:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
@list l11&lt;br /&gt;
	{mso-list-id:1605575881;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-1087757402 67698703 67698713 67698715 67698703 67698713 67698715 67698703 67698713 67698715;}&lt;br /&gt;
@list l11:level1&lt;br /&gt;
	{mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-bidi-font-family:&amp;quot;Times New Roman&amp;quot;;}&lt;br /&gt;
@list l12&lt;br /&gt;
	{mso-list-id:1709915636;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:2042111048 67698689 67698691 67698693 67698689 67698691 67698693 67698689 67698691 67698693;}&lt;br /&gt;
@list l12:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
@list l13&lt;br /&gt;
	{mso-list-id:1745058118;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-1145408414 67698689 67698691 67698693 67698689 67698691 67698693 67698689 67698691 67698693;}&lt;br /&gt;
@list l13:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
@list l14&lt;br /&gt;
	{mso-list-id:1889145465;&lt;br /&gt;
	mso-list-template-ids:1736059350;}&lt;br /&gt;
@list l14:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0A7;&lt;br /&gt;
	mso-level-tab-stop:.5in;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	mso-ansi-font-size:10.0pt;&lt;br /&gt;
	font-family:Wingdings;}&lt;br /&gt;
@list l15&lt;br /&gt;
	{mso-list-id:1992056895;&lt;br /&gt;
	mso-list-type:hybrid;&lt;br /&gt;
	mso-list-template-ids:-1328883150 67698689 67698691 67698693 67698689 67698691 67698693 67698689 67698691 67698693;}&lt;br /&gt;
@list l15:level1&lt;br /&gt;
	{mso-level-number-format:bullet;&lt;br /&gt;
	mso-level-text:F0B7;&lt;br /&gt;
	mso-level-tab-stop:none;&lt;br /&gt;
	mso-level-number-position:left;&lt;br /&gt;
	text-indent:-.25in;&lt;br /&gt;
	font-family:Symbol;}&lt;br /&gt;
ol&lt;br /&gt;
	{margin-bottom:0in;}&lt;br /&gt;
ul&lt;br /&gt;
	{margin-bottom:0in;}&lt;br /&gt;
--&amp;gt;&amp;lt;/style&amp;gt;&amp;lt;!--[if gte mso 10]&amp;gt;&amp;lt;style&amp;gt;&lt;br /&gt;
/* Style Definitions */&lt;br /&gt;
table.MsoNormalTable&lt;br /&gt;
	{mso-style-name:&amp;quot;Table Normal&amp;quot;;&lt;br /&gt;
	mso-tstyle-rowband-size:0;&lt;br /&gt;
	mso-tstyle-colband-size:0;&lt;br /&gt;
	mso-style-noshow:yes;&lt;br /&gt;
	mso-style-parent:&amp;quot;&amp;quot;;&lt;br /&gt;
	mso-padding-alt:0in 5.4pt 0in 5.4pt;&lt;br /&gt;
	mso-para-margin:0in;&lt;br /&gt;
	mso-para-margin-bottom:.0001pt;&lt;br /&gt;
	mso-pagination:widow-orphan;&lt;br /&gt;
	font-size:10.0pt;&lt;br /&gt;
	font-family:Calibri;&lt;br /&gt;
	mso-ansi-language:#0400;&lt;br /&gt;
	mso-fareast-language:#0400;&lt;br /&gt;
	mso-bidi-language:#0400;}&lt;br /&gt;
&amp;lt;/style&amp;gt;&amp;lt;![endif]--&amp;gt;&amp;lt;/head&amp;gt;&amp;lt;body lang=EN-US link=blue vlink=purple style='tab-interval:.5in'&amp;gt;&amp;lt;div class=Section1&amp;gt;&amp;lt;p class=MsoNormal align=center style='text-align:center'&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;ECE633 Independent Study: Architecture of&lt;br /&gt;
Parallel Computers&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal align=center style='text-align:center'&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Karishma Navalakha&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Abstract:&amp;lt;span&lt;br /&gt;
style='mso-tab-count:1'&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;There&lt;br /&gt;
has been tremendous research and development in the field of multi-core&lt;br /&gt;
Architecture in the last decade. In such a dynamic environment it is very&lt;br /&gt;
difficult to have text books covering latest developments in the field. Wiki&lt;br /&gt;
written text books comes as an extremely handy tool for students to get&lt;br /&gt;
acquainted and interested in ongoing research.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;In this independent study we explored an academic learning technique&lt;br /&gt;
where students could learn the fundamental concepts of the subject through the&lt;br /&gt;
text book available to students and lectures delivered by Prof. Gehringer in&lt;br /&gt;
class. They can now build on this foundation and gather latest information from&lt;br /&gt;
the varied online resources and technical papers and summarize their findings&lt;br /&gt;
in the form of wiki pages.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Software is&lt;br /&gt;
also being currently developed to assist the students and was adopted in this&lt;br /&gt;
course. We tried to enhance the quality of student submitted wiki pages through&lt;br /&gt;
peer reviewing. Professor Gehringer and I constantly provided inputs to&lt;br /&gt;
students to improve both their quality of wiki pages as well as quality of&lt;br /&gt;
reviewing. The software being developed under the able guidance of professor&lt;br /&gt;
Gehringer has been vital in overcoming administrative hurdles involved in&lt;br /&gt;
assigning topics to students, maintaining the updates and tracking progress of&lt;br /&gt;
their writings, getting feedbacks through peer reviewing and handling the&lt;br /&gt;
re-submitted work. All this has been managed via the software in an organized&lt;br /&gt;
fashion.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Experience with Wiki&lt;br /&gt;
written text book:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;The&lt;br /&gt;
software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.&lt;br /&gt;
This is a beginning masters-level course that is taken by all Computer&lt;br /&gt;
Engineering masters students. It is optional for Computer Science students, but&lt;br /&gt;
as it is one way to fulfill a core requirement, it is popular with them too.&lt;br /&gt;
The recently adopted textbook for this course is the locally written&lt;br /&gt;
Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems&lt;br /&gt;
[Solihin 2009]. It did not make sense to have the students rewrite this excellent&lt;br /&gt;
text, but the book concentrates on theory and design fundamentals, without&lt;br /&gt;
detailed application to current parallel machines. We felt that students would&lt;br /&gt;
benefit from learning how the principles were applied in current architectures.&lt;br /&gt;
Furthermore, they would learn about the newest machines in this fast-changing&lt;br /&gt;
field.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;After&lt;br /&gt;
every chapter covered in class, two individuals, or pairs of students were&lt;br /&gt;
required to sign up for writing the wiki supplement for that particular&lt;br /&gt;
chapter. (That is, we solicited two supplements for each chapter, each of which&lt;br /&gt;
could be authored by one or two students.) They were asked to add specific&lt;br /&gt;
types of information which was not included in the chapter.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Initially,&lt;br /&gt;
students were not clear about the purpose of their wiki pages. The first pages&lt;br /&gt;
they wrote had substantial duplication of topics covered in the textbook.&lt;br /&gt;
Students were attempting to give a complete coverage of issues discussed in the&lt;br /&gt;
chapter. We wanted them to concentrate instead on recent developments. Upon&lt;br /&gt;
seeing this, we established the practice of having the first two authors of&lt;br /&gt;
this paper (Gehringer and Navalakha) review the student work, along with three&lt;br /&gt;
peer reviews from fellow students. A lot of review time was spent providing&lt;br /&gt;
guidance on how to revise.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;At&lt;br /&gt;
the beginning we gave the students complete freedom to explore resources for&lt;br /&gt;
the topic they had chosen to write on. This was not very successful, as the&lt;br /&gt;
students seemingly chose to read the first few search hits, which tended to&lt;br /&gt;
provide an overview of the topic, rather than in-depth information on&lt;br /&gt;
particular implementations. Sometimes students were not aware that the&lt;br /&gt;
information they found was already covered in the next chapter, which they have&lt;br /&gt;
not read yet. The first review which we gave students was mainly just making&lt;br /&gt;
them aware of topics covered in later chapters. A lot of effort in writing the&lt;br /&gt;
initial draft was thus wasted. After the first two sets of topics, we began to&lt;br /&gt;
provide links for students to material that we wanted the students to pay&lt;br /&gt;
attention to. Gehringer and Navalakha met weekly to discuss what to provide to&lt;br /&gt;
students. We regularly consulted other textbooks, technology news, and Web&lt;br /&gt;
sites of major processor manufacturers, such as Intel and AMD. As the semester&lt;br /&gt;
progressed, the quality of the initial submissions improved, and the students&lt;br /&gt;
realized better returns for their effort.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;The&lt;br /&gt;
quality of work seemed to improve as the semester progressed. A comparison of&lt;br /&gt;
the grades for the wiki pages revealed that the average score for the first&lt;br /&gt;
chapter written by each student was 82.8% while the average for the second&lt;br /&gt;
submission was 82.7%. The quality of wiki pages had improved, but at the same&lt;br /&gt;
time, the peer reviewers became more demanding. Students were given more inputs&lt;br /&gt;
to improve their work via peer reviewing. Thus the improvement was seen in the&lt;br /&gt;
final wiki page produced as against the grades received by students. The&lt;br /&gt;
initial wiki pages provided randomly collected data and was cluttered by&lt;br /&gt;
diagrams and graphs. This information reinstated facts given in the textbook.&lt;br /&gt;
The later wiki pages focused on a comparative study of present-day&lt;br /&gt;
supercomputers produced by Intel, AMD and IBM.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;For&lt;br /&gt;
example while writing the wiki for cache-coherence protocols, the students&lt;br /&gt;
examined which protocol was favored by which company and why. They also&lt;br /&gt;
discussed protocols which have been introduced in recent two years e.g.,&lt;br /&gt;
Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to&lt;br /&gt;
readers. Gehringer and Navalakha provided additional reviews which helped in&lt;br /&gt;
constantly improving the quality of wiki pages. These reviews gave the students&lt;br /&gt;
insight into what was expected expected of them. This led to an increasing&lt;br /&gt;
focus on current developments while peer reviewing. It was observed that later&lt;br /&gt;
versions of reviews included guidance similar to that received from Gehringer&lt;br /&gt;
and Navalakha. The organization of the wiki pages and the volume of relevant&lt;br /&gt;
data collected by students improved as the semester progressed.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Electronic&lt;br /&gt;
peer-review systems have been widely used to review student work, but never&lt;br /&gt;
before, to our knowledge, have they been applied to assignments consisting of&lt;br /&gt;
multiple interrelated parts with precedence constraints. The growing interest&lt;br /&gt;
in large collaborative projects, such as wiki textbooks, has led to a need for&lt;br /&gt;
electronic support for the process, lest the administrative burden on&lt;br /&gt;
instructor and TA grow too large.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Chapter wise learning from&lt;br /&gt;
this independent study:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Chapter 1: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;It&lt;br /&gt;
covered an interesting topic of supercomputer evolution. Wiki pages written for&lt;br /&gt;
this topic included a lot data from literature. Students came up with&lt;br /&gt;
interesting topics which were not covered in the text book such as &amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Timeline of supercomputers&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;, &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;First Supercomputer(ENIAC)&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;, &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Cray History&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;, &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Supercomputer Hierarchal Architecture&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;, &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Supercomputer Operating System&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;, &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Cooling Supercomputer&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt; and&lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic;&lt;br /&gt;
text-decoration:none;text-underline:none'&amp;gt;Processor Family&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=QuoteChar&amp;gt;&amp;lt;span style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;.&lt;br /&gt;
From their research we could see the increase in dominance of Intel’s&lt;br /&gt;
processors in the consumer market. We also conclude that Unix has been the&lt;br /&gt;
platform for most of these super computers. &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;Massive Parallel Processing&lt;br /&gt;
(MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely&lt;br /&gt;
used multiprocessor machine architectures which was replaced by constellation&lt;br /&gt;
computing in the 2000 and currently is dominated by cluster computing.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;References:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;a href=&amp;quot;http://www.top500.org/&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;http://www.top500.org/&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution20supercomputers&amp;amp;f=false&amp;quot;&lt;br /&gt;
title=&amp;quot;http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolu&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;The future of&lt;br /&gt;
supercomputing: an interim report By National Research Council (U.S.).&lt;br /&gt;
Committee on the Future of Supercomputing&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 2: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;Data Parallel Programming: &amp;lt;/span&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;The&lt;br /&gt;
students provided comparisons between data parallelism and task parallelism. &amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#002BB8'&amp;gt;Haveraaen (2000)&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt; notes that data parallel&lt;br /&gt;
codes typically bear a strong resemblance to sequential codes, making them&lt;br /&gt;
easier to read and write. &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;Students noted that the data parallel model may be used with the&lt;br /&gt;
shared memory or the message passing model without conflict. &amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;In their comparisons they concluded combining the data&lt;br /&gt;
parallel and message passing models results in reduction in the amount and&lt;br /&gt;
complexity of communication required relative to a task parallel approach.&lt;br /&gt;
Similarly, combining the data parallel and shared memory models tends to simplify&lt;br /&gt;
and reduce the amount of synchronization required. &amp;lt;span style='mso-bidi-font-style:&lt;br /&gt;
italic'&amp;gt;SIMD (single-instruction-multiple-data)&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;processors&lt;br /&gt;
are specifically designed to run data parallel algorithms.&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt; Modern examples include CUDA processors&lt;br /&gt;
developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and&lt;br /&gt;
IBM).&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l4 level1 lfo8'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;W.&lt;br /&gt;
Daniel Hillis and Guy L. Steele, Jr., &amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://portal.acm.org/citation.cfm?id=7903&amp;quot;&lt;br /&gt;
title=&amp;quot;http://portal.acm.org/citation.cfm?id=7903&amp;quot;&amp;gt;&amp;lt;span style='font-family:&lt;br /&gt;
&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;quot;Data parallel algorithms,&amp;quot;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt; Communications of the&lt;br /&gt;
ACM, 29(12):1170-1183, December 1986.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l4 level1 lfo8'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Alexander&lt;br /&gt;
C. Klaiber and Henry M. Levy, &amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://portal.acm.org/citation.cfm?id=192020&amp;quot;&lt;br /&gt;
title=&amp;quot;http://portal.acm.org/citation.cfm?id=192020&amp;quot;&amp;gt;&amp;lt;span style='font-family:&lt;br /&gt;
&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;quot;A comparison of message passing and shared memory&lt;br /&gt;
architectures for data parallel programs,&amp;quot;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt; in Proceedings of the 21st&lt;br /&gt;
Annual International Symposium on Computer Architecture, April 1994, pp.&lt;br /&gt;
94-105.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;Chapter3: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;In this wiki supplement, the three kinds of parallelisms,&lt;br /&gt;
i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism&lt;br /&gt;
techniques were discussed with examples in the form of Open MP code as&lt;br /&gt;
discussed in the text book. Besides the students provided additional depth in&lt;br /&gt;
this topic by discussing parallel_for, parallel_reduce,&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;parallel_scan,&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;pipeline,&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Reduction,&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;DOALL,&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;DOACROSS, DOPIPE with respect to Intel Thread&lt;br /&gt;
Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads.&lt;br /&gt;
Finally they conclude : Pthreads works for all the parallelism and could&lt;br /&gt;
express functional parallelism easily, but it needs to build specialized&lt;br /&gt;
synchronization primitives and explicitly privatize variables, makes it more&lt;br /&gt;
effort needed to switch a serial program in to parallel mode.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;OpenMP can provide many performance enhancing features,&lt;br /&gt;
such as atomic, barrier and flush synchronization primitives. It is very simple&lt;br /&gt;
to use OpenMP to exploit DOALL parallelism, but the syntax for expressing&lt;br /&gt;
functional parallelism is awkward.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Intel TBB relies on generic programming, it performs better&lt;br /&gt;
with custom iteration spaces or complex reduction operations. Also, it provides&lt;br /&gt;
generic parallel patterns for parallel while-loops, data-flow pipeline models,&lt;br /&gt;
parallel sorts and prefixes, so it's better in cases go beyond loop-based&lt;br /&gt;
parallelism.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l6 level1 lfo9'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https2F2Fmail3Fui26ik26view26th26attid26disp26realattid26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw&amp;quot;&lt;br /&gt;
title=&amp;quot;https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https2F2Fmail3Fui26ik26view26th26attid26disp26realattid%3Df_g602o&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;An Optimal Abtraction Model&lt;br /&gt;
for Hardware Multithreading in Modern Processor Architectures&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l6 level1 lfo9'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.threadingbuildingblocks.org/uploads/81/91/Latest20Source%20Documentation/Reference.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.threadingbuildingblocks.org/uploads/81/91/Latest20Source%20Documentation/Reference.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Intel Threading Building&lt;br /&gt;
Blocks 2.2 for Open Source Reference Manual&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l6 level1 lfo9'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-list:Ignore'&amp;gt;3.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;https://computing.llnl.gov/tutorials/pthreads/#Joining&amp;quot;&lt;br /&gt;
title=&amp;quot;https://computing.llnl.gov/tutorials/pthreads/#Joining&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;POSIX Threads Programming&lt;br /&gt;
by Blaise Barney, Lawrence Livermore National Laboratory&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 6:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt;Cache Structures of&lt;br /&gt;
Multi-Core Architectures: Students added additional insight on this topic by&lt;br /&gt;
discussing Shared Memory Multiprocessors, write policies and replacement&lt;br /&gt;
policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was&lt;br /&gt;
an additional subtopic students threw light on. Students also gave definitions&lt;br /&gt;
about Trace Cache and Smart Cache techniques by Intel.&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;The most important take away from this topic&lt;br /&gt;
was how students discussed WRITE POLICIES used in recent multi core&lt;br /&gt;
architectures. For example, Intel IA 32 IA64 architecture implements Write&lt;br /&gt;
Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No&lt;br /&gt;
Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion&lt;br /&gt;
unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT,&lt;br /&gt;
with allocate on load and noallocate on stores.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l9 level1 lfo10'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://download.intel.com/technology/architecture/sma.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://download.intel.com/technology/architecture/sma.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;http://download.intel.com/technology/architecture/sma.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l9 level1 lfo10'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.intel.com/Assets/PDF/manual/248966.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.intel.com/Assets/PDF/manual/248966.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;http://www.intel.com/Assets/PDF/manual/248966.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l9 level1 lfo10'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;3.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.intel.com/design/intarch/papers/cache6.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.intel.com/design/intarch/papers/cache6.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;http://www.intel.com/design/intarch/papers/cache6.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 7: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Shared-memory multiprocessors run into several problems&lt;br /&gt;
that are more pronounced than their uniprocessor counterparts. The Solihin text&lt;br /&gt;
used in this course goes into detail on three of these issues, that is cache&lt;br /&gt;
coherence, memory consistency and synchronization.&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;The goal of this wiki supplement was to&lt;br /&gt;
discuss these three issues and also what can be done to ensure that&lt;br /&gt;
instructions are handled in both a timely and efficient manner and in a manner&lt;br /&gt;
that is consistent with what the programmer might desire. Memory consistency&lt;br /&gt;
was discussed by comparing ordering on a uniprocessor vs ordering on a&lt;br /&gt;
multiprocessor. They concluded that in a multiprocessor much more care must be&lt;br /&gt;
taken to ensure that all of the loads and stores are committed to memory in a&lt;br /&gt;
valid order. Synchronization was discussed as applicable to Open MP and fence&lt;br /&gt;
insertion. Other methods such as test and set method and direct interrupt to&lt;br /&gt;
another core were also briefly discussed. &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span style='font-family:&lt;br /&gt;
&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;The programmer (or complier) is responsible for&lt;br /&gt;
knowing which synchronization directives are available on a given architecture&lt;br /&gt;
and implementing them in an efficient manner. The students also discussed commonly&lt;br /&gt;
used instructions for synchronization in popular processor architectures. For&lt;br /&gt;
example &amp;lt;span class=apple-style-span&amp;gt;SPARC V8 uses store barrier, Alpha uses&lt;br /&gt;
memory barrier and write memory barrier whereas Intel x86 uses lfence (load)&lt;br /&gt;
sfence (store).&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l11 level1 lfo11'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l11 level1 lfo11'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790&amp;quot;&lt;br /&gt;
title=&amp;quot;http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 8:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Students discussed the existing bus-based cache coherence&lt;br /&gt;
in real machines. They went ahead and classified the cache coherence protocols&lt;br /&gt;
based on the year they were introduced and they processors which uses them. MSI&lt;br /&gt;
protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is&lt;br /&gt;
called D (Dirty) but works the same as MSI protocol works. MSI has a major&lt;br /&gt;
drawback in that each read-write sequence incurs 2 bus transactions&lt;br /&gt;
irrespective of whether the cache line is stored in only one cache or not. The&lt;br /&gt;
Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture&lt;br /&gt;
microprocessor to support SMP and MESI. The MESIF protocol, used in the latest&lt;br /&gt;
Intel multi-core processors was introduced to accommodate the point-to-point&lt;br /&gt;
links used in the QuickPath Interconnect.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;MESI came with the drawback of using much time and bandwidth. MOESI was&lt;br /&gt;
the AMD’s answer to this problem . MOESI' has become one of the most popular&lt;br /&gt;
snoop-based protocols supported in the AMD64 architecture. The AMD dual-core&lt;br /&gt;
Opteron can maintain cache coherence in systems up to 8 processors using this&lt;br /&gt;
protocol. The Dragon Protocol is an update based coherence protocol which does&lt;br /&gt;
not invalidate other cached copies. The Dragon Protocol , was developed by&lt;br /&gt;
Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation.&lt;br /&gt;
This protocol was used in the Xerox PARC Dragon multiprocessor workstation. &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;References:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l7 level1 lfo12'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache20&amp;amp;%20MESI.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache20&amp;amp;%20MESI.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Cache consistency with MESI&lt;br /&gt;
on Intel processor&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l7 level1 lfo12'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://techreport.com/articles.x/8236/2&amp;quot;&lt;br /&gt;
title=&amp;quot;http://techreport.com/articles.x/8236/2&amp;quot;&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:#3366BB'&amp;gt;AMD dual core Architecture&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l7 level1 lfo12'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;3.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913&amp;quot;&lt;br /&gt;
title=&amp;quot;http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Silicon Graphics Computer&lt;br /&gt;
Systems&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l7 level1 lfo12'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;4.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533&amp;quot;&lt;br /&gt;
title=&amp;quot;http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Synapse tightly coupled&lt;br /&gt;
multiprocessors: a new approach to solve old problems&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l7 level1 lfo12'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;5.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://en.wikipedia.org/wiki/Dragon_protocol&amp;quot;&lt;br /&gt;
title=&amp;quot;http://en.wikipedia.org/wiki/Dragon_protocol&amp;quot;&amp;gt;&amp;lt;span style='font-family:&lt;br /&gt;
&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Dragon Protocol&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify'&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 9:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Synchronization: Students classified synchronization&lt;br /&gt;
techniques based on implementation. Hardware synchronization uses locks,&lt;br /&gt;
barriers and mutual exclusion. Software synchronization examples include ticket&lt;br /&gt;
locks and queue-based MCS locks.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Mutex&lt;br /&gt;
implementation uses execution of atomic statements. &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Some common examples include &amp;lt;span style='mso-bidi-font-weight:&lt;br /&gt;
bold'&amp;gt;Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. &amp;lt;/span&amp;gt;Another&lt;br /&gt;
type of lock that was not discussed in the text is known as the&lt;br /&gt;
&amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. They also&lt;br /&gt;
discussed reasons why a programmer should attempt to write programs in such a&lt;br /&gt;
way as to avoid locks. There are API's that exist for parallel architectures&lt;br /&gt;
that provide specific types of synchronization. If the API are used they way&lt;br /&gt;
they were design, performance can be maximized while minimizing overhead.Load&lt;br /&gt;
Locked(LL) and Store Conditional(SC) are a pair of instructions are improved&lt;br /&gt;
hardware primitives that are used for lock-free read-modify-write operation.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Detailed description of &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;Combining Tree Barrier,&lt;br /&gt;
Tournament Barrier and Disseminating Barrier was included. One of the&lt;br /&gt;
interesting topics discussed in this wiki supplement was the performance&lt;br /&gt;
evaluation of different barrier implementations. They showed that&lt;br /&gt;
barrier/centralized blocking barrier does not scale with number of threads and&lt;br /&gt;
the contention increases with increase in number of threads.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;References:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l8 level1 lfo13'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l8 level1 lfo13'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.statemaster.com/encyclopedia/Deadlock&amp;quot;&amp;gt;&amp;lt;span style='font-family:&lt;br /&gt;
&amp;quot;Times New Roman&amp;quot;'&amp;gt;http://www.statemaster.com/encyclopedia/Deadlock&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l8 level1 lfo13'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;3.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.ukhec.ac.uk/publications/reports/synch_java.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;http://www.ukhec.ac.uk/publications/reports/synch_java.pdf&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;Chapter 10: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Students discussed the existing bus-based cache coherence&lt;br /&gt;
in real machines. They went ahead and classified the cache coherence protocols&lt;br /&gt;
based on the year they were introduced and they processors which uses them. MSI&lt;br /&gt;
protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is&lt;br /&gt;
called D (Dirty) but works the same as MSI protocol works. MSI has a major&lt;br /&gt;
drawback in that each read-write sequence incurs 2 bus transactions&lt;br /&gt;
irrespective of whether the cache line is stored in only one cache or not. The&lt;br /&gt;
Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture&lt;br /&gt;
microprocessor to support SMP and MESI. The MESIF protocol, used in the latest&lt;br /&gt;
Intel multi-core processors was introduced to accommodate the point-to-point&lt;br /&gt;
links used in the QuickPath Interconnect.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;MESI came with the drawback of using much time and bandwidth. MOESI was&lt;br /&gt;
the AMD’s answer to this problem . MOESI' has become one of the most popular&lt;br /&gt;
snoop-based protocols supported in the AMD64 architecture. The AMD dual-core&lt;br /&gt;
Opteron can maintain cache coherence in systems up to 8 processors using this&lt;br /&gt;
protocol. The Dragon Protocol is an update based coherence protocol which does&lt;br /&gt;
not invalidate other cached copies. The Dragon Protocol , was developed by&lt;br /&gt;
Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation.&lt;br /&gt;
This protocol was used in the Xerox PARC Dragon multiprocessor workstation.&lt;br /&gt;
References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l5 level1 lfo14'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Shared Memory Consistency&lt;br /&gt;
Models&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpMiddle style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l5 level1 lfo14'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273&amp;quot;&lt;br /&gt;
title=&amp;quot;http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Designing Memory&lt;br /&gt;
Consistency Models For Shared-Memory Multiprocessors&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l5 level1 lfo14'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-list:Ignore'&amp;gt;3.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html&amp;quot;&lt;br /&gt;
title=&amp;quot;http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;Consistency Models&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 11:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;The cache coherence protocol presented in Chapter 11 of&lt;br /&gt;
Solihin 2008 is simpler than most real directory-based protocols. This textbook&lt;br /&gt;
supplement presents the directory-based protocols used by the DASH&lt;br /&gt;
multiprocessor and the Alewife multiprocessor. It concludes with an argument of&lt;br /&gt;
why complexity might be undesirable in cache coherence protocols. The DASH&lt;br /&gt;
multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to&lt;br /&gt;
ensure cache coherence within cluster and a directory-based protocol to ensure&lt;br /&gt;
coherence across clusters. The protocol uses a Remote Access Cache (RAC) at&lt;br /&gt;
each cluster, which essentially consolidates memory blocks from remote clusters&lt;br /&gt;
into a single cache on the local snoopy bus. When a request is issued for a&lt;br /&gt;
block from a remote cluster that is not in the RAC, the request is denied but&lt;br /&gt;
the request is also forwarded to the owner. The owner supplies the block to the&lt;br /&gt;
RAC. Eventually, when the requestor retries, the block will be waiting in the&lt;br /&gt;
RAC. Read and readx operations on a Dash processor were discussed in detail.&lt;br /&gt;
They also discuss two race conditions which mainly arises on a Dash&lt;br /&gt;
processor.The first occurs when a Read from requester R is forwarded from home&lt;br /&gt;
H to owner O, but O sends a Writeback to H before the forwarded Read arrives.&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;Another possible race occurs&lt;br /&gt;
when the home node H replies with data (ReplyD) to a Read from requester R but&lt;br /&gt;
an invalidation (Inv) arrives first.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;LimitLESS is the cache coherence protocol used by the Alewife&lt;br /&gt;
multiprocessor.&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;   &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Unlike the DASH multiprocessor, the Alewife multiprocessor&lt;br /&gt;
is not organized into clusters of nodes with local buses, and therefore cache&lt;br /&gt;
coherence through the system is maintain through the directory. &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l0 level1 lfo15'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;i&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/i&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop&lt;br /&gt;
Gupta, and John Hennessy (1990).&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://doi.acm.org/10.1145/325164.325132&amp;quot;&lt;br /&gt;
title=&amp;quot;http://doi.acm.org/10.1145/325164.325132&amp;quot;&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;&lt;br /&gt;
color:#3366BB;text-decoration:none;text-underline:none'&amp;gt;&amp;quot;The&lt;br /&gt;
directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;In&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;i&amp;gt;&amp;lt;span style='color:black'&amp;gt;Proceedings of the 17th&lt;br /&gt;
Annual International Symposium on Computer Architecture.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/i&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l0 level1 lfo15'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;David Chaiken, John Kubiatowicz, and Anant Agarwal (1991).&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-converted-space&amp;gt;&amp;lt;span style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf&amp;quot;&lt;br /&gt;
title=&amp;quot;http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB;text-decoration:none;&lt;br /&gt;
text-underline:none'&amp;gt;&amp;quot;LimitLESS directories: A scalable cache coherence&lt;br /&gt;
scheme.&amp;quot;&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span class=apple-converted-space&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt; &amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;i&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;ACM SIGPLAN Notices&amp;lt;/span&amp;gt;&amp;lt;/i&amp;gt;&amp;lt;span style='color:black'&amp;gt;.&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;b&lt;br /&gt;
style='mso-bidi-font-weight:normal'&amp;gt;&amp;lt;span style='color:black'&amp;gt;Chapter 12:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;Interconnection Networks: Advances in multiprocessors,&lt;br /&gt;
parallel computing &amp;amp; networking and parallel computer architectures demand&lt;br /&gt;
very high performance from interconnection networks. Due to this,&lt;br /&gt;
interconnection network structure has changed over time, trying to meet higher&lt;br /&gt;
bandwidths and performance. Students discussed criterion to be considered for&lt;br /&gt;
choosing the best Network. It included Performance Requirements, Scalability,&lt;br /&gt;
Incremental expandability, Partitionability, Simplicity, Distance Span,&lt;br /&gt;
Physical Constraints, Reliability and Reparability, Expected Workloads and Cost&lt;br /&gt;
Constraints. They provided in depth discussion on Classification of&lt;br /&gt;
Interconnection networks.&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Shared-Medium&lt;br /&gt;
Networks include Token Ring, Token Bus,&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Backplane Bus. Direct Networks include Mesh, Torus, Hypercube,&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Tree, Cube-Connected Cycles and&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;de Bruijn and Star Graph Networks.&amp;lt;span&lt;br /&gt;
style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Indirect Networks include Regular Topologies&lt;br /&gt;
like Crossbar Network and Multistage Interconnection Network and Hybrid&lt;br /&gt;
Networks such as&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Multiple Backplane&lt;br /&gt;
Buses, Hierarchical Networks,&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;Cluster-Based Networks and Hypergraph Topologies. They also discussed&lt;br /&gt;
routing algorithms and deadlock, starvation and livelock associated with it.&lt;br /&gt;
These topics were covered in in an extremely detailed way. The students&lt;br /&gt;
included a diagrammatic&amp;lt;span style='mso-spacerun:yes'&amp;gt;&amp;lt;/span&amp;gt;representation&lt;br /&gt;
for every topology. &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;References: &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpFirst style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l3 level1 lfo16'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;span style='mso-list:Ignore'&amp;gt;1.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.top500.org/2007_overview_recent_supercomputers/sci&amp;quot;&lt;br /&gt;
title=&amp;quot;http://www.top500.org/2007_overview_recent_supercomputers/sci&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;;color:#3366BB'&amp;gt;http://www.top500.org/2007_overview_recent_supercomputers/sci&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
class=apple-style-span&amp;gt;&amp;lt;span style='color:black'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=ListParagraphCxSpLast style='text-align:justify;text-indent:-.25in;&lt;br /&gt;
mso-list:l3 level1 lfo16'&amp;gt;&amp;lt;![if !supportLists]&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;span&lt;br /&gt;
style='mso-list:Ignore'&amp;gt;2.&amp;lt;span style='font:7.0pt &amp;quot;Times New Roman&amp;quot;'&amp;gt;     &lt;br /&gt;
&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;![endif]&amp;gt;&amp;lt;a&lt;br /&gt;
href=&amp;quot;http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html&amp;quot;&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html&amp;lt;/span&amp;gt;&amp;lt;/a&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-family:&amp;quot;Times New Roman&amp;quot;'&amp;gt;&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;b style='mso-bidi-font-weight:&lt;br /&gt;
normal'&amp;gt;&amp;lt;span style='font-family:&amp;quot;Times New Roman&amp;quot;;color:black'&amp;gt;Conclusion:&amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;This independent study helped me to increase my knowledge&lt;br /&gt;
to a great extent in the field of Architecture of parallel computers. There&lt;br /&gt;
were 4 students working on every chapter and came up with 2 wiki pages per&lt;br /&gt;
group. We collected a total of 18 wiki supplements. The data collected was&lt;br /&gt;
enormous. While reviewing their content I kept updating my knowledge base. I&lt;br /&gt;
also provided the resources from where they can collect data. This helped me to&lt;br /&gt;
come across latest developments in the field. Interacting with students helped&lt;br /&gt;
me to increase my communication skills. Constant discussions with Prof.&lt;br /&gt;
Gehringer helped me to understand key concepts. This idea of writing wiki&lt;br /&gt;
supplements got selected for KU Village presentation. I got an opportunity to&lt;br /&gt;
present this paper along with Prof. Gehringer. &amp;lt;o:p&amp;gt;&amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=apple-style-span&amp;gt;&amp;lt;span&lt;br /&gt;
style='color:black'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=MsoNormal style='text-align:justify'&amp;gt;&amp;lt;span class=QuoteChar&amp;gt;&amp;lt;span&lt;br /&gt;
style='font-style:normal;mso-bidi-font-style:italic'&amp;gt;&amp;lt;o:p&amp;gt; &amp;lt;/o:p&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;/div&amp;gt;&lt;br /&gt;
		&amp;lt;!-- END: body --&amp;gt;&lt;br /&gt;
		&lt;br /&gt;
		&amp;lt;a href=http://www.milonic.com/&amp;gt;&amp;lt;font color=&amp;quot;#FFFFFF&amp;quot;&amp;gt;JavaScript Menu Courtesy of Milonic.com&amp;lt;/font&amp;gt;&amp;lt;/a&amp;gt;&lt;br /&gt;
	&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot; width=&amp;quot;10&amp;quot;&amp;gt;&amp;lt;img src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;10&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot;&amp;gt;&lt;br /&gt;
    	&amp;lt;!-- INIT: left_body --&amp;gt;&lt;br /&gt;
    	&lt;br /&gt;
&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Site Utility&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/printer_version.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&lt;br /&gt;
           &amp;lt;a href=&amp;quot;/en/software/os/i_love_wiki/index.mpl?print=1&amp;amp;&amp;quot;&amp;gt;Printer version&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/flag_italy.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
          &amp;lt;a href=&amp;quot;/it/software/os/i_love_wiki/index.mpl?&amp;quot;&amp;gt;Leggilo in italiano&amp;lt;/a&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/stats.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a target=&amp;quot;_blank&amp;quot; href=&amp;quot;/cgi-bin/perl/awstats/awstats.pl?lang=en&amp;quot;&amp;gt;Site stats&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/stats.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a href=&amp;quot;/phpBB2/index.php&amp;quot;&amp;gt;Read forums&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/sitemap.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a href=&amp;quot;/sitemap.mpl&amp;quot;&amp;gt;Site map&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;&lt;br /&gt;
    &amp;lt;br&amp;gt;&lt;br /&gt;
            &lt;br /&gt;
        &amp;lt;!-- INIT: Google adSense --&amp;gt;&lt;br /&gt;
        &amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Ads by &lt;br /&gt;
  		&amp;lt;font color=&amp;quot;#2168E0&amp;quot;&amp;gt;G&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#D8240C&amp;quot;&amp;gt;o&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#F4C513&amp;quot;&amp;gt;o&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#2168E0&amp;quot;&amp;gt;g&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#2C9F2C&amp;quot;&amp;gt;l&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#D70000&amp;quot;&amp;gt;e&amp;lt;/font&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt; &lt;br /&gt;
        &amp;lt;script type=&amp;quot;text/javascript&amp;quot;&amp;gt;&amp;lt;!--&lt;br /&gt;
google_ad_client = &amp;quot;pub-6275448308915555&amp;quot;;&lt;br /&gt;
google_ad_width = 120;&lt;br /&gt;
google_ad_height = 240;&lt;br /&gt;
google_ad_format = &amp;quot;120x240_as&amp;quot;;&lt;br /&gt;
google_ad_type = &amp;quot;text_image&amp;quot;;&lt;br /&gt;
google_ad_channel =&amp;quot;&amp;quot;;&lt;br /&gt;
google_color_border = &amp;quot;FFFFFF&amp;quot;;&lt;br /&gt;
google_color_bg = &amp;quot;FFFFFF&amp;quot;;&lt;br /&gt;
google_color_link = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_url = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_text = &amp;quot;000000&amp;quot;;&lt;br /&gt;
//--&amp;gt;&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&lt;br /&gt;
  src=&amp;quot;http://pagead2.googlesyndication.com/pagead/show_ads.js&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;/script&amp;gt;&lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;   &lt;br /&gt;
&amp;lt;!-- END: Google adSense --&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
            &amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;&amp;lt;font color=&amp;quot;#dc9529&amp;quot;&amp;gt;&lt;br /&gt;
		&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/chat_with_me.gif&amp;quot; width=&amp;quot;27&amp;quot; height=&amp;quot;16&amp;quot;&amp;gt; Chat with me&amp;lt;/font&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#dc9529&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#dc9529&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt; &lt;br /&gt;
       	&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        	My nickname is &amp;lt;b&amp;gt;emi &amp;lt;/b&amp;gt;but now I am &amp;lt;font color=&amp;quot;#FF00FF&amp;quot;&amp;gt;&amp;lt;b&amp;gt;offline&amp;lt;/b&amp;gt;&amp;lt;/font&amp;gt;.&lt;br /&gt;
			However you can usually find me on these channels of the&lt;br /&gt;
			&amp;lt;a href=&amp;quot;http://www.azzurra.org/&amp;quot;&amp;gt;Azzurra&amp;lt;/a&amp;gt; networks:&amp;lt;/p&amp;gt;&lt;br /&gt;
		&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;span style=&amp;quot;font-size: 7pt&amp;quot;&amp;gt;&amp;lt;a href=&amp;quot;irc://irc.azzurra.net/areanetworking&amp;quot;&amp;gt;areanetworking&amp;lt;/a&amp;gt;, &lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/telug&amp;quot;&amp;gt;telug&amp;lt;/a&amp;gt;,&lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/controguerra&amp;quot;&amp;gt;controguerra&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/webgui&amp;quot;&amp;gt;webgui&amp;lt;/a&amp;gt;, &lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/geeks&amp;quot;&amp;gt;geeks&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/pescaralug&amp;quot;&amp;gt;pescaralug&amp;lt;/a&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt; &lt;br /&gt;
        &lt;br /&gt;
    &lt;br /&gt;
	&amp;lt;br&amp;gt;&lt;br /&gt;
	&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Credits&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo&amp;quot;&amp;gt;&lt;br /&gt;
	&amp;lt;a href=&amp;quot;http://www.masonhq.com/&amp;quot;&amp;gt;Mason&amp;lt;/a&amp;gt;&lt;br /&gt;
&amp;lt;/p&amp;gt;&lt;br /&gt;
&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
	All this website has built using Mason. Without this language derived from Perl,&lt;br /&gt;
	most of the tools used for the manage of this site would never been developed.&lt;br /&gt;
	&amp;lt;/p&amp;gt;&lt;br /&gt;
    &amp;lt;p class=&amp;quot;menu_right_titolo&amp;quot;&amp;gt;&lt;br /&gt;
	&amp;lt;a href=http://www.milonic.com/&amp;gt;JavaScript Menu&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
	&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
	 The great menu at the top of this page.&amp;lt;/p&amp;gt;&lt;br /&gt;
    &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;&lt;br /&gt;
	&lt;br /&gt;
    	&amp;lt;!-- END:  left_body --&amp;gt;&lt;br /&gt;
    &amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot; width=&amp;quot;8&amp;quot;&amp;gt;&amp;lt;img src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;8&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
  &amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&amp;lt;!-- INIT: Comments --&amp;gt;&lt;br /&gt;
		&lt;br /&gt;
		&lt;br /&gt;
&amp;lt;!-- END:  Comments --&amp;gt;&lt;br /&gt;
&amp;lt;p&amp;gt;&lt;br /&gt;
&amp;lt;!-- INIT: Google adSense --&amp;gt;&lt;br /&gt;
&amp;lt;table align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;&lt;br /&gt;
&amp;lt;td align=&amp;quot;bottom&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&amp;gt;&amp;lt;!--&lt;br /&gt;
google_ad_client = &amp;quot;pub-6275448308915555&amp;quot;;&lt;br /&gt;
google_ad_width = 728;&lt;br /&gt;
google_ad_height = 90;&lt;br /&gt;
google_ad_format = &amp;quot;728x90_as&amp;quot;;&lt;br /&gt;
google_ad_type = &amp;quot;text_image&amp;quot;;&lt;br /&gt;
google_ad_channel =&amp;quot;&amp;quot;;&lt;br /&gt;
google_color_border = &amp;quot;4575A3&amp;quot;;&lt;br /&gt;
google_color_bg = &amp;quot;D1E9FF&amp;quot;;&lt;br /&gt;
google_color_link = &amp;quot;0033CC&amp;quot;;&lt;br /&gt;
google_color_url = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_text = &amp;quot;000000&amp;quot;;&lt;br /&gt;
//--&amp;gt;&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&lt;br /&gt;
  src=&amp;quot;http://pagead2.googlesyndication.com/pagead/show_ads.js&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;/td&amp;gt;&lt;br /&gt;
&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&amp;lt;!-- END: Google adSense --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;100%&amp;quot;&amp;gt;&lt;br /&gt;
  &amp;lt;tr&amp;gt;&amp;lt;td colspan=&amp;quot;3&amp;quot; bgcolor=&amp;quot;#4575A3&amp;quot; class=&amp;quot;linea&amp;quot; HEIGHT=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
  &amp;lt;tr&amp;gt;&lt;br /&gt;
    &amp;lt;td&amp;gt; Copyright© 1997-2006 Emiliano Bruni&amp;lt;/td&amp;gt;&lt;br /&gt;
	&amp;lt;td align=&amp;quot;center&amp;quot;&amp;gt;Online from 16/08/1998 with &amp;lt;img src=&amp;quot;/cgi-bin/c/Count.cgi?ft=0&amp;amp;df=ebruni.it.dat&amp;amp;comma=T&amp;amp;md=8&amp;amp;pad=T&amp;amp;dd=verdana&amp;quot; alt=&amp;quot;&amp;quot; align=&amp;quot;bottom&amp;quot;&amp;gt;&lt;br /&gt;
 visitors&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td align=&amp;quot;right&amp;quot;&amp;gt;Write me to:&lt;br /&gt;
    &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/mail.gif&amp;quot; width=&amp;quot;77&amp;quot; height=&amp;quot;10&amp;quot; align=&amp;quot;baseline&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
  &amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&lt;br /&gt;
		&amp;lt;!-- END: body --&amp;gt;&lt;br /&gt;
		&lt;br /&gt;
		&amp;lt;a href=http://www.milonic.com/&amp;gt;&amp;lt;font color=&amp;quot;#FFFFFF&amp;quot;&amp;gt;JavaScript Menu Courtesy of Milonic.com&amp;lt;/font&amp;gt;&amp;lt;/a&amp;gt;&lt;br /&gt;
	&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot; width=&amp;quot;10&amp;quot;&amp;gt;&amp;lt;img src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;10&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot;&amp;gt;&lt;br /&gt;
    	&amp;lt;!-- INIT: left_body --&amp;gt;&lt;br /&gt;
    	&lt;br /&gt;
&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Site Utility&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/printer_version.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&lt;br /&gt;
           &amp;lt;a href=&amp;quot;/en/software/os/i_love_wiki/index.mpl?print=1&amp;amp;&amp;quot;&amp;gt;Printer version&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/flag_italy.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
          &amp;lt;a href=&amp;quot;/it/software/os/i_love_wiki/index.mpl?&amp;quot;&amp;gt;Leggilo in italiano&amp;lt;/a&amp;gt;&amp;lt;/b&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/stats.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a target=&amp;quot;_blank&amp;quot; href=&amp;quot;/cgi-bin/perl/awstats/awstats.pl?lang=en&amp;quot;&amp;gt;Site stats&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/stats.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a href=&amp;quot;/phpBB2/index.php&amp;quot;&amp;gt;Read forums&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/sitemap.gif&amp;quot; width=&amp;quot;14&amp;quot; height=&amp;quot;10&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;a href=&amp;quot;/sitemap.mpl&amp;quot;&amp;gt;Site map&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;&lt;br /&gt;
    &amp;lt;br&amp;gt;&lt;br /&gt;
            &lt;br /&gt;
        &amp;lt;!-- INIT: Google adSense --&amp;gt;&lt;br /&gt;
        &amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Ads by &lt;br /&gt;
  		&amp;lt;font color=&amp;quot;#2168E0&amp;quot;&amp;gt;G&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#D8240C&amp;quot;&amp;gt;o&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#F4C513&amp;quot;&amp;gt;o&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#2168E0&amp;quot;&amp;gt;g&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#2C9F2C&amp;quot;&amp;gt;l&amp;lt;/font&amp;gt;&amp;lt;font color=&amp;quot;#D70000&amp;quot;&amp;gt;e&amp;lt;/font&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt; &lt;br /&gt;
        &amp;lt;script type=&amp;quot;text/javascript&amp;quot;&amp;gt;&amp;lt;!--&lt;br /&gt;
google_ad_client = &amp;quot;pub-6275448308915555&amp;quot;;&lt;br /&gt;
google_ad_width = 120;&lt;br /&gt;
google_ad_height = 240;&lt;br /&gt;
google_ad_format = &amp;quot;120x240_as&amp;quot;;&lt;br /&gt;
google_ad_type = &amp;quot;text_image&amp;quot;;&lt;br /&gt;
google_ad_channel =&amp;quot;&amp;quot;;&lt;br /&gt;
google_color_border = &amp;quot;FFFFFF&amp;quot;;&lt;br /&gt;
google_color_bg = &amp;quot;FFFFFF&amp;quot;;&lt;br /&gt;
google_color_link = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_url = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_text = &amp;quot;000000&amp;quot;;&lt;br /&gt;
//--&amp;gt;&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&lt;br /&gt;
  src=&amp;quot;http://pagead2.googlesyndication.com/pagead/show_ads.js&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;/script&amp;gt;&lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;   &lt;br /&gt;
&amp;lt;!-- END: Google adSense --&amp;gt;&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
            &amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#dc9529&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;&amp;lt;font color=&amp;quot;#dc9529&amp;quot;&amp;gt;&lt;br /&gt;
		&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/chat_with_me.gif&amp;quot; width=&amp;quot;27&amp;quot; height=&amp;quot;16&amp;quot;&amp;gt; Chat with me&amp;lt;/font&amp;gt;&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#dc9529&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#dc9529&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt; &lt;br /&gt;
       	&lt;br /&gt;
        &amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        	My nickname is &amp;lt;b&amp;gt;emi &amp;lt;/b&amp;gt;but now I am &amp;lt;font color=&amp;quot;#FF00FF&amp;quot;&amp;gt;&amp;lt;b&amp;gt;offline&amp;lt;/b&amp;gt;&amp;lt;/font&amp;gt;.&lt;br /&gt;
			However you can usually find me on these channels of the&lt;br /&gt;
			&amp;lt;a href=&amp;quot;http://www.azzurra.org/&amp;quot;&amp;gt;Azzurra&amp;lt;/a&amp;gt; networks:&amp;lt;/p&amp;gt;&lt;br /&gt;
		&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
        &amp;lt;span style=&amp;quot;font-size: 7pt&amp;quot;&amp;gt;&amp;lt;a href=&amp;quot;irc://irc.azzurra.net/areanetworking&amp;quot;&amp;gt;areanetworking&amp;lt;/a&amp;gt;, &lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/telug&amp;quot;&amp;gt;telug&amp;lt;/a&amp;gt;,&lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/controguerra&amp;quot;&amp;gt;controguerra&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/webgui&amp;quot;&amp;gt;webgui&amp;lt;/a&amp;gt;, &lt;br /&gt;
        &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/geeks&amp;quot;&amp;gt;geeks&amp;lt;/a&amp;gt;, &amp;lt;a href=&amp;quot;irc://irc.azzurra.net/pescaralug&amp;quot;&amp;gt;pescaralug&amp;lt;/a&amp;gt;&amp;lt;/span&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
        &lt;br /&gt;
        &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt; &lt;br /&gt;
        &lt;br /&gt;
    &lt;br /&gt;
	&amp;lt;br&amp;gt;&lt;br /&gt;
	&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;148&amp;quot;&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; rowspan=&amp;quot;8&amp;quot; width=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td bgcolor=&amp;quot;#3399FF&amp;quot; height=&amp;quot;2&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo_gruppo&amp;quot;&amp;gt;Credits&amp;lt;/p&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
&lt;br /&gt;
        &amp;lt;td height=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td height=&amp;quot;2&amp;quot; bgcolor=&amp;quot;#3399FF&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
      &amp;lt;tr&amp;gt;&lt;br /&gt;
        &amp;lt;td&amp;gt;&amp;lt;p class=&amp;quot;menu_right_titolo&amp;quot;&amp;gt;&lt;br /&gt;
	&amp;lt;a href=&amp;quot;http://www.masonhq.com/&amp;quot;&amp;gt;Mason&amp;lt;/a&amp;gt;&lt;br /&gt;
&amp;lt;/p&amp;gt;&lt;br /&gt;
&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
	All this website has built using Mason. Without this language derived from Perl,&lt;br /&gt;
	most of the tools used for the manage of this site would never been developed.&lt;br /&gt;
	&amp;lt;/p&amp;gt;&lt;br /&gt;
    &amp;lt;p class=&amp;quot;menu_right_titolo&amp;quot;&amp;gt;&lt;br /&gt;
	&amp;lt;a href=http://www.milonic.com/&amp;gt;JavaScript Menu&amp;lt;/a&amp;gt;&amp;lt;/p&amp;gt;&lt;br /&gt;
	&amp;lt;p class=&amp;quot;menu_right&amp;quot;&amp;gt;&lt;br /&gt;
	 The great menu at the top of this page.&amp;lt;/p&amp;gt;&lt;br /&gt;
    &amp;lt;/td&amp;gt;&lt;br /&gt;
      &amp;lt;/tr&amp;gt;&lt;br /&gt;
    &amp;lt;/table&amp;gt;&lt;br /&gt;
	&lt;br /&gt;
    	&amp;lt;!-- END:  left_body --&amp;gt;&lt;br /&gt;
    &amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td valign=&amp;quot;top&amp;quot; width=&amp;quot;8&amp;quot;&amp;gt;&amp;lt;img src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;8&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
  &amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&amp;lt;!-- INIT: Comments --&amp;gt;&lt;br /&gt;
		&lt;br /&gt;
		&lt;br /&gt;
&amp;lt;!-- END:  Comments --&amp;gt;&lt;br /&gt;
&amp;lt;p&amp;gt;&lt;br /&gt;
&amp;lt;!-- INIT: Google adSense --&amp;gt;&lt;br /&gt;
&amp;lt;table align=&amp;quot;center&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;tr&amp;gt;&amp;lt;td&amp;gt;&lt;br /&gt;
&amp;lt;td align=&amp;quot;bottom&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&amp;gt;&amp;lt;!--&lt;br /&gt;
google_ad_client = &amp;quot;pub-6275448308915555&amp;quot;;&lt;br /&gt;
google_ad_width = 728;&lt;br /&gt;
google_ad_height = 90;&lt;br /&gt;
google_ad_format = &amp;quot;728x90_as&amp;quot;;&lt;br /&gt;
google_ad_type = &amp;quot;text_image&amp;quot;;&lt;br /&gt;
google_ad_channel =&amp;quot;&amp;quot;;&lt;br /&gt;
google_color_border = &amp;quot;4575A3&amp;quot;;&lt;br /&gt;
google_color_bg = &amp;quot;D1E9FF&amp;quot;;&lt;br /&gt;
google_color_link = &amp;quot;0033CC&amp;quot;;&lt;br /&gt;
google_color_url = &amp;quot;CC3300&amp;quot;;&lt;br /&gt;
google_color_text = &amp;quot;000000&amp;quot;;&lt;br /&gt;
//--&amp;gt;&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;script type=&amp;quot;text/javascript&amp;quot;&lt;br /&gt;
  src=&amp;quot;http://pagead2.googlesyndication.com/pagead/show_ads.js&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;/script&amp;gt;&lt;br /&gt;
&amp;lt;/td&amp;gt;&lt;br /&gt;
&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&amp;lt;!-- END: Google adSense --&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;table border=&amp;quot;0&amp;quot; cellpadding=&amp;quot;0&amp;quot; cellspacing=&amp;quot;0&amp;quot; width=&amp;quot;100%&amp;quot;&amp;gt;&lt;br /&gt;
  &amp;lt;tr&amp;gt;&amp;lt;td colspan=&amp;quot;3&amp;quot; bgcolor=&amp;quot;#4575A3&amp;quot; class=&amp;quot;linea&amp;quot; HEIGHT=&amp;quot;1&amp;quot;&amp;gt;&amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/1x1.gif&amp;quot; width=&amp;quot;1&amp;quot; height=&amp;quot;1&amp;quot; alt=&amp;quot;&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&amp;lt;/tr&amp;gt;&lt;br /&gt;
  &amp;lt;tr&amp;gt;&lt;br /&gt;
    &amp;lt;td&amp;gt; Copyright© 1997-2006 Emiliano Bruni&amp;lt;/td&amp;gt;&lt;br /&gt;
	&amp;lt;td align=&amp;quot;center&amp;quot;&amp;gt;Online from 16/08/1998 with &amp;lt;img src=&amp;quot;/cgi-bin/c/Count.cgi?ft=0&amp;amp;df=ebruni.it.dat&amp;amp;comma=T&amp;amp;md=8&amp;amp;pad=T&amp;amp;dd=verdana&amp;quot; alt=&amp;quot;&amp;quot; align=&amp;quot;bottom&amp;quot;&amp;gt;&lt;br /&gt;
 visitors&amp;lt;/td&amp;gt;&lt;br /&gt;
    &amp;lt;td align=&amp;quot;right&amp;quot;&amp;gt;Write me to:&lt;br /&gt;
    &amp;lt;img border=&amp;quot;0&amp;quot; src=&amp;quot;/images/mail.gif&amp;quot; width=&amp;quot;77&amp;quot; height=&amp;quot;10&amp;quot; align=&amp;quot;baseline&amp;quot;&amp;gt;&amp;lt;/td&amp;gt;&lt;br /&gt;
  &amp;lt;/tr&amp;gt;&lt;br /&gt;
&amp;lt;/table&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt;&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;/body&amp;gt;&lt;br /&gt;
&amp;lt;/html&amp;gt;&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40156</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40156"/>
		<updated>2010-11-08T03:05:42Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40155</id>
		<title>CSC/ECE 506 Spring 2010/summary</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/summary&amp;diff=40155"/>
		<updated>2010-11-08T02:59:48Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;ECE633 Independent Study: Architecture of Parallel Computers&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
*Abstract: *&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
*Experience with Wiki written text book:*&lt;br /&gt;
&lt;br /&gt;
The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
*Chapter wise learning from this independent study:*&lt;br /&gt;
&lt;br /&gt;
*Chapter 1: *&lt;br /&gt;
&lt;br /&gt;
It covered an interesting topic of supercomputer evolution. Wiki pages written for this topic included a lot data from literature. Students came up with interesting topics which were not covered in the text book such as [Timeline of supercomputers|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers], [First Supercomputer(ENIAC)|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29], [Cray History|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History], [Supercomputer Hierarchal Architecture|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture], [Supercomputer Operating System|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System], [Cooling Supercomputer|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer] and [Processor Family|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family]. From their research we could see the increase in dominance of Intel’s processors in the consumer market. We also conclude that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely used multiprocessor machine architectures which was replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing|http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false]&lt;br /&gt;
&lt;br /&gt;
*Chapter 2: *&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The students provided comparisons between data parallelism and task parallelism. [Haveraaen (2000)|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. Students noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In their comparisons they concluded combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      W. Daniel Hillis and Guy L. Steele, Jr., [&amp;quot;Data parallel algorithms,&amp;quot;|http://portal.acm.org/citation.cfm?id=7903] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2.      Alexander C. Klaiber and Henry M. Levy, [&amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;|http://portal.acm.org/citation.cfm?id=192020] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
*Chapter3: *&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism techniques were discussed with examples in the form of Open MP code as discussed in the text book. Besides the students provided additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally they conclude : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures|https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw]&lt;br /&gt;
&lt;br /&gt;
2.      [Intel Threading Building Blocks 2.2 for Open Source Reference Manual|http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf]&lt;br /&gt;
&lt;br /&gt;
3.      [POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory|https://computing.llnl.gov/tutorials/pthreads/#Joining]&lt;br /&gt;
&lt;br /&gt;
*Chapter 6:*&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: Students added additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was an additional subtopic students threw light on. Students also gave definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic was how students discussed WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.       [http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2.       [http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3.       [http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
*Chapter 7: *&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement was to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency was discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. They concluded that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization was discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core were also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. The students also discussed commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.       [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2.       [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
*Chapter 8:*&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [Cache consistency with MESI on Intel processor|http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [AMD dual core Architecture|http://techreport.com/articles.x/8236/2]&lt;br /&gt;
&lt;br /&gt;
3.      [Silicon Graphics Computer Systems|http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913]&lt;br /&gt;
&lt;br /&gt;
4.      [Synapse tightly coupled multiprocessors: a new approach to solve old problems|http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533]&lt;br /&gt;
&lt;br /&gt;
5.      [Dragon Protocol|http://en.wikipedia.org/wiki/Dragon_protocol]&lt;br /&gt;
&lt;br /&gt;
*Chapter 9:*&lt;br /&gt;
&lt;br /&gt;
Synchronization: Students classified synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements.&lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. They also discussed reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier was included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. They showed that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3.      [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
*Chapter 10: *&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References:&lt;br /&gt;
&lt;br /&gt;
1.      [Shared Memory Consistency Models|http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [Designing Memory Consistency Models For Shared-Memory Multiprocessors|http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273]&lt;br /&gt;
&lt;br /&gt;
3.      [Consistency Models|http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html]&lt;br /&gt;
&lt;br /&gt;
* *&lt;br /&gt;
&lt;br /&gt;
*Chapter 11:*&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor.   Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
_1.      _Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [&amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;|http://doi.acm.org/10.1145/325164.325132] In _Proceedings of the 17th Annual International Symposium on Computer Architecture._&lt;br /&gt;
&lt;br /&gt;
2.      David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [&amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;|http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf] _ACM SIGPLAN Notices_.&lt;br /&gt;
&lt;br /&gt;
*Chapter 12:*&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. They provided in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. They also discussed routing algorithms and deadlock, starvation and livelock associated with it. These topics were covered in in an extremely detailed way. The students included a diagrammatic representation for every topology.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2.      [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;br /&gt;
&lt;br /&gt;
* *&lt;br /&gt;
&lt;br /&gt;
*Conclusion:*&lt;br /&gt;
&lt;br /&gt;
This independent study helped me to increase my knowledge to a great extent in the field of Architecture of parallel computers. There were 4 students working on every chapter and came up with 2 wiki pages per group. We collected a total of 18 wiki supplements. The data collected was enormous. While reviewing their content I kept updating my knowledge base. I also provided the resources from where they can collect data. This helped me to come across latest developments in the field. Interacting with students helped me to increase my communication skills. Constant discussions with Prof. Gehringer helped me to understand key concepts. This idea of writing wiki supplements got selected for KU Village presentation. I got an opportunity to present this paper along with Prof. Gehringer.&lt;br /&gt;
&lt;br /&gt;
ECE633 Independent Study: Architecture of Parallel Computers&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
*Abstract: *&lt;br /&gt;
&lt;br /&gt;
There has been tremendous research and development in the field of multi-core Architecture in the last decade. In such a dynamic environment it is very difficult to have text books covering latest developments in the field. Wiki written text books comes as an extremely handy tool for students to get acquainted and interested in ongoing research. In this independent study we explored an academic learning technique where students could learn the fundamental concepts of the subject through the text book available to students and lectures delivered by Prof. Gehringer in class. They can now build on this foundation and gather latest information from the varied online resources and technical papers and summarize their findings in the form of wiki pages. Software is also being currently developed to assist the students and was adopted in this course. We tried to enhance the quality of student submitted wiki pages through peer reviewing. Professor Gehringer and I constantly provided inputs to students to improve both their quality of wiki pages as well as quality of reviewing. The software being developed under the able guidance of professor Gehringer has been vital in overcoming administrative hurdles involved in assigning topics to students, maintaining the updates and tracking progress of their writings, getting feedbacks through peer reviewing and handling the re-submitted work. All this has been managed via the software in an organized fashion.&lt;br /&gt;
&lt;br /&gt;
*Experience with Wiki written text book:*&lt;br /&gt;
&lt;br /&gt;
The software was first deployed in CSC/ECE 506, Architecture of Parallel Computers. This is a beginning masters-level course that is taken by all Computer Engineering masters students. It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too. The recently adopted textbook for this course is the locally written Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems [Solihin 2009]. It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines. We felt that students would benefit from learning how the principles were applied in current architectures. Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter. (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.) They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter. We wanted them to concentrate instead on recent developments. Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students. A lot of review time was spent providing guidance on how to revise.&lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on. This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations. Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for students to material that we wanted the students to pay attention to. Gehringer and Navalakha met weekly to discuss what to provide to students. We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD. As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed. A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 82.8% while the average for the second submission was 82.7%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students. The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the textbook. The later wiki pages focused on a comparative study of present-day supercomputers produced by Intel, AMD and IBM.&lt;br /&gt;
&lt;br /&gt;
For example while writing the wiki for cache-coherence protocols, the students examined which protocol was favored by which company and why. They also discussed protocols which have been introduced in recent two years e.g., Intel's MESIF protocol. Such in depth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave the students insight into what was expected expected of them. This led to an increasing focus on current developments while peer reviewing. It was observed that later versions of reviews included guidance similar to that received from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
*Chapter wise learning from this independent study:*&lt;br /&gt;
&lt;br /&gt;
*Chapter 1: *&lt;br /&gt;
&lt;br /&gt;
It covered an interesting topic of supercomputer evolution. Wiki pages written for this topic included a lot data from literature. Students came up with interesting topics which were not covered in the text book such as [Timeline of supercomputers|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Timeline_of_supercomputers], [First Supercomputer(ENIAC)|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#First_Supercomputer_.28_ENIAC_.29], [Cray History|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cray_History], [Supercomputer Hierarchal Architecture|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Supercomputer_Hierarchal_Architecture], [Supercomputer Operating System|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#SuperComputer_Operating_System], [Cooling Supercomputer|http://pg-server.csc.ncsu.edu/mediawiki/index.php/1.1#Cooling_Supercomputer] and [Processor Family|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch1_lm#Processor_Family]. From their research we could see the increase in dominance of Intel’s processors in the consumer market. We also conclude that Unix has been the platform for most of these super computers. Massive Parallel Processing (MPP) and Symmetric Multiprocessing (SMP) were the earliest style of widely used multiprocessor machine architectures which was replaced by constellation computing in the 2000 and currently is dominated by cluster computing.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
[http://www.top500.org/]&lt;br /&gt;
&lt;br /&gt;
[The future of supercomputing: an interim report By National Research Council (U.S.). Committee on the Future of Supercomputing|http://books.google.com/books?id=wx4kNh8ArH8C&amp;amp;pg=PA3&amp;amp;lpg=PA3&amp;amp;dq=evolution+of+supercomputers&amp;amp;source=bl&amp;amp;ots=7DVWaEYsZ4&amp;amp;sig=WKRWRuqtM-UfPoB-Wdka5ZWTgng&amp;amp;hl=en&amp;amp;ei=xAleS-TmDpqutgfcj_2jAg&amp;amp;sa=X&amp;amp;oi=book_result&amp;amp;ct=result&amp;amp;resnum=1&amp;amp;ved=0CAoQ6AEwADgK#v=onepage&amp;amp;q=evolution%20of%20supercomputers&amp;amp;f=false]&lt;br /&gt;
&lt;br /&gt;
*Chapter 2: *&lt;br /&gt;
&lt;br /&gt;
Data Parallel Programming: The students provided comparisons between data parallelism and task parallelism. [Haveraaen (2000)|http://pg-server.csc.ncsu.edu/mediawiki/index.php/CSC/ECE_506_Spring_2010/ch_2_maf#References] notes that data parallel codes typically bear a strong resemblance to sequential codes, making them easier to read and write. Students noted that the data parallel model may be used with the shared memory or the message passing model without conflict. In their comparisons they concluded combining the data parallel and message passing models results in reduction in the amount and complexity of communication required relative to a task parallel approach. Similarly, combining the data parallel and shared memory models tends to simplify and reduce the amount of synchronization required. SIMD (single-instruction-multiple-data) processors are specifically designed to run data parallel algorithms. Modern examples include CUDA processors developed by nVidia and Cell processors developed by STI (Sony, Toshiba, and IBM).&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      W. Daniel Hillis and Guy L. Steele, Jr., [&amp;quot;Data parallel algorithms,&amp;quot;|http://portal.acm.org/citation.cfm?id=7903] Communications of the ACM, 29(12):1170-1183, December 1986.&lt;br /&gt;
&lt;br /&gt;
2.      Alexander C. Klaiber and Henry M. Levy, [&amp;quot;A comparison of message passing and shared memory architectures for data parallel programs,&amp;quot;|http://portal.acm.org/citation.cfm?id=192020] in Proceedings of the 21st Annual International Symposium on Computer Architecture, April 1994, pp. 94-105.&lt;br /&gt;
&lt;br /&gt;
*Chapter3: *&lt;br /&gt;
&lt;br /&gt;
In this wiki supplement, the three kinds of parallelisms, i.e. DOALL, DOACROSS and DOPIPE were discussed. These three parallelism techniques were discussed with examples in the form of Open MP code as discussed in the text book. Besides the students provided additional depth in this topic by discussing parallel_for, parallel_reduce, parallel_scan, pipeline, Reduction, DOALL, DOACROSS, DOPIPE with respect to Intel Thread Building Blocks. They also compared DOPIPE, DOACROSS, DOALL in POSIX Threads. Finally they conclude : Pthreads works for all the parallelism and could express functional parallelism easily, but it needs to build specialized synchronization primitives and explicitly privatize variables, makes it more effort needed to switch a serial program in to parallel mode.&lt;br /&gt;
&lt;br /&gt;
OpenMP can provide many performance enhancing features, such as atomic, barrier and flush synchronization primitives. It is very simple to use OpenMP to exploit DOALL parallelism, but the syntax for expressing functional parallelism is awkward.&lt;br /&gt;
&lt;br /&gt;
Intel TBB relies on generic programming, it performs better with custom iteration spaces or complex reduction operations. Also, it provides generic parallel patterns for parallel while-loops, data-flow pipeline models, parallel sorts and prefixes, so it's better in cases go beyond loop-based parallelism.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [An Optimal Abtraction Model for Hardware Multithreading in Modern Processor Architectures|https://docs.google.com/viewer?a=v&amp;amp;pid=gmail&amp;amp;attid=0.1&amp;amp;thid=126f8a391c11262c&amp;amp;mt=application%2Fpdf&amp;amp;url=https%3A%2F%2Fmail.google.com%2Fmail%2F%3Fui%3D2%26ik%3Dd38b56c94f%26view%3Datt%26th%3D126f8a391c11262c%26attid%3D0.1%26disp%3Dattd%26realattid%3Df_g602ojwk0%26zw&amp;amp;sig=AHIEtbTeQDhK98IswmnVSfrPBMfmPLH5Nw]&lt;br /&gt;
&lt;br /&gt;
2.      [Intel Threading Building Blocks 2.2 for Open Source Reference Manual|http://www.threadingbuildingblocks.org/uploads/81/91/Latest%20Open%20Source%20Documentation/Reference.pdf]&lt;br /&gt;
&lt;br /&gt;
3.      [POSIX Threads Programming by Blaise Barney, Lawrence Livermore National Laboratory|https://computing.llnl.gov/tutorials/pthreads/#Joining]&lt;br /&gt;
&lt;br /&gt;
*Chapter 6:*&lt;br /&gt;
&lt;br /&gt;
Cache Structures of Multi-Core Architectures: Students added additional insight on this topic by discussing Shared Memory Multiprocessors, write policies and replacement policies. Greedy Dual Size (GDS) and Priority Cache(PC) replacement policy was an additional subtopic students threw light on. Students also gave definitions about Trace Cache and Smart Cache techniques by Intel. The most important take away from this topic was how students discussed WRITE POLICIES used in recent multi core architectures. For example, Intel IA 32 IA64 architecture implements Write Combining, Write Collapsing, Weakly Ordered, Uncacheable &amp;amp; Write No Allocate and Non-temporal techniques in its cache. AMD uses cache exclusion unlike Intel’s cache inclusion. Sun's Niagara and SPARC use L1 caches as WT, with allocate on load and noallocate on stores.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.       [http://download.intel.com/technology/architecture/sma.pdf]&lt;br /&gt;
&lt;br /&gt;
2.       [http://www.intel.com/Assets/PDF/manual/248966.pdf]&lt;br /&gt;
&lt;br /&gt;
3.       [http://www.intel.com/design/intarch/papers/cache6.pdf]&lt;br /&gt;
&lt;br /&gt;
*Chapter 7: *&lt;br /&gt;
&lt;br /&gt;
Shared-memory multiprocessors run into several problems that are more pronounced than their uniprocessor counterparts. The Solihin text used in this course goes into detail on three of these issues, that is cache coherence, memory consistency and synchronization. The goal of this wiki supplement was to discuss these three issues and also what can be done to ensure that instructions are handled in both a timely and efficient manner and in a manner that is consistent with what the programmer might desire. Memory consistency was discussed by comparing ordering on a uniprocessor vs ordering on a multiprocessor. They concluded that in a multiprocessor much more care must be taken to ensure that all of the loads and stores are committed to memory in a valid order. Synchronization was discussed as applicable to Open MP and fence insertion. Other methods such as test and set method and direct interrupt to another core were also briefly discussed. The programmer (or complier) is responsible for knowing which synchronization directives are available on a given architecture and implementing them in an efficient manner. The students also discussed commonly used instructions for synchronization in popular processor architectures. For example SPARC V8 uses store barrier, Alpha uses memory barrier and write memory barrier whereas Intel x86 uses lfence (load) sfence (store).&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.       [https://wiki.ittc.ku.edu/ittc/images/0/0f/Loghi.pdf]&lt;br /&gt;
&lt;br /&gt;
2.       [http://portal.acm.org/citation.cfm?id=782854&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84866326&amp;amp;CFTOKEN=84791790]&lt;br /&gt;
&lt;br /&gt;
*Chapter 8:*&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [Cache consistency with MESI on Intel processor|http://www.zak.ict.pwr.wroc.pl/nikodem/ak_materialy/Cache%20consistency%20&amp;amp;%20MESI.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [AMD dual core Architecture|http://techreport.com/articles.x/8236/2]&lt;br /&gt;
&lt;br /&gt;
3.      [Silicon Graphics Computer Systems|http://ieeexplore.ieee.org.www.lib.ncsu.edu:2048/stamp/stamp.jsp?tp=&amp;amp;arnumber=4913]&lt;br /&gt;
&lt;br /&gt;
4.      [Synapse tightly coupled multiprocessors: a new approach to solve old problems|http://portal.acm.org/citation.cfm?id=1499317&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=83027384&amp;amp;CFTOKEN=95680533]&lt;br /&gt;
&lt;br /&gt;
5.      [Dragon Protocol|http://en.wikipedia.org/wiki/Dragon_protocol]&lt;br /&gt;
&lt;br /&gt;
*Chapter 9:*&lt;br /&gt;
&lt;br /&gt;
Synchronization: Students classified synchronization techniques based on implementation. Hardware synchronization uses locks, barriers and mutual exclusion. Software synchronization examples include ticket locks and queue-based MCS locks. Mutex implementation uses execution of atomic statements.&lt;br /&gt;
&lt;br /&gt;
Some common examples include Test-and-Set, Fetch-and-Increment, Exchange, Compare-and-Swap. Another type of lock that was not discussed in the text is known as the &amp;quot;Hand-off&amp;quot; lock was discussed in detail by the students. They also discussed reasons why a programmer should attempt to write programs in such a way as to avoid locks. There are API's that exist for parallel architectures that provide specific types of synchronization. If the API are used they way they were design, performance can be maximized while minimizing overhead.Load Locked(LL) and Store Conditional(SC) are a pair of instructions are improved hardware primitives that are used for lock-free read-modify-write operation.&lt;br /&gt;
&lt;br /&gt;
Detailed description of Combining Tree Barrier, Tournament Barrier and Disseminating Barrier was included. One of the interesting topics discussed in this wiki supplement was the performance evaluation of different barrier implementations. They showed that barrier/centralized blocking barrier does not scale with number of threads and the contention increases with increase in number of threads.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [http://www2.cs.uh.edu/~hpctools/pub/iwomp-barrier.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [http://www.statemaster.com/encyclopedia/Deadlock]&lt;br /&gt;
&lt;br /&gt;
3.      [http://www.ukhec.ac.uk/publications/reports/synch_java.pdf]&lt;br /&gt;
&lt;br /&gt;
*Chapter 10: *&lt;br /&gt;
&lt;br /&gt;
Students discussed the existing bus-based cache coherence in real machines. They went ahead and classified the cache coherence protocols based on the year they were introduced and they processors which uses them. MSI protocol was first used in SGI IRIS 4D series. In Synapse protocol M state is called D (Dirty) but works the same as MSI protocol works. MSI has a major drawback in that each read-write sequence incurs 2 bus transactions irrespective of whether the cache line is stored in only one cache or not. The Pentium Pro microprocessor, introduced in 1992 was the first Intel architecture microprocessor to support SMP and MESI. The MESIF protocol, used in the latest Intel multi-core processors was introduced to accommodate the point-to-point links used in the QuickPath Interconnect. MESI came with the drawback of using much time and bandwidth. MOESI was the AMD’s answer to this problem . MOESI' has become one of the most popular snoop-based protocols supported in the AMD64 architecture. The AMD dual-core Opteron can maintain cache coherence in systems up to 8 processors using this protocol. The Dragon Protocol is an update based coherence protocol which does not invalidate other cached copies. The Dragon Protocol , was developed by Xerox Palo Alto Research Center(Xerox PARC), a subsidiary of Xerox Corporation. This protocol was used in the Xerox PARC Dragon multiprocessor workstation. References:&lt;br /&gt;
&lt;br /&gt;
1.      [Shared Memory Consistency Models|http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-95-7.pdf]&lt;br /&gt;
&lt;br /&gt;
2.      [Designing Memory Consistency Models For Shared-Memory Multiprocessors|http://portal.acm.org/citation.cfm?id=193889&amp;amp;dl=GUIDE&amp;amp;coll=GUIDE&amp;amp;CFID=84028355&amp;amp;CFTOKEN=32262273]&lt;br /&gt;
&lt;br /&gt;
3.      [Consistency Models|http://cs.gmu.edu/cne/modules/dsm/green/memcohe.html]&lt;br /&gt;
&lt;br /&gt;
* *&lt;br /&gt;
&lt;br /&gt;
*Chapter 11:*&lt;br /&gt;
&lt;br /&gt;
The cache coherence protocol presented in Chapter 11 of Solihin 2008 is simpler than most real directory-based protocols. This textbook supplement presents the directory-based protocols used by the DASH multiprocessor and the Alewife multiprocessor. It concludes with an argument of why complexity might be undesirable in cache coherence protocols. The DASH multiprocessor uses a two-level coherence protocol, relying on a snoopy bus to ensure cache coherence within cluster and a directory-based protocol to ensure coherence across clusters. The protocol uses a Remote Access Cache (RAC) at each cluster, which essentially consolidates memory blocks from remote clusters into a single cache on the local snoopy bus. When a request is issued for a block from a remote cluster that is not in the RAC, the request is denied but the request is also forwarded to the owner. The owner supplies the block to the RAC. Eventually, when the requestor retries, the block will be waiting in the RAC. Read and readx operations on a Dash processor were discussed in detail. They also discuss two race conditions which mainly arises on a Dash processor.The first occurs when a Read from requester R is forwarded from home H to owner O, but O sends a Writeback to H before the forwarded Read arrives. Another possible race occurs when the home node H replies with data (ReplyD) to a Read from requester R but an invalidation (Inv) arrives first. LimitLESS is the cache coherence protocol used by the Alewife multiprocessor.   Unlike the DASH multiprocessor, the Alewife multiprocessor is not organized into clusters of nodes with local buses, and therefore cache coherence through the system is maintain through the directory.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
_1.      _Daniel Lenoski, James Laudon, Kourosh Gharachorloo, Anoop Gupta, and John Hennessy (1990). [&amp;quot;The directory-based cache coherence protocol for the DASH multiprocessor.&amp;quot;|http://doi.acm.org/10.1145/325164.325132] In _Proceedings of the 17th Annual International Symposium on Computer Architecture._&lt;br /&gt;
&lt;br /&gt;
2.      David Chaiken, John Kubiatowicz, and Anant Agarwal (1991). [&amp;quot;LimitLESS directories: A scalable cache coherence scheme.&amp;quot;|http://groups.csail.mit.edu/cag/papers/pdf/asplos4.pdf] _ACM SIGPLAN Notices_.&lt;br /&gt;
&lt;br /&gt;
*Chapter 12:*&lt;br /&gt;
&lt;br /&gt;
Interconnection Networks: Advances in multiprocessors, parallel computing &amp;amp; networking and parallel computer architectures demand very high performance from interconnection networks. Due to this, interconnection network structure has changed over time, trying to meet higher bandwidths and performance. Students discussed criterion to be considered for choosing the best Network. It included Performance Requirements, Scalability, Incremental expandability, Partitionability, Simplicity, Distance Span, Physical Constraints, Reliability and Reparability, Expected Workloads and Cost Constraints. They provided in depth discussion on Classification of Interconnection networks. Shared-Medium Networks include Token Ring, Token Bus, Backplane Bus. Direct Networks include Mesh, Torus, Hypercube, Tree, Cube-Connected Cycles and de Bruijn and Star Graph Networks. Indirect Networks include Regular Topologies like Crossbar Network and Multistage Interconnection Network and Hybrid Networks such as Multiple Backplane Buses, Hierarchical Networks, Cluster-Based Networks and Hypergraph Topologies. They also discussed routing algorithms and deadlock, starvation and livelock associated with it. These topics were covered in in an extremely detailed way. The students included a diagrammatic representation for every topology.&lt;br /&gt;
&lt;br /&gt;
References:&lt;br /&gt;
&lt;br /&gt;
1.      [http://www.top500.org/2007_overview_recent_supercomputers/sci]&lt;br /&gt;
&lt;br /&gt;
2.      [http://www.cs.nmsu.edu/~pfeiffer/classes/573/notes/topology.html]&lt;br /&gt;
&lt;br /&gt;
* *&lt;br /&gt;
&lt;br /&gt;
*Conclusion:*&lt;br /&gt;
&lt;br /&gt;
This independent study helped me to increase my knowledge to a great extent in the field of Architecture of parallel computers. There were 4 students working on every chapter and came up with 2 wiki pages per group. We collected a total of 18 wiki supplements. The data collected was enormous. While reviewing their content I kept updating my knowledge base. I also provided the resources from where they can collect data. This helped me to come across latest developments in the field. Interacting with students helped me to increase my communication skills. Constant discussions with Prof. Gehringer helped me to understand key concepts. This idea of writing wiki supplements got selected for KU Village presentation. I got an opportunity to present this paper along with Prof. Gehringer.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32462</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32462"/>
		<updated>2010-07-02T16:10:27Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second submission was only slightly greater. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave an insight to students about what is expected and thus they focussed on current developments while peer reviewing. It was observed that later versions of reviews included similar guidance received by them from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students constantly improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
Thus for large projects, this is an efficient way where the students have to sign up for the topics they want to work on. This can be&lt;br /&gt;
achieved using the electronic signup system. Students can suggest topics using the “New Suggestion” form and the instructor can approve topics which are reasonable. If the topics need to be done in a particular order the instructor can set the precedence and the system would set the deadlines appropriately for the topics. The system would also see to it that the students are notified of their deadlines in accordance with the topic they selected.&lt;br /&gt;
&lt;br /&gt;
Electronic peer-review systems have been widely used to review student work, but never before, to our knowledge, have they been applied to assignments consisting of multiple interrelated parts with precedence constraints. The growing interest in large collaborative projects, such as wiki textbooks, has led to a need for electronic support for the process, lest the administrative burden on instructor and TA grow too large.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32461</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32461"/>
		<updated>2010-07-02T16:02:25Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second submission was only slightly greater. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. These reviews gave an insight to students about what is expected and thus they focussed on current developments while peer reviewing. It was observed that later versions of reviews included similar guidance received by them from Gehringer and Navalakha. The organization of the wiki pages and the volume of relevant data collected by students constantly improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32460</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32460"/>
		<updated>2010-07-02T15:57:43Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second submission was only slightly greater. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. The organization of the wiki pages and the volume of relevant data collected by students constantly improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32459</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32459"/>
		<updated>2010-07-02T15:57:15Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second submission was only slighly greater. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. The organization of the wiki pages and the volume of relevant data collected by students constantly improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32458</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32458"/>
		<updated>2010-07-02T15:55:07Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second was 82.73%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages. The organization of the wiki pages and the volume of relevant data collected by students constantly improved as the semester progressed.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32457</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32457"/>
		<updated>2010-07-02T15:53:04Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Reys et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
Our software was first deployed in CSC/ECE 506, Architecture of Parallel Computers.  This is a beginning masters-level course that is taken by all Computer Engineering masters students.  It is optional for Computer Science students, but as it is one way to fulfill a core requirement, it is popular with them too.  The recently adopted textbook for this course is the locally written ''Fundamentals of Parallel Computer Architecture: Multichip and Multicore Systems'' [Solihin 2009].  It did not make sense to have the students rewrite this excellent text, but the book concentrates on theory and design fundamentals, without detailed application to current parallel machines.  We felt that students would benefit from learning how the principles were applied in current architectures.  Furthermore, they would learn about the newest machines in this fast-changing field.&lt;br /&gt;
&lt;br /&gt;
After every chapter covered in class, two individuals, or pairs of students were required to sign up for writing the wiki supplement for that particular chapter.  (That is, we solicited two supplements for each chapter, each of which could be authored by one or two students.)  They were asked to add specific types of information which was not included in the chapter.&lt;br /&gt;
&lt;br /&gt;
Initially, students were not clear about the purpose of their wiki pages. The first pages they wrote had substantial duplication of topics covered in the textbook. Students were attempting to give a complete coverage of issues discussed in the chapter.  We wanted them to concentrate instead on recent developments.  Upon seeing this, we established the practice of having the first two authors of this paper (Gehringer and Navalakha) review the student work, along with three peer reviews from fellow students.  A lot of review time was spent providing guidance on how to revise. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources for the topic they had chosen to write on.  This was not very successful, as the students seemingly chose to read the first few search hits, which tended to provide an overview of the topic, rather than in-depth information on particular implementations.  Sometimes students were not aware that the information they found was already covered in the next chapter, which they have not read yet. The first review which we gave students was mainly just making them aware of topics covered in later chapters. A lot of effort in writing the initial draft was thus wasted. After the first two sets of topics, we began to provide links for  students to material that we wanted the students to pay attention to.  Gehringer and Navalakha met weekly to discuss what to provide to students.  We regularly consulted other textbooks, technology news, and Web sites of major processor manufacturers, such as Intel and AMD.  As the semester progressed, the quality of the initial submissions improved, and the students realized better returns for their effort.&lt;br /&gt;
&lt;br /&gt;
The quality of work seemed to improve as the semester progressed.  A comparison of the grades for the wiki pages revealed that the average score for the first chapter written by each student was 83.37% while the average for the second was 82.73%. The quality of wiki pages had improved, but at the same time, the peer reviewers became more demanding.  Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.The initial wiki pages provided randomly collected data and was cluttered by diagrams and graphs. This information reinstated facts given in the text book. The later wiki pages focussed on a comparative study of present day supercomputers produced by Intel, AMD and IBM. For example while writing the wiki for Cache coherence protocols the students compared which protocol was favoured by which company and why. They also discussed protocols which have been introduced in recent two years eg. MESIF protocol by Intel. Such in deapth analysis made the wiki more appealing to readers. Gehringer and Navalakha provided additional reviews which helped in constantly improving the quality of wiki pages.&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
[Barnett and Blumner 2008]	Barnett, R.W. and Blumner, J. S.&lt;br /&gt;
Writing Centers and Writing Across the Curriculum Programs: Building Interdisciplinary Partnerships, IAP, 2008 &lt;br /&gt;
&lt;br /&gt;
[Bednar et al. 1991]	Bednar, A. E., Cunningham, D.D., Thomas, M., &amp;amp; Perry, D. [1991]. Theory into practice: How do we link? In G. Anglin [Ed.] Instructional technology: Past, present, and future. Denver, CO: Libraries Unlimited&lt;br /&gt;
&lt;br /&gt;
[Carpenter et al. 2006]	Carpenter, P., Bullock, A., &amp;amp; Potter, J. [2006] Textbooks in teaching and learning: The views of students and their teachers. Brookes eJournal of Learning and Teaching 2[1], 1-9. &lt;br /&gt;
&lt;br /&gt;
[Cunningham et al. 2000]	Cunningham, D.J., Duffy, T. M. &amp;amp; Knuft, R.A. [2000]. The textbook of the future [CRLT Technical Report No. 14-00]. Center for Research on Learning and Technology: Bloomington, IN. &lt;br /&gt;
&lt;br /&gt;
[Gehringer et al. 2007]	Gehringer, E.F., Ehresman, L.M., Conger, S.G.,  and Wagle, P.A. &amp;quot;Reusable learning objects through peer review: The Expertiza approach,&amp;quot; Innovate-Journal of Online Education 3:6 (August/September 2007).  &lt;br /&gt;
&lt;br /&gt;
[Gehringer 2009]	Gehringer, E.F. &amp;quot;Expertiza: information management for collaborative learning,&amp;quot; in Monitoring and Assessment in Online Collaborative Environments: Emergent Computational Technologies for E-Learning Support, A. A. Juan Perez [ed.], IGI Global Press, 2009.&lt;br /&gt;
&lt;br /&gt;
[Gehringer, Kadanjoth and Kidd 2010]	Gehringer, E.F., Kadanjoth, R., and Kidd, J. &amp;quot;Software Support for Peer-Reviewing Wiki Textbooks and Other Large Projects,&amp;quot; Proceedings of the Workshop on Computer-Supported Peer Review in Education, June 14, 2010.&lt;br /&gt;
&lt;br /&gt;
[NRC 2005]	National Research Council [NRC] of the National Academies [2005]. How students learn: History, Mathematics and Science in the classroom. Washington, D.C.: The National Academies Press. Retrieved May 16, 2009 from http://www.nap.edu/books/0309074339/html/&lt;br /&gt;
&lt;br /&gt;
[Rainie 2007]	Rainie, L. [2007]. Wikipedia: When in doubt, multitudes seek it out. [Pew Internet &amp;amp; American Life Project]. Pew Research Center: Washington, D.C. Accessed online on October 12, 2007 from http://pewresearch.org/pubs/460/wikipedia&lt;br /&gt;
&lt;br /&gt;
[Reys et al. 2004]	Reys B. J., Reys R. E. &amp;amp; Chavez, O. [2004]. Why mathematics textbooks matter. Educational Leadership 61[5], 61-66.&lt;br /&gt;
&lt;br /&gt;
[Solihin 2009]	Solihin, Y.  Fundamentals of Paralell Computer Architecture: Multichip and Multicore Systems, Solihin Books, 2008, 2009.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32448</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32448"/>
		<updated>2010-06-29T15:10:51Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37% while the average for the second was 82.73%. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32447</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32447"/>
		<updated>2010-06-29T15:10:40Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37% while the average for the second was 82.73%. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32446</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32446"/>
		<updated>2010-06-29T15:08:08Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37% while the average for the second was 82.73%. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32445</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32445"/>
		<updated>2010-06-29T15:07:16Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37 while the average for the second was 82.73. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32444</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32444"/>
		<updated>2010-06-29T15:06:45Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37 while the average for the second was 82.73. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32443</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32443"/>
		<updated>2010-06-29T15:01:18Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.A comparison of the grades for the wiki pages revealed that the average score for the first submission was 83.37 while the average for the second was 82.73. The quality of wiki pages had improved, but at the same time peer reviewing quality also improved. Students were given more inputs to improve their work via peer reviewing. Thus the improvement was seen in the final wiki page produced as against the grades received by students.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32442</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32442"/>
		<updated>2010-06-29T14:42:11Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as for students who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32441</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32441"/>
		<updated>2010-06-29T14:41:25Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in a group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as student who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32440</id>
		<title>CSC/ECE 506 Spring 2010/KU Village</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Spring_2010/KU_Village&amp;diff=32440"/>
		<updated>2010-06-29T14:39:55Z</updated>

		<summary type="html">&lt;p&gt;Knnavala: /* Experience with wiki-textbook writing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=== Experience with a student-written wiki textbook supplement ===&lt;br /&gt;
Edward F. Gehringer&lt;br /&gt;
&lt;br /&gt;
Karishma Navalakha&lt;br /&gt;
&lt;br /&gt;
Reejesh Kadanjoth&lt;br /&gt;
&lt;br /&gt;
North Carolina State University&lt;br /&gt;
&lt;br /&gt;
{efg, knnavala, rkadanj}@ncsu.edu&lt;br /&gt;
&lt;br /&gt;
== Abstract ==&lt;br /&gt;
&lt;br /&gt;
As wiki usage becomes common in educational settings, instructors are beginning to experiment with student-authored wiki textbooks.  Instead of reading textbooks selected by the instructor, students are challenged to read the primary literature and organize it for consumption by the other members of the class.  This has important pedagogical advantages, as students are stimulated to take responsibility for their own learning and perform tasks similar to those in the real world.  These benefits, however, come with an array of administrative challenges, including sequencing the material to be covered, and assigning other students to peer-review the submitted work.  We are developing software to assist in this effort.  This presentation discusses our experience with the process and the software in an advanced course on parallel computer architecture, where students were assigned to write supplements for each textbook chapter, describing how the theory covered in class was realized in state-of-the-art multicore processors.&lt;br /&gt;
&lt;br /&gt;
== Introduction ==&lt;br /&gt;
&lt;br /&gt;
In the last half-dozen years, the wiki has emerged as one of the leading collaborative tools on the Web.  It has the advantage that editing is done in place, without the need to pass copies around by e-mail.  This eases collaboration, by making it obvious which version is the most current.  Moreover, changes become visible instantly to anyone who accesses a page, which means that no intervention by the instructor is needed to disseminate new versions to the rest of the class.  These characteristics make it possible for students to work together to write text that is intended to be read by their fellow students.&lt;br /&gt;
&lt;br /&gt;
Forward-looking instructors were quick to apply wiki-based collaboration to a task that would heretofore have been intractable: having students write their own peer-reviewed textbook for the class.  The advantages are many: Rather than simply consume what is fed to them by the instructor and textbook author(s), students now have to take responsibility for their own learning [NRC 2005], determining what is worthy of being taught to the class.  In so doing, the students are &amp;quot;constructing&amp;quot; their own learning.  This meshes well with constructivism [Bednar et al. 1991]--the theory that in order to assimilate knowledge thoroughly, students must &amp;quot;build&amp;quot; it in their own minds rather than simply receive it from an external source.  Assigned textbooks deprive students of the motivation to decide what topics are relevant and remove the need to evaluate different points of view.  For this reason, they have been called &amp;quot;inconsistent with constructivist principles&amp;quot; [Cunningham et al. 2000].&lt;br /&gt;
&lt;br /&gt;
Researching a wiki textbook forces students to read the primary literature--a skill that is very necessary in the outside world, and one that is rarely given thorough attention in undergraduate courses.  If left to their own devices, students favor secondary research sources like Wikipedia [Rainie 2007].  Ironically, this testifies both to the attractiveness of the wiki for constructing reference works, and the need to encourage students to do their own research.&lt;br /&gt;
&lt;br /&gt;
Writing a textbook article is beneficial because it is expository writing.  A broad body of knowledge supports the idea of &amp;quot;writing across the curriculum&amp;quot; [Barnett and Blumner 2008], which says that writing experience should be integrated into every academic discipline, rather than confined to writing courses.  By contributing to the textbook, students gain experience writing up their thoughts for an audience of their peers.  Feedback from their classmates helps them learn from their mistakes and improve their writing skills.&lt;br /&gt;
&lt;br /&gt;
Finally, in a world where textbook prices are a significant component of the cost of education, student-authored textbooks have the ability to save students money.  Surprisingly, there is little evidence that students benefit from what they pay for their textbooks.  A four-year old English study found &amp;quot;no correlation between textbook purchase and the grade received&amp;quot; [Carpenter et al. 2006].  While there is a developing body of research on student-authored wiki textbooks, little research has been done on the efficacy of most commercial textbooks, either before or after publication [Revs et al. 2004].&lt;br /&gt;
&lt;br /&gt;
== The administrative burden ==&lt;br /&gt;
&lt;br /&gt;
Wikis take care of version control and dissemination of student writing, but many administrative issues remain.  Writing a textbook is a series of different projects, which usually need to be spaced out throughout the semester.  One must arrange for at least one student to choose each of the chapters or topics that need to be included.  In a face-to-face class, this can be arranged by passing around a signup sheet, but in a distance-education class, software support is needed.  Even in a face-to-face class, software support is useful, because choices made by students are registered automatically in the system, and students have an equal ability to sign up while there are still many topics available.&lt;br /&gt;
&lt;br /&gt;
Appropriate deadlines must be assigned for each topic or chapter.  Peer review requires separate deadlines for submission and review ... and, if authors are to revise their work in response to peer comments, there must be resubmission and final review deadlines as well.  There is usually a precedence relationship between topics: Some topics must be learned before other topics can be presented.  This means that the same four deadlines (submission, initial review, etc.) are applied to different work at different times during the semester.  A topic may not be written on before all prerequisite topics have been completed.  Getting all of these deadlines set is time consuming, and sending reminders to the students involved makes it more complex.  Software support is clearly desirable.  In the Expertiza system [Gehringer et al. 2007, Gehringer 2009], we have implemented support for signup sheets and staggered deadlines [Gehringer, Kadanjoth and Kidd 2010].&lt;br /&gt;
&lt;br /&gt;
== Experience with wiki-textbook writing ==&lt;br /&gt;
&lt;br /&gt;
While designing the structure for the course ECE 506 Architecture of Parallel Computer, a graduate level program, we discussed many ideas to make the course relevant to the industry which has shown important ramifications in recent years in the field of Parallel Computing. We wanted to make sure that the students learn about the fundamental basics which are described well in text books and at the same time update themselves with the recent advancements in the field. By making the students write wiki pages, which not only served as a learning tool but also created supplements to previous text books which needed constant revisions due to the number of supercomputers designed every year, we achieved multi-fold goal. &lt;br /&gt;
&lt;br /&gt;
Initially the idea of wiki pages was not clear to students. The initial wiki pages had a lot of duplication of the text book. Students wanted to give a complete coverage of the chapter. We wanted them to concentrate on recent developments. A lot of reviewing time was required for the initial releases. &lt;br /&gt;
&lt;br /&gt;
At the beginning we gave the students complete freedom to explore resources and come up with wiki pages. This initial approach did not prove very useful. After every chapter covered in class, two groups of students were required to sign up for writing the wiki page for that particular chapter. They were asked to add additional information which is not included in the chapter. The chapters covered in the class followed a certain sequence and had a logical flow with respect to the previous chapter. We realized that students were not aware that the information they find is already present in the next chapter which they have not studied yet. The first review which we gave students was mainly about making them aware of topics covered in later chapters. A lot of effort put in by students was unnecessarily wasted. After the first two rounds we changed gears and provided links to students which included information which we wanted the students to pay more attention to. We had weekly meetings regarding the topics which we would like to incorporate for every chapter. To come up with a list of topics we referred the previous text book, technology news as well as websites of every major processor giants. We realized the quality of wiki pages improved and energy of students was more channelized.&lt;br /&gt;
&lt;br /&gt;
The first half of the wiki pages were a group assignment. Two students came up with a wiki page. As the topic was new it made sense for students to work in group. We also assigned 4 reviewers for every wiki page. This was beneficial for the students who were writing the wiki as well as student who were reviewing it. We assigned individual wiki pages to students during the later half of the semester. Students had already learnt from their mistakes as well as others mistakes while peer reviewing. We noticed an improving quality of work as the semester progressed.&lt;/div&gt;</summary>
		<author><name>Knnavala</name></author>
	</entry>
</feed>