<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Crbarile</id>
	<title>Expertiza_Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Crbarile"/>
	<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Special:Contributions/Crbarile"/>
	<updated>2026-09-11T16:04:03Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62615</id>
		<title>CSC 456 Spring 2012/10b CP</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62615"/>
		<updated>2012-04-23T17:37:09Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Performance */ added markup for sections&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Romanescu, Lebeck, and Sorin make a great point that &amp;quot;The most important feature of a computer is correct execution.&amp;quot; Computers are expected to produce correct output consistently. Memory consistency -- the intentional ordering of all reads and writes to memory addresses (Solihin) -- plays a crucial role in guaranteeing that the results of running a program are the results intended by the programmer. Maintaining memory consistency is a problem on all multiprocessor machines (Solihin).&lt;br /&gt;
&lt;br /&gt;
== Models in Use ==&lt;br /&gt;
===Strict Consistency===&lt;br /&gt;
===Sequential Consistency===&lt;br /&gt;
===Weak Consistency===&lt;br /&gt;
&lt;br /&gt;
== Performance ==&lt;br /&gt;
===Weak Consistency===&lt;br /&gt;
===Sequential Consistency===&lt;br /&gt;
===Strict Consistency===&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62614</id>
		<title>CSC 456 Spring 2012/10b CP</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62614"/>
		<updated>2012-04-23T17:36:38Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Models in Use */ Added markup for sections&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Romanescu, Lebeck, and Sorin make a great point that &amp;quot;The most important feature of a computer is correct execution.&amp;quot; Computers are expected to produce correct output consistently. Memory consistency -- the intentional ordering of all reads and writes to memory addresses (Solihin) -- plays a crucial role in guaranteeing that the results of running a program are the results intended by the programmer. Maintaining memory consistency is a problem on all multiprocessor machines (Solihin).&lt;br /&gt;
&lt;br /&gt;
== Models in Use ==&lt;br /&gt;
===Strict Consistency===&lt;br /&gt;
===Sequential Consistency===&lt;br /&gt;
===Weak Consistency===&lt;br /&gt;
&lt;br /&gt;
== Performance ==&lt;br /&gt;
From best to worst:&lt;br /&gt;
Weak Consistency&lt;br /&gt;
Sequential Consistency&lt;br /&gt;
Strict Consistency&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62453</id>
		<title>CSC 456 Spring 2012/10b CP</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2012/10b_CP&amp;diff=62453"/>
		<updated>2012-04-17T05:22:12Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: pass 1 at introduction&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Romanescu, Lebeck, and Sorin make a great point that &amp;quot;The most important feature of a computer is correct execution.&amp;quot; Computers are expected to produce correct output consistently. Memory consistency -- the intentional ordering of all reads and writes to memory addresses (Solihin) -- plays a crucial role in guaranteeing that the results of running a program are the results intended by the programmer. Maintaining memory consistency is a problem on all multiprocessor machines (Solihin).&lt;br /&gt;
&lt;br /&gt;
== Models in Use ==&lt;br /&gt;
&lt;br /&gt;
== Performance ==&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=62452</id>
		<title>CSC 456 Spring 2011</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=62452"/>
		<updated>2012-04-17T05:16:22Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Adding our page link&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Chapter 1: Nick Nicholls, Albert Chu]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 4a: Brandon Chisholm, Chris Barile]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 6: Joshua Mohundro, Patrick Wong]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 6: Allison Hamann, Chris Barile]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch1 BC]] &amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch7 MN]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch7 AA]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch4b]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/10a AJ]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/10b CP]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/11b AB]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/11a NC]]&amp;lt;br/&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60581</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60581"/>
		<updated>2012-03-26T17:21:37Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Polyhedral Transformation */ added references&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
This technique involves running the program while recording each memory access in detail. The memory accesses are then looked at to determine what code can be parallelized. This technique is simple, but does not cover all the possible input combinations. These may lead to false positives in which code segments can be labelled as parallelizeable when they actually aren't.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
Polyhedral transformation involves representing loop iterations from the source code as lattice points in something called a polytope. This polytope is then transformed into a more optimized form. This approach is limited when pointers are found inside loops. The following pictures represent a polytope before and after the transformation.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
[[File:polytope_model.png|thumb|center|200px|Polytope model, unskewed&amp;lt;ref name=&amp;quot;polytope&amp;quot;/&amp;gt;]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|center|200px|Polytope model, skewed&amp;lt;ref name=&amp;quot;polytope&amp;quot;/&amp;gt;]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
Also known as &amp;quot;concolic execution&amp;quot;, &amp;quot;mixed concrete and symbolic execution&amp;quot;, or &amp;quot;symbolic execution&amp;quot;, automatic program exploration is originally a bug finding technique that provides full coverage of the program. The first step in this approach is to symbolically represent the inputs to the program. Then, each statement that has a symbolic variable is executed by symbolic manipulation. If a branch condition is found with a symbolic variable, then both paths are conceptually taken. The last step is using the path constraints to find the concrete values that allow the program to go down both paths. This is done until there are no more possible inputs. &amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
Scalar and Array Analysis are actually two different forms of analysis that are often used together. Scalar Analysis checks for dependencies between scalar variables in code. A dependency can be defined as &amp;quot;when a memory location written on one iteration of a loop is accessed (read or write) on a different iteration.&amp;quot;&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Array Analysis is then used on code that could not be parallelized using Scalar Analysis. Array Analysis looks for arrays that can be privatized.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt; Privatization means to allow each thread its own private copy of the array in question, in whole or in part. This is only possible if there are no dependencies in the array.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
One problem in parallelizing code is that synchronization may be required in some programs to prevent errors in computation. Commutativity Analysis is used to determine which operations can be performed out of order without affecting the results of the code. It is based on the mathematical concept of commutativity, where if the result of running two operations is the same, regardless of which order they are performed in.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
LLVM works by taking C/C++ code and compiling it into LLVM IR bytecode. Optimizations are done at the IR level and then a code generator brings it back into native code to be executed.&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;polytope&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Polytope_model&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60580</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60580"/>
		<updated>2012-03-26T17:20:20Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* References */ added polytope ref&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
This technique involves running the program while recording each memory access in detail. The memory accesses are then looked at to determine what code can be parallelized. This technique is simple, but does not cover all the possible input combinations. These may lead to false positives in which code segments can be labelled as parallelizeable when they actually aren't.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
Polyhedral transformation involves representing loop iterations from the source code as lattice points in something called a polytope. This polytope is then transformed into a more optimized form. This approach is limited when pointers are found inside loops. The following pictures represent a polytope before and after the transformation.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
[[File:polytope_model.png|thumb|center|200px|Polytope model, unskewed]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|center|200px|Polytope model, skewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
Also known as &amp;quot;concolic execution&amp;quot;, &amp;quot;mixed concrete and symbolic execution&amp;quot;, or &amp;quot;symbolic execution&amp;quot;, automatic program exploration is originally a bug finding technique that provides full coverage of the program. The first step in this approach is to symbolically represent the inputs to the program. Then, each statement that has a symbolic variable is executed by symbolic manipulation. If a branch condition is found with a symbolic variable, then both paths are conceptually taken. The last step is using the path constraints to find the concrete values that allow the program to go down both paths. This is done until there are no more possible inputs. &amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
Scalar and Array Analysis are actually two different forms of analysis that are often used together. Scalar Analysis checks for dependencies between scalar variables in code. A dependency can be defined as &amp;quot;when a memory location written on one iteration of a loop is accessed (read or write) on a different iteration.&amp;quot;&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Array Analysis is then used on code that could not be parallelized using Scalar Analysis. Array Analysis looks for arrays that can be privatized.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt; Privatization means to allow each thread its own private copy of the array in question, in whole or in part. This is only possible if there are no dependencies in the array.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
One problem in parallelizing code is that synchronization may be required in some programs to prevent errors in computation. Commutativity Analysis is used to determine which operations can be performed out of order without affecting the results of the code. It is based on the mathematical concept of commutativity, where if the result of running two operations is the same, regardless of which order they are performed in.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
LLVM works by taking C/C++ code and compiling it into LLVM IR bytecode. Optimizations are done at the IR level and then a code generator brings it back into native code to be executed.&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;polytope&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Polytope_model&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60579</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60579"/>
		<updated>2012-03-26T17:09:23Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Automatic Program Exploration */  Fixed spelling error.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
This technique involves running the program while recording each memory access in detail. The memory accesses are then looked at to determine what code can be parallelized. This technique is simple, but does not cover all the possible input combinations. These may lead to false positives in which code segments can be labelled as parallelizeable when they actually aren't.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
Polyhedral transformation involves representing loop iterations from the source code as lattice points in something called a polytope. This polytope is then transformed into a more optimized form. This approach is limited when pointers are found inside loops. The following pictures represent a polytope before and after the transformation.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
[[File:polytope_model.png|thumb|center|200px|Polytope model, unskewed]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|center|200px|Polytope model, skewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
Also known as &amp;quot;concolic execution&amp;quot;, &amp;quot;mixed concrete and symbolic execution&amp;quot;, or &amp;quot;symbolic execution&amp;quot;, automatic program exploration is originally a bug finding technique that provides full coverage of the program. The first step in this approach is to symbolically represent the inputs to the program. Then, each statement that has a symbolic variable is executed by symbolic manipulation. If a branch condition is found with a symbolic variable, then both paths are conceptually taken. The last step is using the path constraints to find the concrete values that allow the program to go down both paths. This is done until there are no more possible inputs. &amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
Scalar and Array Analysis are actually two different forms of analysis that are often used together. Scalar Analysis checks for dependencies between scalar variables in code. A dependency can be defined as &amp;quot;when a memory location written on one iteration of a loop is accessed (read or write) on a different iteration.&amp;quot;&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Array Analysis is then used on code that could not be parallelized using Scalar Analysis. Array Analysis looks for arrays that can be privatized.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt; Privatization means to allow each thread its own private copy of the array in question, in whole or in part. This is only possible if there are no dependencies in the array.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
One problem in parallelizing code is that synchronization may be required in some programs to prevent errors in computation. Commutativity Analysis is used to determine which operations can be performed out of order without affecting the results of the code. It is based on the mathematical concept of commutativity, where if the result of running two operations is the same, regardless of which order they are performed in.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
LLVM works by taking C/C++ code and compiling it into LLVM IR bytecode. Optimizations are done at the IR level and then a code generator brings it back into native code to be executed.&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60560</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60560"/>
		<updated>2012-03-22T06:45:48Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Scalar and Array Analysis */  Added section on Scalar and Array analysis.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
This technique involves running the program while recording each memory access in detail. The memory accesses are then looked at to determine what code can be parallelized. This technique is simple, but does not cover all the possible input combinations. These may lead to false positives in which code segments can be labelled as parallelizeable when they actually aren't.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
Polyhedral transformation involves representing loop iterations from the source code as lattice points in something called a polytope. This polytope is then transformed into a more optimized form. This approach is limited when pointers are found inside loops. The following pictures represent a polytope before and after the transformation.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
[[File:polytope_model.png|thumb|center|200px|Polytope model, unskewed]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|center|200px|Polytope model, skewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
Also known as &amp;quot;concolic execution&amp;quot;, &amp;quot;mixed concrete and symbolic execution&amp;quot;, or &amp;quot;symbolic execution&amp;quot;, automatic program exploration is originally a bug finding technique that provides full coverage of the program. The firs0 tstep in this approach is to symbolically represent the inputs to the program. Then, each statement that has a symbolic variable is executed by symbolic manipulation. If a branch condition is found with a symbolic variable, then both paths are conceptually taken. The last step is using the path constraints to find the concrete values that allow the program to go down both paths. This is done until there are no more possible inputs. &amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
Scalar and Array Analysis are actually two different forms of analysis that are often used together. Scalar Analysis checks for dependencies between scalar variables in code. A dependency can be defined as &amp;quot;when a memory location written on one iteration of a loop is accessed (read or write) on a different iteration.&amp;quot;&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Array Analysis is then used on code that could not be parallelized using Scalar Analysis. Array Analysis looks for arrays that can be privatized.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt; Privatization means to allow each thread its own private copy of the array in question, in whole or in part. This is only possible if there are no dependencies in the array.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
One problem in parallelizing code is that synchronization may be required in some programs to prevent errors in computation. Commutativity Analysis is used to determine which operations can be performed out of order without affecting the results of the code. It is based on the mathematical concept of commutativity, where if the result of running two operations is the same, regardless of which order they are performed in.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
LLVM works by taking C/C++ code and compiling it into LLVM IR bytecode. Optimizations are done at the IR level and then a code generator brings it back into native code to be executed.&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60559</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60559"/>
		<updated>2012-03-22T06:19:49Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Commutativity Analysis */  Added section on commutativity analysis.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
This technique involves running the program while recording each memory access in detail. The memory accesses are then looked at to determine what code can be parallelized. This technique is simple, but does not cover all the possible input combinations. These may lead to false positives in which code segments can be labelled as parallelizeable when they actually aren't.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
Polyhedral transformation involves representing loop iterations from the source code as lattice points in something called a polytope. This polytope is then transformed into a more optimized form. This approach is limited when pointers are found inside loops. The following pictures represent a polytope before and after the transformation.&amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
[[File:polytope_model.png|thumb|center|200px|Polytope model, unskewed]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|center|200px|Polytope model, skewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
Also known as &amp;quot;concolic execution&amp;quot;, &amp;quot;mixed concrete and symbolic execution&amp;quot;, or &amp;quot;symbolic execution&amp;quot;, automatic program exploration is originally a bug finding technique that provides full coverage of the program. The firs0 tstep in this approach is to symbolically represent the inputs to the program. Then, each statement that has a symbolic variable is executed by symbolic manipulation. If a branch condition is found with a symbolic variable, then both paths are conceptually taken. The last step is using the path constraints to find the concrete values that allow the program to go down both paths. This is done until there are no more possible inputs. &amp;lt;ref name=&amp;quot;chia&amp;quot; /&amp;gt;&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
One problem in parallelizing code is that synchronization may be required in some programs to prevent errors in computation. Commutativity Analysis is used to determine which operations can be performed out of order without affecting the results of the code. It is based on the mathematical concept of commutativity, where if the result of running two operations is the same, regardless of which order they are performed in.&amp;lt;ref name=&amp;quot;dipa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
LLVM works by taking C/C++ code and compiling it into LLVM IR bytecode. Optimizations are done at the IR level and then a code generator brings it back into native code to be executed.&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=File:Polytope_model_skewed.png&amp;diff=60054</id>
		<title>File:Polytope model skewed.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=File:Polytope_model_skewed.png&amp;diff=60054"/>
		<updated>2012-03-19T18:06:04Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60053</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60053"/>
		<updated>2012-03-19T18:05:53Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Polyhedral Transformation */  add skewed img&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
[[File:polytope_model.png|thumb|right|200px|Polytope model, unskewed]]&lt;br /&gt;
[[File:polytope_model_skewed.png|thumb|right|200px|Polytope model, skewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60042</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60042"/>
		<updated>2012-03-19T18:01:02Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Polyhedral Transformation */ resized polytope image&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
[[File:polytope_model.png|thumb|right|200px|Polytope model, unskewed]]&lt;br /&gt;
&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=File:Polytope_model.png&amp;diff=60035</id>
		<title>File:Polytope model.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=File:Polytope_model.png&amp;diff=60035"/>
		<updated>2012-03-19T17:58:03Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60033</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60033"/>
		<updated>2012-03-19T17:57:43Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: add polytope image&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
[[File:polytope_model.png]]&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60016</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60016"/>
		<updated>2012-03-19T17:49:44Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: removed notes section&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60015</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60015"/>
		<updated>2012-03-19T17:48:52Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: added reference for ncsa blog and fixed ref tag in limitations.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;ncsa&amp;quot; /&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;ncsa&amp;quot;&amp;gt;http://www.ncsa.illinois.edu/extremeideas/site/on_the_limits_of_automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60011</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=60011"/>
		<updated>2012-03-19T17:47:41Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Limitations */  Started limitations section&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
One of the major limitations of automatic parallelization is that a computer lacks the insight into the overall intention of a program that a human would have. The programmer understands what the program must do, and can use that to determine if there are alternate approaches or algorithms for parallelizing the code. Even when parallelizing manually, a programmer may not have enough insight into parallel programming, and will need the assistance of an expert to improve code performance.&amp;lt;ref name=&amp;quot;chai&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59998</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59998"/>
		<updated>2012-03-19T17:34:07Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added wiki ref for introduction&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually.&amp;lt;ref name=&amp;quot;wiki&amp;quot; /&amp;gt; There are several techniques that have been created for parallelizing code, but each has limitations.&lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59994</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59994"/>
		<updated>2012-03-19T17:31:56Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added Introduction&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Automatic Parallelism is the process of automatically converting sequential code into code that will make use of multiple processors. One main reason for implementing automatic parallelism is to save time and energy compared to converting the code manually. There are several techniques that have been created for parallelizing code, but each has limitations. &lt;br /&gt;
&lt;br /&gt;
== Techniques ==&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59556</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59556"/>
		<updated>2012-03-14T17:46:01Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* References */ Added wiki page to reflist&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Techniques ==&lt;br /&gt;
&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;wiki&amp;quot;&amp;gt;http://en.wikipedia.org/wiki/Automatic_parallelization&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59555</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59555"/>
		<updated>2012-03-14T17:45:08Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* References */ added dipasquale and yaun cho papers to reflist&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Techniques ==&lt;br /&gt;
&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;chia&amp;quot;&amp;gt;http://www.eecs.berkeley.edu/%7Echiayuan/cs262a/cs262a_parallel.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;dipa&amp;quot;&amp;gt;http://www.csc.villanova.edu/%7Etway/publications/DiPasquale_Masplas05_Paper5.pdf&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59552</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59552"/>
		<updated>2012-03-14T17:39:46Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* References */ add sample reference&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Techniques ==&lt;br /&gt;
&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;br /&gt;
&amp;lt;references&amp;gt;&lt;br /&gt;
&amp;lt;ref name=&amp;quot;sample&amp;quot;&amp;gt;Sample Ref, www.sample.com&amp;lt;/ref&amp;gt;&lt;br /&gt;
&amp;lt;/references&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59549</id>
		<title>Chapter 4a: Brandon Chisholm, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_4a:_Brandon_Chisholm,_Chris_Barile&amp;diff=59549"/>
		<updated>2012-03-14T17:38:05Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added Page structure&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Techniques ==&lt;br /&gt;
&lt;br /&gt;
===Profile-Driven Parallelism===&lt;br /&gt;
===Polyhedral Transformation===&lt;br /&gt;
===Automatic Program Exploration===&lt;br /&gt;
===Scalar and Array Analysis===&lt;br /&gt;
===Commutativity Analysis===&lt;br /&gt;
===Low Level Virtual Machine===&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Limitations ==&lt;br /&gt;
&lt;br /&gt;
== Notes ==&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
== References ==&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=59548</id>
		<title>CSC 456 Spring 2011</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=59548"/>
		<updated>2012-03-14T17:32:31Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added Link to our page for chapter 4a&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Chapter 1: Nick Nicholls, Albert Chu]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 4a: Brandon Chisholm, Chris Barile]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 6: Joshua Mohundro, Patrick Wong]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 6: Allison Hamann, Chris Barile]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch1 BC]] &amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch7 MN]]&amp;lt;br/&amp;gt;&lt;br /&gt;
[[CSC 456 Spring 2012/ch7 AA]]&amp;lt;br/&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59017</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59017"/>
		<updated>2012-02-22T05:51:49Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Organization */ Added image of cache to section&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Victim Cache==&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
===Handling Misses===&lt;br /&gt;
[[File:Jouppi_victimcaching_fig7.jpg|thumb|Jouppi's illustration of victim caching.]]&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
===Implementation===&lt;br /&gt;
Victim caching can be found in AMD's Opteron processor series produced specifically for servers and workstations. &amp;lt;ref&amp;gt;http://en.wikipedia.org/wiki/Opteron&amp;lt;/ref&amp;gt; Opteron processors use a victim cache that is capable of holding eight victim blocks. &amp;lt;ref&amp;gt; Hennessy, John L., and David A. Patterson.  ''Computer Architecture: A Quantitative Approach''. Elsevier, Inc., 2012, p. B-14.&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Sector Cache==&lt;br /&gt;
One early cache organization technique was sector cache. It was used on the IBM 360/85, which was one of the earliest commercial CPUs&amp;lt;ref name=&amp;quot;rot99&amp;quot;&amp;gt;Rothman, Jeffrey B. and Alan Jay Smith. Sector Cache Design and Performance&amp;lt;/ref&amp;gt;. Sector cache allowed for smaller tag sizes, which made searching in cache easier and quicker, without requiring cache lines to be excessively long.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; It was also easier to build with the circuit technology of the time.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; In testing during the design of the IBM 360/85, it was found to run at 81% of the maximum ideal efficiency calculated by IBM's designers.&amp;lt;ref name=&amp;quot;lip68&amp;quot;&amp;gt;Liptay, J. S. Structural Aspects of the System/360 Model 85, Part II: The Cache.&amp;lt;/ref&amp;gt; Sector caches were later replaced by set associative caches, which were found to be more efficient.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Organization===&lt;br /&gt;
[[Image:Sectorcache.png|200px|thumb|right|Sectors in cache are made of subsectors. Each subsector has a validity bit indicating whether or not it has been loaded with data.]]The cache is divided into sectors, which correspond to logical sectors on the main storage device.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Subsectors are similar to the lines in a direct mapped cache, and are only loaded into cache when needed. This prevents large amounts of data from having to be transferred for every memory reference. Subsectors have a validity bit that indicates whether they are loaded with data or not.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Miss Handling===&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=File:Sectorcache.png&amp;diff=59016</id>
		<title>File:Sectorcache.png</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=File:Sectorcache.png&amp;diff=59016"/>
		<updated>2012-02-22T05:44:18Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Depiction of sector cache with 3 subsectors. For each subsector, there is a validity bit (V) that tells if the subsector has been loaded with the correct data from disk.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Depiction of sector cache with 3 subsectors. For each subsector, there is a validity bit (V) that tells if the subsector has been loaded with the correct data from disk.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59015</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59015"/>
		<updated>2012-02-22T05:23:10Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */ Fixed missing text for Liptay reference&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Victim Cache==&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
===Handling Misses===&lt;br /&gt;
[[File:Jouppi_victimcaching_fig7.jpg|thumb|Jouppi's illustration of victim caching.]]&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
===Implementation===&lt;br /&gt;
Victim caching can be found in AMD's Opteron processor series produced specifically for servers and workstations. &amp;lt;ref&amp;gt;http://en.wikipedia.org/wiki/Opteron&amp;lt;/ref&amp;gt; Opteron processors use a victim cache that is capable of holding eight victim blocks. &amp;lt;ref&amp;gt; Hennessy, John L., and David A. Patterson.  ''Computer Architecture: A Quantitative Approach''. Elsevier, Inc., 2012, p. B-14.&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Sector Cache==&lt;br /&gt;
One early cache organization technique was sector cache. It was used on the IBM 360/85, which was one of the earliest commercial CPUs&amp;lt;ref name=&amp;quot;rot99&amp;quot;&amp;gt;Rothman, Jeffrey B. and Alan Jay Smith. Sector Cache Design and Performance&amp;lt;/ref&amp;gt;. Sector cache allowed for smaller tag sizes, which made searching in cache easier and quicker, without requiring cache lines to be excessively long.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; It was also easier to build with the circuit technology of the time.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; In testing during the design of the IBM 360/85, it was found to run at 81% of the maximum ideal efficiency calculated by IBM's designers.&amp;lt;ref name=&amp;quot;lip68&amp;quot;&amp;gt;Liptay, J. S. Structural Aspects of the System/360 Model 85, Part II: The Cache.&amp;lt;/ref&amp;gt; Sector caches were later replaced by set associative caches, which were found to be more efficient.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Organization===&lt;br /&gt;
The cache is divided into sectors, which correspond to logical sectors on the main storage device.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Subsectors are similar to the lines in a direct mapped cache, and are only loaded into cache when needed. This prevents large amounts of data from having to be transferred for every memory reference. Subsectors have a validity bit that indicates whether they are loaded with data or not.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Miss Handling===&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59014</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59014"/>
		<updated>2012-02-22T05:21:42Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  Added references&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Victim Cache==&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
===Handling Misses===&lt;br /&gt;
[[File:Jouppi_victimcaching_fig7.jpg|thumb|Jouppi's illustration of victim caching.]]&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
===Implementation===&lt;br /&gt;
Victim caching can be found in AMD's Opteron processor series produced specifically for servers and workstations. &amp;lt;ref&amp;gt;http://en.wikipedia.org/wiki/Opteron&amp;lt;/ref&amp;gt; Opteron processors use a victim cache that is capable of holding eight victim blocks. &amp;lt;ref&amp;gt; Hennessy, John L., and David A. Patterson.  ''Computer Architecture: A Quantitative Approach''. Elsevier, Inc., 2012, p. B-14.&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Sector Cache==&lt;br /&gt;
One early cache organization technique was sector cache. It was used on the IBM 360/85, which was one of the earliest commercial CPUs&amp;lt;ref name=&amp;quot;rot99&amp;quot;&amp;gt;Rothman, Jeffrey B. and Alan Jay Smith. Sector Cache Design and Performance&amp;lt;/ref&amp;gt;. Sector cache allowed for smaller tag sizes, which made searching in cache easier and quicker, without requiring cache lines to be excessively long.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; It was also easier to build with the circuit technology of the time.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt; In testing during the design of the IBM 360/85, it was found to run at 81% of the maximum ideal efficiency calculated by IBM's designers.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;Liptay, J. S. Structural Aspects of the System/360 Model 85, Part II: The Cache.&amp;lt;/ref&amp;gt; Sector caches were later replaced by set associative caches, which were found to be more efficient.&amp;lt;ref name=&amp;quot;rot99&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Organization===&lt;br /&gt;
The cache is divided into sectors, which correspond to logical sectors on the main storage device.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Subsectors are similar to the lines in a direct mapped cache, and are only loaded into cache when needed. This prevents large amounts of data from having to be transferred for every memory reference. Subsectors have a validity bit that indicates whether they are loaded with data or not.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Miss Handling===&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt; Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&amp;lt;ref name=&amp;quot;lip68&amp;quot;/&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59013</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=59013"/>
		<updated>2012-02-22T05:15:16Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */ Added background information on sector cache.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Victim Cache==&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
===Handling Misses===&lt;br /&gt;
[[File:Jouppi_victimcaching_fig7.jpg|thumb|Jouppi's illustration of victim caching.]]&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
===Implementation===&lt;br /&gt;
Victim caching can be found in AMD's Opteron processor series produced specifically for servers and workstations. &amp;lt;ref&amp;gt;http://en.wikipedia.org/wiki/Opteron&amp;lt;/ref&amp;gt; Opteron processors use a victim cache that is capable of holding eight victim blocks. &amp;lt;ref&amp;gt; Hennessy, John L., and David A. Patterson.  ''Computer Architecture: A Quantitative Approach''. Elsevier, Inc., 2012, p. B-14.&amp;lt;/ref&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==Sector Cache==&lt;br /&gt;
One early cache organization technique was sector cache. It was used on the IBM 360/85, which was one of the earliest commercial CPUs. Sector cache allowed for smaller tag sizes, which made searching in cache easier and quicker, without requiring cache lines to be excessively long. It was also easier to build with the circuit technology of the time. In testing during the design of the IBM 360/85, it was found to run at 81% of the maximum ideal efficiency calculated by IBM's designers. Sector caches were later replaced by set associative caches, which were found to be more efficient.&lt;br /&gt;
&lt;br /&gt;
===Organization===&lt;br /&gt;
The cache is divided into sectors, which correspond to logical sectors on the main storage device. When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors. Subsectors are similar to the lines in a direct mapped cache, and are only loaded into cache when needed. This prevents large amounts of data from having to be transferred for every memory reference. Subsectors have a validity bit that indicates whether they are loaded with data or not.&lt;br /&gt;
&lt;br /&gt;
===Miss Handling===&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits. Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58934</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58934"/>
		<updated>2012-02-20T18:16:43Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  Added sub headers that will have the content expanded on.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Victim Cache==&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
===Handling Misses===&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
==Sector Cache==&lt;br /&gt;
Sector cache is where the cache is divided into sectors. &lt;br /&gt;
&lt;br /&gt;
===Sectors and Subsectors===&lt;br /&gt;
These sectors correspond to a logical sector on the main storage device. When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors. Subsectors are similar to the lines in a direct mapped cache. &lt;br /&gt;
&lt;br /&gt;
===Miss Handling===&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits. Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&lt;br /&gt;
&lt;br /&gt;
===Theory===&lt;br /&gt;
One reason for this approach is that programs are generally organized in contiguous blocks on disk, . Another is that data is first looked up by sector, and then by subsector, which means that it can be found much quicker, and the hardware to do the simultaneous comparison of tags is less expensive.&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58931</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58931"/>
		<updated>2012-02-20T18:07:13Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added References list at the bottom of the page.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Victim Cache=&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
==Handling Misses==&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.&amp;lt;ref&amp;gt;http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547&amp;lt;/ref&amp;gt;&amp;lt;ref&amp;gt;http://dl.acm.org/citation.cfm?id=1101876&amp;lt;/ref&amp;gt; Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
Sector cache is where the cache is divided into sectors. These sectors correspond to a logical sector on the main storage device. When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors. Subsectors are similar to the lines in a direct mapped cache. &lt;br /&gt;
&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits. Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&lt;br /&gt;
&lt;br /&gt;
One reason for this approach is that programs are generally organized in contiguous blocks on disk, . Another is that data is first looked up by sector, and then by subsector, which means that it can be found much quicker, and the hardware to do the simultaneous comparison of tags is less expensive.&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&amp;lt;references /&amp;gt;&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58787</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58787"/>
		<updated>2012-02-15T07:03:13Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  Revising content. Editing for clarity.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=Victim Cache=&lt;br /&gt;
Victim caches were first proposed by Norman P. Jouppi in 1990. Victim caching implements a small, fully-associative cache between direct-mapped L1 memory and the next level of memory. The cache allows lines evicted from the L1 cache a “second-chance” by loading them into the victim cache. Victim caches decrease the overall conflict miss rate (Jouppi).&lt;br /&gt;
&lt;br /&gt;
Direct-mapped caches can especially benefit from victim caching due to their large miss rates. Victim caching allows direct-mapped caches to still be used in order to take advantage of their speed while decreasing the miss rate to an even lower rate than the miss rate found in set-associative caches (Jouppi).&lt;br /&gt;
&lt;br /&gt;
==Handling Misses==&lt;br /&gt;
The proposed victim cache is fully-associative and lies between the L1 memory and the next level of memory. While Jouppi proposed a victim cache with 1 to 5 entries, Naz et al. proposed that the victim caches should be 4 to 16 cache lines.[http://ieeexplore.ieee.org/xpl/freeabs_all.jsp?arnumber=134547][http://dl.acm.org/citation.cfm?id=1101876] Regardless of the size, when a miss occurs in the L1 cache, the victim cache is then scanned for the wanted line. If a miss occurs in both the L1 and victim cache, the needed line is then pulled from the next level, and the line evicted from the L1 cache is then placed in the victim cache. If a miss occurs in the L1 cache but hits in the victim cache, the two lines are swapped between the two caches. Thus, this eliminates the majority of conflict misses that occur due to temporal locality.&lt;br /&gt;
&lt;br /&gt;
Sector cache is where the cache is divided into sectors. These sectors correspond to a logical sector on the main storage device. When sectors are needed, they are not loaded into cache all at once, but in smaller pieces known as subsectors. Subsectors are similar to the lines in a direct mapped cache. &lt;br /&gt;
&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then the subsector where the data is located is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When data from other subsectors within a loaded sector are requested, the system loads those subsectors into the cache sector and sets their validity bits. Sectors are not removed from the cache until the system needs to reclaim the space to process another request.&lt;br /&gt;
&lt;br /&gt;
One reason for this approach is that programs are generally organized in contiguous blocks on disk, . Another is that data is first looked up by sector, and then by subsector, which means that it can be found much quicker, and the hardware to do the simultaneous comparison of tags is less expensive.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58061</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58061"/>
		<updated>2012-02-06T18:53:41Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  drafting content&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;br /&gt;
&lt;br /&gt;
=Victim Cache=&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
A. Organziation&lt;br /&gt;
 1) Sectors&lt;br /&gt;
 2) Subsectors&lt;br /&gt;
 3) Validity Bit&lt;br /&gt;
B. Load Procedure&lt;br /&gt;
 1) Sector miss&lt;br /&gt;
 2) Subsector miss&lt;br /&gt;
C. Advantages&lt;br /&gt;
&lt;br /&gt;
Sector cache is where the cache is divided into sectors. These sectors correspond to a logical sector on the main storage device. Sectors are not loaded into the cache all at once, but in smaller pieces known as subsectors, which are similar to the cache lines in direct mapped cache. &lt;br /&gt;
&lt;br /&gt;
When a process requests data from a disk sector that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then a portion of that sector, otherwise known as the subsector, is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When data from other subsectors are requested, the system loads those subsectors into the cache sector and sets their validity bits. The sector is not removed from the cache until it is needed for another program.&lt;br /&gt;
&lt;br /&gt;
One reason for this approach is that programs are generally organized in contiguous blocks on disk. Another is that data is first looked up by sector, and then by subsector, which means that it can be found much quicker, and the hardware to do the simultaneous comparison of tags is less expensive.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58054</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58054"/>
		<updated>2012-02-06T18:39:51Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Change to Sector Cache, accidentally edited whole page&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;br /&gt;
&lt;br /&gt;
=Victim Cache=&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
A. Organziation&lt;br /&gt;
 1) Sectors&lt;br /&gt;
 2) Subsectors&lt;br /&gt;
 3) Validity Bit&lt;br /&gt;
B. Load Procedure&lt;br /&gt;
 1) Sector miss&lt;br /&gt;
 2) Subsector miss&lt;br /&gt;
C. Advantages&lt;br /&gt;
&lt;br /&gt;
Sector Caches are organized into sectors, which correspond to sectors of the main storage. Sectors are then organized into subsectors, which are similar to cache lines. Each subsector has a &amp;quot;validity bit&amp;quot; which tells whether or not that subsector has been loaded into cache from the main storage. &lt;br /&gt;
&lt;br /&gt;
When a process requests data that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then a portion of that sector, known as the subsector, is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. &lt;br /&gt;
&lt;br /&gt;
One reason for this approach is that programs are generally organized in contiguous blocks on disk. Another is that data is first looked up by sector, and then by subsector, which means that it can be found much quicker, and the hardware to do the simultaneous comparison of tags is less expensive.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58037</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=58037"/>
		<updated>2012-02-06T18:27:45Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  Adding paragraph about cache loading&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;br /&gt;
&lt;br /&gt;
=Victim Cache=&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
A. Organziation&lt;br /&gt;
 1) Sectors&lt;br /&gt;
 2) Subsectors&lt;br /&gt;
 3) Validity Bit&lt;br /&gt;
B. Load Procedure&lt;br /&gt;
 1) Sector miss&lt;br /&gt;
 2) Subsector miss&lt;br /&gt;
C. Advantages&lt;br /&gt;
&lt;br /&gt;
Sector Caches are organized into sectors and subsectors. A sector is a collection of smaller subsectors. Subsectors are similar to lines in a direct-mapped cache. Each subsector has a &amp;quot;validity bit&amp;quot; which tells whether or not that subsector has been loaded into cache from the main storage. &lt;br /&gt;
&lt;br /&gt;
When a process requests data that is not in the cache, a cache sector is assigned to the sector on the main storage device where the requested data is stored. Then a portion of that sector, known as the subsector, is loaded into the cache. The subsector's validity bit is then set to reflect that it has been loaded from the main storage. When a process requests data from the loaded sector that is not already in the cache, the subsector it is in becomes loaded into memory.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=57331</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=57331"/>
		<updated>2012-01-30T19:08:20Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /* Sector Cache */  Started drafting content. Need to add reference section to page and possibly archive papers.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;br /&gt;
&lt;br /&gt;
=Victim Cache=&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
A. Organziation&lt;br /&gt;
 1) Sectors&lt;br /&gt;
 2) Subsectors&lt;br /&gt;
 3) Validity Bit&lt;br /&gt;
B. Load Procedure&lt;br /&gt;
 1) Sector miss&lt;br /&gt;
 2) Subsector miss&lt;br /&gt;
C. Advantages&lt;br /&gt;
&lt;br /&gt;
Sector Caches are organized into sectors and subsectors. A sector is a collection of smaller subsectors. Subsectors are similar to lines in a direct-mapped cache. Each subsector has a &amp;quot;validity bit&amp;quot; which tells whether or not that subsector has been loaded into cache from the main memory.&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=57326</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=57326"/>
		<updated>2012-01-30T19:00:01Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: /*Sector Cache*/  Starting to outline content for this page. Will replace with content.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;br /&gt;
&lt;br /&gt;
=Victim Cache=&lt;br /&gt;
&lt;br /&gt;
=Sector Cache=&lt;br /&gt;
A. Organziation&lt;br /&gt;
 1) Sectors&lt;br /&gt;
 2) Subsectors&lt;br /&gt;
 3) Validity Bit&lt;br /&gt;
B. Load Procedure&lt;br /&gt;
 1) Sector miss&lt;br /&gt;
 2) Subsector miss&lt;br /&gt;
C. Advantages&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=56609</id>
		<title>Chapter 6: Allison Hamann, Chris Barile</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Chapter_6:_Allison_Hamann,_Chris_Barile&amp;diff=56609"/>
		<updated>2012-01-23T18:54:57Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Adding reference to ieee paper&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=876437&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=56605</id>
		<title>CSC 456 Spring 2011</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=56605"/>
		<updated>2012-01-23T18:48:00Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Added missing line break&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Chapter 6: Joshua Mohundro, Patrick Wong]]&amp;lt;br /&amp;gt;&lt;br /&gt;
[[Chapter 6: Allison Hamann, Chris Barile]]&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=56604</id>
		<title>CSC 456 Spring 2011</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC_456_Spring_2011&amp;diff=56604"/>
		<updated>2012-01-23T18:47:23Z</updated>

		<summary type="html">&lt;p&gt;Crbarile: Adding Page for Chapter 6&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Chapter 6: Joshua Mohundro, Patrick Wong]]&lt;br /&gt;
[[Chapter 6: Allison Hamann, Chris Barile]]&lt;/div&gt;</summary>
		<author><name>Crbarile</name></author>
	</entry>
</feed>