<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Vvijaya</id>
	<title>Expertiza_Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.expertiza.ncsu.edu/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Vvijaya"/>
	<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=Special:Contributions/Vvijaya"/>
	<updated>2026-10-01T10:09:38Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.41.0</generator>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4700</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4700"/>
		<updated>2007-09-27T22:56:40Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
''Mapping'' assigns the processes to different processors with an intent of minimizing the communication between them. Considering the ''grep'' problem, the master maps the document parts to processors executing map/reduce function. It is fair to assume that much of the time is spent in transferring the data across the network, from the master to the workers. One way of solving the problem is to, store replicas of the splits at different worker locations. The master can now assign the M splits in such a way that the replicas of the input document are local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Summary=&lt;br /&gt;
As seen from the ''grep'' example, ''MapReduce'' is a highly scalable and customizable algorithm for processing large data sets. In cases where, the data set is very large compared to the the number of available workers, the problem can be solved recursively. The final output can be fed back to the ''MapReduce'' engine by further refining the key/value pairs. Also, because of the user-customizable nature of the inputs, data and/or keys, any ''MapReduce'' solution is easily portable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4699</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4699"/>
		<updated>2007-09-27T22:55:54Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
''Mapping'' assigns the processes to different processors with an intent of minimizing the communication between them. Considering the ''grep'' problem, the master maps the document parts to processors executing map/reduce function. It is fair to assume that much of the time is spent in transferring the data across the network, from the master to the workers. One way of solving the problem is to, store replicas of the splits at different worker locations. The master can now assign the M splits in such a way that the replicas of the input document are local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=Summary=&lt;br /&gt;
As seen from the ''grep'' example, ''MapReduce'' is a highly scalable and customizable algorithm for processing large data sets. In cases where, the data set is very large compared to the the number of available workers, the problem can be solved recursively. The final output can be fed back to the ''MapReduce'' engine by further refining the key/value pairs. Also, because of the user-customizable nature of the inputs, data and/or keys, any ''MapReduce'' solution is easily portable.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4698</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4698"/>
		<updated>2007-09-27T22:55:37Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Summary */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
''Mapping'' assigns the processes to different processors with an intent of minimizing the communication between them. Considering the ''grep'' problem, the master maps the document parts to processors executing map/reduce function. It is fair to assume that much of the time is spent in transferring the data across the network, from the master to the workers. One way of solving the problem is to, store replicas of the splits at different worker locations. The master can now assign the M splits in such a way that the replicas of the input document are local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=Summary=&lt;br /&gt;
As seen from the ''grep'' example, ''MapReduce'' is a highly scalable and customizable algorithm for processing large data sets. In cases where, the data set is very large compared to the the number of available workers, the problem can be solved recursively. The final output can be fed back to the ''MapReduce'' engine by further refining the key/value pairs. Also because of the user-customizable nature of the inputs, data and/or keys, any ''MapReduce'' solution is easily portable.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4697</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4697"/>
		<updated>2007-09-27T22:54:46Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
''Mapping'' assigns the processes to different processors with an intent of minimizing the communication between them. Considering the ''grep'' problem, the master maps the document parts to processors executing map/reduce function. It is fair to assume that much of the time is spent in transferring the data across the network, from the master to the workers. One way of solving the problem is to, store replicas of the splits at different worker locations. The master can now assign the M splits in such a way that the replicas of the input document are local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=Summary=&lt;br /&gt;
As seen from the ''grep'' example, ''MapReduce'' is a highly scalable and customizable algorithm for processing large data sets. In cases where, the data set is very large compared to the the number of available workers, the problem can be solved recursively. The final output can be fed back to a ''MapReduce'' engine by further refining the key/value pairs. Also because of the user-customizable nature of the inputs, data and/or keys, any ''MapReduce'' solution is easily portable.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4696</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4696"/>
		<updated>2007-09-27T22:38:44Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Mapping */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
''Mapping'' assigns the processes to different processors with an intent of minimizing the communication between them. Considering the ''grep'' problem, the master maps the document parts to processors executing map/reduce function. It is fair to assume that much of the time is spent in transferring the data across the network, from the master to the workers. One way of solving the problem is to, store replicas of the splits at different worker locations. The master can now assign the M splits in such a way that the replicas of the input document are local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4695</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4695"/>
		<updated>2007-09-27T22:32:22Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Orchestration */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represent an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4694</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4694"/>
		<updated>2007-09-27T22:31:31Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Orchestration */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration'' is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represents an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
By providing an intermediate step for reduce ''MapReduce'' algorithm ensures that data locality is maintained.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4693</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4693"/>
		<updated>2007-09-27T22:30:24Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Orchestration */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
''Orchestration&amp;quot; is a step which groups the processes based on their interdependencies. The main purpose of this step is to reduce the synchronization and communication overhead between the processes. In the ''grep'' problem, the intermediate counts generated by the map workers represents an unsorted data set. By alphabetically sorting the intermediate word counts the reduce workers can be assigned with a coherent data set. This increases the probability that each reduce worker may work on a single or a small group of word counts.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4692</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4692"/>
		<updated>2007-09-27T22:19:19Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4691</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4691"/>
		<updated>2007-09-27T22:18:59Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. These partial counts from map workers are reduced into a final count by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides for load based task assignment which is beneficial in a scalable multiprocessor architecture&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4690</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4690"/>
		<updated>2007-09-27T22:17:06Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is a step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched.  The word(s) to be searched represents the key/value pair. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. The partial counts form map workers are reduced into a final count of occurrences by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers.&lt;br /&gt;
MapReduce provides for load based task assignment which is beneficial in a scalable multiprocessor architecture&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4689</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4689"/>
		<updated>2007-09-27T22:16:48Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is  step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result. In the ''grep'' problem, each worker is assigned a part of the document along with the word(s) to be searched.  The word(s) to be searched represents the key/value pair. Each worker searches the part of the document assigned to it and generates a partial count of the occurrences. The partial counts form map workers are reduced into a final count of occurrences by the reduce workers. The master is responsible for assigning M parts of the document and the R intermediate counts to idle workers.&lt;br /&gt;
MapReduce provides for load based task assignment which is beneficial in a scalable multiprocessor architecture&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4688</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4688"/>
		<updated>2007-09-27T22:07:24Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, the ''grep'' problem, which counts the number of occurrences of a specific word or a group of words in a document, the document is split into M equal parts. The size of each part is based on the number of distinct word(s)to be searched.  Hence searching for larger group of words requires the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is  step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4687</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4687"/>
		<updated>2007-09-27T22:05:13Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is  step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate R partial/complete outputs. These partial outputs are consolidated into a final result&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4686</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4686"/>
		<updated>2007-09-27T22:03:34Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Assignment */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
''Assignment'' is  step that assigns the related decomposed tasks to processes.&lt;br /&gt;
In ''MapReduce'' splits are assigned to idle map workers that generate ''R'' partial/complete outputs. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4685</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4685"/>
		<updated>2007-09-27T21:55:47Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides an adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4684</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4684"/>
		<updated>2007-09-27T21:55:33Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts. &amp;lt;br&amp;gt;&lt;br /&gt;
Thus ''MapReduce'' provides a adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4683</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4683"/>
		<updated>2007-09-27T21:55:15Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts.&lt;br /&gt;
Thus ''MapReduce'' provides a adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4682</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4682"/>
		<updated>2007-09-27T21:55:06Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts.&lt;br /&gt;
&lt;br /&gt;
Thus ''MapReduce'' provides a adaptive method for decomposing a problem.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4681</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4681"/>
		<updated>2007-09-27T21:49:53Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
''Decomposition'' is a step that divides the problem into manageable tasks. In ''MapReduce'' the input data set is divided into M splits. Typically a master process takes care of the decomposition. The ''MapReduce'' algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4680</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4680"/>
		<updated>2007-09-27T21:48:31Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Decomposition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;=MapReduce=&lt;br /&gt;
One of the major problems faced in the network world, is the need to classify huge amounts of data into easily understandable smaller data sets. It might be the problem of classifying hits for a website from different sources or classifying the frequency of occurrence of words in text, the problem involves analyzing large amounts of data into small understandable data sets.&lt;br /&gt;
&lt;br /&gt;
''MapReduce'' is a programming model for processing large data sets. This was developed as a scalable, multiprocessing model by [http://www.google.com/about.html Google]. This model boasts of high scalability and high fault tolerance. ''MapReduce'' helps in ''mapping'' large data sets to smaller number of keys, thus ''reducing'' the data set itself to a more manageable proportion. More specifically, users specify a ''map'' function to process a key/value pair and generate a set of intermediate key/value pairs. The ''reduce'' function merges these intermediate key/values and generates the consolidated output.&lt;br /&gt;
&lt;br /&gt;
Some of the examples, where ''MapReduce'' can be employed are:&lt;br /&gt;
*Distributed Grep&lt;br /&gt;
*Count of URL Access Frequency&lt;br /&gt;
*ReverseWeb-Link Graph&lt;br /&gt;
*Term-Vector per Host&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=Steps of Parallelization=&lt;br /&gt;
The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:MapReduce03.JPG]]&amp;lt;br&amp;gt;[http://209.85.163.132/papers/mapreduce-osdi04.pdf Execution Overview]&lt;br /&gt;
&lt;br /&gt;
===Decomposition===&lt;br /&gt;
'Decomposition' is a step that divides the problem into manageable tasks. In MapReduce the input data set is divided into M splits. Typically a master process takes care of the decomposition. The MapReduce algorithm tries to map a large dataset into a smaller number of key/value pairs.&lt;br /&gt;
The number of key/value pairs to be arrived at, will determine the size of the split. For example, in a problem of counting the number of occurrences of a specific word or a group of words in a document, the document would be split into M equal parts. The size of each part would be based on the number of distinct word(s)to be searched.  Thus searching for larger group of words would require the document to be divided into smaller parts.&lt;br /&gt;
&lt;br /&gt;
Each part would be searched in parallel for the specific pattern. The number of parts, M, depends on the size of the entire document(data set) and the number of different words(keys) searched.&lt;br /&gt;
&lt;br /&gt;
===Assignment===&lt;br /&gt;
The tasks are assigned to various processes that execute them to generate partial/complete output. For the problem explained above, each process would be assigned a part of the document with a word(s) to be searched. The word(s) to be searched represents the key/value pair. Each process searches the part of the document assigned to it and generates a partial count of occurrences. The output, again will be generated in the form of key/value pair, where key represents the word(s) searched and value represents the number of occurrences, for the part of the document assigned.&lt;br /&gt;
&lt;br /&gt;
===Orchestration=== &lt;br /&gt;
The main purpose of orchestration is to reduce the synchronization and communication overhead between the processes.  The high I/O traffic generated may become a crucial problem. In the above example, the intermediate values generated by the processes are consolidated into a final output. The reduce workers read the data from the map workers through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key values.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
[[Image:MapReduce02.gif]]&amp;lt;br&amp;gt;[http://labs.google.com/papers/mapreduce-osdi04-slides/index-auto-0008.html Parallel Execution]&amp;lt;br&amp;gt;&lt;br /&gt;
===Mapping===&lt;br /&gt;
This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document parts are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. document parts and intermediate key/value pairs) across the network. One way of solving the problem is to, store replicas of the document parts at different worker locations. The master can now assign the M parts in such a way that the replicas of the input document part is local to the processor on which it is processed. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the communication.&lt;br /&gt;
&lt;br /&gt;
=References=&lt;br /&gt;
*[http://labs.google.com/papers/mapreduce.html Google Research publications]&lt;br /&gt;
*[http://en.wikipedia.org/wiki/MapReduce Wikipage on MapReduce]]&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4406</id>
		<title>CSC/ECE 506 Fall 2007/wiki2 4 BV</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki2_4_BV&amp;diff=4406"/>
		<updated>2007-09-24T22:44:08Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Google uses &amp;quot;Map Reduce&amp;quot;, a programming model for processing and generating large data sets. This model has been known for the high scalability and fault tolerant nature. At a bird's view, users specify a &amp;quot;map&amp;quot; function to process a key/value pair and generate a set of intermediate key/value pairs. The reduce function merges these intermediate values associated with the same intermediate key. The programs written based on this model can be parallelized and executed concurrently on a number of systems. The model provides the capability for partitioning of data, assignment to different processes and merging the results.&lt;br /&gt;
&lt;br /&gt;
The steps of parallelization are explained as below.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Decomposition  The data set is decomposed into N sets. Typically a master process assumes the responsibility of it. In the context of a problem like counting the number of occurrences of a specific word or a group of words in a document. The document would be split into N pieces. The number of pieces depend upon the size of the entire document and the number of different words. This phase will give the number of concurrent tasks.&lt;br /&gt;
&lt;br /&gt;
  &lt;br /&gt;
Assignment The tasks are assigned to various processes that execute them to generate partial/complete output. For the example explained above, a process would be assigned a split with a key/value pair which the represent the group of words to be looked for. Each process reads the key/value pair and generates an intermediate result in the form of a key/intermediate number of occurrences. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Orchestration The main purpose of orchestration is to reduced the synchronization and communication overhead between the tasks.  The high I/O traffic generated may become a crucial problem as the network bottle neck. In the above example, the intermediate values generated by the processes are reduced to generated a  final output. The reduce workers read the data from the map workes through remote procedure calls. The data read is sorted based on the intermediate keys so that all occurrences of the same key are grouped together. This increases the probability that each reduce process may work on the same key for intermediate key/value pairs.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Mapping  This phase will assign the processes to different processors with an intent of minimizing the communication between them. Considering the above mentioned problem, the document splits are mapped to different map workers (i.e processors) by the master. It is quite possible that much of the time is spent on transferring the data ( i.e. splits and intermediate key/value pairs) across the network. One way of solving the problem is to assign the splits in such a way that the replica of the input split is local to the processor on which it is processor. This will eliminate the actual file transfer from the master to the map worker and thereby reducing the comminication.&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3548</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3548"/>
		<updated>2007-09-11T03:57:18Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* New Trends in Vector and Array Processing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like :-&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
*http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3547</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3547"/>
		<updated>2007-09-11T03:57:09Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* New Trends in Vector and Array Processing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like :-&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
*http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3545</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3545"/>
		<updated>2007-09-11T03:56:49Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like :-&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
*http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3357</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3357"/>
		<updated>2007-09-10T23:27:55Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* New Trends in Vector and Array Processing */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like :-&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
*http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3351</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3351"/>
		<updated>2007-09-10T23:18:22Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Bibliography */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
*http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3349</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3349"/>
		<updated>2007-09-10T23:17:45Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY1 supercomputer at http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3348</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3348"/>
		<updated>2007-09-10T23:17:30Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Block diagram of CRAY 1 supercomputer http://research.microsoft.com/users/gbell/CrayTalk/sld063.htm&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3345</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3345"/>
		<updated>2007-09-10T23:12:15Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3344</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3344"/>
		<updated>2007-09-10T23:12:05Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3343</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3343"/>
		<updated>2007-09-10T23:11:39Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Bibliography */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
Array and Vector Procesing&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;br /&gt;
&lt;br /&gt;
Recent Trends&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/PowerPC_G4&lt;br /&gt;
&lt;br /&gt;
http://en.wikipedia.org/wiki/Cell_microprocessor&lt;br /&gt;
&lt;br /&gt;
http://www.ausairpower.net/OSR-0600.html&lt;br /&gt;
&lt;br /&gt;
http://www.cs.ucsc.edu/~mslater/papers/VectorProcessingG4.pdf&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3341</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3341"/>
		<updated>2007-09-10T23:07:42Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Importance and Trends */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== New Trends in Vector and Array Processing ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the traditional problems of supercomputing, vector processing is finding applications in various other domains like&lt;br /&gt;
&lt;br /&gt;
a. Gaming&lt;br /&gt;
&lt;br /&gt;
Most of the code in gaming is a mixture of integer, floating point and vector calculations. This is best handled by a CPU with a vector unit. Dot &lt;br /&gt;
products are very critical to games as they are used to find out vector lengths, projections and transformations. Vector processing is best suited &lt;br /&gt;
for them.&lt;br /&gt;
    &lt;br /&gt;
Sony's third generation Play Stations called PS3 use Power PC based Cell processor having AltiVec vector processing units.&lt;br /&gt;
&lt;br /&gt;
b. Image Processing&lt;br /&gt;
&lt;br /&gt;
Image processing is a challenging domain: high computing power is required to calculate image adjustments in real time.  These typically include &lt;br /&gt;
processing of pixel arrays and performing mathematical transforms (like Fast Four Transforms) on them. Vector processing helps in accelerating the &lt;br /&gt;
performance for such applications.&lt;br /&gt;
   &lt;br /&gt;
Apple Computers use the fourth generation of Power PC processors developed by Motorola ( MPC7400 series ) which incorporates the AltiVec model and is &lt;br /&gt;
one of the widely used embedded processors in this realm.&lt;br /&gt;
&lt;br /&gt;
c. Signal processing applications like RADAR and SONAR&lt;br /&gt;
&lt;br /&gt;
These processors can be harnessed well for signal processing applications. SONAR and RADAR are computationally intensive embedded applications&lt;br /&gt;
Vector processing architectures like AltiVec( from IBM, Motorola and Apple) deliver very well in the areas of graphics and multimedia which can stretch &lt;br /&gt;
even the current super scalar CPUs.&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3036</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3036"/>
		<updated>2007-09-06T02:47:05Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Description */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and storing the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more throughput than serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements rather than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. &lt;br /&gt;
&lt;br /&gt;
Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research on applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3007</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=3007"/>
		<updated>2007-09-06T02:34:49Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Importance and Trends */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org. Horst D Simon, the Director of NERSC Center, gave a presentation on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. &lt;br /&gt;
&lt;br /&gt;
Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processors and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research on applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2997</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2997"/>
		<updated>2007-09-06T02:28:41Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Future */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges like understanding, detecting and predicting the human influence on climate and modeling the earth system including atmosphere, ocean, land and their interactions require the aid of supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2991</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2991"/>
		<updated>2007-09-06T02:20:17Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* History */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length. These elements communicated with each other through an interconnect that resembled a ring. Each element was provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2985</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2985"/>
		<updated>2007-09-06T02:14:40Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Definition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described as an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width. The pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2981</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2981"/>
		<updated>2007-09-06T02:13:27Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Definition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array Processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2976</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2976"/>
		<updated>2007-09-06T02:12:14Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Definition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting Instruction Level Parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2973</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2973"/>
		<updated>2007-09-06T02:11:31Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Definition */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items at the same time. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting instruction level parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2970</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2970"/>
		<updated>2007-09-06T02:10:41Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* Bibliography */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting instruction level parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://en.wikipedia.org/wiki/Vector_processor&lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2968</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2968"/>
		<updated>2007-09-06T02:09:08Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: /* History */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting instruction level parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD (Single Instruction Multiple Data), which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2966</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2966"/>
		<updated>2007-09-06T02:07:08Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting instruction level parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD i.e. single instruction multiple data, which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2964</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2964"/>
		<updated>2007-09-06T02:04:12Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory which execute the instruction. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting instruction level parallelism (ILP) where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of  n-bit data elements each of a definite width and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
Flynn’s Taxonomy classifies computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD i.e. single instruction multiple data, which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory capacity for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with a memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element 63)is directly connected to PE 0(Processing Element 0).&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation. It used vector registers to hold multiple data elements. It had a high startup time and relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched by the control unit will be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyzes, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.lib.ncsu.edu:2162/citation.cfm?id=808415&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses vectors i.e. a series of values or elements than a scalar i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional Units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect which is used for communication.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The pipelined functional units perform arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. Cray Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
More information on the same at http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_3.html.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). The listings are available at http://www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches. Each node consisted of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It could run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions have streaming data and perform same operation on multiple elements.The functional units are deeply pipelined to exploit ILP (Instruction Level Parallelism). Multiple load/store units take advantage of this nature of inputs. Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at http://www.lib.ncsu.edu:2162/citation.cfm?id=956540&amp;amp;coll=portal&amp;amp;dl=ACM&amp;amp;CFID=28753767&amp;amp;CFTOKEN=45906295&lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
&lt;br /&gt;
Computational simulation is one of the areas which require supercomputing. The scientific challenges including understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions can only be done through supercomputers. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta. &lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
	<entry>
		<id>https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2859</id>
		<title>CSC/ECE 506 Fall 2007/wiki1 9 vr</title>
		<link rel="alternate" type="text/html" href="https://wiki.expertiza.ncsu.edu/index.php?title=CSC/ECE_506_Fall_2007/wiki1_9_vr&amp;diff=2859"/>
		<updated>2007-09-06T00:44:44Z</updated>

		<summary type="html">&lt;p&gt;Vvijaya: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Definition ==&lt;br /&gt;
Array processing is a CPU design concept which uses multiple interconnected processing elements to execute the same instruction on different data items. A control processor dispatches a single instruction stream to each of these processing elements containing a processor and a memory which execute the instruction. Communication between the processing elements is achieved by interconnecting the nodes.&lt;br /&gt;
&lt;br /&gt;
Vector Processing can be described an alternative model for exploiting ILP where multiple data elements contained in vector registers are processed using pipelined vector functional units. A vector register is a linear array of vector length n-bit data elements and the pipelined units perform arithmetic operations on the data elements in parallel.&lt;br /&gt;
&lt;br /&gt;
== History == &lt;br /&gt;
&lt;br /&gt;
Flynn’s Taxonomy classifies the computers based on the number of concurrent instructions and data streams available for execution. Array and Vector Processing come under a category called SIMD i.e. single instruction multiple data, which means multiple data streams are processed for the same instruction.&lt;br /&gt;
&lt;br /&gt;
During 1960’s, fear of performance stagnation pushed computer architects to look for smarter alternatives to increase the throughput.  Daniel Slotnick, a Professor from Computer Science Department of University of Illinois, proposed a conceptual SIMD machine called “SOLOMON”, with 1024 1-bit processing elements each having a memory for 128 32–bit values.  The machine was never built, but the design was a starting point for the advanced computer called ILLIAC-IV.  It had 64 processing elements each with memory capacity of 2,048 words of 64 bit length.  These elements communicated with each through an interconnect that resembled a ring. Each element was is provided a direct data path to four other elements, its immediate right and left neighbors and the neighbors spaced eight elements away. This interconnection structure is wrapped around, so that PE 63(Processing Element) is directly connected to PE 0.&lt;br /&gt;
&lt;br /&gt;
The first successful implementation of vector processing architecture was CDC STAR- 100, from Central Data Corporation.  It used vector registers to hold multiple data elements.  It had a high startup time and very relatively slow. Vector architectures exhibited SIMD behavior by having operations that applied to all elements in a vector register.&lt;br /&gt;
&lt;br /&gt;
These parallel computing architectures tried to exploit inherent data parallelism in programs.&lt;br /&gt;
&lt;br /&gt;
== Description ==&lt;br /&gt;
&lt;br /&gt;
An array processor, usually has multiple processing elements each capable of performing arithmetic/logical operations and store the result in its memory. Parallelism is achieved by operating on a stream of data rather than a single element. A control processor is responsible for fetching and broadcasting the instruction which will be executed by the PEs (Processing Elements). Array processing provides more performance than a serial computing.&lt;br /&gt;
&lt;br /&gt;
For example a computation that takes an array of elements and performs some operation on it would require a serial computer to process one element at a time. However an array processor does this by distributing the array elements among the PEs. Each PE may be assigned an element in an array or a set of rows. The instruction dispatched would be executed by the PE’s which communicate with each other. It is needless to say that this computing technique is suited for matrix multiplications and array operations which are extensively used in statistical analyses, numerical linear algebra, numerical solution of partial differential equations and digital signal processing calculations.&lt;br /&gt;
&lt;br /&gt;
More information on the same at &lt;br /&gt;
&lt;br /&gt;
Vector Processing as the name itself suggests uses VECTORS i.e. a series of values or elements than a SCALAR i.e. single value or an element. Vector Processors typically have &lt;br /&gt;
&lt;br /&gt;
•	Vector Registers&lt;br /&gt;
&lt;br /&gt;
•	Vector Functional units&lt;br /&gt;
&lt;br /&gt;
•	Scalar Units with registers and data paths,&lt;br /&gt;
&lt;br /&gt;
•	Vector Load Store Units &lt;br /&gt;
&lt;br /&gt;
•	Interconnect with is used for communication between them.  &lt;br /&gt;
&lt;br /&gt;
Each vector register is capable of holding multiple data elements of a definite width.  A typical system would have a number of such registers. The load/store units are responsible for fetching operands and writing the results into the memory. The functional units are pipelined and they perform the arithmetic and logical operations. All of these operate on a series of values which are either residing in the main memory or in the registers. CRAY Y-MP, a supercomputer built by Cray Inc used vector processing to increase the performance.&lt;br /&gt;
&lt;br /&gt;
Consider an example of multiplication of two arrays on a vector processor. The operands which are elements of the input arrays are loaded into two vector registers. The first elements of each of these vector registers are fed into a pipelined multiplication unit which performs the operation and stores the result in another vector register. All functional units are pipelined so that the overall execution time is low. The results are written back to the main memory using the load/store unit.&lt;br /&gt;
&lt;br /&gt;
== Importance and Trends ==&lt;br /&gt;
&lt;br /&gt;
National Energy Research Scientific Computing Center (NERSC) at Berkeley, California has collaborations with the computer and computational science departments for several universities. This organization also ranks the most powerful supercomputers in the world based on Rmax (a benchmark from Linpack). &lt;br /&gt;
The listings are available at www.top500.org.&lt;br /&gt;
&lt;br /&gt;
Horst D Simon, the Director of NERSC Center, presented on the trends in supercomputing in December 2003. As indicated in his presentations global climate modeling and earth simulators are few of the computationally intensive examples which need supercomputers. Earth Simulator, developed for Japan Aerospace Exploration Agency is a highly parallel vector supercomputer with 640 processor nodes connected by 640x640 single-stage crossbar switches.&lt;br /&gt;
Each node consists of 8 vector type arithmetic processor and 16 GB memory with a peak performance of 8Gflops per vector processor. It can run holistic simulations of atmosphere and oceans down to the resolution of 10 km.&lt;br /&gt;
&lt;br /&gt;
Apart from the domain of supercomputers, vector processors find application in multimedia processing which is computationally intensive and places large demands on portable devices. These functions are inherently data streaming and perform same operation on multiple elements. Multiple load/store units can take advantage of the nature of inputs. The functional units are deeply pipelined to exploit ILP 9instruction level parallelism.&lt;br /&gt;
Motorola has done a lot of research of applying these techniques for multimedia processing.&lt;br /&gt;
&lt;br /&gt;
More information can be obtained at &lt;br /&gt;
&lt;br /&gt;
== Future ==&lt;br /&gt;
Computational simulation is one of the areas which requires supercomputing. These scientific challenges include understanding, detecting and predicting the human influence on climate and modeling the full earth system including atmosphere, ocean, land and their interactions to name a few. &lt;br /&gt;
&lt;br /&gt;
Several companies like IBM, Cray Inc and SGI are doing pioneering research in supercomputing which continues to scale towards new heights.&lt;br /&gt;
&lt;br /&gt;
==  Bibliography  ==&lt;br /&gt;
&lt;br /&gt;
*Parallel Computer Architecture: A Hardware/Software Approach by David Culler, J.P. Singh, Anoop Gupta &lt;br /&gt;
&lt;br /&gt;
*http://www.llnl.gov/computing/tutorials/parallel_comp/&lt;br /&gt;
&lt;br /&gt;
*http://www.pcc.qub.ac.uk/tec/courses/cray/ohp/CRAY-slides_1.html&lt;br /&gt;
&lt;br /&gt;
*http://www.cs.berkeley.edu/~pattrsn/252S98/Lec06-vector.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www-ugrad.cs.colorado.edu/~csci4576/VectorArch/VectorArch.html&lt;br /&gt;
&lt;br /&gt;
*http://www.nersc.gov/~simon/Talks/Five_Trends.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.cray.com/downloads/science/climate_modeling_flyer.pdf&lt;br /&gt;
&lt;br /&gt;
*http://www.kuro5hin.org/story/2002/2/10/145957/917&lt;/div&gt;</summary>
		<author><name>Vvijaya</name></author>
	</entry>
</feed>