Notes
Outline
Supercomputers(t)
Gordon Bell
Bay Area Research Center
Microsoft Corp.
http://research.microsoft.com/~gbell
Photos courtesy of
The Computer Museum History Center
Please only copy with credit!
http://www.computerhistory.org
Supercomputer
Largest computer at a given time
Technical use for science and engineering calculations
Large government defense, weather, aero laboratories are first buyers
Price is no object
Market size is 3-5
Growth in Computational Resources
Used for UK Weather Forecasting
10T •
1T •
100G •
10G •
1G •
100M •
10M •
1M •
100K •
10K •
1K •
100 •
10 •
What a difference 25 years and spending >10x more makes!
Harvard Mark I
aka IBM ASCC
I think there is a world market for maybe five computers.
The scientific market is still about that size… 3 computers
When scientific processing was 100% of the industry a good predictor
$3 Billion: 6 vendors, 7 architectures
DOE buys 3 very big ($100-$200 M) machines every 3-4 years
Supercomputer price (t)
Time $M structure example
1950 1 mainframes many...
1960 3 instruction //sm IBM / CDC
mainframe SMP
1970 10 pipelining 7600 / Cray 1
1980 30 vectors; SCI “Crays”
1990 250 MIMDs: mC, SMP, DSM “Crays”/MPP
2000 1,000 ASCI, COTS MPP Grid, Legion
Supercomputing:
speed at any price, using parallelism
Intra processor
Memory overlap & instruction lookahead
Functional parallelism (2-4)
Pipelining (10)
SIMD ala ILLIAC 2d array of 64 pe vs vectors
Wide instruction word (2-4)
MTA (10-20)
MIMDs… processor replication
SMP (4-64)
Distributed Shared Memory SMPs 100
MIMD… computer replication
Multicomputers aka MPP aka clusters (10K)
Grid: 100K
High performance architectures timeline
1950 . 1960 . 1970 . 1980 . 1990 . 2000
Vtubes Trans. MSI(mini) Micro RISC nMicr
Processor overlap, lookahead “killer micros”
Cray era 6600 7600 Cray1 X Y C  T
Vector-----SMP---------------->
SMP mainframes---> “multis”----------->
DSM KSR   SGI---->
Clusters Tandm VAX IBM     UNIX->
MPP if n>1000   Ncube  Intel IBM->
Networks n>10,000 NOW Grid
High performance architectures timeline
1950 . 1960 . 1970 . 1980 . 1990 . 2000
Vtubes Trans. MSI(mini) Micro RISC nMicr
Sequential programming---->------------------------------
<SIMD Vector--//---------------
  Parallelization---
Parallel programming     <---------------
multicomputers     <--MPP era------
ultracomputers 10X in price 10xMPP
“in situ” resources 100x in //sm    NOW VLC
Grid
Time line of  hpcc contributions
Time line of hpcc contributions
Lehmer  UC/Berkeley pre-computer number sieves
Eniac c1946
Manchester: the first computer. Baby, Mark I, and Atlas
von Neumann computers: Rand Johniac
Gene Amdahl’s Dissertation and first computer
IBM
IBM Stretch c1961 & 360/91 c1965
consoles!
IBM Terabit Photodigital Store c1967
STC Terabytes of storage c1999
Amdahl aka Fujitsu version of the 360 c1975
IBM ASCI Red @ LLNL
CDC, ETA, Cray Research, Cray Computer
Cray
1925
-1996
Circuits and Packaging,
Plumbing (bits and atoms) & Parallelism… plus Programming and Problems
Packaging, including heat removal
High level bit plumbing… getting the bits from I/O, into memory through a processor and back to memory and to I/O
Parallelism
Programming: O/S and compiler
Problems being solved
Seymour Cray Computers
1951: ERA 1103 control circuits
1957: Sperry Rand NTDS; to CDC
1959: Little Character to test transistor ckts
1960: CDC 1604 (3600, 3800) & 160/160A
1964: CDC 6600 (6xxx series)
1969: CDC 7600
Cray Research, Cray Computer Corp. and SRC Computer Corp.
1976: Cray 1...
(1/M, 1/S, XMP, YMP, C90, T90)
1985: Cray Computer
Cray 2 from Cray Research;
GaAs: Cray 3 (1993), Cray 4
1999: SRC Company large scale,
shared memory multiprocessor using x86 microprocessors
Cray contributions…
Creative and productive during his entire career 1951-1996.
Creator and un-disputed designer of supers from c1960 1604 to Cray 1, 1s, 1m c1977… basis for
SMPvector: XMP, YMP, T90, C90, 2, 3
Circuits, packaging, and cooling…
“the mini” as a peripheral computer
Use I/O computers versus I/O processors
Use the main processor and interrupt it for I/O versus I/O processors aka IBM Channels
Cray Contributions
Multi-theaded processor (6600 PPUs)
CDC 6600 functional parallelism leading to RISC… software control
Pipelining in the 7600 leading to...
Use of vector registers: adopted by 10+ companies.
Mainstream for technical computing
Established the template for vector supercomputer architecture
SRC Company use of x86 micro in 1986 that could lead to largest, smP?
“Cray” Clock speed (Mhz),
no. of processors, peak power (Mflops)
CDC 1604 & 6600
CDC 7600: pipelining
CDC 8600
Prototype:
SMP, scalar,
discrete circuits,
 
failed to achieve clock speed
CDC STAR… ETA10
CDC 7600 & Cray 1 at Livermore
Cray 1 #6 from LLNL.
Located at The Computer Museum History Center, Moffett Field
Cray 1 150 Kw. MG set & heat exchanger
Cray XMP/4
Proc.
c1984
Cray 2 from NERSC/LBL
Cray 3 c1995 processor
500 MHz
32 modules 1K GaAs
  ic’s/module
8 proc.
c1970: Beginning the search for parallelism
SIMDs
Illiac IV
CDC Star
Cray 1
Iliac IV: first SIMD c 1970s
SCI
(Strategic Computing Initiative)
funded by DARPA and aimed at a Teraflops!

Era of State computers and many efforts to build high speed computers… lead to HPCC

Thinking Machines, Intel supers,
Cray T3 series
Minisupercomputers: a market whose time never came. 
Alliant,  Convex,
Ardent+Stellar= Stardent = 0,
Cydrome and Multiflow:
prelude to wide word parallelism in Merced
Minisupers with VLIW attack the market
Like the minisupers, they are repelled
It’s software, software, and software
Was it a basically good idea that will now work as Merced?
MasPar...
A less costly, CM 1/2 done in silicon chips
It is repelled.
S is the fatal flaw
Thinking Machines:
Thinking Machines: CM1 & CM5 c1983-1993
In Dec. 1995 computers with 1,000 processors will do most of the  scientific processing.
Danny Hillis
1990 (1 paper or 1 company)
The Bell-Hillis Bet
Massive Parallelism in 1995
Bell-Hillis Bet: wasn’t paid off!
My goal was not necessarily to just win the bet!
Hennessey and Patterson were to evaluate what was really happening…
Wanted to understand degree of MPP progress and programmability
KSR 1:
first commercial DSM
NUMA
(non-uniform memory access) aka
COMA
(cache-only memory architecture)
SCI (c1980s):
Strategic Computing Initiative funded
ATT/Columbia (Non Von), BBN Labs,
Bell Labs/Columbia (DADO),
CMU Warp (GE & Honeywell),
CMU (Production Systems), Encore, ESL, GE (like connection machine), Georgia Tech, Hughes (dataflow), IBM (RP3), MIT/Harris, MIT/Motorola (Dataflow), MIT Lincoln Labs, Princeton (MMMP), Schlumberger (FAIM-1), SDC/Burroughs, SRI (Eazyflow),
University of Texas,
Thinking Machines (Connection Machine),
Those who gave their lives in the search for parallellism
Alliant, American Supercomputer, Ametek, AMT, Astronautics, BBN Supercomputer, Biin, CDC, Chen Systems, CHOPP, Cogent, Convex (now HP), Culler, Cray Computers, Cydrome, Dennelcor, Elexsi, ETA, E & S Supercomputers, Flexible, Floating Point Systems, Gould/SEL, IPM, Key, KSR, MasPar, Multiflow, Myrias, Ncube, Pixar, Prisma, SAXPY, SCS, SDSA, Supertek (now Cray), Suprenum, Stardent (Ardent+Stellar), Supercomputer Systems Inc., Synapse, Thinking Machines, Vitec, Vitesse, Wavetracer.
NCSA Cluster of 8 x 128 processors SGI Origin c1999
Humble
beginning:
In 1981…
would you have predicted this would be the basis of supers?
Intel’s ipsc 1 & Touchstone Delta
Intel Sandia Cluster 9K PII: 1.8 TF
GB with NT, Compaq, HP cluster
The Alliance LES NT Supercluster
Our Tax Dollars At Work
ASCI for Stockpile Stewardship
Intel/Sandia: 
9000x1 node Ppro
LLNL/IBM:
512x8 PowerPC (SP2)
LANL/Cray:
6144 CPUs
Maui Supercomputer Center
512x1 SP2
ASCI Blue Mountain 3.1
Tflops SGI Origin 2000
12,000 sq. ft. of floor space
1.6 MWatts of power
530 tons of cooling
384 cabinets to house 6144 CPU’s with 1536 GB (32GB / 128 CPUs)
48 cabinets for metarouters
96 cabinets for 76 TB of raid disks
36 x HIPPI-800 switch Cluster Interconnect
9 cabinets for 36 HIPPI switches
 about 348 miles of fiber cable
Half of SGI ASCI Computer at LASL c1999
LASL ASCI Cluster Interconnect
LASL ASCI Cluster Interconnect
3 TeraOps makes a difference!
LLNL Architecture
I/O Hardware Architecture
Fujitsu VPP5000 multicomputer:
(not available in the U.S.)
Computing nodes
speed: 9.6 Gflops vector, 1.2 Gflops scalar
primary memory: 4-16 GB
memory bandwidth: 76 GB/s (9.6 x 64 Gb/s)
inter-processor comm: 1.6 GB/s non-blocking
with global addressing among all nodes
I/O: 3 GB/s to scsi, hippi, gigabit ethernet, etc.
1-128 computers deliver 1.22 Tflops
NEC SX 5: clustered SMPv
(not available in the U.S.)
SMPv computing nodes
4 - 8 processors/computer
Processor pap: 8 Gflops
Memory
I/O speed
Cluster
NEC Supers
High Performance COTS
Raceway and (RACE++) Busses
ANSI Standardized
Mapped Memory, Message Passing, ‘Planned Direct’ Transfers
Circuit Switched;  Basic Bus Interface Unit Is a 6 (8) Port Bidirectional Switch at 40MB/s (66MB/s) Per Port.
Scales to » 4000 Processors
Skychannel
ANSI Standardized
320mb/sec; Crossbar backplane supports up to 1.6 GB/s Throughput Non-blocking
Heart of Air Force $3M /  256 Gflops System
Mercury & Sky Computers - & $
Rugged System With 10 Modules ~ $100K; $1K /#
Scalable to several K processors; ~1-10 Gflop / Ft3
10 9U Boards * 4 Ppc750’s » 440 Specfp95 in
1 Ft3 (18.5 * 8 * 10.75”)
Sky 384 Signal Processor, #20 on ‘Top 500’, $3M
Brookhaven/Columbia QCD c1999
(1999 Bell Prize for performance/$)
Brookhaven/Columbia QCD board
HT-MT: What’s 0.55? c1999
HT-MT…
Mechanical: cooling and signals
Chips: design tools, fabrication
Chips: memory, PIM
Architecture: mta on steroids
Storage material
HTMT challenges the heuristics for a successful computer
Mead 11 year rule: time between lab appearance and commercial use
Requires >2 break throughs
Team’s first computer or super
It’s government funded…
albeit at a university