From Linus.Torvalds@cs.Helsinki.FI  Wed Jan 11 11:20:19 1995
Return-Path: Linus.Torvalds@cs.Helsinki.FI
Received: from keos.Helsinki.FI (keos.Helsinki.FI [128.214.4.83]) by jurix.jura.uni-sb.de (8.6.9/8.6.9) with ESMTP id LAA03237 for <florian@jurix.jura.uni-sb.de>; Wed, 11 Jan 1995 11:20:19 +0100
Received: (torvalds@localhost) by keos.Helsinki.FI (8.6.9/H46) id LAA29620; Wed, 11 Jan 1995 11:59:46 +0200
Date: Wed, 11 Jan 1995 11:59:46 +0200
From: Linus Torvalds <Linus.Torvalds@cs.Helsinki.FI>
Message-Id: <199501110959.LAA29620@keos.Helsinki.FI>
In-Reply-To: "Linux Activists"'s message as of Jan 11,  8:17
X-Mn-Key: KERNEL
X-Mailer: Mail User's Shell (7.2.0 10/31/90)
To: "Linux-Activists" <linux-activists@Niksula.hut.fi>,
        linux-alpha@vger.rutgers.edu,
        Florian La Roche <florian@jurix.jura.uni-sb.de>
Subject: Re: Alpha opcodes
Status: RO

Florian La Roche <florian@jurix.jura.uni-sb.de>: "Alpha opcodes":
> 
> (Is niksula still the correct address for KERNEL submissions??)

It works, but..  I'd suggest using the "vger.rutgers.edu" mailing lists
set up by davem: this one is going to the linux-alpha list as well. 

> This text contains some information about the Alpha processor.
> All that info is from reading the Alpha specific part of
> binutils-2.5.2, gcc 2.6.2 and glibc 1.0.9.

DEC has a few nice alpha manuals for general alpha information: the
"Alpha Architecture Handbook" being a small and thin thing which
explains the register set and stuff like that. 

> The integer registers are numbered $0 - $31. (For gas, greater numbers can be
> used for labels. E.g. Linus uses $200 and onwards for labels...)

The register naming schemes are rather strange.  DEC uses its own
naming, giving symbolic names depending on use (so "as" is register 28,
short for "assembly temporary" etc).  The $xx naming is ugly, in my
opinion, and I'd much rather see "%rxx" or something like that instead. 

As to labels: to avoid confusion I think I'll stop using the $XXX
scheme, that one was what I first saw in code from DEC, and it only
helps make the assembly even more confusing.  I'm using numbered labels
(wiohout the "$") in the division routines now, for example. 

> $31 is a special reg. It always contains zero and can never be changed.
> Reg $16 through $21 are used to store parameters before calling a function
> and reg $0 is used to store the function return value.
> Reg $9 to $14 should be saved, if they are used during a function call.
> Several further registers are used for special purposes. Please look at
> the description of the assembler opcodes for further details.
> Gcc source mentions that the stack pointer should only be changed in multiple
> of 16 bytes. This is probably only a convention and not a hardware limitation.

8-byte multiples are a hardware limitation, otherwise you'll get
unaligned exceptions when trying to use the stack for quad-word stuff. 
The 16-byte alignment is probably just to make it better suited to the
cache. 

> The 32 floating point registers are written as $f0 - $f31. $f31 is again
> a special reg and contains the constant zero.
> You are not allowed to load unaligned memory addresses into a fp reg.
> (Must be aligned to 32 bit.)

You can't do unaligned load/stores at all.  If you want to load a byte
at address 0x12345, you need to load the quadword at 0x12340 and then
extract byte #5 (or alternatively load a long-word at 0x12344 and
extract byte #1).  Only 32- and 64-bit load/stores are allowed. 

There are "load/store unaligned" instructions, but they will not
actually do an unaligned load/store: they'll just ignore the low bits in
the address and do an aligned load/store that way. 

All in all: interesting.  The code for loading a byte can be rather
complex. 

Anyway, the above may make the alpha sound terrible to program.  It's
not.  It's a very nice chip, and a very clean architecture, but what it
*is* is different.  Some high-ligts:

 - no status register.  None at all.  When you do a compare instruction,
   the result of the compare goes into another normal register.

 - conditional branches don't look at the status register, as that
   doesn't exist.  Instead, they look at a normal integer register, so
   you can say:

	bne $12,address

   to jump to "address" if register $12 is non-zero.

 - there is no "carry".  Again, with no status register, where would you
   put it? You have to check by hand if an overflow if you want to
   carry.  If you want an exception, that's ok too (but that's probably
   useful only for things like FORTRAM which are supposed to complain on
   integer arithmetic). 

The reason for a lack of status register is simple: it makes it more
likely that you can do things in parallell, when instructions don't all
interfere with each others writing of the status reg.. 

 - no locked read-modify-write memory cycles: they are bad for SMP, and
   aren't that general anyway.  Instead the alpha has the load-locked +
   store-conditional that's also found in MIPS machines.  You can
   essentially do anything between the load and the store (very nice),
   and then if the store fails when somebody else got there in between,
   you just repeat.  This gives very good generality, plus no repeats or
   locks for the normal case. 

 - move conditional ("cmove").  One of those nice things that not
   everybody has, and which makes many of those simple tests quicker
   without a jump. 

A few downsides:

 - no integer divide.  It can easily be done in software, and I know why
   they didn't want it in hardware, but I still am not too happy about
   it.  Everybody has some gripe. 

 - the DEC standard assembly syntax is rather horrid.  For example,
   there is no "compugt" instruction to compare if one reg is greater
   than another, as "compult" can be used instead by changing the
   register order.  But you'd think that the assembler could do it for
   you...

> A short overview about all 64 registers:
> 32 floting point registers, written as $f0 - $f31
> 	0	return value of functions (nonsaved)
> 	1	nonsaved
> 	2-9	saved
> 	10-15	nonsaved
> 	16-21	input args to functions (nonsaved)
> 	22-30	nonsaved
> 	31	always 0 (constant reg) (unallocated), also called FP
> 		used as frame pointer by gcc, but will be
> 		replaced by the hardware frame pointer or the stack pointer
> 32 interger registers, written as $0 - $31, can hold any mode
> 	0	return value (nonsaved)
> 	1-8	nonsaved
> 	9-14	saved, (not saved for int handler in Linus code ???)

I don't save them in my current handler, just because C functions will
never modify them (or rather, they'll save-restore them if they do).

 
>	15 frame pointer  
>	16-21 input args (nonsaved)  
>	22-25 nonsaved  
>	26 return PC for jsr (saved) (jsr = jump service routine)  
>	27 address of function call  
>	28 nonsaved  

28 is "assembly temporary": the assembler will use this without warning
you to create a few of the "complex" instructions.  Don't use unless you
know what you do (but see divide.S in the next release: "don't do this
at home without a parent" ;-)

> 29 global pointer (unallocated)  
> 30 stack pointer (unallocated)  
> 31 always 0 (unallocated), also called AP  

Note that none of the register numbers are set by hardware (except 31,
which is zero in hardware).  The "return address" register aka "ra" aka
register 26, for example, is nothing but a convention, and you can use
your own registers for your own subroutine calls if you want to. 

The "gp" aka "global pointer", ie 29, is used to get memory addresses,
and again is just a convention (although one assumed by the C runtime
system).  Similarly the stack/frame pointer. 

> One feature of the Alpha processor is the ability to define new opcodes  
> on the running system (called PAL codes - Programmable Assembler Language??).   
> You have to load those definitions of new opcodes into the processor.  They are  
> then always available and needn't be loaded every time they are used.   
> Those new opcodes are probably defined by a sequence of the already existing  
> opcodes, but I haven't seen any such...   

PAL-code also has special privileges and some special processor-specific
instructions, but you aren't supposed to know about them. 

> Those newly available assembler sequences are always garantueed to by executed  
> as a whole and not stopped by a hardware exception/signal.   
> One possible usage for those PALs are the atomic bit operatioons in the  
> kernel.   
> Another common usage of PALs is for often generated code sequences generated  
> by a C compiler.  So for code reduction.   

Not really supported now: I won't write my own PAL-code until much
later.  Essentially the PAL-code only allows to do some things that need
special privs (return-from-interrupt, page tabe management, low-level
interrupt stuff). 

> Such a PAL can be also used like a very short function call, but will then  
> need much less overhead for calling that function.   

Actually, the overhead for a PAL call is more than a normal function
call: before executing the PALcode, the processor will have to do an
implicit trap barrier and set up various internal states, so it's only
used for things that really need to be atomic with respect from
interrupts etc. 

> Just define the assembler sequences and assign them to a new opcode.   
> Then define the registers to be used as parameters, the regs that are  
> clobbered during function call and the return value.   
> (This can be done, using the special GNU assembler syntax within CC.  So GNU  
> CC knows which regs he has to save and which not.  Code will be optimized -  
> a bit like inline functions without copying the assembler sequences all the  
> time.   

I think the current gcc inline assembly doesn't allow you to specify
which registers you want as inputs/outputs.  Major bummer.  You have to
move the regs around yourself. 

> Just start reading the alpha version of head.S from the kernel and read  
> my comments in the next paragraph...   
> - "bis x,y,z"  

"bis" is DEC-speak for "or".  Don't ask me why.  It's something like
"BIt Set". 

> If used with "bis $30,$30,$15", this will store reg $30 in reg $15.  So it  
> is some kind of a move command.  If used like "bis $31,$31,$31", this is  
> a nop - it won't do anything (no operation).   
> Linus uses it also with "bis $30,1,$15" - dunno about that one...   

There is no direct register move command: to move a register to another
one, you usually "or" itself to itself and tell that the results goes to
the other register.  "bis $31,1,$15" will make reg 15 contain 1: it does
an "or" of the zero-register and the constant 1.  You can do it other
ways too, but that's one of them. 

The logical operations are: "and" and "xor" (surprisingly called
something logical), the afore-mentioned "bis" (aka "or"), and the forms
that do the same operations on negated values: "bic" is an and-not ("BIt
Clear"), "ornot" is or-not (don't ask me why they have "bis" but then
"ornot"), and "eqv" is xor-not. 

To do a simple "not", you do either: "ornot $31,src,dst" or "eqv
src,$31,dst". 

> - "br $1, $200" branch to label $200 and save the PC (process counter = IP =  
> instruction pointer) of the next instruction in reg $1.   
> - ".long x" will just store the 32 bit value x at that location  
> - "ldq $30,0($1)" will load the 64 bit int at address $1 into reg $30, which is  
> used as stack pointer.  ldq = load quad.   
> - "lda $2,-8($1)" will load the address of "-8($1)" into $2.  As the first two  
> assembler instructions are both 4 bytes, this will load the address offset  
> at which "__start" is currently loaded into memory, into reg $2  
> - "subq $3,$2,$6" will do C-style "$6 = $3 - $2" (or the other way round??)  
> - "stq $5,0($3)" will store the quad at memory location "0($3)"  
> - "bne $4,$201" will branch to label $201, if reg $4 is not zero (branch not  
> zero)  
> - "addq $1,6,$2" will compute $1+6 and store it into reg $2  
> - "jmp $31,($1),$203" will store the offset of the next instruction in $31  
> (so it is lost...).  I duno why there are two further arguments.   
> We just need one destination address??  

($1) is the real destination address, ie jump to register $1, and save
the old IP in $31 (ie throw it away).  The $203 is just an instruction
hint for the instruction pre-fetch (so that it can possibly pre-fetch
even before having the $1 register calculated). 

Don't ask.

> - "br $27,$100" will branch (jump) to label $100, which is just the next  
> instruction, and will save the offset of the otherwise next instruction in  
> reg $27...  (This opcode uses a relative address for the second parameter;  
> we have already been moved to another location in memory.)  
> - "ldgp $29,0($27)" is probably some kind of a special instruction reserved  
> for kernel level code - has probably something got to do with memory  
> offsets/addressing.  gp = general pointer, is always in reg $29.   

This just primes the general pointer - magic.  It's one of the composite
assembly instructions, and the assembler and linker will do the right
thing to make sure the gp points to the base of program memory. 

> - There is a special code sequence for doing function calls.   
> "lda $27,fun_name" loads the address (memory offset) of function fun_name  
> into reg $27.  (load address = lda)  
> "jsr $26,($27),start_kernel" will store the address of the next instruction  
> in reg $26 (so it contains the return value).  I dunno about the next two  
> parameters.  Normally we need only once the destination address.   
> Maybe the third parameter isn't used at all?  

The third parameter is again the prefetch hint. 

> - "and $1,$2,$3" will do C-style "$3 = $2 & $1"  
> - "stq_c" must do something like store the first parameter at the address  
> given by the second parameter and return zero in reg $0, if that memory  
> address has already changed (in the last 2 instructions?? can't be!!)  

Yup.  That's the "store conditional".  This needs hardware support (and
has it), but is very practical.  It usually goes through normally, but
just in case an interrupt happened and we stored something else, the
store will fail, and we'll just try again all over. 

> - cmpbge - compare byte greater equal  

Does a 8-byte compare in one go.  The use in the bitops isn't
necessarily completely obvious, but it was fun doing it that way. 

> Linus has written (1.1.78) an assembler routines to be used by GNU gcc for  
> dividing 2 integers (return div or qoutient for signed/unsigned integers).   
> Looking at the exactly same-pupose functions included in glibc 1.09 for  
> OSF/1, you can find in those assembler files ready to use "div"-instructions.   
> Dunno what's going on here...   

The "div" instruction exists in assembly language, and gcc will emit
them, but the assembler will change them into a function call.  The
functions normally exist in the DEC C library.  I don't know the
copyright status of those (even though I have copies of sources), and
besides, the DEC routines use a lookup table that makes the division
routines take about 6kB of memory.  My routines are slightly slower for
some cases (faster for others, but I'm afraid the DEC ones are better in
the "usual case") but are much smaller and don't punish the cache as
much. 

> Something else that strikes me: As of 1.1.78 in the file entry.S:  
> "entInt" is the routine that is called, if an hardware interrupt occurs.   
> It must save all registers before calling the C function that does the rest  
> of the job.  That assembler file subtracts 144 from the stack pointer and  
> then stores all registers in that space.  Reg $8 *and* reg $19 are both  
> stored at the same location.   

Bug.  That stuff is by no means the final stack layout anyway (it
disagrees with my "ptrace.h" ;-)

> If more then three people are interested to learn more, I will put an entry  
> in my /etc/aliases for discussion about this.  If more than 15 people are  
> interested, I will set up a majordomo mailing list.   

Already done:

	...

	Again, the lists are all majordomo like the rest of the new
	linux-activists lists. So a simple way to join is:

	echo "subscribe linux-alpha" | mail majordomo@vger.rutgers.edu

	just replace linux-alpha with the list you want to join. If you would
	like a simple help file explaining how to use majordomo just send
	...

Anyway, while I'm about I'll just give a short status-report on my alpha
(Jim probably is much further along, but I haven't seen his code):

 - mainly "general portability" stuff done: you've seen the source tree
   layout changes.  The kernel compiles on my alpha except for the
   memory management. 

 - I have another source tree that actually boots up and does hardware
   tests (a small console driver etc).  I'm working on getting the code
   from that one into the standard kernel and have a bootable kernel
   this week (note that "booting" and "working" are completely separate
   things, although the booting step is rather important). 

 - currently it's rather EISA AlphaPC-specific: the PCI stuff needs to
   change mainly the io.h include and possibly memory layout slightly. 

I'll try to remember to send in reports here every once in a while,

		Linus

