Developer (??) Information about the Alpha processor and Linux: 95/01/12 =============================================================== Info is taken from reading the Alpha specific source code in binutils 2.5.2, gcc 2.6.3 and glibc 1.09. Also Linus wrote a long reply to the first version of this text. (Long portions of this text are just his words.) This text is for people who already know about an assmbler language and now just need a short kick to learn the Alpha way of doing it... I also want to review the different processors of DEC, but not all their different bundles of complete computers... I wasn't sure, if writing this text is a good thing. But probably some people will take a quick look at this and learn more about the Alpha architecture. Not everybody wants to buy a book from DEC for that... Hardware information about the different Alpha processors: ========================================================== 21064 166/233 MHz SpecInt SpecFloat Cache 21066 Some definition of commonly used words: ======================================= reg = register General notes about (GNU) assembly language: ============================================ Read also the GNU info files from binutils 2.5.2 about the GNU assembler. I have uploaded further information on the ftp server mentioned at the end of this text. Linus now uses normal numbers as labels. E.g. "1:" at the beginning of a line marks that position with the label "1". You normally reference it with "1b" (last label named "1" backwards in the assembler text) or with "1f" (next label named "1" further down in the text; f=forward). General information about the Alpha processor: ============================================== You can't do unaligned load/stores at all. If you want to load a byte at address 0x12345, you need to load the quadword at 0x12340 and then extract byte #5 (or alternatively load a long-word at 0x12344 and extract byte #1). Only 32- and 64-bit load/stores are allowed. There are "load/store unaligned" instructions, but they will not actually do an unaligned load/store: they'll just ignore the low bits in the address and do an aligned load/store that way. No status register. When you do a compare instruction, the result goes into another normal register. And conditional branches look at normal integer registers like "bne $12,address" to jump to "address" if register $12 is non-zero. As you don't have a carry flag in a status register, you have to check by hand for an overflow/carry. The processor would also allow to generate exceptions on overflow, but that would only be useful for things like fortran which are supposed to complain on integer arithmetic. The reason for a lack of status register is simple: it makes it more likely that you can do things in parallell, when instructions don't all interfere with each others writing of the status reg.. - no locked read-modify-write memory cycles: they are bad for SMP, and aren't that general anyway. Instead the alpha has the load-locked + store-conditional that's also found in MIPS machines. You can essentially do anything between the load and the store (very nice), and then if the store fails when somebody else got there in between, you just repeat. This gives very good generality, plus no repeats or locks for the normal case. - move conditional ("cmove"). One of those nice things that not everybody has, and which makes many of those simple tests quicker without a jump. A few downsides: - no integer divide. It can easily be done in software, and I know why they didn't want it in hardware, but I still am not too happy about it. Everybody has some gripe. - the DEC standard assembly syntax is rather horrid. For example, there is no "compugt" instruction to compare if one reg is greater than another, as "compult" can be used instead by changing the register order. But you'd think that the assembler could do it for you... The Alpha processor has 64 registers. 32 floating point and 32 integer registers. Even if the type of cpu doesn't support floating point instructions, the floating point registers must still be present, as they may be used for 32 or 64 bit integer instructions. Here are some definitions about the various registers: - nonsaved = This register needn't be saved, if this register is used in a function. So if you call a function, this reg may be changed to anything on function return. - saved = You do not have to save this reg, if you make a call to another function. - unallocated = This reg is normally not used by gcc to store variables in it, since it probably is some kind of a special purpose reg. The integer registers are named $0 through $31. (DEC uses symbolic names depending on use of the different regs. Like e.g. "as" for the assembly temorary reg $28. doesn't just give the regs numbers, but (For gas, greater numbers can be used for labels. E.g. Linus uses $200 and onwards for labels...) $31 is a special reg. It always contains zero and can never be changed. Reg $16 through $21 are used to store parameters before calling a function and reg $0 is used to store the function return value. Reg $9 to $14 should be saved, if they are used during a function call. Several further registers are used for special purposes. Please look at the description of the assembler opcodes for further details. Reg $28 is also called assembly temporary ("as" by DEC naming convention) and some complex assembler instructions use it as temorary register. So do not store anything in it, if you do not have detailed knowledge about the other used assembler instructions. Gcc source mentions that the stack pointer should only be changed in multiple of 16 bytes. This is probably only a convention and not a hardware limitation. If you access the stack for quad-word stuff, you must use 8-byte aligned addresses. Otherwise, you get "unaligned exceptions". GNU cc will align stack usage (e.g. local variables) to multiplies of 16 bytes to (maybe) better use the cache. The 32 floating point registers are written as $f0 - $f31. $f31 is again a special reg and contains the constant zero. You are not allowed to load unaligned memory addresses into a fp reg. (Must be aligned to 32 bit.) A short overview about all 64 registers: 32 floting point registers, written as $f0 - $f31 0 return value of functions (nonsaved) 1 nonsaved 2-9 saved 10-15 nonsaved 16-21 input args to functions (nonsaved) 22-30 nonsaved 31 always 0 (constant reg) (unallocated), also called FP used as frame pointer by gcc, but will be replaced by the hardware frame pointer or the stack pointer 32 interger registers, written as $0 - $31, can hold any mode 0 return value (nonsaved) 1-8 nonsaved 9-14 saved 15 frame pointer 16-21 input args (nonsaved) 22-25 nonsaved 26 return adresse (ra) (saved) (jsr = jump service routine) 27 address of function call 28 nonsaved 29 global pointer (unallocated) used to get memory addresses in the C runtime system 30 stack pointer (unallocated) 31 always 0 (unallocated), also called AP used as argument pointer by gcc, but will be replaced by the stack pointer or the hardware frame pointer One feature of the Alpha processor is the ability to define new opcodes on the running system (called PAL codes - Programmable Assembler Language??). You have to load those definitions of new opcodes into the processor. They are then always available and needn't be loaded every time they are used. Those new opcodes are probably defined by a sequence of the already existing opcodes, but I haven't seen any such... Those newly available assembler sequences are always garantueed to by executed as a whole and not stopped by a hardware exception/signal. One possible usage for those PALs are the atomic bit operatioons in the kernel. Another common usage of PALs is for often generated code sequences generated by a C compiler. So for code reduction. Such a PAL can be also used like a very short function call, but will then need much less overhead for calling that function. Just define the assembler sequences and assign them to a new opcode. Then define the registers to be used as parameters, the regs that are clobbered during function call and the return value. (This can be done, using the special GNU assembler syntax within CC. So GNU CC knows which regs he has to save and which not. Code will be optimized - a bit like inline functions without copying the assembler sequences all the time. I will describe some assembler opcodes in the following text. Most often, they are taken from Linus already written code. The end of an assembler opcode often describes the type of parameters used. - 's' stands for signed (integers) - 'u' for unsigned - 'l' for long, a 32 bit integer - 'q' for quad, a 64 bit integer Just start reading the alpha version of head.S from the kernel and read my comments in the next paragraph... - "bis x,y,z" DEC language for "or" (BIt Set) If used with "bis $30,$30,$15", this will store reg $30 in reg $15. So it is some kind of a move command. If used like "bis $31,$31,$31", this is a nop - it won't do anything (no operation). Linus uses it also with "bis $30,1,$15" - dunno about that one... - "br $1, $200" branch to label $200 and save the PC (process counter = IP = instruction pointer) of the next instruction in reg $1. - ".long x" will just store the 32 bit value x at that location - "ldq $30,0($1)" will load the 64 bit int at address $1 into reg $30, which is used as stack pointer. ldq = load quad. - "lda $2,-8($1)" will load the address of "-8($1)" into $2. As the first two assembler instructions are both 4 bytes, this will load the address offset at which "__start" is currently loaded into memory, into reg $2 - "subq $3,$2,$6" will do C-style "$6 = $3 - $2" (or the other way round??) - "stq $5,0($3)" will store the quad at memory location "0($3)" - "bne $4,$201" will branch to label $201, if reg $4 is not zero (branch not zero) - "addq $1,6,$2" will compute $1+6 and store it into reg $2 - "jmp $31,($1),$203" will store the offset of the next instruction in $31 (so it is lost...). I duno why there are two further arguments. We just need one destination address?? - "br $27,$100" will branch (jump) to label $100, which is just the next instruction, and will save the offset of the otherwise next instruction in reg $27... (This opcode uses a relative address for the second parameter; we have already been moved to another location in memory.) - "ldgp $29,0($27)" is probably some kind of a special instruction reserved for kernel level code - has probably something got to do with memory offsets/addressing. gp = general pointer, is always in reg $29. - There is a special code sequence for doing function calls. "lda $27,fun_name" loads the address (memory offset) of function fun_name into reg $27. (load address = lda) "jsr $26,($27),start_kernel" will store the address of the next instruction in reg $26 (so it contains the return value). I dunno about the next two parameters. Normally we need only once the destination address. Maybe the third parameter isn't used at all? - "and $1,$2,$3" will do C-style "$3 = $2 & $1" - "stq_c" must do something like store the first parameter at the address given by the second parameter and return zero in reg $0, if that memory address has already changed (in the last 2 instructions?? can't be!!) - cmpbge - compare byte greater equal Linus has written (1.1.78) an assembler routines to be used by GNU gcc for dividing 2 integers (return div or qoutient for signed/unsigned integers). Looking at the exactly same-pupose functions included in glibc 1.09 for OSF/1, you can find in those assembler files ready to use "div"-instructions. Dunno what's going on here... Something else that strikes me: As of 1.1.78 in the file entry.S: "entInt" is the routine that is called, if an hardware interrupt occurs. It must save all registers before calling the C function that does the rest of the job. That assembler file subtracts 144 from the stack pointer and then stores all registers in that space. Reg $8 *and* reg $19 are both stored at the same location. That is probably still a very early test function and Linus for sure knows about this... It also seems, like not enough registers are saved here. (He probably added one reg and didn't bother to change the offset in the following lines. Lazy Linus... As long as he *eventually* will be coding it The Right Way (is this a trademark?). Small note about the forthcoming "64-bit Linux" version: long = 64 bits int = 32 bits short = 16 bits char = 8 bits I haven't taken too much time to write this small essay. As soon as I find out more about it, I will maybe rewrite it... If more then three people are interested to learn more, I will put an entry in my /etc/aliases for discussion about this. If more than 15 people are interested, I will set up a majordomo mailing list. (This is probably just for the next 2 or 3 weeks, until we have all necessary information together. Someone else will have to do a mailing list for "how does Linux run on my home alpha machine?"...) I will for sure print end results on the KERNEL channel or at least post an ftp site, where you could pick up updated information... Cross-Compiling for Alpha on a Linux box: ========================================= Compile and install binutils-2.5.2(.6): - ./configure --prefix=/usr --target=alpha-osf i486-linux - make "CFLAGS=-O2 -pipe" LDFLAGS=-s - make "CFLAGS=-O2 -pipe" LDFLAGS=-s install (as root) Compile and install gcc 2.6.3: - ./configure --prefix=/usr --local-prefix=/usr/alpha-osf/local \ --target=alpha-osf --host=i486-linux --with-gnu-as --with-gnu-ld \ --with-stabs - change in config/alpha/alpha.h: #undef SDB_DEBUGGING_INFO #undef MIPS_DEBUGGING_INFO - make "CFLAGS=-O2 -pipe" LDFLAGS=-s LANGUAGES=C - make "CFLAGS=-O2 -pipe" LDFLAGS=-s LANGUAGES=C install (as root) Binaries are on jurix.jura.uni-sb.de:/pub/linux/alpha (for libc 4.6.27). I'd really like to have that Alpha-Emulator running on my Linux box. The DEC person promised that one for the end of last year. Let's hope to see it soon. What about the Linux C library? The DEC-Person said, that someone has already done work on that. But what about a 64-bit version?? glibc 1.09 includes a DEC OSF/1 port, which is also 64 bit. (Is that one complete?) Further possibilities of information: ===================================== The "Alpha Architecture Handbook" from DEC explains the register set and stuff like that for an Alpha processor. DEC WWW server: http://www.dec.com or http://www.digital.com things that I make available: ftp://jurix.jura.uni-sb.de:/pub/linux/alpha/ (on that ftp site should also be the newest version of this text) The book about the Alpha processor... (What's it exactly called??) Feedback to: ============ I'd be happy to include some further information in here. So please feel free to email me. (Even speeling corrections are appreciated, since English is not my first language.) Florian La Roche florian@jurix.jura.uni-sb.de