
This is the third public release of MD, which stands for Multiple
Devices. Its main goal is to group several disks or partitions
together, making them look like a single block device.  Furthermore,
it is interface independent, so it is possible to mix IDE
(MFM/RLL/ESDI/AT-BUS), SCSI, and even old XT-like disks.

**WARNING** : This is some **VERY ALPHA** software. Don't use it
unless you exactly know what you're doing. This patch works for me,
and I trust my data to it, but I do very frequent backups. So, if this
damned thing goes wild, I won't loose a lot. Please backup your own
system BEFORE using this software.

--

Theory of operations :
	The kernel patch included in the package create a new block
device driver called "md" (device major=9, which is now the official
and registered major number). The main difference between this driver
and the other ones like "hd" and "sd" (among others...) is that "md"
never access the disks itself. The trick is to find the real device
when the kernel tries to queue the requests (in ll_rw_block.c). With a
little help from a hash table (a friend of mine...), you shouldn't
notice too much difference about speed.

	There's currently two ways to manage such devices. The first
is called 'linear', which means real devices are appended to each
other. This kind of device is easily expandable (see 'Future
extensions' section), but gives no or very small speed
improvement. The second is called 'stripped', which uses 'stacked'
devices and spread contiguous blocks across those devices.

--

Using it :
	Once the kernel is patched and running, you can use a small
command set to manage your md-devices :
- mdadd md-dev block-dev1 block-dev2 ... : add blocks devices to
md-dev.
- mdrun md-dev : make the md device useable as a block device.
- mdstop md-dev : stop the device (if started) and cancel the group.

For example, that's what I used to have in my /etc/rc :
	/sbin/mdadd /dev/md0 /dev/sdb1 /dev/sdc2
	/sbin/mdrun -l /dev/md0

It groups /dev/sdb1 and /dev/sdc2 in a single device called /dev/md0,
and starts it in linear mode. Note that it could also be written :
        /sbin/mdadd /dev/md0 /dev/sdb1
        /sbin/mdadd /dev/md0 /dev/sdc2
        /sbin/mdrun -l /dev/md0

BE CAREFUL ! :
        /sbin/mdadd /dev/md0 /dev/sdb1 /dev/sdc2
is NOT equivalent to
        /sbin/mdadd /dev/md0 /dev/sdc2 /dev/sdb1

It produces a device that have the same size, but that will be very
different, specially if you have already put a filesystem on it. So
once you've configured a md-device with an arbitrary order, ALWAYS USE
THAT ORDER, or you won't retreive your data. You been warned !!

To create a 'stripped' device, you should have used
        /sbin/mdrun -s0 /dev/md0
which runs the device in stripped mode, with a 0 factor. The factor
indicate the size of a chunk on a real device, according to the
following formula :
	chunk_size = PAGE_SIZE << factor

So, on a 386, a 0 factor indicate a chunk size of 4096 bytes, a 1
factor indicate a chunk size of 8192 bytes, a 2 factor indicate a
chunk size of 16384 bytes...

It is also a good idea to create a /etc/mtab file that contains
entries to run mdadd on. For exemple, here is my own /etc/mdtab :

# mdtab for wild-wind #1 & #2

/dev/md0	stripped,0 /dev/sdb1	/dev/sdc2	# /usr/local
/dev/md1        linear	   /dev/hda6	/dev/hda7	# /mnt for swapfile

So I can simply do
	/sbin/mdadd -a
	/sbin/mdrun -a

You can even use the '-r' flag to automagically start the device once
it has been successfully completed. In this case, the above becomes :
	/sbin/mdrun -ar

To use a particular device, use the syntax
	/sbin/mdadd [-r] md_device

mdstop is not very useful, but it's a good way of being sure that the
md-device has been sync-ed.

You can see the exact state of your md-devices by looking at
/proc/mdstat which looks like this on a test machine running in linear
mode :
	$ cat /proc/mdstat
	md0 : active linear hda6 hda7 10360 blocks
	      [hda6] [hda6/hda7] [hda7]
	md1 : inactive
	md2 : inactive
	md3 : inactive

First, you have the state of each device the kernel is configured for.
Here, /dev/md0 is active and /dev/md[1-3] are stopped. Then you see
the devices used for the current md-device and the total size in 1024
bytes blocks. The second line shows the hash table state, which is
only printed for debugging purpose, and may go away in a near future
(at least I hope so ;-).

Now, the same machine running in stripped mode with a 0 factor :

$ cat /proc/mdstat
md0 : active stripped(0) hda6 hda7 10360 blocks
      [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0]
[z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0] [z0/z1] [z1]
      z0=[hda6/hda7] zo=0 do=0 s=9976
      z1=[hda6] zo=9976 do=4988 s=384
md1 : inactive
md2 : inactive
md3 : inactive

You can now do whatever you want with an active md-device (create a
filesystem, a swapfile, or even a swap partition, but the last one
isn't really useful, since the kernel already handles multiple swap
partitions).

--

Install :
	In the package, you should have found some files :

- README   : this file,
- Changelog: history of the project,
- mdpatch  : patch to kernel against version 1.1.92 (works from 1.1.83)
- mdadd.c  : source for mdadd, mdrun and mdstop,
- mdparse.c: parser for mdadd & co
- mdtab	   : Example for /etc/mdtab
- mdadd.8  : man page for mdadd, mdstop and mdrun,
- md.lsm   : lsm entry for md,
- Makefile : guess what...

As you're used to, do a
		$ patch -p0 -s < wherever_it_lives/mdpatch
from the directory that contains your linux tree (supposed to be clean...).
You can now compile your brand new (bugged ? ;-) kernel (don't forget
to do a 'make config' first and to answer 'yes' to the question about
md-device).

In the meantime (Yes sir, that's what I call multitasking !), get back
to the package, edit Makefile to suit your taste, and do a
		$ make
		$ make install
(the latest may require from you to be logged as root).

You shouldn't get too much warnings during both of the
compilations. It also creates entries in /dev called md[0-3]. If you
need some more, add them and change the MAX_MD_DEV constant in
md.h. Also, if you want more than 8 real devices per md-device, adjust
the MAX_REAL constant to suit your needs.

--

Bugs :
	There is at least one known bug with mkdosfs, which asks the
device about physical geometry. This is, of course, not relevant with
a md-dev. So the max size of a DOSFS on such a device is sometimes
limited to a fraction of the avaible size. If there's a DOSFS guru out
there that can help me, please have a look at the HDIO_GETGEO ioctl in
md.c. But who will use DOSFS on such a device anyway ?

	The code heavily depends on 1024 bytes blocks and 512 bytes
sectors. I hope to change it soon.

	You CANNOT use floppies as real devices for md, since they are
not managed the same way hard-disks are. Maybe some day, I'll dig into
that, but don't be too sure, since it has a rather low priority on my
'Things-to-do-one-day-when-I-have-time-to-give-a-look' list ;-).

	In stripped mode, chunk size depends on hardware page size.
So a stripped device with a 0 factor on an i386 (PAGE_SIZE=4k) cannot
be used on an Alpha (PAGE_SIZE=8k).

	If you find a bug, report as fast a possible. Please include :
- the kernel version you're running,
- the kind of real devices you're using,
- a copy of your partition tables (using fdisk),
- a copy of your /proc/mdtab,
- if the bug results in a crash, kernel info as described in
/usr/src/linux/README, since this will help a lot
- any application message that shows that there's a problem.

	This file itself is a known bug ;-).

--

Future extensions :
	Well, what I'd like to have now is a way to extend an existing
file system by just adding a new device to a md-dev running in linear
mode (of course, no data loss !!). If Remy Card read this, I'd be glad
to have his own opinion about it (Yes, I can read french ;-).

	Some people asked me about RAID support. I'm thinking hard
about it, but don't have much time yet. Anyway, stripped personality
(which is in fact RAID-0) is a first step in that direction. The main
problem is to generate multiple requests for a single buffer, since a
block is supposed to be an atomic entity. So real RAID (say, 1 and up)
would introduce major changes in the kernel, and I'm not yet ready to
handle such a task alone.

Also feel free to send me any idea about what you'd like to see in a
future version. Any positive feedback would be very nice too !

--

Thanks to :
	Linus Torvalds and others for writing such a beast,
	pc, jpl, arthur & jmm for everything, including missed train ;-)
	BB for understanding (sometimes).

--

Please send any idea, opinion or bug report to
<zyngier@amertume.ufr-info-p7.ibp.fr> (preferred) or to
<maz@gloups.fdn.fr>

Please don't expect any immediate answer, as I can only access my mail
account once a week :-(, but mailing on Friday gives you a chance to
receive an answer during week-end !

			Marc ZYNGIER

