I've patched the latest (that was 2.0.12) kernel release to allow swapping via 
NFS. Yes, I know, it is slow. But there is little sense in allowing the root 
file system on NFS and not supporting swapping via NFS, except maybe for 
administrative reasons.

There is a patch for linux-1.3.15 contained in the etherboot-2.0 package
written by Markus Gutschke (gutschk@math.uni-muenster.de) to add support
for nfs-swapping. However, this patch doesn't apply without modification 
to the 2.0.** kernels as the nfs code has changed a lot.

However, I used it as a starting point. Some notes on the implementation:

* to prevent deadlocks because of insufficient memory during nfs-swapping,
  every task that is involved in nfs transfer has implicitly the GFP_NFS
  priority for memory allocation. This is achieved by a global flag,
  "int nfs_swap_active", which is incremented every time a task enteres
  rw_swap_page() and tries to swap via nfs, and is decremented when that
  finished it's nfs swapping. Together with a new flag in the task structure,
  i.e. "task->doing_nfs" this allows for the priority override for memory
  allocations. "current->doing_nfs" is set to one whenever a task performs
  nfs transfers. 

  Thus the priority override takes precedence over the priority argument to
  get_free_pages() only when some task is currently swapping via nfs and
  if the task requiring memory actually is involved in nfs transfer at the
  moment.

* The nfs code calls sometimes kmalloc() to allocate some pages. When it does
  so, this may lead to a call to try_to_swap_out(). Then it could happen the
  task would recursively call the nfs code to swap some pages out via nfs.
  That MUST be prohibited as this could lead to a loop that would use up the
  nfs request queues. Therefore, if "current->doing_nfs" is set to 1 then
  try_to_swap_out() skips swap files that are located on nfs mounted volumes.
	
  Maybe one could do this a bit more fine grained, but it seems to work.

* The routine rw_swap_page() calls "nfs_proc_write()" and "nfs_proc_read()"
  directly, rather then calling "file->f_op->write()" etc.
  First, this reduces the overhead a little bit, and secondly, one has to
  fake root permissions when accessing the swap files (swap files really
  should be read and writable ONLY be the superuser) and I felt uncomfortable
  with doin it. However, it might be quite a good thing to use the f_ops.

However, there is something more against using the fops:

* asynchronous access to nfs mounted volumes, that contain active swap-files
  must be prohibited.

  I'm not quite sure why, but I think this is mainly because 
  generic_file_read_ahead() uses up memory very quickly because the nfs code
  simply marks pages as not uptodate and returns when doing asynchronous 
  reads. Also, the nfs request queues are filled very quickly when doing 
  asychronous reads AND swapping on the same nfs volume. Of course, the 
  reason is that there is an infinite loop somewhere, because one does 
  async io, runs out of memory and recursively enteres the nfs code again
  calling try_to_swap_out(). Also, the task doing async IO via nfs is NOT 
  locked via "task->doing_nfs", as it actually doesn't do any nfs tranfers,
  but leave this to the nfsiods.
  
  To come to an end, I prevent async IO on nfs volumes containing active
  swap files by adding a new field to the nfs server struct, namely
  "no_async" that is incremented at each call to sys_swapon() and decremented
  by sys_swapoff().


I've tested the thing on a quite paranoid configuration, only 4M of memory
and all disk, include root fs, were located on nfs. 
  
I tried to break the system by overloading it, but it seems that the code
works now. I started a couple of large tasks that were trying to eat up
about 12M at the same time, but the system still didn't hang, though it
took, of course, ages until all the programs were loaded.
	
Also, I used a modificated version of the program "swapd" by 
Nick Holloway <Nick.Holloway@alfie.demon.co.uk>, that allocates swap files
on the file-system as needed (I modified it to be able to lock itself
in memory and to acquire real time priority, if somebody is interested)
This means that the system survived repeated addon's and removals of swap
space, as well as multiple nfs swap files.

Warning: it is still possible to lock up the system due to memory fragmentation.
Unluckily, the page alloc code can't guarantee the allocation of memory chunks
of more than 1 page. This means that really save setups are only possible
with the nfs rsize and wsize below 4096 byte, at least for the nfs volumes that
contain swap files. Of course, the preformance is diminuished by a small rsize
and wsize. Also, it might be necessary to adjust the freepages parameters in
/proc/sys/vm/freepages, at least for systems with few physical memory. This 
might be necessary as the nfs code is allowed to allocate reserved pages to 
prevent deadlock. But when no reserved pages remain, one looses.

/proc/sys/vm/freepages containes 3 numbers:
min_free_pages free_pages_low free_pages_high

Linux tried to set these values to "16 #somevalue #somevalue" for the 
4Mb machine. I used a bit higher values. Hint: one can set these parameters
with 

echo "30 50 80" > /proc/sys/vm/freepages

when the proc fs is mounted rw.

One last note: there is a bug in the kswapd code, that can cause the kswapd
to sleep forever and never wake up again. This is because the kswapd timer 
as well as the kswapd daemon itself try to set the variable "kswapd_awake"
if the daemon wakes up, which can lead to a deadlock for the kswapd daemon.
I'll send mail on this bug to Linus separately (I'm quite sure that it 
almost never happens except of course in my testing environment that was
so low with RAM).
