


    This is yet another release of the Linux NFS client with enhanced
    async read-ahead functionality.  This release replaces the
    simple-minded nfsiod scheme.

    Rather than using a fixed number of nfsiod processes, we now use
    a single rpciod daemon that takes care of async RPC requests. This
    avoids unnecessary context switches. With an nfsiod scheme, control
    has to switch to the nfsiod process before a reply can be received.

    With the new scheme, the readahead RPC request is enqueued by the
    user process as before, but there's no need to wake up rpciod. Instead,
    incoming replies trigger the rpc_data_ready callback, which looks up
    the client for whom this reply is meant, and wakes up the handler
    of the first request on the pending list. If the request happens to be
    an async RPC call, rpciod is woken up to dispatch it, otherwise a
    normal user process will do it.

    It also incoporates other small enhancements; for a full list please
    see the ChangeLog.

    NOTE

    With the current scheme, you may occasionally see messages in your
    syslog saying the IP fragment reassembly failed. This may be annoying,
    but it's nothing to worry about. This problem is still pending, though.


    HOW TO USE

    This stuff compiles as a loadable module (the last version I tried it
    on was 1.3.99).  Simply type mkmodule, and insmod nfs.o. When mounting
    your first NFS volume, this will start the rpciod daemon; and after
    unmounting the last volume, the daemon will exit. This should help
    kerneld users a lot.

    Note that readahead will not be enabled unless rsize is at greater
    or equal PAGE_SIZE (4096 at the moment).

    Alternatively, you can put it right into the kernel: remove everything
    from fs/nfs, move the Makefile and all *.c to this directory, and
    copy all *.h files to include/linux.

    After mounting, you should be able to watch (with tcpdump) several
    RPC READ calls being placed simultaneously.


    HOW IT WORKS

    When a process reads from a file on an NFS volume, the following
    happens:

     *	nfs_file_read calls generic_file_read.

     *	generic_file_read requests one ore more pages via
    	nfs_readpage.

     *	nfs_readpage allocates an RPC request struct, fills in the READ
	request, and registers a callback function with the request.
	It then sends out the RPC call, enqueues the request, and returns.
    	If the server is congested, nfs_readpage places the call
    	directly, waiting for the reply (sync readpage).

     *  When the reply arrives, the first process on clnt->pending is
	woken up, and receives and dispatches the reply. For synchronous
	requests, the calling process is woken up. For async requests,
	the callback is executed, which marks the page uptodate and
	wakes up all processes waiting on the page.

    This is the rough outline only. There are a few things to note:

     *	Async RPC will not be tried when server->rsize < PAGE_SIZE.

     *	When an error occurs, rpciod has no way of returning
    	the error code to the user process. Therefore, it flags
    	page->error and wakes up all processes waiting on that
    	page (they usually do so from withing generic_readpage).

    	generic_readpage finds that the page is still not
    	uptodate, and calls nfs_readpage again. This time around,
    	nfs_readpage notices that page->error is set and
    	unconditionally does a synchronous RPC call.

    	This area needs a lot of improvement, since read errors
    	are not that uncommon (e.g. we have to retransmit calls
    	if the fsuid is different from the ruid in order to
    	cope with root squashing and stuff like that).

	Retransmits with fsuid/ruid change should be handled by the BIO
	code, but this doesn't come easily (a more general nfs_call
	routine that does all this may be useful...)

     *	To save some time on readaheads, we save one data copy
    	by frobbing the page into the iovec passed to the
	RPC code so that the networking layer copies the
    	data into the page directly.

    	This needs to be adjustable (different authentication
    	flavors; AUTH_NULL versus AUTH_SHORT verifiers).

     *	Currently, the maximum number of outstanding RPC requests
	is limited to twelve. I found that if the limit is too high
	(e.g. 32), the congestion window value starts to oscillate,
	and effective throughput drops.

	However, this is not a good solution. There must be an algorithm
	that adapts more easily to varying server load (by watching RTT
	deviation?) and client load (e.g., low memory may lead IP
	defragmentation failures).


    WISH LIST

    After giving this thing some testing, I'd like to add some more
    features:

     *	Some sort of async write handling. True write-back doesn't
	work with the current kernel (I think), because invalidate_pages
	kills all pages, regardless of whether they're dirty or not.
	Besides, this may require special bdflush treatment because
	write caching on clients is really hairy.

	Alternatively, a write-through scheme might be useful where
	the client enqueues the request, but leaves collecting the
	results to nfsiod. Again, we need a way to pass RPC errors
	back to the application.

     *	Support for different authentication flavors.

     *	/proc/net/nfsclnt (for nfsstat, etc.).

May 14, 1996
Olaf Kirch <okir@monad.swb.de>
