
------------------------------------------------------------------------------

A license is hereby granted to reproduce this software source code and
to create executable versions from this source code for personal,
non-commercial use.  The copyright notice included with the software
must be maintained in all copies produced.

THIS PROGRAM IS PROVIDED "AS IS". THE AUTHOR PROVIDES NO WARRANTIES
WHATSOEVER, EXPRESSED OR IMPLIED, INCLUDING WARRANTIES OF
MERCHANTABILITY, TITLE, OR FITNESS FOR ANY PARTICULAR PURPOSE.  THE
AUTHOR DOES NOT WARRANT THAT USE OF THIS PROGRAM DOES NOT INFRINGE THE
INTELLECTUAL PROPERTY RIGHTS OF ANY THIRD PARTY IN ANY COUNTRY.

Copyright (c) 1995-1999, John Conover, All Rights Reserved.

Comments and/or bug reports should be addressed to:

    john@johncon.com (John Conover)

------------------------------------------------------------------------------

Rel is a suite of programs and tools for building wide area full text
information retrieval systems over the Internet. The search mechanisms
are capable of sorting documents by relevance to keyword search
criteria. Boolean operations (and, or, not, and grouping operators,)
on multiple keywords are fully supported and the programs are capable
of phonetic keyword search. The programs are also find application in
enterprise wide area information retrieval systems.

The rel.zip tape archive contains the C sources for the programs and
applications.  The archive consists of the programs:

    rel which orders the relevance of text documents to a search
    criteria.

    rels which orders the relevance of text documents to a simple
    phonetic search criteria.

    relx which orders the relevance of text documents to a phonetic
    search criteria.

    htmlrel which orders the relevance of html text documents to a
    search criteria.

    htmlrels which orders the relevance of html text documents to a
    simple phonetic search criteria.

    htmlrelx which orders the relevance of html text documents to a
    phonetic search criteria.

    wgetrel which searches Internet Web pages for documents that are
    relevant to a search criteria.

    wgetrels which searches Internet Web pages for documents that are
    relevant to a simple phonetic search criteria.

    wgetrelx which searches Internet Web pages for documents that are
    relevant to a phonetic search criteria.

The wgetrel suite of programs use the htmlrel programs to control the
direction of search across the Internet as determined by the relevance
of documents found at a site. The rel suite of programs is useful for
building enterprise wide email repositories that are searchable by
relevance. The rel suite of programs are Open Source software.

The concept of document relevance searching is not new. It was first
proposed by Vannevar Bush (Bush, V. (1941) "Memorandum regarding
Memex," [Vannevar Bush Papers, Library of Congress], Box 50, General
Correspondence File, Eric Hodgins.)

MORE ABOUT THE PROGRAMS

Rel is a program that determines the relevance of text documents to a
set of keywords expressed in boolean infix notation. The list of file
names that are relevant are printed to the standard output, in order
of relevance.

Rels is a program that determines the relevance of text documents to a
set of keywords expressed in boolean infix notation. The relevance is
determined by comparing the phonetic representation of the keywords
with the phonetic representation of every word in a
document. (Phonetic searching has some degree of tolerance to
misspelled words.) The list of file names that are relevant are
printed to the standard output, in order of relevance.

Relx is a program that determines the relevance of text documents to a
set of keywords expressed in boolean infix notation. The relevance is
determined by comparing the phonetic representation of the keywords
with the phonetic representation of every word in a
document. (Phonetic searching has some degree of tolerance to
misspelled words.) The list of file names that are relevant are
printed to the standard output, in order of relevance. The phonetic
algorithm of the rels(1) program is modified-see the man page for
rels(1) and relx(1).

    There is a companion set of procmail/smartlist, (both programs are
    available via anonymous ftp from:
    ftp://ftp.informatik.rwth-aachen.de in
    /pub/packages/procmail/procmail.tar.gz and
    /pub/packages/procmail/SmartList.tar.gz,) scripts in the
    applications directory to construct an enterprise wide, full text
    information retrieval system that uses the Unix MTA (Message
    Transfer Agent,) as a delivery, query, and distribution
    system. The programs rel(1), rels(1), and relx(1) are used for
    searching the enterprise wide archive.

Htmlrel is a program that determines the relevance of text documents
to a set of keywords expressed in boolean infix notation. The list of
file names that are relevant are printed to the standard output, in
order of relevance. The program is identical to the rel(1) program,
except the output file list format has been altered to be compatible
with Netscape's level 1 bookmark file syntax.

    There is a companion shell script, wgetrel, in the applications
    directory which is an Internet Web page search engine, using the
    programs htmlrel(1) and wget(1). (The program wget(1) is available
    via anonymous ftp from ftp://prep.ai.mit.edu/pub/gnu/wget.tar.gz.)
    The direction of the search is controlled through determination of
    relevance of the documents to a search criteria.

Htmlrels is a program that determines the relevance of text documents
to a set of keywords expressed in boolean infix notation. The
relevance is determined by comparing the phonetic representation of
the keywords with the phonetic representation of every word in a
document. (Phonetic searching has some degree of tolerance to
misspelled words.) The list of file names that are relevant are
printed to the standard output, in order of relevance.  The program is
identical to the rel(1) program, except the output file list format
has been altered to be compatible with Netscape's level 1 bookmark
file syntax.

    There is a companion shell script, wgetrels, in the applications
    directory which is an Internet Web page search engine, using the
    programs htmlrels(1) and wget(1). (The program wget(1) is
    available via anonymous ftp from
    ftp://prep.ai.mit.edu/pub/gnu/wget.tar.gz.) The direction of the
    search is controlled through determination of relevance of the
    documents to a search criteria.

Htmlrelx is a program that determines the relevance of text documents
to a set of keywords expressed in boolean infix notation. The
relevance is determined by comparing the phonetic representation of
the keywords with the phonetic representation of every word in a
document. (Phonetic searching has some degree of tolerance to
misspelled words.) The list of file names that are relevant are
printed to the standard output, in order of relevance. The phonetic
algorithm of the htmlrels(1) program is modified-see the man page for
htmlrels(1) and htmlrelx(1).  The program is identical to the rel(1)
program, except the output file list format has been altered to be
compatible with Netscape's level 1 bookmark file syntax.

    There is a companion shell script, wgetrelx, in the applications
    directory which is an Internet Web page search engine, using the
    programs htmlrelx(1) and wget(1). (The program wget(1) is
    available via anonymous ftp from
    ftp://prep.ai.mit.edu/pub/gnu/wget.tar.gz.) The direction of the
    search is controlled through determination of relevance of the
    documents to a search criteria.

john@johncon.com (John Conover)
Campbell, California, USA
March, 1999
