Commit 01d7839a authored by Chris Pepper's avatar Chris Pepper
Browse files

Finish cleanup.

	Remove <em> around e.g. & i.e.


git-svn-id: https://svn.apache.org/repos/asf/httpd/httpd/branches/2.0.x@650012 13f79535-47bb-0310-9956-ffa450edef68
parent 10acf73b
Loading
Loading
Loading
Loading
+81 −75
Changes for docs/manual/rewrite/rewrite_guide_advanced.xml: 81 added lines, 75 removed lines.
Original line number Diff line number Diff line
@@ -36,7 +36,7 @@

    <note type="warning">ATTENTION: Depending on your server configuration
    it may be necessary to adjust the examples for your
    situation, <em>e.g.,</em> adding the <code>[PT]</code> flag if
    situation, e.g., adding the <code>[PT]</code> flag if
    using <module>mod_alias</module> and
    <module>mod_userdir</module>, etc. Or rewriting a ruleset
    to work in <code>.htaccess</code> context instead
@@ -61,7 +61,7 @@ introduction</a></seealso>

        <dd>
          <p>We want to create a homogeneous and consistent URL
          layout across all WWW servers on an Intranet web cluster, <em>i.e.,</em>
          layout across all WWW servers on an Intranet web cluster, i.e.,
          all URLs (by definition server-local and thus
          server-dependent!) become server <em>independent</em>!
          What we want is to give the WWW namespace a single consistent
@@ -310,7 +310,7 @@ RewriteRule (.*) netsw-lsdir.cgi/$1

    <section id="redirect404">

      <title>Redirect Failing URLs to Another Webserver</title>
      <title>Redirect Failing URLs to Another Web Server</title>

      <dl>
        <dt>Description:</dt>
@@ -373,10 +373,10 @@ RewriteRule ^(.+) http://<strong>webserverB</strong>.dom/$1
          <p>Do you know the great CPAN (Comprehensive Perl Archive
          Network) under <a href="http://www.perl.com/CPAN"
          >http://www.perl.com/CPAN</a>?
          This does a redirect to one of several FTP servers around
          the world which each carry a CPAN mirror and (theoretically)
          near the requesting client. Actually this
          can be called an FTP access multiplexing service.
          CPAN automatically redirects browsers to one of many FTP
          servers around the world (generally one near the requesting
          client); each server carries a full CPAN mirror. This is
          effectively an FTP access multiplexing service.
          CPAN runs via CGI scripts, but how could a similar approach
          be implemented via <module>mod_rewrite</module>?</p>
        </dd>
@@ -430,7 +430,7 @@ com ftp://ftp.cxan.com/CxAN/
        <dd>
          <p>At least for important top-level pages it is sometimes
          necessary to provide the optimum of browser dependent
          content, <em>i.e.,</em> one has to provide one version for
          content, i.e., one has to provide one version for
          current browsers, a different version for the Lynx and text-mode
          browsers, and another for other browsers.</p>
        </dd>
@@ -478,12 +478,12 @@ RewriteRule ^foo\.html$ foo.<strong>32</strong>.html [<strong>L
          explicit up-to-date copy of the remote data on the local
          machine. For a web server we could use the program
          <code>webcopy</code> which runs via HTTP. But both
          techniques have one major drawback: The local copy is
          always just as up-to-date as the last time we ran the program. It
          would be much better if the mirror is not a static one we
          techniques have a major drawback: The local copy is
          always only as up-to-date as the last time we ran the program. It
          would be much better if the mirror was not a static one we
          have to establish explicitly. Instead we want a dynamic
          mirror with data which gets updated automatically when
          there is need (updated on the remote host).</p>
          mirror with data which gets updated automatically on the
          as needed on the remote host(s).</p>
        </dd>

        <dt>Solution:</dt>
@@ -543,20 +543,20 @@ RewriteRule ^http://www\.remotesite\.com/(.*)$ /mirror/of/remotesite/$1
          <p>This is a tricky way of virtually running a corporate
          (external) Internet web server
          (<code>www.quux-corp.dom</code>), while actually keeping
          and maintaining its data on a (internal) Intranet webserver
          and maintaining its data on an (internal) Intranet web server
          (<code>www2.quux-corp.dom</code>) which is protected by a
          firewall. The trick is that on the external webserver we
          retrieve the requested data on-the-fly from the internal
          firewall. The trick is that the external web server retrieves
          the requested data on-the-fly from the internal
          one.</p>
        </dd>

        <dt>Solution:</dt>

        <dd>
          <p>First, we have to make sure that our firewall still
          protects the internal webserver and that only the
          <p>First, we must make sure that our firewall still
          protects the internal web server and only the
          external web server is allowed to retrieve data from it.
          For a packet-filtering firewall we could for instance
          On a packet-filtering firewall, for instance, we could
          configure a firewall ruleset like the following:</p>

<example><pre>
@@ -596,18 +596,18 @@ RewriteRule ^/home/([^/]+)/.www/?(.*) http://<strong>www2</strong>.quux-corp.dom
        <dt>Solution:</dt>

        <dd>
          <p>There are a lot of possible solutions for this problem.
          We will discuss first a commonly known DNS-based variant
          and then the special one with <module>mod_rewrite</module>:</p>
          <p>There are many possible solutions for this problem.
          We will first discuss a common DNS-based method,
          and then one based on <module>mod_rewrite</module>:</p>

          <ol>
            <li>
              <strong>DNS Round-Robin</strong>

              <p>The simplest method for load-balancing is to use
              the DNS round-robin feature of <code>BIND</code>.
              DNS round-robin.
              Here you just configure <code>www[0-9].foo.com</code>
              as usual in your DNS with A(address) records, <em>e.g.,</em></p>
              as usual in your DNS with A (address) records, e.g.,</p>

<example><pre>
www0   IN  A       1.2.3.1
@@ -618,7 +618,7 @@ www4 IN A 1.2.3.5
www5   IN  A       1.2.3.6
</pre></example>

              <p>Then you additionally add the following entry:</p>
              <p>Then you additionally add the following entries:</p>

<example><pre>
www   IN  A       1.2.3.1
@@ -630,15 +630,17 @@ www IN A 1.2.3.5

              <p>Now when <code>www.foo.com</code> gets
              resolved, <code>BIND</code> gives out <code>www0-www5</code>
              - but in a slightly permutated/rotated order every time.
              - but in a permutated (rotated) order every time.
              This way the clients are spread over the various
              servers. But notice that this is not a perfect load
              balancing scheme, because DNS resolution information
              gets cached by the other nameservers on the net, so
              balancing scheme, because DNS resolutions are
              cached by clients and other nameservers, so
              once a client has resolved <code>www.foo.com</code>
              to a particular <code>wwwN.foo.com</code>, all its
              subsequent requests also go to this particular name
              <code>wwwN.foo.com</code>. But the final result is
              subsequent requests will continue to go to the same
              IP (and thus a single server), rather than being
              distributed across the other available servers. But the
              over result is
              okay, because the requests are collectively
              spread over the various web servers.</p>
            </li>
@@ -651,8 +653,8 @@ www IN A 1.2.3.5
              <code>lbnamed</code> which can be found at <a
              href="http://www.stanford.edu/~schemers/docs/lbnamed/lbnamed.html">
              http://www.stanford.edu/~schemers/docs/lbnamed/lbnamed.html</a>.
              It is a Perl 5 program in conjunction with auxilliary
              tools which provides a real load-balancing for
              It is a Perl 5 program which, in conjunction with auxilliary
              tools, provides real load-balancing via
              DNS.</p>
            </li>

@@ -670,8 +672,8 @@ www IN CNAME www0.foo.com.

              <p>entry in the DNS. Then we convert
              <code>www0.foo.com</code> to a proxy-only server,
              <em>i.e.,</em> we configure this machine so all arriving URLs
              are just pushed through the internal proxy to one of
              i.e., we configure this machine so all arriving URLs
              are simply passed through its internal proxy to one of
              the 5 other servers (<code>www1-www5</code>). To
              accomplish this we first establish a ruleset which
              contacts a load balancing script <code>lb.pl</code>
@@ -712,19 +714,24 @@ while (&lt;STDIN&gt;) {
              <code>www0.foo.com</code> still is overloaded? The
              answer is yes, it is overloaded, but with plain proxy
              throughput requests, only! All SSI, CGI, ePerl, etc.
              processing is completely done on the other machines.
              This is the essential point.</note>
              processing is handled done on the other machines.
              For a complicated site, this may work well. The biggest
              risk here is that www0 is now a single point of failure --
              if it crashes, the other servers are inaccessible.</note>
            </li>

            <li>
              <strong>Hardware/TCP Round-Robin</strong>

              <p>There is a hardware solution available, too. Cisco
              has a beast called LocalDirector which does a load
              balancing at the TCP/IP level. Actually this is some
              sort of a circuit level gateway in front of a
              webcluster. If you have enough money and really need
              a solution with high performance, use this one.</p>
              <strong>Dedicated Load Balancers</strong>

              <p>There are more sophisticated solutions, as well. Cisco,
              F5, and several other companies sell hardware load
              balancers (typically used in pairs for redundancy), which
              offer sophisticated load balancing and auto-failover
              features. There are software packages which offer similar
              features on commodity hardware, as well. If you have
              enough money or need, check these out. The <a
              href="http://vegan.net/lb/">lb-l mailing list</a> is a
              good place to research.</p>
            </li>
          </ol>
        </dd>
@@ -740,8 +747,8 @@ while (&lt;STDIN&gt;) {
        <dt>Description:</dt>

        <dd>
          <p>On the net there are a lot of nifty CGI programs. But
          their usage is usually boring, so a lot of webmaster
          <p>On the net there are many nifty CGI programs. But
          their usage is usually boring, so a lot of webmasters
          don't use them. Even Apache's Action handler feature for
          MIME-types is only appropriate when the CGI programs
          don't need special URLs (actually <code>PATH_INFO</code>
@@ -750,9 +757,9 @@ while (&lt;STDIN&gt;) {
          <code>.scgi</code> (for secure CGI) which will be processed
          by the popular <code>cgiwrap</code> program. The problem
          here is that for instance if we use a Homogeneous URL Layout
          (see above) a file inside the user homedirs has the URL
          <code>/u/user/foo/bar.scgi</code>. But
          <code>cgiwrap</code> needs the URL in the form
          (see above) a file inside the user homedirs might have a URL
          like <code>/u/user/foo/bar.scgi</code>, but
          <code>cgiwrap</code> needs URLs in the form
          <code>/~user/foo/bar.scgi/</code>. The following rule
          solves the problem:</p>

@@ -766,9 +773,9 @@ RewriteRule ^/[uge]/<strong>([^/]+)</strong>/\.www/(.+)\.scgi(.*) ...
          <code>access.log</code> for a URL subtree) and
          <code>wwwidx</code> (which runs Glimpse on a URL
          subtree). We have to provide the URL area to these
          programs so they know on which area they have to act on.
          But usually this is ugly, because they are all the times
          still requested from that areas, <em>i.e.,</em> typically we would
          programs so they know which area they are really working with.
          But usually this is complicated, because they may still be
          requested by the alternate URL form, i.e., typically we would
          run the <code>swwidx</code> program from within
          <code>/u/user/foo/</code> via hyperlink to</p>

@@ -776,10 +783,10 @@ RewriteRule ^/[uge]/<strong>([^/]+)</strong>/\.www/(.+)\.scgi(.*) ...
/internal/cgi/user/swwidx?i=/u/user/foo/
</pre></example>

          <p>which is ugly. Because we have to hard-code
          <p>which is ugly, because we have to hard-code
          <strong>both</strong> the location of the area
          <strong>and</strong> the location of the CGI inside the
          hyperlink. When we have to reorganize the area, we spend a
          hyperlink. When we have to reorganize, we spend a
          lot of time changing the various hyperlinks.</p>
        </dd>

@@ -825,12 +832,12 @@ HREF="*"

        <dd>
          <p>Here comes a really esoteric feature: Dynamically
          generated but statically served pages, <em>i.e.,</em> pages should be
          generated but statically served pages, i.e., pages should be
          delivered as pure static pages (read from the filesystem
          and just passed through), but they have to be generated
          dynamically by the web server if missing. This way you can
          have CGI-generated pages which are statically served unless
          one (or a cronjob) removes the static contents. Then the
          have CGI-generated pages which are statically served unless an
          admin (or a <code>cron</code> job) removes the static contents. Then the
          contents gets refreshed.</p>
        </dd>

@@ -844,16 +851,16 @@ RewriteCond %{REQUEST_FILENAME} <strong>!-s</strong>
RewriteRule ^page\.<strong>html</strong>$          page.<strong>cgi</strong>   [T=application/x-httpd-cgi,L]
</pre></example>

          <p>Here a request to <code>page.html</code> leads to a
          <p>Here a request for <code>page.html</code> leads to an
          internal run of a corresponding <code>page.cgi</code> if
          <code>page.html</code> is still missing or has filesize
          <code>page.html</code> is missing or has filesize
          null. The trick here is that <code>page.cgi</code> is a
          usual CGI script which (additionally to its <code>STDOUT</code>)
          CGI script which (additionally to its <code>STDOUT</code>)
          writes its output to the file <code>page.html</code>.
          Once it was run, the server sends out the data of
          Once it has completed, the server sends out
          <code>page.html</code>. When the webmaster wants to force
          a refresh the contents, he just removes
          <code>page.html</code> (usually done by a cronjob).</p>
          a refresh of the contents, he just removes
          <code>page.html</code> (typically from <code>cron</code>).</p>
        </dd>
      </dl>

@@ -867,9 +874,9 @@ RewriteRule ^page\.<strong>html</strong>$ page.<strong>cgi</strong> [
        <dt>Description:</dt>

        <dd>
          <p>Wouldn't it be nice while creating a complex webpage if
          <p>Wouldn't it be nice, while creating a complex web page, if
          the web browser would automatically refresh the page every
          time we write a new version from within our editor?
          time we save a new version from within our editor?
          Impossible?</p>
        </dd>

@@ -877,10 +884,10 @@ RewriteRule ^page\.<strong>html</strong>$ page.<strong>cgi</strong> [

        <dd>
          <p>No! We just combine the MIME multipart feature, the
          webserver NPH feature and the URL manipulation power of
          web server NPH feature, and the URL manipulation power of
          <module>mod_rewrite</module>. First, we establish a new
          URL feature: Adding just <code>:refresh</code> to any
          URL causes this to be refreshed every time it gets
          URL causes the 'page' to be refreshed every time it is
          updated on the filesystem.</p>

<example><pre>
@@ -1021,18 +1028,17 @@ exit(0);
        <dd>
          <p>The <directive type="section" module="core"
          >VirtualHost</directive> feature of Apache is nice
          and works great when you just have a few dozens
          and works great when you just have a few dozen
          virtual hosts. But when you are an ISP and have hundreds of
          virtual hosts to provide this feature is not the best
          choice.</p>
          virtual hosts, this feature is suboptimal.</p>
        </dd>

        <dt>Solution:</dt>

        <dd>
          <p>To provide this feature we map the remote web page or even
          the complete remote webarea to our namespace by the use
          of the <dfn>Proxy Throughput</dfn> feature (flag <code>[P]</code>):</p>
          the complete remote web area to our namespace using the
          <dfn>Proxy Throughput</dfn> feature (flag <code>[P]</code>):</p>

<example><pre>
##
@@ -1204,11 +1210,11 @@ RewriteRule !^http://[^/.]\.mydomain.com.* - [F]
        <dt>Description:</dt>

        <dd>
          <p>Sometimes a very special authentication is needed, for
          instance a authentication which checks for a set of
          <p>Sometimes very special authentication is needed, for
          instance authentication which checks for a set of
          explicitly configured users. Only these should receive
          access and without explicit prompting (which would occur
          when using the Basic Auth via <module>mod_auth</module>).</p>
          when using Basic Auth via <module>mod_auth</module>).</p>
        </dd>

        <dt>Solution:</dt>