|
||||||||||
PREV CLASS NEXT CLASS | FRAMES NO FRAMES | |||||||||
SUMMARY: NESTED | FIELD | CONSTR | METHOD | DETAIL: FIELD | CONSTR | METHOD |
java.lang.Object dk.netarkivet.wayback.batch.DeduplicateToCDXAdapter
public class DeduplicateToCDXAdapter
Class containing methods for turning duplicate entries in a crawl log into lines in a CDX index file.
Field Summary | |
---|---|
(package private) org.archive.wayback.UrlCanonicalizer |
canonicalizer
canonicalizer used to canonicalize urls. |
Constructor Summary | |
---|---|
DeduplicateToCDXAdapter()
Default constructor. |
Method Summary | |
---|---|
java.lang.String |
adaptLine(java.lang.String line)
If the input line is a crawl log entry representing a duplicate then a CDX entry is written to the output. |
void |
adaptStream(java.io.InputStream is,
java.io.OutputStream os)
Reads an input stream representing a crawl log line by line and converts any lines representing duplicate entries to wayback-compliant cdx lines. |
Methods inherited from class java.lang.Object |
---|
clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait |
Field Detail |
---|
org.archive.wayback.UrlCanonicalizer canonicalizer
Constructor Detail |
---|
public DeduplicateToCDXAdapter()
Method Detail |
---|
public java.lang.String adaptLine(java.lang.String line)
adaptLine
in interface DeduplicateToCDXAdapterInterface
line
- the crawl-log line to be analysed
public void adaptStream(java.io.InputStream is, java.io.OutputStream os)
adaptStream
in interface DeduplicateToCDXAdapterInterface
is
- The input stream from which data is read.os
- The output stream to which the cdx lines are written.
|
||||||||||
PREV CLASS NEXT CLASS | FRAMES NO FRAMES | |||||||||
SUMMARY: NESTED | FIELD | CONSTR | METHOD | DETAIL: FIELD | CONSTR | METHOD |