edu.umd.cloud9.collection
Class XMLInputFormat2
java.lang.Object
org.apache.hadoop.mapreduce.InputFormat<K,V>
org.apache.hadoop.mapreduce.lib.input.FileInputFormat<LongWritable,Text>
org.apache.hadoop.mapreduce.lib.input.TextInputFormat
edu.umd.cloud9.collection.XMLInputFormat2
public class XMLInputFormat2
- extends TextInputFormat
A simple InputFormat for XML documents (org.apache.hadoop.mapreduce API). The class recognizes begin-of-document and end-of-document
tags only: everything between those delimiting tags is returned in an uninterpreted Text
object.
- Author:
- Jimmy Lin
| Methods inherited from class org.apache.hadoop.mapreduce.lib.input.FileInputFormat |
addInputPath, addInputPaths, getInputPathFilter, getInputPaths, getMaxSplitSize, getMinSplitSize, getSplits, setInputPathFilter, setInputPaths, setInputPaths, setMaxInputSplitSize, setMinInputSplitSize |
START_TAG_KEY
public static final String START_TAG_KEY
- See Also:
- Constant Field Values
END_TAG_KEY
public static final String END_TAG_KEY
- See Also:
- Constant Field Values
XMLInputFormat2
public XMLInputFormat2()
createRecordReader
public RecordReader<LongWritable,Text> createRecordReader(InputSplit split,
TaskAttemptContext context)
- Create a record reader for a given split. The framework will call
XMLInputFormat2.XMLRecordReader.initialize(InputSplit, TaskAttemptContext) before
the split is used.
- Overrides:
createRecordReader in class TextInputFormat
- Parameters:
split - the split to be readcontext - the information about the task
- Returns:
- a new record reader
- Throws:
IOException
InterruptedException