initTableSnapshotMapperJob
public static void initTableSnapshotMapperJob(String snapshotName,
Scan scan,
Class<? extends TableMapper> mapper,
Class<?> outputKeyClass,
Class<?> outputValueClass,
org.apache.hadoop.mapreduce.Job job,
boolean addDependencyJars,
org.apache.hadoop.fs.Path tmpRestoreDir,
RegionSplitter.SplitAlgorithm splitAlgo,
int numSplitsPerRegion)
throws IOException
Sets up the job for reading from a table snapshot. It bypasses hbase servers and read directly
from snapshot files.
- Parameters:
snapshotName - The name of the snapshot (of a table) to read from.
scan - The scan instance with the columns, time range etc.
mapper - The mapper class to use.
outputKeyClass - The class of the output key.
outputValueClass - The class of the output value.
job - The current job to adjust. Make sure the passed job is carrying all necessary HBase
configuration.
addDependencyJars - upload HBase jars and jars for any of the configured job classes via
the distributed cache (tmpjars).
tmpRestoreDir - a temporary directory to copy the snapshot files into. Current user should
have write permissions to this directory, and this should not be a subdirectory of rootdir.
After the job is finished, restore directory can be deleted.
splitAlgo - algorithm to split
numSplitsPerRegion - how many input splits to generate per one region
- Throws:
IOException - When setting up the details fails.
- See Also:
TableSnapshotInputFormat