Package org.apache.tika.ml.chardetect
Class MojibusterEncodingDetector
java.lang.Object
org.apache.tika.ml.chardetect.MojibusterEncodingDetector
- All Implemented Interfaces:
Serializable,org.apache.tika.config.SelfConfiguring,org.apache.tika.detect.EncodingDetector
public class MojibusterEncodingDetector
extends Object
implements org.apache.tika.detect.EncodingDetector
Naive-Bayes pipeline detector: structural checks for wide Unicode
+ BOMs before falling through to the bigram NB classifier for
everything else.
Order of operations:
- UTF-32 codepoint validity via
WideUnicodeDetector. 4-byte-aligned probes with valid Unicode codepoints in exactly one endian order are deterministically UTF-32. - UTF-16 column-asymmetry specialist. Stride-2 column histograms reliably distinguish UTF-16-LE / BE from other content — a question bigram NB fundamentally can't answer (LE and BE produce the same bigram multiset).
- Naive-Bayes bigram classifier. Handles the single-byte and multi-byte CJK classes where byte-bigrams are the natural discriminative signal.
BOM detection is NOT handled here. The canonical
location is org.apache.tika.detect.BOMDetector (tika-core),
SPI-registered, runs first in DefaultEncodingDetector's
chain and emits a DECLARATIVE candidate. This pipeline
composes with that detector externally, not internally.
Each prefix layer short-circuits when it produces a confident candidate. Conservative: only return at a layer when that layer's structural check is clean.
- See Also:
-
Field Summary
Fields -
Constructor Summary
ConstructorsConstructorDescriptionDefault SPI constructor: load the NB bigram model from the classpath atDEFAULT_MODEL_RESOURCE.MojibusterEncodingDetector(Path nbModelPath) -
Method Summary
Modifier and TypeMethodDescriptionList<org.apache.tika.detect.EncodingResult>detect(byte[] probe) Byte-array entry point without metadata — same as passingnull.List<org.apache.tika.detect.EncodingResult>detect(byte[] probe, org.apache.tika.metadata.Metadata metadata) Byte-array entry point with optional metadata.List<org.apache.tika.detect.EncodingResult>detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext)
-
Field Details
-
DEFAULT_MODEL_RESOURCE
Default NB bigram model on the classpath.- See Also:
-
-
Constructor Details
-
MojibusterEncodingDetector
Default SPI constructor: load the NB bigram model from the classpath atDEFAULT_MODEL_RESOURCE. The UTF-16 specialist loads its own model the same way.- Throws:
IOException
-
MojibusterEncodingDetector
- Throws:
IOException
-
-
Method Details
-
detect
public List<org.apache.tika.detect.EncodingResult> detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext) throws IOException - Specified by:
detectin interfaceorg.apache.tika.detect.EncodingDetector- Throws:
IOException
-
detect
Byte-array entry point without metadata — same as passingnull. -
detect
public List<org.apache.tika.detect.EncodingResult> detect(byte[] probe, org.apache.tika.metadata.Metadata metadata) Byte-array entry point with optional metadata. If metadata's content-type suggests HTML/XML (or is absent), HTML is stripped before the NB stage — but never before the wide-Unicode structural checks, which need byte alignment intact.
-