Class MojibusterEncodingDetector

java.lang.Object
org.apache.tika.ml.chardetect.MojibusterEncodingDetector
All Implemented Interfaces:
Serializable, org.apache.tika.config.SelfConfiguring, org.apache.tika.detect.EncodingDetector

public class MojibusterEncodingDetector extends Object implements org.apache.tika.detect.EncodingDetector
Naive-Bayes pipeline detector: structural checks for wide Unicode + BOMs before falling through to the bigram NB classifier for everything else.

Order of operations:

  1. UTF-32 codepoint validity via WideUnicodeDetector. 4-byte-aligned probes with valid Unicode codepoints in exactly one endian order are deterministically UTF-32.
  2. UTF-16 column-asymmetry specialist. Stride-2 column histograms reliably distinguish UTF-16-LE / BE from other content — a question bigram NB fundamentally can't answer (LE and BE produce the same bigram multiset).
  3. Naive-Bayes bigram classifier. Handles the single-byte and multi-byte CJK classes where byte-bigrams are the natural discriminative signal.

BOM detection is NOT handled here. The canonical location is org.apache.tika.detect.BOMDetector (tika-core), SPI-registered, runs first in DefaultEncodingDetector's chain and emits a DECLARATIVE candidate. This pipeline composes with that detector externally, not internally.

Each prefix layer short-circuits when it produces a confident candidate. Conservative: only return at a layer when that layer's structural check is clean.

See Also:
  • Field Summary

    Fields
    Modifier and Type
    Field
    Description
    static final String
    Default NB bigram model on the classpath.
  • Constructor Summary

    Constructors
    Constructor
    Description
    Default SPI constructor: load the NB bigram model from the classpath at DEFAULT_MODEL_RESOURCE.
     
  • Method Summary

    Modifier and Type
    Method
    Description
    List<org.apache.tika.detect.EncodingResult>
    detect(byte[] probe)
    Byte-array entry point without metadata — same as passing null.
    List<org.apache.tika.detect.EncodingResult>
    detect(byte[] probe, org.apache.tika.metadata.Metadata metadata)
    Byte-array entry point with optional metadata.
    List<org.apache.tika.detect.EncodingResult>
    detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext)
     

    Methods inherited from class java.lang.Object

    clone, equals, finalize, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
  • Field Details

    • DEFAULT_MODEL_RESOURCE

      public static final String DEFAULT_MODEL_RESOURCE
      Default NB bigram model on the classpath.
      See Also:
  • Constructor Details

    • MojibusterEncodingDetector

      public MojibusterEncodingDetector() throws IOException
      Default SPI constructor: load the NB bigram model from the classpath at DEFAULT_MODEL_RESOURCE. The UTF-16 specialist loads its own model the same way.
      Throws:
      IOException
    • MojibusterEncodingDetector

      public MojibusterEncodingDetector(Path nbModelPath) throws IOException
      Throws:
      IOException
  • Method Details

    • detect

      public List<org.apache.tika.detect.EncodingResult> detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext) throws IOException
      Specified by:
      detect in interface org.apache.tika.detect.EncodingDetector
      Throws:
      IOException
    • detect

      public List<org.apache.tika.detect.EncodingResult> detect(byte[] probe)
      Byte-array entry point without metadata — same as passing null.
    • detect

      public List<org.apache.tika.detect.EncodingResult> detect(byte[] probe, org.apache.tika.metadata.Metadata metadata)
      Byte-array entry point with optional metadata. If metadata's content-type suggests HTML/XML (or is absent), HTML is stripped before the NB stage — but never before the wide-Unicode structural checks, which need byte alignment intact.