Class Utf16SpecialistEncodingDetector
- All Implemented Interfaces:
Serializable,org.apache.tika.config.SelfConfiguring,org.apache.tika.detect.EncodingDetector,StatisticalSpecialist
Utf16ColumnFeatureExtractor to produce a column-asymmetry-based
judgment of UTF-16-LE vs UTF-16-BE.
HTML-immune by construction
The feature set the model consumes (12 per-column byte-range counts)
captures the 2-byte alignment asymmetry that UTF-16 content produces and
HTML content cannot — HTML has no 2-byte alignment, so any byte range
appears with equal expected frequency at even vs odd positions. No
amount of HTML markup can fire this specialist. See
Utf16ColumnFeatureExtractor for the detailed argument.
Stage 1 of the MoE migration
Runs alongside the existing MojibusterEncodingDetector
rather than replacing any piece of it. Emits a single
EncodingResult.ResultType.STATISTICAL candidate for the meta
arbiter (JunkFilterEncodingDetector) to weigh against the other
detectors in the chain. The existing WideUnicodeDetector-based
structural UTF-16 detection inside Mojibuster is not removed yet — both
can operate in parallel during Stage 1 validation.
Model loading
The default constructor loads a trained model from the classpath at
DEFAULT_MODEL_RESOURCE. If the resource is absent or
malformed, construction throws IOException — the detector
never operates in a no-op state because silent no-ops produce wrong
answers without any indication that something's wrong. Deploy the
detector only when a trained model is bundled; remove it from the
chain otherwise.
Probe size
Reads up to MAX_PROBE_BYTES bytes. UTF-16 column-asymmetry
signal stabilises quickly — even ~100 bytes is usually enough for a
strong call. Default 512 is generous.
- See Also:
-
Field Summary
FieldsModifier and TypeFieldDescriptionstatic final StringDefault classpath resource for the trained UTF-16 specialist model.static final intDefault number of probe bytes read.static final StringSpecialist name used inSpecialistOutputfor provenance. -
Constructor Summary
ConstructorsConstructorDescriptionLoad the model from the default classpath location. -
Method Summary
Modifier and TypeMethodDescriptionList<org.apache.tika.detect.EncodingResult>detect(byte[] probe) Byte-array entry point for callers that already hold a probe (e.g.List<org.apache.tika.detect.EncodingResult>detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext) getName()Short name:"utf16","sbcs", etc.provider()ServiceLoader-compatible provider method.score(byte[] probe) StatisticalSpecialistentry point: raw per-class logits, ornullfor a probe too short to evaluate (fewer than 2 bytes) or missing a model.score(org.apache.tika.io.TikaInputStream tis) Convenience: mark/reset the stream, read a probe, and score it.scoreBytes(byte[] probe) Deprecated.
-
Field Details
-
DEFAULT_MODEL_RESOURCE
Default classpath resource for the trained UTF-16 specialist model. Missing resource → detector is a noop (logged once at construction).- See Also:
-
MAX_PROBE_BYTES
public static final int MAX_PROBE_BYTESDefault number of probe bytes read.- See Also:
-
SPECIALIST_NAME
Specialist name used inSpecialistOutputfor provenance.- See Also:
-
-
Constructor Details
-
Utf16SpecialistEncodingDetector
Load the model from the default classpath location.- Throws:
IOException- if the model resource is missing or malformed — the detector does not operate in a no-op state.
-
-
Method Details
-
provider
ServiceLoader-compatible provider method. Wraps the checkedIOExceptionfrom the no-arg constructor in aServiceConfigurationErrorso the arbiter can catch it and skip a specialist whose model is not bundled — without hiding the cause. -
getName
Description copied from interface:StatisticalSpecialistShort name:"utf16","sbcs", etc.- Specified by:
getNamein interfaceStatisticalSpecialist
-
score
StatisticalSpecialistentry point: raw per-class logits, ornullfor a probe too short to evaluate (fewer than 2 bytes) or missing a model. Returningnulldeclines to contribute; an all-low logit vector would muddy the combiner.Unlike
detect(org.apache.tika.io.TikaInputStream, org.apache.tika.metadata.Metadata, org.apache.tika.parser.ParseContext), this method does not apply a margin threshold — downstream pooling sees raw logits for both classes.- Specified by:
scorein interfaceStatisticalSpecialist
-
score
Convenience: mark/reset the stream, read a probe, and score it. Returnsnullif the probe is too short.- Throws:
IOException
-
scoreBytes
Deprecated.usescore(byte[]). Kept for existing tests. -
detect
public List<org.apache.tika.detect.EncodingResult> detect(org.apache.tika.io.TikaInputStream tis, org.apache.tika.metadata.Metadata metadata, org.apache.tika.parser.ParseContext parseContext) throws IOException - Specified by:
detectin interfaceorg.apache.tika.detect.EncodingDetector- Throws:
IOException
-
detect
Byte-array entry point for callers that already hold a probe (e.g.MojibusterEncodingDetector's pipeline). Returns an empty list for probes belowMIN_PROBE_BYTESor when the winning class has margin <MIN_LOGIT_MARGIN.
-
score(byte[]).