Index
All Classes and Interfaces|All Packages|Constant Field Values|Serialized Form
A
- AdaptiveProbe - Class in org.apache.tika.ml.chardetect
-
Reads an encoding-detection probe sized by content, not raw bytes.
- AMBIGUOUS - Enum constant in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Sample is structurally valid but contains no complete multi-byte sequence (pure ASCII, or only a truncated lead at probe-end).
- analyzeBigrams(byte[], int, int) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
For each scored bigram in the probe (same skip rules as
NaiveBayesBigramEncodingDetector.scoreClasses(byte[])), compute and return its dequantized contribution to two specified classes' scores. - appliesTo(String) - Static method in class org.apache.tika.ml.chardetect.CjkDecodeValidator
-
True for the legacy multi-byte CJK charsets this veto applies to (the decode-failure signal is meaningful only for these; ISO-2022 is handled structurally and single-byte charsets don't apply).
- ARABIC - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
B
- bigram - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.BigramContrib
- BigramContrib(int, double, double) - Constructor for class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.BigramContrib
- buildGroupIndices(String[]) - Static method in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Build a per-class group-index array from a label array (e.g. from a
LinearModel), usingCharsetConfusables.GROUPS(both symmetric and superset chains) for probability collapsing in inference. - byteIdenticalOnProbe(byte[], Charset, Charset) - Static method in class org.apache.tika.ml.chardetect.DecodeEquivalence
-
Returns
trueif decodingprobeunder charsetsaandbproduces bit-identical character sequences.
C
- CAP_PER_BIGRAM_NATS - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Per-distinct-bigram cap: top-scoring class's contribution is clipped to the best cross-cohort class's contribution + this many nats.
- CharsetConfusables - Class in org.apache.tika.ml.chardetect
-
Charset relationships used for lenient (lenient) evaluation of charset detectors.
- checkAscii(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueifbytescontains no bytes with value >= 0x80 (i.e. pure 7-bit ASCII, which is a strict subset of UTF-8). - checkAscii(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- checkHz(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif HZ-GB-2312 switching sequences are present. - checkHz(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- checkIbm424(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Detects IBM424 (EBCDIC Hebrew) by examining the sub-0x80 byte landscape.
- checkIbm424(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- checkIbm500(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Detects IBM500 (International EBCDIC / EBCDIC-500) by looking for the combination of the EBCDIC space byte and high-byte Latin letter density.
- checkIbm500(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- checkIso2022Jp(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Deprecated.
- checkUtf8(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Validates the UTF-8 byte grammar of the sample and returns one of three outcomes:
StructuralEncodingRules.Utf8Result.LIKELY_UTF8: all multi-byte sequences are valid and the sample contains enough high bytes to be informative. - checkUtf8(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- CJK - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- CjkDecodeValidator - Class in org.apache.tika.ml.chardetect
-
Structural false-CJK veto: measures how badly a probe fails to decode under a legacy multi-byte CJK charset, robustly against embedded UTF-8.
- collapseGroups(float[], int[][]) - Static method in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Collapse confusable group probabilities: within each group, sum all members' probabilities and assign the total to the highest-scoring member; the other members get 0.
- contribA - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.BigramContrib
- contribB - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.BigramContrib
- countUtf8Errors(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- countUtf8Errors(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- countUtf8Sequences(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- countUtf8Sequences(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- CYRILLIC - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
D
- DecodeEquivalence - Class in org.apache.tika.ml.chardetect
-
Cheap byte-wise decode-equivalence check for single-byte charsets.
- DEFAULT_CONTENT_TARGET - Static variable in class org.apache.tika.ml.chardetect.AdaptiveProbe
-
Default body-content target.
- DEFAULT_MODEL_RESOURCE - Static variable in class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
-
Default NB bigram model on the classpath.
- DEFAULT_MODEL_RESOURCE - Static variable in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Default classpath resource for the trained UTF-16 specialist model.
- DEFAULT_RAW_CAP - Static variable in class org.apache.tika.ml.chardetect.AdaptiveProbe
-
Default hard ceiling on raw bytes read.
- detect(byte[]) - Method in class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
-
Byte-array entry point without metadata — same as passing
null. - detect(byte[]) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- detect(byte[]) - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Byte-array entry point for callers that already hold a probe (e.g.
- detect(byte[], Metadata) - Method in class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
-
Byte-array entry point with optional metadata.
- detect(TikaInputStream, Metadata, ParseContext) - Method in class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
- detect(TikaInputStream, Metadata, ParseContext) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- detect(TikaInputStream, Metadata, ParseContext) - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
- detectIso2022(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Detects ISO-2022-JP, ISO-2022-KR, and ISO-2022-CN by scanning for their characteristic ESC designation sequences.
- detectIso2022(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- diff() - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.BigramContrib
E
- EBCDIC - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- entityCount - Variable in class org.apache.tika.ml.chardetect.HtmlByteStripper.Result
-
Number of well-formed HTML entities stripped from TEXT.
- equals(Object) - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Indicates whether some other object is "equal to" this one.
- errors() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Returns the value of the
errorsrecord component. - extract(byte[]) - Method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
- extract(byte[], int, int) - Method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Extract from a sub-range of a byte array.
- extractSparseInto(byte[], int[], int[]) - Method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Sparse extraction into caller-owned, reusable buffers.
F
- featureLabel(int) - Static method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Human-readable label for feature index
i(for debugging).
G
- getCharsets() - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- getClassLogits() - Method in class org.apache.tika.ml.chardetect.SpecialistOutput
- getCoveredLabels() - Method in class org.apache.tika.ml.chardetect.SpecialistOutput
- getLabel(int) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- getLabels() - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- getLogit(String) - Method in class org.apache.tika.ml.chardetect.SpecialistOutput
-
Raw logit for
label, ornullif not covered. - getName() - Method in interface org.apache.tika.ml.chardetect.StatisticalSpecialist
-
Short name:
"utf16","sbcs", etc. - getName() - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
- getNumBuckets() - Method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
- getNumClasses() - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- getSpecialistName() - Method in class org.apache.tika.ml.chardetect.SpecialistOutput
- GREEK - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- GROUPS - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
All confusable groups (both symmetric and superset chains), used for probability collapsing during inference via
CharsetConfusables.collapseGroups(float[], int[][]).
H
- has2ByteColumnAsymmetry(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif the probe's byte distribution across stride-2 columns is sufficiently asymmetric to be plausible UTF-16 of some script. - has2ByteColumnAsymmetryEvidence(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Evidence-based variant of
StructuralEncodingRules.has2ByteColumnAsymmetry(byte[])with no conservative short-probe default: returnstrueonly when the bytes themselves demonstrate column asymmetry, regardless of probe length. - hasC1Bytes(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif the probe contains any byte in the C1 control range0x80–0x9F. - hasC1Bytes(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- hasCrlfBytes(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif the probe contains at least one CRLF pair (0x0D 0x0A). - hasCrlfBytes(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- hasGb18030FourByteSequence(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif the probe contains at least one GB18030-specific 4-byte sequence. - hasGb18030FourByteSequence(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- hashCode() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Returns a hash code value for this object.
- hasWideUtf8Sequence(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
True if the sample has a COMPLETE 3-/4-byte UTF-8 sequence (lead
0xE0–0xF4, continuations0x80–0xBF). - hasWideUtf8Sequence(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- HEBREW - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- HtmlByteStripper - Class in org.apache.tika.ml.chardetect
-
Byte-level HTML tag stripper used as a preprocess for charset detection.
- HtmlByteStripper.Result - Class in org.apache.tika.ml.chardetect
-
Result of a strip operation: new content length and the number of well-formed tags (including comments) successfully parsed.
I
- isDecisive() - Method in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Returns true when the grammar check produced a directional answer (either LIKELY_UTF8 or NOT_UTF8).
- isEbcdicLikely(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Returns
trueif the probe is plausibly EBCDIC based on the word-separator distribution. - isLenientMatch(String, String) - Static method in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Return
trueif predictingpredictedwhen the true charset isactualis an acceptable ("lenient") result. - ISO_TO_WINDOWS - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Maps each ISO-8859-X charset to its Windows-12XX equivalent.
L
- LATIN - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- length - Variable in class org.apache.tika.ml.chardetect.HtmlByteStripper.Result
-
Content byte count written into the destination.
- LIKELY_UTF8 - Enum constant in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Sample is grammatically valid UTF-8 and contains at least one complete multi-byte sequence.
M
- MARGIN_THRESHOLD_NATS_PER_BIGRAM - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Per-scored-bigram log-score margin (in nats) that defines "model is reliably right" vs "model is genuinely uncertain between candidates."
- MAX_PROBE_BYTES - Static variable in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Default number of probe bytes read.
- MIN_BIGRAMS_FOR_DIVERSITY_GATE - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Minimum scored bigrams required before the diversity gate applies.
- MIN_COLUMN_ASYMMETRY_PROBE - Static variable in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Minimum probe length before
StructuralEncodingRules.has2ByteColumnAsymmetry(byte[])produces meaningful diversity counts. - MIN_DISTINCT_FOR_CAP - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Minimum distinct bigrams required before the per-bigram cap applies.
- MIN_DIVERSITY_RATIO - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Minimum distinct-bigram fraction of total-scored-bigrams.
- MIN_HIGH_BYTES - Static variable in class org.apache.tika.ml.chardetect.CjkDecodeValidator
-
Minimum legacy (non-UTF-8) high bytes required before the rate is trusted.
- MojibusterEncodingDetector - Class in org.apache.tika.ml.chardetect
-
Naive-Bayes pipeline detector: structural checks for wide Unicode + BOMs before falling through to the bigram NB classifier for everything else.
- MojibusterEncodingDetector() - Constructor for class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
-
Default SPI constructor: load the NB bigram model from the classpath at
MojibusterEncodingDetector.DEFAULT_MODEL_RESOURCE. - MojibusterEncodingDetector(Path) - Constructor for class org.apache.tika.ml.chardetect.MojibusterEncodingDetector
N
- NaiveBayesBigramEncodingDetector - Class in org.apache.tika.ml.chardetect
-
Naive-Bayes byte-bigram charset classifier.
- NaiveBayesBigramEncodingDetector(InputStream) - Constructor for class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- NaiveBayesBigramEncodingDetector(Path) - Constructor for class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
- NaiveBayesBigramEncodingDetector.BigramContrib - Class in org.apache.tika.ml.chardetect
-
Per-bigram contribution to the per-class score, used for diagnostic tools that want to understand why a probe scores one class over another.
- NaiveBayesBigramEncodingDetector.Cohort - Enum Class in org.apache.tika.ml.chardetect
-
Script / writing-system family used by
NaiveBayesBigramEncodingDetector.CAP_PER_BIGRAM_NATS. - NaiveBayesBigramEncodingDetector.ScoreResult - Class in org.apache.tika.ml.chardetect
-
Score result returned by
NaiveBayesBigramEncodingDetector.scoreClassesAndCount(byte[]). - NOT_UTF8 - Enum constant in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Sample contains at least one invalid UTF-8 sequence.
- NUM_COLUMNS - Static variable in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Number of columns (even-offset vs odd-offset).
- NUM_FEATURES - Static variable in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Total feature-vector dimension: ranges * columns.
- NUM_RANGES - Static variable in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
-
Number of byte-value ranges tracked.
O
- org.apache.tika.ml.chardetect - package org.apache.tika.ml.chardetect
P
- provider() - Static method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
ServiceLoader-compatible provider method.
R
- read(TikaInputStream, int, int) - Static method in class org.apache.tika.ml.chardetect.AdaptiveProbe
-
Reads from
tis(mark/reset preserved) until tag-stripped content reachescontentTarget, the raw read reachesrawCap, or EOF — whichever first. - Result(int, int, int) - Constructor for class org.apache.tika.ml.chardetect.HtmlByteStripper.Result
S
- SBCS_LATIN_FAMILY - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Single-byte Latin-family charsets that may decode byte-identically to windows-1252 on sparse probes (where the only high bytes present fall in positions the family agrees on — e.g. 0xE4='ä' in every member).
- score(byte[]) - Method in interface org.apache.tika.ml.chardetect.StatisticalSpecialist
-
Per-class logits for the probe, or
nullto decline (probe too short, hard-gated, etc.). - score(byte[]) - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
StatisticalSpecialistentry point: raw per-class logits, ornullfor a probe too short to evaluate (fewer than 2 bytes) or missing a model. - score(TikaInputStream) - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Convenience: mark/reset the stream, read a probe, and score it.
- scoreBytes(byte[]) - Method in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Deprecated.use
Utf16SpecialistEncodingDetector.score(byte[]). Kept for existing tests. - scoreClasses(byte[]) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Compute the raw per-class score vector for a probe, without top-K extraction or softmax.
- scoreClassesAndCount(byte[]) - Method in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Like
NaiveBayesBigramEncodingDetector.scoreClasses(byte[])but also reports the number of bigrams that contributed to the dot product vs the total scored region. - scoredBigrams - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.ScoreResult
- ScoreResult(double[], int, int) - Constructor for class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.ScoreResult
- scores - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.ScoreResult
- sequences() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Returns the value of the
sequencesrecord component. - SPECIALIST_NAME - Static variable in class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Specialist name used in
SpecialistOutputfor provenance. - SpecialistOutput - Class in org.apache.tika.ml.chardetect
-
Raw per-class logits from a single MoE specialist.
- SpecialistOutput(String, Map<String, Float>) - Constructor for class org.apache.tika.ml.chardetect.SpecialistOutput
- StatisticalSpecialist - Interface in org.apache.tika.ml.chardetect
-
SPI contract for an MoE charset-detection specialist.
- strip(byte[], int, int, byte[], int) - Static method in class org.apache.tika.ml.chardetect.HtmlByteStripper
-
Strip HTML/XML tags, comments, and the bodies of
<script>and<style>elements fromsrc[srcOffset .. srcOffset+srcLen)intodststarting atdstOffset. - strippedFailureRate(byte[], Charset) - Static method in class org.apache.tika.ml.chardetect.CjkDecodeValidator
-
Failure rate of
bytesundercjkCharset's vendor superset, counting only legacy high bytes (embedded UTF-8 is skipped, not counted). - stripTags(byte[], int, int, byte[], int) - Static method in class org.apache.tika.ml.chardetect.HtmlByteStripper
-
Strip tags only; entities pass through unchanged.
- stripTagsAndEntities(byte[], int, int, byte[], int) - Static method in class org.apache.tika.ml.chardetect.HtmlByteStripper
-
Strip tags and well-formed HTML entities.
- StructuralEncodingRules - Class in org.apache.tika.ml.chardetect
-
Fast, rule-based encoding checks that run before the statistical model.
- StructuralEncodingRules.Utf8Result - Enum Class in org.apache.tika.ml.chardetect
-
Outcome of the UTF-8 structural check.
- StructuralEncodingRules.Utf8Stats - Record Class in org.apache.tika.ml.chardetect
-
Single-pass tally of UTF-8 structure over the sample: malformed-sequence events and complete valid multi-byte sequences.
- SUBLINEAR_COUNT - Static variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector
-
Sublinear count weighting ("count clipping").
- SUPERSET_OF - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Directional superset relationships: key is a charset, value is its immediate superset.
- SYMMETRIC_GROUPS - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Symmetric-only confusable groups.
- symmetricPeersOf(String) - Static method in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Return the set of charsets that are symmetrically confusable with
charset, not includingcharsetitself.
T
- tagCount - Variable in class org.apache.tika.ml.chardetect.HtmlByteStripper.Result
-
Number of well-formed tags parsed (including comments).
- THAI - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- toCharset() - Method in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
- toResult() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Collapse to the
StructuralEncodingRules.checkUtf8(byte[])tri-state: any invalidity → NOT_UTF8; else ≥1 complete multi-byte sequence → LIKELY_UTF8 (a lone truncated lead is no structural evidence); else AMBIGUOUS (pure ASCII, or truncated-lead-only). - toString() - Method in class org.apache.tika.ml.chardetect.SpecialistOutput
- toString() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Returns a string representation of this record class.
- toString() - Method in class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
- totalBigrams - Variable in class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.ScoreResult
- truncatedTailInvalid() - Method in record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Returns the value of the
truncatedTailInvalidrecord component.
U
- UTF - Enum constant in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
- Utf16ColumnFeatureExtractor - Class in org.apache.tika.ml.chardetect
-
Feature extractor for the UTF-16 specialist of the mixture-of-experts charset detector.
- Utf16ColumnFeatureExtractor() - Constructor for class org.apache.tika.ml.chardetect.Utf16ColumnFeatureExtractor
- Utf16SpecialistEncodingDetector - Class in org.apache.tika.ml.chardetect
-
UTF-16 specialist detector of the mixture-of-experts charset detection architecture.
- Utf16SpecialistEncodingDetector() - Constructor for class org.apache.tika.ml.chardetect.Utf16SpecialistEncodingDetector
-
Load the model from the default classpath location.
- utf8SequenceLength(byte[], int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
-
Length (2/3/4) of a grammatically valid UTF-8 multi-byte sequence starting at
i, or 0 if none. - utf8Stats(byte[]) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- utf8Stats(byte[], int, int) - Static method in class org.apache.tika.ml.chardetect.StructuralEncodingRules
- Utf8Stats(int, int, boolean) - Constructor for record class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Stats
-
Creates an instance of a
Utf8Statsrecord class.
V
- valueOf(String) - Static method in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
-
Returns the enum constant of this class with the specified name.
- valueOf(String) - Static method in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Returns the enum constant of this class with the specified name.
- values() - Static method in enum class org.apache.tika.ml.chardetect.NaiveBayesBigramEncodingDetector.Cohort
-
Returns an array containing the constants of this enum class, in the order they are declared.
- values() - Static method in enum class org.apache.tika.ml.chardetect.StructuralEncodingRules.Utf8Result
-
Returns an array containing the constants of this enum class, in the order they are declared.
W
- WESTERN_LATIN_FAMILY - Static variable in class org.apache.tika.ml.chardetect.CharsetConfusables
-
Strict subset of
CharsetConfusables.SBCS_LATIN_FAMILYcontaining only the Western European Latin members.
All Classes and Interfaces|All Packages|Constant Field Values|Serialized Form
StructuralEncodingRules.detectIso2022(byte[])which distinguishes JP/KR/CN.