CVE-2026-88049
HighCVSS 8.6Summary
Tesseract 5.5.3 and earlier fails to validate dimension consistency in NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart. A crafted .traineddata file with an NT_LSTM layer can cause a heap out-of-bounds write during the first recognition step on the default LSTM engine. No fixed release is available.
Risk Assessment
Processing a malicious .traineddata file can lead to heap corruption, a crash, or potentially controlled corruption. Any environment loading untrusted Tesseract models is at risk.
Recommendation
Do not load untrusted .traineddata files and restrict access to the tessdata directory. Track Tesseract releases and apply the fix as soon as it becomes available.
Other vulnerabilities in Tesseract
See all- CVE-2026-88050Medium
In Tesseract version 5.5.3 and earlier, RecodedCharID::DeSerialize in src/ccutil/unicharcompress.h validates length_ but accepts negative code_ values from a crafted .traineddata recoder component. UnicharCompress::ComputeCodeRange can consequently produce code_range_ equal to zero, after which SetupDecoder indexes is_valid_start_ with the negative code on a size-zero vector, causing an out-of-bounds bit write and a crash or allocation failure. No fixed release is available.
- CVE-2026-88048High
Tesseract 5.5.3 and earlier does not validate the deserialized scalars ni_ and no_ against weight-matrix dimensions in FullyConnected::DeSerialize. A crafted .traineddata file with an NT_SOFTMAX layer can cause a heap out-of-bounds write and read on the default LSTM engine. No fixed release is available.
- CVE-2026-88047High
Tesseract 5.5.3 and earlier, in Classify::ReadNormProtos, extracts a whitespace-delimited token from a .traineddata file into a fixed 61-byte stack buffer without setting stream width. A token longer than 60 characters writes up to 39 attacker-controlled bytes past the buffer during TessBaseAPI::Init of the legacy engine. No fixed release is available.
- CVE-2026-73067Medium
Tesseract is an open source OCR engine. Prior to 5.5.3, a crafted .traineddata model loaded through TessBaseAPI::Init can cause SquishedDawg::read_squished_dawg in src/dict/dawg.cpp to accept an unterminated forward-edge run, after which SquishedDawg::Load calls num_forward_edges(0) and last_edge in src/dict/dawg.h reads beyond edges_, causing a heap out-of-bounds read and process crash before image processing. This issue is fixed in version 5.5.3.
- CVE-2026-73066Medium
Tesseract is an open source OCR engine. Prior to 5.5.3, a crafted .traineddata LSTM model component loaded through Tesseract's deserializer can cause an unchecked signed integer multiplication in Convolve::DeSerialize in src/lstm/convolve.cpp to wrap the convolution output-channel count, undersizing the forward-pass output buffer while writes use the unwrapped element count and causing a heap out-of-bounds write during OCR recognition. This issue is fixed in version 5.5.3.
Original NVD description (English source)
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, prior .traineddata hardening added bounds checks to NetworkIO::CopyTimeStepGeneral and NetworkIO::Randomize in src/lstm/networkio.cpp but left NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart unchecked. In LSTM::Forward in src/lstm/lstm.cpp, source_ is sized from the independently deserialized na_ field while the WriteTimeStepPart count is ns_, which comes from the CI gate WeightMatrix dim1() value. A crafted NT_LSTM layer can make ns_ much larger than na_, causing a heap out-of-bounds write during the first recognition step on the default LSTM engine and resulting in heap corruption, a crash, or potentially controlled corruption. No fixed release is available as of this review.

