CVE-2026-88047
HighCVSS 8.6Summary
Tesseract 5.5.3 and earlier, in Classify::ReadNormProtos, extracts a whitespace-delimited token from a .traineddata file into a fixed 61-byte stack buffer without setting stream width. A token longer than 60 characters writes up to 39 attacker-controlled bytes past the buffer during TessBaseAPI::Init of the legacy engine. No fixed release is available.
Risk Assessment
This can cause stack corruption, denial of service, and potentially control-flow hijacking. Typical libstdc++ builds are affected, while Apple libc++ C++20 bounded array builds are incidentally protected.
Recommendation
Do not load untrusted .traineddata files and restrict access to the tessdata directory. Track Tesseract releases and apply the fix as soon as it is available.
Other vulnerabilities in Tesseract
See all- CVE-2026-88050Medium
In Tesseract version 5.5.3 and earlier, RecodedCharID::DeSerialize in src/ccutil/unicharcompress.h validates length_ but accepts negative code_ values from a crafted .traineddata recoder component. UnicharCompress::ComputeCodeRange can consequently produce code_range_ equal to zero, after which SetupDecoder indexes is_valid_start_ with the negative code on a size-zero vector, causing an out-of-bounds bit write and a crash or allocation failure. No fixed release is available.
- CVE-2026-88049High
Tesseract 5.5.3 and earlier fails to validate dimension consistency in NetworkIO::WriteTimeStepPart and NetworkIO::AddTimeStepPart. A crafted .traineddata file with an NT_LSTM layer can cause a heap out-of-bounds write during the first recognition step on the default LSTM engine. No fixed release is available.
- CVE-2026-88048High
Tesseract 5.5.3 and earlier does not validate the deserialized scalars ni_ and no_ against weight-matrix dimensions in FullyConnected::DeSerialize. A crafted .traineddata file with an NT_SOFTMAX layer can cause a heap out-of-bounds write and read on the default LSTM engine. No fixed release is available.
- CVE-2026-73067Medium
Tesseract is an open source OCR engine. Prior to 5.5.3, a crafted .traineddata model loaded through TessBaseAPI::Init can cause SquishedDawg::read_squished_dawg in src/dict/dawg.cpp to accept an unterminated forward-edge run, after which SquishedDawg::Load calls num_forward_edges(0) and last_edge in src/dict/dawg.h reads beyond edges_, causing a heap out-of-bounds read and process crash before image processing. This issue is fixed in version 5.5.3.
- CVE-2026-73066Medium
Tesseract is an open source OCR engine. Prior to 5.5.3, a crafted .traineddata LSTM model component loaded through Tesseract's deserializer can cause an unchecked signed integer multiplication in Convolve::DeSerialize in src/lstm/convolve.cpp to wrap the convolution output-channel count, undersizing the forward-pass output buffer while writes use the unwrapped element count and causing a heap out-of-bounds write during OCR recognition. This issue is fixed in version 5.5.3.
Original NVD description (English source)
Tesseract is an open source OCR engine. In version 5.5.3 and earlier, Classify::ReadNormProtos in src/classify/normmatch.cpp parses the NORMPROTO component of a .traineddata file and uses std::istream::operator>>(char*) to extract a whitespace-delimited token into a fixed 61-byte stack buffer without setting a stream width. The 100-byte line buffer can carry a token of up to 99 characters, so a token longer than 60 characters writes up to 39 attacker-controlled bytes past the buffer during TessBaseAPI::Init of the legacy engine, causing stack corruption, denial of service, and potentially control-flow hijacking on affected standard-library implementations. Builds using Apple's libc++ C++20 bounded array overload are incidentally protected, while typical libstdc++ builds remain affected. No fixed release is available as of this review.

