At Extreme Compression, OCR Accuracy Can Hide Contextual Guessing
DeepSeek-OCR turns documents into compressed visual tokens and then reconstructs their text, a design DeepSeek presents as a way to process long documents with less context and computing cost. In his account of the paper, Two Minute Papers presenter Károly Zsolnai-Fehér highlights both the system’s strong OCR benchmark scores at high compression and a limit to what those scores prove: the decoder may infer missing words from context rather than recover them from the page. The paper also presents the pipeline as a way to generate training data at scale.

DeepSeek trades document detail for a smaller visual representation
The longer an AI system’s conversation or attached documents become, the more information it has to keep in context—and the more expensive it is to run. Károly Zsolnai-Fehér describes DeepSeek-OCR’s response as a form of artificial forgetting: turn a document into an image, compress that image into a small number of vision tokens, then reconstruct text from those tokens.
The analogy to JPEG is useful but incomplete. JPEG compresses pixels. DeepSeek-OCR first represents a document visually, then compresses that representation into tokens the model can process. The goal is to use substantially less space while recovering enough text to remain useful.
The benchmark chart shown in the presentation reports results on the Fox benchmark with 64 vision tokens per page. At 10.5× compression, precision is 96.5%; at 11.8× it is 93.8%. As compression increases, the results vary rather than declining in a straight line: precision is 83.3% at 13.2× and 85.8% at 15.1×, before falling to 79.3% at 16.5×, 70.3% at 17.7× and 59.1% at 19.7×.
Those figures show that the system can produce text that scores well against the benchmark even when its visual representation is much smaller than the original page. They do not establish whether each output word was read from the compressed visual evidence or inferred from context. That distinction becomes important at the most aggressive compression levels.
Local processing and compression make high-resolution pages manageable
DeepSeek-OCR is formally an OCR system: it takes a document—including, in the source’s description, a handwritten note—and produces text. Its unusual feature is the route it takes: a relatively small number of compressed vision tokens are used to generate a much larger number of text tokens.
The flowchart shows a document divided into patches, then processed through several stages before reaching the DeepSeek-3B decoder. First, a SAM ViTDet component uses local attention. Next, a convolutional stage downsamples by a factor of 16. CLIP then performs global attention, with an embedding layer connecting the visual representation to the decoder.
The order is central to the design. Rather than applying full attention to every pixel at the outset, the encoder processes local regions, compresses them, and then handles the document globally. Zsolnai-Fehér characterizes the process as skimming, squeezing, and only then reading. The local stage and downsampling reduce the representation before the global visual processing and text generation.
This is why the method is more than JPEG compression plugged into an AI system. The compression is learned and token-based, and the encoder is engineered to handle high-resolution documents without the cost of treating every pixel with full attention. The resulting tokens are not simply a smaller copy of the page; they are the input from which the decoder generates text.
At extreme compression, recognition and inference can blur
The paper itself says OCR alone is insufficient to validate “true context optical compression.” A stress test cited in the presentation raises a more specific concern: when visual tokens are severely limited, the decoder may use language priors to fill gaps, guessing missing or ambiguous text from context rather than visual evidence.
If the compressed image no longer carries a word clearly, a decoder may still produce a plausible one. The output can look like successful recovery even when the model has inferred what is likely to be on the page. Zsolnai-Fehér compares this to autocomplete, “but better.”
That concern does not erase the benchmark results. It changes what they can tell us. Text accuracy measures how well the output matches the expected text; by itself, it does not distinguish visual recognition from contextual completion. The limitation is therefore not simply that compression loses information. It is that a language-capable decoder can make the loss harder to see in the final text.
The presentation does not treat this as proof that the system is unusable. Rather, it marks a boundary on the claim: strong OCR scores at high compression are evidence of effective text production, but they do not, alone, show how much of the page’s content survives as visual evidence.
The pipeline is also presented as a way to generate training data
The same document-processing pipeline is presented as useful beyond retaining long documents more cheaply. According to the paper as summarized by Zsolnai-Fehér, DeepSeek-OCR can generate training data for language models and vision-language models at more than 200,000 pages per day on a single A100 GPU.
That throughput gives the approach a second role: processing documents to produce material that can help train other models. It follows from the same effort to make document handling cheaper, though it is distinct from proving that every generated transcription faithfully reflects visual evidence. The source’s account holds both points together: the system can produce text and training data at substantial scale, while severe compression may leave the decoder relying on contextual guesses.
