New OCR model processes complex layouts, tables, handwriting, and math across 100+ languages with frontier-grade accuracy at a fraction of the size and cost
Elastic (NYSE:ESTC) today announced the launch of jina-ocr-v1, a new optical character recognition (OCR) model for end-to-end document processing. At 574M active parameters, it delivers frontier-grade accuracy in a model roughly one tenth the size of the benchmark leader. Jina-ocr-v1 accurately converts complex visual documents into structured, machine-readable text, such as Markdown, in a single pass, making it easy to search, train models, and build agentic applications using the data from scanned documents.
While traditional OCR works well on clean text and simple layouts, complex documents with highly visual content often require separate processing steps, such as page segmentation, element classification, text recognition and reassembly. Each step introduces potential for errors that can accumulate through the fragile processing pipeline. When inaccurate or incomplete data is passed downstream, agents can return incomplete facts, and RAG pipelines can return answers that don't accurately reflect the source documents.
jina-ocr-v1 handles the entire process end to end in a single model. It uses a mixture-of-experts architecture with 3.4B total parameters and 574M active at inference, running at the speed and cost of a sub-600M model. jina-ocr-v1 also adds FastMTP technology, which improves multi-token prediction to accelerate inference.
Login to comment