Skip to main content

Mistral Launches OCR 4 for Document Intelligence and Enterprise RAG

24 JUNE 2026·2 MIN READ·5 SOURCES·Official source plus independent coverage

Mistral AI has released OCR 4, a document intelligence model featuring paragraph-level bounding boxes and structural block classification. The system supports fully self-hosted enterprise deployments and serves as an ingestion engine for retrieval-augmented generation pipelines.

Mistral Launches OCR 4 for Document Intelligence and Enterprise RAG

Key takeaways · 3

  • 01

    Mistral OCR 4 features native paragraph-level bounding box extraction and structural block classification.

  • 02

    The model supports single-container deployments for strict data residency and compliance requirements.

  • 03

    Developers can process documents at $4 per 1,000 pages or $5 per 1,000 annotated pages.

Release and Core Features

Mistral AI released OCR 4 on June 23, 2026, a small, focused model that runs in a single container for fully self-hosted deployments. [3] The service powers Mistral's Document AI stack with native paragraph-level bounding box extraction. [5] Alongside extracted text, the model returns bounding boxes, inline confidence scores, and block classification for elements including titles, tables, equations, and signatures. [3] Self-managed deployment is available to enterprise customers, keeping document data within their environments for residency, sovereignty, and compliance. [3] Bounding boxes localize text for in-context highlighting and reliable data pipelines. [3]

Performance and Multilingual Support

On public benchmarks, OCR 4 achieved the top overall score of 85.20 on OlmOCRBench. [3] Mistral states that independent annotators prefer OCR 4 over every leading OCR and document-AI system tested, with win rates averaging 72 percent. [3] The model supports 170 languages across 10 language groups, delivering measurable gains on specialized and low-resource languages where several competing systems degrade. [3] These block types and inline confidence scores drive source-grounded citations, redactions, and human-in-the-loop verification. [3] The structured output supplies citation-ready inputs to ingestion, retrieval, and evaluation workflows for RAG and enterprise search. [3]

Integration and Pricing

OCR 4 operates as an ingestion component of Search Toolkit, which is Mistral's open-source, composable search framework that was announced at the AI Now Summit. [3] For developers and enterprises, the OCR 4 model is priced at $4 per 1,000 pages, while the cost increases to $5 per 1,000 annotated pages. [5] Users can run OCR at scale via Mistral's Batch API and access official cookbooks for tasks like healthcare document processing, document chunking, and extracting data via annotations. [4] Additionally, the model supports cost-efficient, high-throughput batch processing. [3]

What it means

Mistral’s launch of OCR 4 emphasizes a clear transition from basic text extraction to citation-ready data pipelines tailored for enterprise search and RAG workflows. By offering self-managed, single-container deployments alongside native bounding box capabilities, the company addresses strict corporate requirements for data residency and compliance. The model’s competitive benchmark scores suggest it can challenge existing document intelligence solutions on complex document layouts. The pricing model separates standard OCR from structural annotations, giving developers cost flexibility based on specific operational needs. What the sources don't address: How OCR 4’s latency scales under heavy concurrent request loads in large enterprise environments.

Mistral OCR 4 provides developers with a dedicated, self-hostable tool for structured document ingestion in RAG pipelines. Its integrated bounding box and block classification capabilities streamline the creation of citation-grounded applications.

Why it matters
Story quiz

Test yourself on this story — 4 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 24 June 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.