← Back to feed
2026-07-22visiondatacode

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Pouria Mahdi, Haq Nawaz Malik

PDF preview for Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Read on arXiv →

Key claim

Persian Pixel addresses OCR challenges for Persian script.

In plain English

Imagine you're trying to build an OCR system that can read Persian documents, which is particularly challenging due to the unique characteristics of the Perso-Arabic script. Current OCR solutions struggle with this because they often rely on limited datasets that don't capture the complexities of Persian writing, such as cursive connections and various glyph shapes. This leads to failures in recognizing text accurately, especially in diverse contexts and styles, which is what's called a data bottleneck. Without enough high-quality annotated data, developing effective OCR systems for Persian is a slow and costly process. To tackle this issue, the authors created Persian Pixel, a synthetic dataset that includes over 343,000 image-text pairs generated from a large Persian corpus. This dataset not only simulates the intricacies of Persian script but also incorporates realistic degradation models to mimic real-world document conditions. By providing a scalable and openly available resource, Persian Pixel enables the training of modern OCR architectures, potentially accelerating advancements in Persian document analysis and digitization. This represents a significant step forward compared to previous efforts, as it offers a practical solution to the data scarcity problem that has hindered progress in this area.

Novelty
8.0/10

The introduction of a synthetic dataset specifically for Persian OCR addresses a significant gap in the field.

Reliability
7.5/10

The dataset is large and well-constructed, but its effectiveness in real-world applications needs further validation.

Deep reliability assessment

The methodology supports the creation of a large-scale synthetic dataset for Persian OCR, addressing data scarcity and complexity challenges. However, the effectiveness of synthetic data in fully bridging the domain gap with real-world data may be overclaimed.

Reproducibility

Yes, the dataset is publicly available at https://huggingface.co/datasets/Omarrran/Persian_Pixel.

Key figure

The paper does not provide a specific figure or architectural diagram description.

Codelink
huggingface.co/datasets/Omarrran/Persian_PixelOfficial
Persian Pixel: A large-scale synthetic OCR dataset for Persian language — Frontier Papers