Skip to content
#

arabic-ocr

Here are 45 public repositories matching this topic...

This research aims to fine-tune an Arabic OCR model using Tesseract 5.0, enhancing text recognition accuracy through extensive data collection, preprocessing, and image generation. By leveraging advanced training techniques and data augmentation, we achieve significant improvements in word error rates (WER).

  • Updated Apr 4, 2025
  • Jupyter Notebook

The largest publicly released line-level dataset of historical Arabic manuscripts — 14 books, 3,043 pages, 28,600 lines, with margin/insertion-anchor annotations for non-linear reading order.

  • Updated Sep 1, 2026
  • Python

Optical Character Recognition, OCR pipeline, Arabic OCR, Deep Learning OCR, Computer Vision text extraction, Text recognition system, AI document processing, Multilingual OCR, Transformer OCR, OCR benchmarking, Bounding box detection, Ground truth evaluation.

  • Updated May 20, 2026
  • Python

Add this topic to your repo

To associate your repository with the arabic-ocr topic, visit your repo's landing page and select "manage topics."

Learn more