# Optical Mark Recognition (OMR) Evaluation Pipeline
A robust, computer vision-driven Optical Mark Recognition (OMR) pipeline built with Python and OpenCV. This system processes batch images of bubble sheets, corrects perspective distortions using spatial anchor detection, extracts filled bubble data, evaluates student answers against a master configuration file, and routes graded records into segmented data sheets based on department verification.
---
## Key Features
* **Intelligent Document Alignment:** Utilizes contour analysis, Canny edge detection, and bilinear interpolation to correct skewed or rotated input sheets.
* **Double-Mark Detection:** Robust bubble-fill algorithms detect if a student marks more than one option per question, automatically flagging the response as incorrect to prevent grading ambiguities.
* **Inline Dynamic Evaluation:** Automatically parses, scores, and tracks correct answers, incorrect selections, unattempted questions, and final marks against a centralized master JSON configuration.
* **Automated Data Routing:** Automatically filters records based on allowed department listings, generating separate structured data dumps (`valid_department.csv` and `invalid_department.csv`) instantly.
* **Fault-Tolerant Batch Tracking:** Includes high-precision execution timers tracking per-image and entire batch run performance metrics, while isolating failing files cleanly inside an overwritten runtime log.
---
## Directory Structure
```text
├── config.py # Operational thresholds, anchor ratios, and layout specifications
├── main.py # Pipeline execution engine and batch supervisor
├── failed.json # Freshly overwritten log documenting execution failure exceptions
├── input/ # Target folder for raw OMR input sheets
├── models/
│ └── valid.json # Master verification exam answer key and validation criteria
└── pipeline/
├── crop_sheet.py # Perspective warp and binary preprocessing operations
├── anchor_detection.py # Corner point geometric tracking algorithms
├── grid_mapper.py # Sub-grid extraction arrays and coordinate mapping
├── extractor.py # Pixel density calculation and bubble recognition
└── evaluator.py # Core evaluation metrics, scoring engine, and CSV output tools
Ensure you have your virtual environment activated, then install the mandatory runtime packages:
pip install -r requirements.txt
The pipeline relies on a master criteria blueprint located at models/valid.json. Ensure your configuration structure matches the pattern below:
{
"metadata": {
"exam_name": "Departmental Entrance & Evaluation Test",
"total_questions": 20,
"passing_marks": 10
},
"valid_departments": [
"Computer Engineering",
"Information Technology",
"Electronics & Telecommunication",
"PYTHON"
],
"scoring_rules": {
"correct_marks": 5,
"negative_marks": 0,
"unattempted_marks": 0
},
"answer_key": {
"1": "A",
"2": "B",
"3": "C"
}
}
Note: Ensure metadata.total_questions exactly matches the highest sequence number mapped inside the answer_key dictionary to prevent evaluation truncation bugs.
Drop your target image assets into your designated scan folder (defaulting to the input directory) and run the primary batch pipeline execution command:
python main.py
- Geometric Registration: The sheet undergoes a spatial search tracking defined anchors. If obscured, it triggers adaptive threshold steps alongside bilateral smoothing filters to locate corners.
- Coordinate Matrix Interpolation: The document maps individual bubble groups using linear ratios specified inside the layout variables.
- Density Analysis: OpenCV pixel counters evaluate the darkness value inside individual bubble boundaries.
- Grading & Segregation: Data points are sent to the evaluator, verified against the permitted department string set, scored, and saved to the disk.
Upon completing execution, the pipeline populates data fields within the following system outputs:
output/valid_department.csv: Contains full metrics of processed individuals belonging to a certified department found insidevalid.json.output/invalid_department.csv: Isolates individuals whose documents processed cleanly but carried an unlisted or unrecognized department mapping.
Both files contain the following structured layout header elements:
testid, name, contact, department, status, correct, incorrect, unanswered, score
failed.json: Overwritten from scratch on every run. If an image features an unreadable bubble profile, invalid orientation, or broken anchor setup, it catalogs the file signature along with the stack trace exception reason.