AI engineer developing document recognition systems at a workstation overlooking Tehran Back to experience

AI Engineer

Roshan

Jul 2022 — Mar 2025

Tehran-Iran

Key responsibilities

Advanced OCR Development

Spearheaded the development of the "Alefba" Persian OCR system. Improved handwriting recognition, Arabic diacritics support, and robustness against noise and complex layouts using Attention mechanisms and synthetic data generation.

Advanced Training Techniques

Implemented unsupervised representation learning for segmentation via contrastive and reconstruction methods. Leveraged Teacher-Student pseudo-labeling with large open-source models to continuously bootstrap and expand training datasets.

Model Optimization & Training

Maximized GPU utilization and accelerated training speeds by implementing TAR dataloaders, dynamic learning rates (OneCycleLR, ReduceLROnPlateau), Label Smoothing, and mixed augmentation strategies (random padding, cropping).

Document Layout Analysis

Enhanced structural extraction by replacing heuristic methods with DeepLabV3 segmentation, transitioning towards polygon-based detection, and developing advanced logic for dynamic newspaper column parsing and table extraction.

Algorithmic Problem Solving

Resolved LSTM memory limitations for processing long/continuous text lines by engineering a fixed-chunking and overlapping algorithm combined with CTC loss.

Production Scalability

Improved system resilience by redesigning the document queuing architecture to support fair dynamic resource allocation for long vs. short documents, and implemented Triton inference server health-checks to prevent request failures during crashes.

Inference Acceleration

Accelerated inference speeds using TensorRT and implemented inference on video files via frame-by-frame extraction.