Senior CV Engineer
ZestIoT
Hyderabad, Telangana, India
Job Description
Role: AI / Computer Vision Engineer – Flight Turnaround Monitoring System
Key Responsibilities
- Model development: Design, train and optimise deep learning models for object detection, multi-object tracking, segmentation and video-based activity recognition on airport apron footage.
- Activity-detection pipeline: Build the end-to-end pipeline that converts detections and tracks into turnaround events – combining detectors, trackers, stand-zone (ROI) geometry, temporal action models and a rule/state-machine layer with robust start/end timestamping.
- Accuracy ownership: Define the evaluation protocol with the product team, build and maintain a golden test set, and drive per-activity precision, recall and timestamp accuracy to the agreed targets.
- Dataset engineering: Own the data strategy – camera-view coverage, annotation guidelines and QA, active learning, hard-negative mining, class-imbalance handling and synthetic/augmented data for rare or safety-critical events.
- Robustness: Make models reliable across night, rain, fog, glare, shadows, heavy occlusion by aircraft and vehicles, small/distant objects, differing camera angles and different airports or stands.
- Error analysis: Run systematic failure analysis (by activity, camera, time of day and weather) and convert findings into data, model or logic improvements in fast iteration cycles.
- Production inference: Optimise multi-stream, low-latency inference with GPU acceleration (TensorRT / ONNX Runtime / DeepStream) on data-centre or edge hardware.
- MLOps: Set up experiment tracking, dataset and model versioning, automated evaluation, shadow-mode validation, drift monitoring and a retraining loop with human-in-the-loop correction.
- Integration: Work with backend, product and operations teams to publish activity events via Redis / RabbitMQ / SQL to dashboards, alerts and turnaround-time (TAT) analytics.
- Leadership: Mentor junior engineers, contribute to architecture decisions, technical reviews and roadmap planning, and track emerging research in video understanding and edge AI. Accuracy & Performance Expectations Targets below are indicative and will be finalised per activity during the pilot. The engineer is expected to define how they are measured and to reach them in production conditions, not only on a held-out dataset.
- Event-level detection: Target ≥ 95% precision and ≥ 90–95% recall per activity (event level F1), measured against a manually annotated golden set; safety-critical activities (e.g. refuelling, chocks, pushback) held to the higher bar.
- Timestamp accuracy: Start/end times of activities within an agreed tolerance (e.g. ± 5–10 seconds) of ground truth; reported using temporal IoU and mean absolute timestamp error.
- Object detection / tracking: Strong mAP on apron object classes (aircraft, GSE, personnel, doors, cones, chocks, hoses) with stable track IDs through occlusion (high IDF1 / low ID-switch rate).
- Robustness: Performance reported per slice (day/night, weather, camera, stand) with no slice falling materially below the overall target.
- Operational: Real-time processing across all configured camera streams per GPU, with a false-alarm rate that operations teams can act on.
Required Skills and Qualifications Core computer vision and deep learning
- Expert-level Python with strong software engineering practices (testing, packaging, code review).
- Strong foundation in machine learning, neural networks and image/video processing.
- Extensive hands-on experience with PyTorch for training, fine-tuning and deploying models.
- Proficiency in OpenCV and video I/O (FFmpeg, GStreamer, RTSP/H.264/H.265 streams).
- Practical experience with modern detectors and segmenters (YOLO family, RT-DETR / DETR, Faster R-CNN, Mask R-CNN, U-Net) including fine-tuning on custom datasets. Video understanding and activity recognition (critical for this role)
- Multi-object tracking: ByteTrack, BoT-SORT, DeepSORT or similar, including re identification and handling of long occlusions.
- Action recognition / temporal action detection: SlowFast, X3D, TSM, TimeSformer, VideoMAE, Video Swin or similar; experience with temporal modelling (LSTM / GRU / Transformers / temporal convolution) on top of frame-level features.
- Spatio-temporal reasoning: zone/ROI-based logic, object–object and object–person interaction (e.g. vehicle docked at aircraft door, hose connected), and rule or state-machine event engines.
- Pose estimation and small-object detection techniques (tiling / SAHI, high-resolution inference, multi-scale features). Data, training and accuracy engineering
- Building annotation pipelines with CVAT / Label Studio or similar; writing labelling guidelines and running inter-annotator QA.
- Handling class imbalance and rare events: hard-negative mining, active learning, focal / weighted losses, targeted augmentation, semi-/self-supervised learning and synthetic data.
- Rigorous evaluation: mAP, precision/recall/F1, confusion matrices, temporal IoU, IDF1/MOTA, threshold and confidence calibration, and sliced analysis by lighting, weather and camera view.
- Robustness to domain shift: day/night, rain, fog, glare, camera placement and cross-site generalisation; low-light enhancement and, where applicable, thermal/IR imagery.
- Experiment tracking and reproducibility: MLflow or Weights & Biases, DVC or equivalent for dataset/model versioning. Deployment and integration
- Real-time inference optimisation with ONNX Runtime and TensorRT (FP16/INT8 quantisation, batching); NVIDIA DeepStream or Triton Inference Server is a strong plus.
- Experience running multi-stream video analytics on GPU servers and/or edge devices (e.g. NVIDIA Jetson).
- Proficiency in Docker and CUDA for containerised, GPU-accelerated deployment.
- Experience integrating with Redis, RabbitMQ (or Kafka) and SQL databases for event streaming and storage.
- Strong command of Linux, including scripting, debugging and performance profiling. Experience and soft skills
- 2+ years in applied computer vision, with at least 2 years delivering video-based detection or activity-recognition systems to production.
- Proven record of taking a model from a prototype to meeting a measured accuracy target in real-world conditions.
- Ability to solve complex problems independently, lead technical initiatives and mentor others.
- Clear communication for cross-functional collaboration and technical documentation. Good to Have
- Prior work in aviation, airport operations, logistics, industrial safety or other CCTV-based activity monitoring.
- Experience with vision-language / foundation models (e.g. CLIP, SAM, Grounding DINO, VLMs) for auto-labelling or open-vocabulary detection.
- Knowledge of A-CDM / turnaround milestone standards and airport ground-handling processes.
- Experience with multi-camera calibration, camera-to-stand mapping, and model compression (pruning, distillation).
- Publications or open-source contributions in video understanding or object tracking.