
WHAT YOU'LL DO - Own the design, training, and deployment of computer vision models for hygiene and food safety detection: PPE compliance, handwashing and food-handling behaviours, cross-contamination risk, and related tasks - Improve accuracy and reduce false positives on real production footage, where the failure modes are messy and the data is imperfect - Build and improve the data pipeline: collection strategy, annotation standards and QA, dataset versioning, and evaluation methodology - Take models from prototype to production on the edge — optimisation, quantisation, and deployment to on-site devices and on-prem servers, with hard constraints on latency, cost, and reliability - Design evaluation frameworks that reflect real operational risk, not just benchmark metrics - Set the technical direction for the CV stack and raise the engineering standard of the team - Mentor junior engineers and contribute to hiring as the team grows WHAT WE'RE LOOKING FOR REQUIRED - 5+ years of professional experience in computer vision or applied deep learning, with meaningful time spent on systems that ran in production - Strong command of modern CV: object detection, tracking, segmentation, and video/temporal action recognition - Hands-on experience with edge deployment and inference optimisation — TensorRT, ONNX Runtime, OpenVINO, quantisation and pruning — on constrained hardware such as NVIDIA Jetson (Orin/Xavier/Nano), Coral, or equivalent. You have shipped models that had to run within a real latency and power budget, not only in the cloud. - Direct experience with human activity or behaviour recognition, pose estimation, or multi-object tracking — for example SlowFast, X3D, VideoMAE, TimeSformer for action recognition; MediaPipe, OpenPose, MMPose, YOLO-Pose, ViTPose for pose; ByteTrack, BoT-SORT, DeepSORT, or StrongSORT for tracking - Practical depth with modern detection frameworks and tooling: the YOLO family (Ultralytics YOLOv8/v11, YOLO-World), RT-DETR, DETR/Deformable DETR, Detectron2, MMDetection, Grounding DINO or SAM for open-vocabulary and segmentation work - Deep practical fluency in Python and PyTorch (or TensorFlow), plus OpenCV, and the discipline that comes from debugging models that failed in the field - Comfort with video ingestion and streaming pipelines — GStreamer, FFmpeg, RTSP, DeepStream — not just working from static datasets - Demonstrated experience shipping models to production — you have owned the path from dataset to deployed inference, not just the training script - Experience working with real-world video and imperfect data: low light, occlusion, motion blur, camera variation, class imbalance - Sound judgment on data strategy: what to label, how to label it, how to evaluate, and when more data is not the answer - Clear technical communication in English, including with non-technical stakeholders STRONGLY PREFERRED - Experience in a safety-critical, regulated, or compliance-driven domain — food safety, industrial safety, healthcare, or similar - Cloud pipeline experience — building and operating the backend behind a vision product: AWS/GCP/Azure, containerised inference services (Docker, Kubernetes), event and message queues (Kafka, SQS, MQTT), object storage, and video/data pipelines at scale - Edge-to-cloud architecture and optimisation — deciding what runs on-device versus in the cloud, and designing for the trade-off: bandwidth and storage cost, selective upload and event-triggered clips, intermittent connectivity and store-and-forward, device fleet management and OTA model updates, and continuous retraining loops that feed edge data back into cloud training - Multi-camera work: calibration, re-identification, and tracking people across overlapping or disjoint views - Experience training custom detectors on self-collected data, including annotation tooling (CVAT, Label Studio, Roboflow) and active learning loops - Exposure to MLOps practice: experiment tracking, reproducible training, monitoring models after deployment - Early-stage or small-team experience, where scope is wide and ambiguity is normal - Arabic language ability (useful, not required)
Posted
Apply