Computer Vision Engineer
Company Description Kavach X is an AI-powered security, automation, and smart infrastructure platform transforming how homes, businesses, and enterprises manage safety and connected technologies. The company is building a unified, technology-first ecosystem that integrates physical security, intelligent automation, customer experience, and digital service management. Its platform is being developed to support CCTV surveillance, building automation, access control, smart gate systems, video door communication, fire safety, networking, and AI-powered monitoring services. With a scalable SaaS-driven architecture and predictive maintenance capabilities, Kavach X serves residential, commercial, industrial, educational, hospitality, and healthcare environments. The culture emphasizes innovation, ownership, execution, and customer success.
Role Description This is a full-time Computer Vision Engineer role based in Ludhiana with a hybrid work arrangement, allowing some work from home. The Computer Vision Engineer will design, develop, and optimize computer vision algorithms for security, surveillance, and automation applications across Kavach X’s platform. Day-to-day responsibilities include building and training models for object detection, tracking, and pattern recognition, integrating these models with edge devices and cloud services, and collaborating with data science and robotics teams to deploy solutions in real-world environments. The role involves prototyping vision pipelines, evaluating model performance, improving accuracy and latency, and working closely with product and engineering stakeholders to align solutions with platform requirements. The engineer will also contribute to documentation, code reviews, and continuous improvement of the AI-powered monitoring and support ecosystem.
Qualifications
We are hiring an engineer to build an edge AI system that runs on NVIDIA Jetson devices and makes existing IP cameras searchable by voice.
The system reads RTSP streams from cameras of any brand, detects and tracks people and vehicles, recognises enrolled faces, and stores time-stamped events with snapshots and clips. Users ask questions in a mobile app, such as "Who came between 5 and 6 PM?", and get a spoken answer with the matching clips.
A working starter codebase exists. You will take it from prototype to a reliable product running at real customer sites.
Responsibilities
- Deploy and optimise detection and tracking models (YOLO, ByteTrack) on Jetson Orin with TensorRT
- Handle RTSP/ONVIF streams from many camera and NVR brands, including reconnects and hardware decoding
- Build and tune face recognition for enrolled people, keeping false matches low
- Improve the event pipeline: snapshots, clips, local storage, offline-first sync to cloud
- Build the cloud backend that receives events and clips from many sites
- Extend the natural-language query layer (English, Hindi, Hinglish) and integrate LLM APIs
- Work with the mobile app developer on APIs, clip playback and voice
- Test at live sites, measure accuracy and performance, and fix what breaks in the field
- Add new analytics over time: people counting, attendance, number plates, zone alerts
Must-Have Skills
- Strong Python; comfortable reading and extending an existing codebase
- Computer vision with OpenCV and object detection (YOLO or similar), including training or fine-tuning on custom data
- Multi-object tracking (ByteTrack, DeepSORT or similar)
- Model deployment on NVIDIA Jetson or other GPU edge devices: TensorRT, FP16/INT8, measuring FPS
- Video streams: RTSP, FFmpeg and/or GStreamer, H.264/H.265
- Linux, systemd, Docker, basic networking (IP cameras, LAN, VLANs)
- REST API design and a SQL database (SQLite/PostgreSQL)
- Git, and writing tests
Good-to-Have Skills
- NVIDIA DeepStream
- Face recognition with InsightFace/ArcFace, and setting match thresholds
- LLM APIs with function/tool calling
- Speech-to-text and text-to-speech, especially Hindi and Hinglish
- Cloud: AWS/GCP/Azure, S3-style storage, MQTT
- Number-plate recognition (ANPR) or people counting
- Open-source NVR/VMS such as Frigate
- Hands-on experience with CCTV installations or security products
- Privacy-aware design: encryption, data retention, consent
Experience & Role Details
Item
Requirement
Experience
3+ years in computer vision or ML engineering, with at least one project deployed on real cameras or edge devices
Education
B.Tech/M.Tech in CS, ECE or related field, or equivalent proven work
Type
Full-time (contract-to-hire also considered)
Location
[City] / hybrid; occasional site visits for testing
Languages
English for technical work; Hindi helpful for the voice layer
Compensation
[Range] — based on experience
All work is under NDA, and code and IP belong to the company.