Senior AI/ML Engineer – Personalization, Bandits & AI Agents

OWOW
India

We’re a founder-led AI company building an autonomous growth platform that decides, for every customer, which offer, message, channel and timing will work best, and learns from the results. We’re looking for a hands-on senior engineer who has built these decision systems in production.

Location: Remote (3–4 hrs daily overlap with US Pacific time)

Experience: 6+ years of software or ML engineering

Reports to: Founder & CEO

You must have:

  • Built and run a production system that chooses between options for each user and learns from outcomes, using multi-armed bandits or reinforcement learning on live traffic (e.g., recommendations, ranking, ad or offer selection, pricing, or send-time/channel optimization)
  • 3+ years shipping LLM features or AI agents to real users (not proofs of concept)
  • 6+ years of hands-on engineering, excluding internships, teaching and study periods
  • Writing code most of the week today, with experience as the technical lead on a product
  • Strong Python (TypeScript is a plus)

This role is not a fit if your RL experience is only:

RLHF/DPO fine-tuning, A/B testing of prompts or models, building RL environments or training data for AI labs, or coursework/research projects.

You’ll likely be a great fit if you’ve worked as:

Applied Scientist, ML Engineer (Personalization, Recommendations, Ranking, Ads) or Decision Scientist at a consumer product company in e-commerce, food delivery, ride-hailing, fintech, gaming, adtech or martech.

What you’ll do:

  • Own the bandit/RL decision loops at the core of our platform, including reward design and evaluation
  • Build and ship AI agents that plan and optimize campaigns
  • Write most of the critical code yourself while leading architecture

Why join us: Small team, direct work with our CEO, and real ownership of the product.

Score my resume against this job, free →

Get your ATS score for this role — free. Score my resume free →