AI Engineer (AI labs)

Weekday AI

This role is for one of Weekday’s clients
Salary range: Rs 3000000 - Rs 20000000 (ie INR 30 - 200 LPA)


Min Experience: 3+ years
Location: Bengaluru, Karnataka, India
JobType: full-time

This role involves building and improving AI agents that generate and validate real-world tasks. You will be responsible for enhancing these agents using traces and evaluation results, turning enterprise workflows into executable tasks for AI training and evaluation, and ensuring their quality and behavior as frontier models evolve.

Requirements

Key Responsibilities

  • Build and improve agents, prompts, tool interfaces, context handling, and orchestration across task generation and validation.
  • Inspect model calls, tool use, trajectories, screenshots, state changes, and grader outputs for trace-driven debugging.
  • Turn recurring failures into targeted code and agent changes.
  • Review pipeline outputs, form hypotheses, implement fixes, and test on fresh runs and regression suites before rolling out improvements.
  • Choose and instrument tracing, debugging, and experiment tools.
  • Track quality, latency, token usage, and cost by agent and pipeline stage to improve model and tool choices.
  • Improve validation logic and reward criteria for task and grader quality, ensuring shipped tasks are solvable and evaluations reflect genuine completion.
  • Partner with Platform Engineers to turn agent improvements into modular components that can be evaluated and released reliably.
  • Own a workflow from output review through debugging, evaluation, and release to find the actual failure.
  • Build repeatable checks to distinguish agent failures from task or grader defects, making evaluation signals trustworthy.
  • Fix recurring failure patterns and prove gains across task families, fresh batches, and regression suites.
  • Establish a repeatable way to compare quality, latency, and cost across versions for fast iteration and measurable improvements.

Qualifications

  • Strong software engineering fundamentals and hands-on Python experience.
  • Ability to read unfamiliar code and debug a system across component boundaries.
  • Experience building LLM agents or multi-step AI workflows with tool use, structured outputs, and state management.
  • Comfort designing evaluations and using traces and data to distinguish a real improvement from a lucky run.
  • Demonstrated ownership through implementation and verification, with clear communication about failures, tradeoffs, and evidence-supported decisions.
  • Useful additional experience includes computer-use agents, RL environments, reward design, browser automation, or tracing and evaluation platforms.

Must-have skills

LLM Agents, Multi-step AI workflows, Reinforced Learning environments

Good-to-have skills

Browser automation, State management, Reward design

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.