Bengaluru
Location: Bengaluru (In-person only)
Duration: 6-months (Jul-Dec 2026)
Stipend: 40-50K per month
Full Time, Mon-Sat (6 days a week)
We’re building a horizontal voice-agent platform, so new use cases come up constantly, and each one needs a thoughtful way to measure quality. Your job is to turn business problems into evals: understand what “good” means for a use case, then design and build the evaluation that measures it. You’ll own the work that helps us move to cheaper and faster models without losing quality.
The day-to-day is practical and iterative: writing judge prompts, curating and sampling data, designing test plans, implementing them, reviewing results, and refining. It is detailed work, especially early on, and it matters because it shapes how reliably the product improves. If that kind of ownership energizes you, you’ll do well here.
Work with the team to understand each use case, define what quality means, and build evals that capture it.
Write and refine LLM-as-judge prompts and rubrics.
Curate datasets and design sampling strategies that evaluate models against real usage.
Build and run test pipelines, analyse results, and translate them into clear recommendations.
Keep iterating as use cases and models change.
Learns fast. New domains come up regularly, and you can ramp on them quickly.
Comfortable with careful iteration. Eval quality comes from focused, repetitive improvement. You find that satisfying, especially through the build phase.
Strong pull toward the business problem, not just the technical task — you ask “what is this use case actually trying to achieve?” before writing a metric.
Hands-on. Comfortable writing code to implement pipelines and clear in written communication.
Sharp and rigorous. Notices when a metric is measuring the wrong thing and pushes back constructively.
Bonus: voice/conversational AI exposure, Hindi-English code-switching familiarity, or experience with eval tooling (promptfoo, Langfuse, DeepEval, Inspect).
This may not be the right fit for someone seeking only novel or research work, or for someone who wants to avoid hands-on data and prompt iteration.