TymblHub

© 2026 TymblHub

Machine Learning Engineer - Evaluation

Hackerrank
Posted on
Hackerrank logo

Experience
1 - 4 yrs
Job Location
Bengaluru, India
Vacancy
1
Designation
Machine Learning Engineer
Job Type
ONSITE

Job Description

HackerRank sits at the center of this problem with a rare combination of scale, longitudinal data, and direct relationships with the companies making hiring decisions. The opportunity here is to define what rigorous, fair, and meaningful skill evaluation looks like in the agentic era. That methodology does not exist yet. This role exists to build it.
What you will do
  • Build LLM-powered evaluation pipelines that assess AI usage skills consistently, fairly, and at production scale
  • Own the evaluation methodology end to end: what the rubric is, how the model applies it, how you measure whether it is being applied correctly, and how you audit for bias
  • Design and run experiments to determine what good evaluation actually looks like. The answer is not known. You will be finding it
  • Build RAG pipelines and fine-tuning workflows that make evaluation models adhere reliably to the rules we set for them
  • Define the benchmarking infrastructure: how we know when our evaluation quality has improved, and how we catch regressions before candidates do
  • Translate model behavior into outcomes that product managers, enterprise customers, and candidates can understand and trust
Who you are
  • You have shipped LLM-powered systems in production where consistency and reliability were hard constraints, not nice-to-haves
  • You think as rigorously about how you measure your model as about the model itself. A poorly constructed eval is a worse outcome than a weaker model
  • You have a research mindset. You are comfortable operating in a space where the right methodology does not exist yet and needs to be invented
  • You think in systems. The data pipeline, the model, the serving layer, and the rubric it enforces are one problem to you
  • You can defend ML judgment in plain language to people who are not ML engineers, because the translation layer is part of the job
Even better if you have
  • Experience building evaluation frameworks for generative or conversational AI systems
  • Background in educational assessment, psychometrics, or human-in-the-loop evaluation at scale
  • Publications or open-source contributions in LLM evaluation, benchmarking, or alignment
  • Prior work at the interface of research and product, where you had to ship science, not just publish it
You will thrive here if
  • You find the measurement problem as interesting as the model problem, maybe more interesting.
  • You hold evaluation methodology to the same standard as model performance, and you are uncomfortable shipping something you cannot explain.
  • You want your work to define what good looks like in a field that is just now figuring that out.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.