TymblHub

Β© 2026 TymblHub

AI Safety & Evaluation Lead

Smart Source
Posted on
Smart Source logo

Experience
4 - 9 yrs
Salary (CTC)
β‚Ή40L - β‚Ή55L
Job Location
Noida, India
Vacancy
1
Designation
Safety Lead
Job Type
Not specified

Job Description

Key Responsibilities

  • Design and own the AI evaluation framework: benchmark construction, automated and human-in-the-loop eval pipelines.
  • Lead red-teaming and adversarial testing across deployed models: prompt injection, jailbreaking,

edge case identification.

  • Define AI safety standards, responsible AI policies, and child-safe content guardrails for

deployment at scale.

  • Build and maintain evaluation tooling: RAGAS, LangSmith, PromptBench, EleutherAI lmevaluation-

harness.

  • Conduct systematic bias detection and fairness measurement across outputs, especially for

regional and linguistic variation.

  • Manage AI risk assessments and translate safety findings into actionable model or prompt

changes.

  • Establish AI governance documentation: model cards, audit trails, safety reports for internal and

regulatory stakeholders.

  • Embed evaluation and safety checkpoints into the model release pipeline in collaboration with

engineering.

  • Track and apply evolving regulatory requirements: EU AI Act, NIST AI RMF, India AI policy

frameworks.

  • Measure and report model accuracy, grounding fidelity, and response quality across deployed AI

systems through systematic evaluation.

  • Build and maintain knowledge of AI governance and compliance standards, operationalising

responsible AI practices into daily engineering workflows.

Must-Have Skills

  • Strong experience in AI/LLM safety, evaluation, and risk assessment
  • Expertise in designing and executing model evaluation frameworks and test methodologies
  • Experience in red-teaming, adversarial testing, and identifying model vulnerabilities
  • Strong understanding of bias, fairness, hallucination, and content safety in AI systems
  • Experience measuring model accuracy, grounding, and response quality
  • Knowledge of AI governance, compliance, and responsible AI practices
  • Familiarity with frameworks such as NIST AI RMF and AI regulatory standards; understanding of student psychology and age-appropriate AI output design
  • Strong analytical and problem-solving skills with experience working on production AI systems