TymblHub

© 2026 TymblHub

AI QA Engineer

QualityKiosk Technologies
Posted on
QualityKiosk Technologies logo

Experience
2 - 4 yrs
Salary (CTC)
₹8L - ₹12L
Job Location
Pune, India
Vacancy
1
Designation
QA Engineer
Job Type
ONSITE

Job Description

L1 AI Evaluation Analyst


Role Summary

We are looking for an AI Evaluation Analyst to design, execute, and analyze supervised evaluation workstreams for AI voice bots, chatbots, and LLM-based systems. The role requires hands-on involvement in test case creation, dataset preparation, AI response evaluation, defect analysis, and client-ready reporting.

Experience Required

2 to 4 years of total experience.

Relevant experience may include AI testing, chatbot testing, voice bot testing, QA automation, data analysis, NLP testing, prompt evaluation, or software quality engineering.

Educational Qualification

B.E. / B.Tech / M.Tech / M.S. in Artificial Intelligence, Computer Science, Computer Engineering, Information Technology, or related Computer discipline.

Key Responsibilities
Create AI evaluation test cases based on business flows, use cases, intents, personas, and expected outcomes.
Prepare evaluation datasets, golden test sets, regression packs, and prompt variations.
Execute AI evaluations for voice bots, chatbots, and LLM applications.
Analyze AI responses for hallucination, correctness, completeness, safety, policy compliance, tone, robustness, and fallback handling.
Operate AI evaluation platforms for test execution and result analysis.
Review LLM-as-judge outputs and perform human calibration where required.
Identify failure patterns and prepare structured defect reports.
Support evaluation coverage across multiple languages, personas, user journeys, and edge cases.
Prepare client-ready summaries, dashboards, and evaluation reports under senior guidance.
Participate in client discussions to understand evaluation scope and clarify test requirements.
Maintain traceability between use cases, test cases, datasets, defects, and evaluation outcomes
.
Required Knowledge and Skills
Strong understanding of software testing and QA concepts.
Working knowledge of AI, NLP, LLMs, chatbots, voice bots, and conversational AI systems.
Ability to design test cases for AI behavior, conversation flows, and business scenarios.
Hands-on experience with Python and SQL for data handling and analysis.
Ability to work with JSON, APIs, CSV files, and evaluation datasets.
Understanding of common AI failure modes such as hallucination, ambiguity handling, incorrect intent detection, unsafe response, and inconsistent answers.
Good analytical and problem-solving skills.
Strong written communication and reporting skills.
Ability to work with clients and internal stakeholders in a structured manner.

Preferred Skills
Experience with prompt evaluation platforms, annotation tools, or AI testing tools.
Exposure to LLM-as-judge evaluation methods.
Experience in voice bot testing, IVR testing, ASR/TTS validation, or contact center testing.
Exposure to regulated domains such as healthcare, banking, insurance, or automotive.
Experience preparing test coverage matrices and evaluation dashboards.

No Referrers Available

There are currently no referrers available for this job. You can still apply, will let you know once there is any referrer available.