Skip to content Skip to footer

Data Scientist – LLM Evaluation & AI Quality

Location: Hyderabad
Work Mode: Work from Office (WFO)
Experience: 5+ Years
Employment Type: Full-Time

ABOUT NSTARX

NStarX is an AI-first, Cloud-first engineering services company focused on AI, ML, Cloud, Automation, and Software Engineering. We help organizations accelerate business value through modern technology and engineering solutions.

ABOUT THE ROLE

We are looking for a Data Scientist – LLM Evaluation & AI Quality to define, measure, and improve the quality and reliability of AI-generated outputs. This is an evaluation and quality-focused role, involving LLM evaluation, RAG quality assessment, evidence ranking, and fail-closed validation rather than traditional model development.

KEY RESPONSIBILITIES
  • Design evidence tiering frameworks to rank and weight sources based on authority and recency.
  • Define fail-closed quality checks to determine whether AI-generated outputs should be released or withheld.
  • Design and execute evaluation frameworks covering groundedness, citation accuracy, completeness, and coverage.
  • Develop measurable evaluation rubrics in collaboration with client Subject Matter Experts (SMEs).
  • Analyze LLM and RAG failure modes and provide actionable feedback to retrieval and prompt engineering teams.
  • Establish quality metrics, evaluation datasets, and reporting mechanisms.
  • Support AI quality sign-off and delivery gate reviews with evidence-based results.
REQUIRED SKILLS & EXPERIENCE
  • 5+ years of experience in Data Science / Applied ML with hands-on LLM evaluation experience.
  • Strong Python, SQL, and analytical skills.
  • Experience designing evaluation methodologies for Generative AI / LLM systems.
  • Strong understanding of RAG architectures and failure modes.
  • Experience developing testable evaluation rubrics from expert/SME judgment.
  • Strong understanding of AI quality, reliability, and evaluation metrics.
GOOD TO HAVE
  • Experience with RAGAS, DeepEval, LangSmith, or similar evaluation frameworks.
  • Experience working with scientific, clinical, legal, or other high-authority literature.
  • Experience with human-in-the-loop evaluation and annotation.
  • Knowledge of inter-rater reliability and statistical evaluation methods.
KEY COMPETENCIES

Python | SQL | LLM Evaluation | Generative AI | RAG | RAGAS | DeepEval | LangSmith | AI Quality | Evaluation Frameworks | Prompt Evaluation | Groundedness | Citation Accuracy | Data Analysis

To apply for this job email your details to recruiting@nstarxinc.com

Privacy Overview
NStarX Logo

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.

Necessary

Strictly Necessary Cookie should be enabled at all times so that we can save your preferences for cookie settings.