Employer Login | Partner Login | Candidate Login
WhatsApp Logo Chat us Call Logo Call us

AI Quality Engineer, Safety and RAG - Assistant Vice President

Citi Bank — Chennai, CHENNAI

Apply Now
Experience 0-5 Years yrs
Age 21-35 yrs
Gender Both
Salary Not Disclosed LPA
Department Sales Loans
Product Loans

Job Description

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of approximately 219,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

We are seeking a highly motivated and experienced AI Quality Engineer to join our Retail and Wealth Risk Engineering team under the Enterprise Risk Technology platform.This role spans the full spectrum of modern AI quality engineering — fromAgentic AI flow testing andRAG pipeline validation toAI safety, test automation, andperformance & reliability engineering.

You will be the quality pillar for complex autonomous AI systems, ensuring they aresafe, accurate, explainable, resilient, and production-ready at scale. This is a high-impact, highly technical role that requires both depth in AI/ML and breadth across testing disciplines.

Responsibilities

Agentic AI Testing

  • Design and executeend-to-end test strategies for Agentic AI pipelines, including single-agent and multi-agent workflows.
  • Validateagent reasoning, planning, and decision-making chains (e.g., ReAct, Chain-of-Thought, Plan-and-Execute, Reflexion).
  • Testtool-use correctness — ensuring agents invoke the right tools, with correct parameters, at the right time.
  • Evaluateagent memory systems (short-term, long-term, episodic) for accuracy and context retention across sessions.
  • Validateagent handoff and delegation logic in multi-agent orchestration frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • Testtermination conditions, loop detection, andinfinite loop prevention in autonomous agent loops.

RAG (Retrieval-Augmented Generation) Testing

  • Design comprehensive test strategies forend-to-end RAG pipelines — covering ingestion, chunking, embedding, retrieval, reranking, and generation stages.
  • Validateretrieval accuracy and relevance — ensuring the correct context chunks are retrieved for a given query.
  • Testembedding model quality and vector similarity thresholds across different document corpora.
  • Evaluatefaithfulness, groundedness, and answer relevance of generated responses using frameworks likeRAGAS, TruLens, DeepEval.
  • Testchunking strategies (fixed, semantic, hierarchical) for their impact on retrieval quality.
  • Validatecontext window management — ensuring retrieved context does not exceed token limits or degrade generation quality.
  • Conductend-to-end regression testing when the underlying knowledge base, embedding model, or LLM changes.
  • Testmulti-turn conversational RAG for context coherence and citation accuracy across turns.

Test Automation

  • Build and maintainautomated test harnesses for Agentic and RAG systems, including agent trajectory replay, tool mock injection, and prompt simulation.
  • Developautomated evaluation pipelines integrated into CI/CD workflows for continuous model and agent validation.
  • Createdata validation and data quality frameworks (using Great Expectations, Deequ, or custom tooling) for training, retrieval, and inference data.
  • Buildprompt regression suites to detect behavioral drift across LLM versions or prompt changes.
  • Implementdeterminism and reproducibility tests for stochastic LLM-based decisions.
  • Automatevector database validation — index integrity, embedding drift, and retrieval consistency checks.
AI Safety & Security Testing
  • Conductred-teaming and adversarial testing to uncover jailbreaks, prompt injection vulnerabilities, and goal misalignment in LLM-based systems.
  • Testoutput guardrails and content filters for unsafe, biased, toxic, or out-of-scope model behavior.
  • Validateprivilege escalation controls — ensuring agents do not exceed permitted actions or access unauthorized resources.
  • Performdata poisoning and backdoor attack simulations to assess model robustness.
  • Evaluate models forbias, fairness, and discrimination using frameworks such as AI Fairness 360 and Aequitas.
  • TestPII leakage and data privacy controls in RAG and agent pipelines in accordance with GDPR, CCPA, and internal data governance policies.
  • Conduct security testing aligned with theOWASP Top 10 for LLM Applications, including:
    • Prompt Injection (Direct & Indirect)
    • Insecure Output Handling
    • Training Data Poisoning
    • Insecure Plugin / Tool Design
    • Sensitive Information Disclosure
  • Validateconstitutional AI constraints, RLHF-aligned behavior boundaries, and system prompt integrity.
  • Collaborate with cybersecurity teams onAI-specific threat modeling and vulnerability management.
  • Maintainsafety testing playbooks and document red-team findings with severity ratings and remediation recommendations.

Performance & Reliability Testing

  • Define and executeload, stress, soak, and spike testing for AI-powered APIs, inference endpoints, and agent orchestration services.
  • Measure and optimizeend-to-end latency across RAG and agentic pipelines — from query to final response.
  • BenchmarkLLM inference throughput (tokens/second) and identify bottlenecks across model serving infrastructure.
  • Testauto-scaling behavior of AI services under variable load conditions.
  • Validatecircuit breaker, retry, and fallback mechanisms in agentic and RAG systems for graceful degradation.
  • Testvector database performance — query latency, index build time, and retrieval accuracy under high concurrency.
  • Conductcost efficiency analysis — measuring token consumption, API call costs, and infrastructure spend per agent task.
  • EstablishSLOs (Service Level Objectives) andSLAs for AI system availability, latency percentiles (P50, P95, P99), and error rates.
  • Collaborate with MLOps teams to set upobservability dashboards, monitoring alerts, and automated anomaly detection for production AI systems.
  • Performchaos engineering experiments to validate agent and RAG system resilience under infrastructure failures.

Domain Knowledge

  • Deep understanding ofRAG architecture patterns — naive RAG, advanced RAG, modular RAG.
  • Solid grasp ofagent design patterns: ReAct, Plan-and-Execute, Reflexion, MRKL, Mixture-of-Agents.
  • Familiarity withAI safety and alignment principles (RLHF, Constitutional AI, guardrail layers).
  • Knowledge oftoken economics, context management, and LLM cost optimization.
  • Proficiency inperformance engineering methodologies for distributed AI systems.

Preferred Qualifications

  • Experience withMCP (Model Context Protocol) or similar agentic communication standards.
  • Exposure tomulti-modal agent testing (agents handling text, images, code, documents).
  • Experience inregulated industries (banking, finance, healthcare) with strict compliance requirements.
  • Familiarity withchaos engineering tools (Chaos Monkey, Gremlin, LitmusChaos).

Education

  • Bachelor’s degree in Computer Science, Engineering, or a related field.

  • Master’s degree is a plus.

Experience
  • 8+ years of experience in software or AI/ML quality engineering.
  • 3+ years of hands-on experience withRAG systems, or Agentic AI.
  • Proven experience buildingautomated test frameworks for non-deterministic AI systems.
  • Strong background inperformance testing andAI safety/security assessments.

------------------------------------------------------

Job Family Group:

Technology

------------------------------------------------------

Job Family:

Technology Quality

------------------------------------------------------

Time Type:

Full time

------------------------------------------------------

Most Relevant Skills

Please see the requirements listed above.

------------------------------------------------------

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

------------------------------------------------------

Citi is an equal opportunity employer, and qualified candidates will receive consideration without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other characteristic protected by law.

If you are a person with a disability and need a reasonable accommodation to use our search tools and/or apply for a career opportunity review Accessibility at Citi.

View Citi’s EEO Policy Statement and the Know Your Rights poster.

Frequently Asked Questions

What are the eligibility criteria for the AI Quality Engineer, Safety and RAG - Assistant Vice President job at Citi Bank?

Eligibility typically includes the qualifications and experience outlined in the job description above, with around 0-5 Years years of relevant experience expected for this role.

What are the primary responsibilities of a AI Quality Engineer, Safety and RAG - Assistant Vice President at Citi Bank?

The key responsibilities for this role are detailed in the Key Responsibilities section above, covering the core duties expected of a AI Quality Engineer, Safety and RAG - Assistant Vice President at Citi Bank.

Is prior experience required for this position?

This role requires around 0-5 Years years of relevant experience, as specified in the job listing. Please refer to the Qualifications & Experience section above for full details.

What skills are important for this role?

Skills relevant to this position are outlined in the Qualifications & Experience section above. In general, strong communication, domain knowledge, and the ability to meet role-specific targets are valued across similar BFSI positions.

How do I apply for the AI Quality Engineer, Safety and RAG - Assistant Vice President role at Citi Bank?

You can apply directly using the Apply Now button on this page, which will take you to Citi Bank's application process for this role.