Skip to main content
Top expert badge
Recommended expert
Profile header background

Arya D.-Quality Analyst & AI Evaluation Specialist

Arya D. - Quality Analyst & AI Evaluation Specialist - profile avatar
Profile header overlay
Available
Surat, India

Check rate

Experience

Mar 2026 - Present
India

Quality Analyst

Alignerr

Position summary
Quality Analyst at Alignerr
Industries
Information Technology
Business areas
Information Technology
Quality Assurance
Research and Development
Explore QA Engineer
  • Completed AI evaluation and data-focused assessment projects across text, code, image, and structured data modalities while consistently meeting platform quality benchmarks and review standards.
  • Applied RLHF-based evaluation workflows including preference ranking, rubric scoring, adversarial testing, and instruction-following assessment across multi-turn LLM interactions.
  • Reviewed model outputs for hallucinations, instruction drift, contextual inconsistencies, logical contradictions, and constraint violations using structured annotation and evaluation frameworks.
  • Performed annotation quality reviews and provided evaluator feedback to improve scoring consistency and adherence to evaluation guidelines.
  • Contributed to AI evaluation, generalist, and data science-oriented projects involving reasoning validation, quality assessment, and model behavior analysis.
Mar 2026 - Present

LLM Failure Mode Taxonomy & Severity Classification System

Turing

Position summary
LLM Failure Mode Taxonomy & Severity Classification System at Turing
Industries
Information Technology
Business areas
Product Development
Quality Assurance
Research and Development
  • Production LLM evaluation lacked a shared framework for categorising and prioritising failure types across evaluators, so built a structured taxonomy covering six patterns — hallucination, constraint leakage, instruction drift, tone inconsistency, logical contradiction, and contextual memory failure — each with documented triggering conditions and annotated real-output examples.
  • Assigned severity tiers (critical / major / minor) to each failure type based on user-impact potential, using a consistent rating logic that could be applied uniformly across different evaluators reviewing different output batches.
  • Applied the framework in Turing's weekly evaluation workflow to produce severity-graded engineering reports, reducing ambiguity in failure communication and giving the client engineering team a reproducible prioritisation signal.
Jan 2026 - Jul 2026

Rubric-Based Multi-Turn Adversarial Evaluation Framework

Turing

Position summary
Rubric-Based Multi-Turn Adversarial Evaluation Framework at Turing
Industries
Information Technology
Business areas
Information Technology
Quality Assurance
Research and Development
  • Inconsistent scoring across evaluators under adversarial prompting conditions pointed to a missing rubric structure — designed a modular evaluation framework across four quality dimensions: instruction adherence, factual consistency, tone regulation, and contextual continuity, each with annotated pass/fail thresholds and example outputs calibrated to evaluator judgment.
  • Incorporated progressive difficulty tiers from single-turn edge cases through to multi-turn jailbreak attempts, allowing the rubric to scale difficulty without changing the underlying scoring logic.
  • Applied at Turing to support model release readiness evaluations and subsequently adapted for text and code modality annotation tasks on Alignerr and RWS platforms.
Dec 2025 - Mar 2026
India

Business Analyst – AI Evaluation

Turing

Position summary
Business Analyst – AI Evaluation at Turing
Industries
Information Technology
Business areas
Information Technology
Quality Assurance
Research and Development
  • Evaluated 150+ production-grade LLM outputs weekly using structured rubrics, adversarial prompting scenarios, and multi-turn response analysis to assess model quality, reasoning, and instruction adherence.
  • Identified recurring model failure modes including hallucinations, contextual memory failures, instruction drift, logical contradictions, tone inconsistencies, and constraint violations through systematic output review.
  • Reviewed annotation quality and supported evaluator calibration efforts to improve scoring consistency across evaluation workflows.
  • Developed severity-based reporting frameworks that standardized failure categorization and improved prioritization of model issues for downstream review.
  • Translated client quality requirements into rubric criteria, evaluation dimensions, and QA documentation to support release-readiness assessments and model behavior evaluation.
Jul 2025 - Dec 2025
Surat, India

Business Analyst

Mountain Monk Consulting

Position summary
Business Analyst at Mountain Monk Consulting
Industries
Energy
Food and Beverage
Retail
Business areas
Audit
Operations
Project Management
  • Supported client engagements across retail, solar energy, F&B, interior design, and experiential retail domains through requirements gathering, sprint planning, and process documentation.
  • Created SOWs, process documentation, stakeholder presentations, and business process models aligned with client delivery objectives.
  • Audited 20+ client deliverables monthly for logical consistency, semantic accuracy, and business alignment while coordinating cross-functional stakeholder communication.
Feb 2025 - Jul 2025

LLM Content Evaluation & Annotation Framework

Independent Practice

Position summary
LLM Content Evaluation & Annotation Framework at Independent Practice
Industries
Information Technology
Business areas
Information Technology
Quality Assurance
Research and Development
  • Developed a structured annotation and evaluation pipeline to assess content quality, semantic equivalence, contextual accuracy, intent preservation, and tonal consistency across LLM-generated outputs.
  • Evaluated LLM-generated text pairs using taxonomy-driven labels for sentiment classification, intent detection, contextual correctness, and quality assessment; documented ambiguous and edge-case scenarios using structured justification notes.
  • Produced a documented annotation dataset with defined quality thresholds and a reusable labeling schema that later informed rubric design approaches used in professional AI evaluation workflows.
Sep 2024 - Jul 2025
Surat, India

Analyst

AUM Electric Engineering Pvt. Ltd.

Position summary
Analyst at AUM Electric Engineering Pvt. Ltd.
Industries
Construction
Energy
Utilities
Business areas
Operations
Project Management
Research and Development
Strategy
  • Gathered and structured client requirements across industrial electrification and infrastructure projects, translating technical inputs into operational planning frameworks.
  • Conducted market research, competitor analysis, technical documentation, and stakeholder reporting to support business planning and project execution.
Feb 2024 - Jul 2024

Driver Drowsiness Detection System

NMIMS

Position summary
Driver Drowsiness Detection System at NMIMS
Industries
Automotive
Information Technology
Business areas
Information Technology
Product Development
Research and Development
  • Driver fatigue is difficult to catch without continuous real-time monitoring — built a computer vision system to classify drowsiness from facial video, applying ML-based behavioural analysis to a safety-critical problem where detection latency has direct real-world consequences.
  • Used OpenCV for eye-blink tracking and facial landmark detection from live frames; implemented a feature extraction pipeline isolating drowsiness indicators; trained a Keras neural network classifier on labelled fatigue-state data captured under controlled recording sessions.
  • Achieved approximately 85% classification accuracy with real-time inference performance suitable for deployment on standard consumer hardware.
Jan 2024 - Jul 2024

ML-Based Sign Language Recognition Model

NMIMS

Position summary
ML-Based Sign Language Recognition Model at NMIMS
Industries
Education
Information Technology
Business areas
Information Technology
Product Development
Research and Development
  • Accessibility tooling for sign language users typically requires expensive hardware or cloud-dependent APIs — built a standalone end-to-end system to translate American Sign Language gestures into live text output using only a standard webcam.
  • Used OpenCV for gesture capture and frame preprocessing, TensorFlow to train a classifier across 26 ASL hand signs, and Flask to serve live predictions via a lightweight web interface; designed capture, inference, and serving as independent modules to support extension to other gesture sets without restructuring the pipeline.
  • Delivered a working real-time gesture-to-text pipeline with consistent single-sign recognition and a Flask interface accessible without local installation, supporting educational and public-access deployment.
Aug 2023 - Jul 2025
Mumbai, India

Scriptwriter

MadLads Studios

Position summary
Scriptwriter at MadLads Studios
Industries
Advertising
Media and Entertainment
Business areas
Marketing
  • Produced 50+ English-language scripts across podcasts, advertising campaigns, and narrative content formats.
  • Developed structured approaches for evaluating narrative coherence, instruction consistency, and tone alignment that later informed AI evaluation workflows.

Industry experience

See where this freelancer has spent most of their professional time.

Experienced in Advertising, Media and Entertainment, Information Technology, Energy, Construction, and Utilities.

Advertising
Media and Entertainment
Information Technology
Energy
Construction
Utilities
Profile match chart

Business area experience

See which departments and functions this freelancer has contributed to most.

Experienced in Research and Development, Marketing, Information Technology, Operations, Project Management, and Quality Assurance.

Research and Development
Marketing
Information Technology
Operations
Project Management
Quality Assurance
Profile match chart

Summary

Quality Analyst & AI Evaluation Specialist with experience in LLM evaluation, RLHF preference ranking, annotation quality assurance, and reviewer workflows across Turing, Alignerr, and RWS. Evaluated 150+ production-grade LLM outputs weekly using structured rubrics, adversarial testing, and multi-turn assessment frameworks to identify hallucinations, instruction drift, logical inconsistencies, and contextual failures.

Experienced in annotation review, evaluator calibration, rubric design, failure-mode analysis, and severity-based reporting to support model alignment and release-readiness evaluations.

Skills

  • Ai Evaluation & Llm Assessment: Llm Evaluation, Ai Model Evaluation, Rlhf Preference Ranking, Annotation Quality Assurance, Adversarial Testing, Reasoning Evaluation, Rubric Design, Failure-Mode Analysis, Hallucination Detection, Multi-Turn Evaluation

  • Business Analysis & Quality Review: Requirements Gathering, Stakeholder Management, Process Analysis, Technical Documentation, Uat, Data Validation, Kpi Tracking, Agile Workflows, Sow Writing

  • Programming & Machine Learning: Python, Sql, Pandas, Numpy, Tensorflow, Keras, Opencv, Flask

  • Tools & Platforms: Power Bi, Github, Google Workspace, Microsoft Office Suite, Alignerr, Rws, Labelbox, Feather Openai

Languages

English
Native
Gujarati
Native
Hindi
Native

Education

Aug 2020 - Jul 2024

Mukesh Patel School of Technology Management & Engineering, NMIMS

Bachelor of Technology, Specialization: Artificial Intelligence · Computer Science Engineering · Mumbai, India

Statistics

Experience

Total positions 10
Experience in Advertising 2 y
Avg length 8 m
Longest experience 1 y 11 m

Global experience

Countries worked in 1 (India)
Primary country India

Expertise

Recent roles Quality Analyst, LLM Failure Mode Taxonomy & Severity Classification System, Rubric-Based Multi-Turn Adversarial Evaluation Framework
Main industries Advertising, Media and Entertainment, Information Technology
Main business areas Research and Development, Marketing, Information Technology

Qualifications

Highest degree Bachelor

Profile

Member since
Need a freelancer? Find your match in seconds.
Try FRATCH GPT
More actions

Frequently asked questions

Have questions? Find more information here.

Arya is based in Surat, India.

Arya speaks the following languages: English (Native), Gujarati (Native), Hindi (Native).

Arya has at least 3 years of experience. During this time, Arya has worked in at least 10 different roles and for 7 different companies. The average length of individual experience is 4 months. Note that Arya may not have shared all experience and actually has more experience.

Based on recent experience, Arya would be well-suited for roles such as: Quality Analyst, LLM Failure Mode Taxonomy & Severity Classification System, Rubric-Based Multi-Turn Adversarial Evaluation Framework.

Arya's most recent position is Quality Analyst at Alignerr.

In recent years, Arya has worked for Alignerr, Turing, Mountain Monk Consulting, Independent Practice, and AUM Electric Engineering Pvt. Ltd..

Arya is most experienced in industries like Advertising, Media and Entertainment, and Information Technology. Arya also has some experience in Energy, Construction, and Utilities.

Arya is most experienced in business areas like Research and Development, Marketing, and Information Technology. Arya also has some experience in Operations, Project Management, and Quality Assurance.

Arya holds a Bachelor in Computer Science Engineering from Mukesh Patel School of Technology Management & Engineering, NMIMS.

Arya is immediately available for suitable projects.

Daily rate distribution

0 1 2 3 4
<€320 €320-​480 €480-​640 €640-​800 €960+

The rates shown represent the typical market range for freelancers in this position based on recent contracts on our platform.

Average rates for similar positions

Rates are based on recent contracts and do not include FRATCH margin.

600
450
300
150
Rate comparison chart
Daily rate avg. 446 €

The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.

600
450
300
150
Rate comparison chart
Median rate 400 €

The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.

Calculated based on our freelancers’ daily rates as of 9 Sep 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.