Arya D.-Quality Analyst & AI Evaluation Specialist

Check rate
Experience
Quality Analyst
Alignerr
- Completed AI evaluation and data-focused assessment projects across text, code, image, and structured data modalities while consistently meeting platform quality benchmarks and review standards.
- Applied RLHF-based evaluation workflows including preference ranking, rubric scoring, adversarial testing, and instruction-following assessment across multi-turn LLM interactions.
- Reviewed model outputs for hallucinations, instruction drift, contextual inconsistencies, logical contradictions, and constraint violations using structured annotation and evaluation frameworks.
- Performed annotation quality reviews and provided evaluator feedback to improve scoring consistency and adherence to evaluation guidelines.
- Contributed to AI evaluation, generalist, and data science-oriented projects involving reasoning validation, quality assessment, and model behavior analysis.
LLM Failure Mode Taxonomy & Severity Classification System
Turing
- Production LLM evaluation lacked a shared framework for categorising and prioritising failure types across evaluators, so built a structured taxonomy covering six patterns — hallucination, constraint leakage, instruction drift, tone inconsistency, logical contradiction, and contextual memory failure — each with documented triggering conditions and annotated real-output examples.
- Assigned severity tiers (critical / major / minor) to each failure type based on user-impact potential, using a consistent rating logic that could be applied uniformly across different evaluators reviewing different output batches.
- Applied the framework in Turing's weekly evaluation workflow to produce severity-graded engineering reports, reducing ambiguity in failure communication and giving the client engineering team a reproducible prioritisation signal.
Rubric-Based Multi-Turn Adversarial Evaluation Framework
Turing
- Inconsistent scoring across evaluators under adversarial prompting conditions pointed to a missing rubric structure — designed a modular evaluation framework across four quality dimensions: instruction adherence, factual consistency, tone regulation, and contextual continuity, each with annotated pass/fail thresholds and example outputs calibrated to evaluator judgment.
- Incorporated progressive difficulty tiers from single-turn edge cases through to multi-turn jailbreak attempts, allowing the rubric to scale difficulty without changing the underlying scoring logic.
- Applied at Turing to support model release readiness evaluations and subsequently adapted for text and code modality annotation tasks on Alignerr and RWS platforms.
Business Analyst – AI Evaluation
Turing
- Evaluated 150+ production-grade LLM outputs weekly using structured rubrics, adversarial prompting scenarios, and multi-turn response analysis to assess model quality, reasoning, and instruction adherence.
- Identified recurring model failure modes including hallucinations, contextual memory failures, instruction drift, logical contradictions, tone inconsistencies, and constraint violations through systematic output review.
- Reviewed annotation quality and supported evaluator calibration efforts to improve scoring consistency across evaluation workflows.
- Developed severity-based reporting frameworks that standardized failure categorization and improved prioritization of model issues for downstream review.
- Translated client quality requirements into rubric criteria, evaluation dimensions, and QA documentation to support release-readiness assessments and model behavior evaluation.
Business Analyst
Mountain Monk Consulting
- Supported client engagements across retail, solar energy, F&B, interior design, and experiential retail domains through requirements gathering, sprint planning, and process documentation.
- Created SOWs, process documentation, stakeholder presentations, and business process models aligned with client delivery objectives.
- Audited 20+ client deliverables monthly for logical consistency, semantic accuracy, and business alignment while coordinating cross-functional stakeholder communication.
LLM Content Evaluation & Annotation Framework
Independent Practice
- Developed a structured annotation and evaluation pipeline to assess content quality, semantic equivalence, contextual accuracy, intent preservation, and tonal consistency across LLM-generated outputs.
- Evaluated LLM-generated text pairs using taxonomy-driven labels for sentiment classification, intent detection, contextual correctness, and quality assessment; documented ambiguous and edge-case scenarios using structured justification notes.
- Produced a documented annotation dataset with defined quality thresholds and a reusable labeling schema that later informed rubric design approaches used in professional AI evaluation workflows.
Analyst
AUM Electric Engineering Pvt. Ltd.
- Gathered and structured client requirements across industrial electrification and infrastructure projects, translating technical inputs into operational planning frameworks.
- Conducted market research, competitor analysis, technical documentation, and stakeholder reporting to support business planning and project execution.
Driver Drowsiness Detection System
NMIMS
- Driver fatigue is difficult to catch without continuous real-time monitoring — built a computer vision system to classify drowsiness from facial video, applying ML-based behavioural analysis to a safety-critical problem where detection latency has direct real-world consequences.
- Used OpenCV for eye-blink tracking and facial landmark detection from live frames; implemented a feature extraction pipeline isolating drowsiness indicators; trained a Keras neural network classifier on labelled fatigue-state data captured under controlled recording sessions.
- Achieved approximately 85% classification accuracy with real-time inference performance suitable for deployment on standard consumer hardware.
ML-Based Sign Language Recognition Model
NMIMS
- Accessibility tooling for sign language users typically requires expensive hardware or cloud-dependent APIs — built a standalone end-to-end system to translate American Sign Language gestures into live text output using only a standard webcam.
- Used OpenCV for gesture capture and frame preprocessing, TensorFlow to train a classifier across 26 ASL hand signs, and Flask to serve live predictions via a lightweight web interface; designed capture, inference, and serving as independent modules to support extension to other gesture sets without restructuring the pipeline.
- Delivered a working real-time gesture-to-text pipeline with consistent single-sign recognition and a Flask interface accessible without local installation, supporting educational and public-access deployment.
Scriptwriter
MadLads Studios
- Produced 50+ English-language scripts across podcasts, advertising campaigns, and narrative content formats.
- Developed structured approaches for evaluating narrative coherence, instruction consistency, and tone alignment that later informed AI evaluation workflows.
Industry experience
See where this freelancer has spent most of their professional time.
Experienced in Advertising, Media and Entertainment, Information Technology, Energy, Construction, and Utilities.
Business area experience
See which departments and functions this freelancer has contributed to most.
Experienced in Research and Development, Marketing, Information Technology, Operations, Project Management, and Quality Assurance.
Summary
Quality Analyst & AI Evaluation Specialist with experience in LLM evaluation, RLHF preference ranking, annotation quality assurance, and reviewer workflows across Turing, Alignerr, and RWS. Evaluated 150+ production-grade LLM outputs weekly using structured rubrics, adversarial testing, and multi-turn assessment frameworks to identify hallucinations, instruction drift, logical inconsistencies, and contextual failures.
Experienced in annotation review, evaluator calibration, rubric design, failure-mode analysis, and severity-based reporting to support model alignment and release-readiness evaluations.
Skills
Ai Evaluation & Llm Assessment: Llm Evaluation, Ai Model Evaluation, Rlhf Preference Ranking, Annotation Quality Assurance, Adversarial Testing, Reasoning Evaluation, Rubric Design, Failure-Mode Analysis, Hallucination Detection, Multi-Turn Evaluation
Business Analysis & Quality Review: Requirements Gathering, Stakeholder Management, Process Analysis, Technical Documentation, Uat, Data Validation, Kpi Tracking, Agile Workflows, Sow Writing
Programming & Machine Learning: Python, Sql, Pandas, Numpy, Tensorflow, Keras, Opencv, Flask
Tools & Platforms: Power Bi, Github, Google Workspace, Microsoft Office Suite, Alignerr, Rws, Labelbox, Feather Openai
Languages
Education
Mukesh Patel School of Technology Management & Engineering, NMIMS
Bachelor of Technology, Specialization: Artificial Intelligence · Computer Science Engineering · Mumbai, India
Statistics
Experience
Global experience
Expertise
Qualifications
Profile
Frequently asked questions
Have questions? Find more information here.
Arya is based in Surat, India.
Arya speaks the following languages: English (Native), Gujarati (Native), Hindi (Native).
Arya has at least 3 years of experience. During this time, Arya has worked in at least 10 different roles and for 7 different companies. The average length of individual experience is 4 months. Note that Arya may not have shared all experience and actually has more experience.
Based on recent experience, Arya would be well-suited for roles such as: Quality Analyst, LLM Failure Mode Taxonomy & Severity Classification System, Rubric-Based Multi-Turn Adversarial Evaluation Framework.
Arya's most recent position is Quality Analyst at Alignerr.
In recent years, Arya has worked for Alignerr, Turing, Mountain Monk Consulting, Independent Practice, and AUM Electric Engineering Pvt. Ltd..
Arya is most experienced in industries like Advertising, Media and Entertainment, and Information Technology. Arya also has some experience in Energy, Construction, and Utilities.
Arya is most experienced in business areas like Research and Development, Marketing, and Information Technology. Arya also has some experience in Operations, Project Management, and Quality Assurance.
Arya holds a Bachelor in Computer Science Engineering from Mukesh Patel School of Technology Management & Engineering, NMIMS.
Arya is immediately available for suitable projects.
Daily rate distribution
The rates shown represent the typical market range for freelancers in this position based on recent contracts on our platform.
Average rates for similar positions
Rates are based on recent contracts and do not include FRATCH margin.
The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.
The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.
Calculated based on our freelancers’ daily rates as of 9 Sep 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.
Similar freelancers
Discover other experts with similar qualifications and experience
Experts recently working on similar projects
Freelancers with hands-on experience in comparable project as a Quality Analyst
Nearby freelancers
Professionals working in or nearby Surat, India
