Ilya A.-Clinical AI & Medical Safety Specialist | Clinical SME (DMD) | Healthcare LLM Evaluation

Check rate
Experience
General Dentist
Private Practice
- Provided general dentistry with responsibilities spanning diagnosis, restorative care, endodontics, and oral surgery within a private-practice setting.
- Performed comprehensive examinations, evidence-based treatment planning, clinical documentation, and interpretation of CBCT and 2D radiographic studies, including complex dental and maxillofacial findings.
- Applied clinical reasoning and risk-based decision-making across routine and complex dental cases, including referral and escalation when specialist care was indicated.
Oral & Maxillofacial Surgery Assistant
Clinical Hospital
- Joined an oral and maxillofacial surgery department during dental school and accumulated approximately four years of hospital-based clinical exposure under specialist supervision.
- Assisted in trauma, reconstructive, and other operative cases, including periods of high-acuity wartime clinical workload involving complex and critically ill patients.
- Supported patient preparation, surgical workflow, perioperative safety, intraoperative assistance, postoperative care, and selected supervised clinical tasks.
- Worked in demanding multidisciplinary environments requiring close adherence to surgical protocols, rapid prioritization, and clear communication across the care team.
LLM Response Evaluation & Comparative Ranking
General AI Evaluation Projects
- Compared multiple AI-generated responses to the same expert-level prompt and ranked them using structured quality criteria.
- Assigned individual quality scores and produced concise ranking rationales covering factual accuracy, logical coherence, instruction following, user-intent alignment, usefulness, specificity, naturalness, and organization.
- Identified hallucinations, logical inconsistencies, redundancy, off-topic content, weak formatting, and lack of user specificity while applying project-specific rubrics and discard rules.
AI Evaluation Specialist & Quality Reviewer
Global AI Evaluation Projects
- Evaluated large volumes of AI-generated content across ASR, machine translation, LLM, and other AI evaluation tasks, applying detailed project-specific guidelines and quality standards.
- Reviewed 10,000+ machine-generated ASR transcriptions, identifying semantic, grammatical, formatting, pronunciation, and meaning-related errors.
- Evaluated machine-generated translations across English, Ukrainian, and Russian, including English → Ukrainian, Ukrainian → English, Russian → English, and other multilingual language pairs.
- Performed complete source-target comparisons to identify mistranslations, omissions, additions, repetitions, terminology inconsistencies, fluency problems, stylistic issues, and locale-specific errors.
- Assigned Acceptability Scores on a 0–100 scale, assessing whether translations accurately conveyed source meaning and met linguistic quality requirements.
- Applied structured translation error taxonomies, including Accuracy, Fluency, Terminology, Omission, Addition, Repetition, Style, and Locale-related issues.
- Validated other evaluators' Score Justifications, determining whether comments accurately reflected the source–translation relationship and whether assigned scores were reasonable.
- Reviewed individual evaluation findings for factual validity, category accuracy, and severity, including Critical, Major, Minor, and informational classifications.
- Identified incorrectly classified, overstated, understated, or overlooked translation errors and provided evidence-based rationales with corrected scores or classifications when required.
- Distinguished objective linguistic errors from acceptable alternative translations and subjective stylistic preferences, avoiding unjustified penalties.
- Performed secondary quality audits to improve annotation consistency and evaluator agreement across multilingual projects.
- Conducted large-scale audio segmentation for voice-assistant interactions, identifying trigger phrases, interaction boundaries, overlapping speech, background audio, and edge-case scenarios.
- Applied semantic consistency scoring, evaluating keyword preservation, meaning retention, natural language quality, and overall comprehensibility.
- Categorized transcription errors involving Proper Nouns, Hot Words, Function Words, Affixes, punctuation, formatting, pronunciation, and semantic consistency.
- Followed project-specific SOPs, evaluation rubrics, quality standards, and delivery requirements.
- Contributed to quality improvement through reviewer feedback, consistency checks, and secondary quality assurance.
Clinical Safety & Adversarial Medical AI Evaluation
Healthcare AI Safety Tasks
- Stress-tested medical AI outputs for unsafe treatment guidance, missed contraindications, failed emergency escalation, and other safety-critical clinical errors.
- Assessed robustness under adversarial or context-manipulated prompts designed to surface misleading, overconfident, or insufficiently risk-aware behavior.
- Classified safety failures by type, severity, and potential clinical impact, including hallucination, reasoning failure, missed red flags, unsupported treatment advice, and inappropriate reassurance.
- Reviewed whether safety-sensitive responses preserved uncertainty, stated limitations, and avoided presenting speculative conclusions as established clinical facts.
Medical AI Subject Matter Expert / Healthcare LLM Evaluator
Healthcare LLM Evaluation & AI Validation Projects
- Evaluated 5,000+ AI-generated medical responses across clinical reasoning and healthcare AI evaluation tasks.
- Assessed AI-generated differential diagnoses, treatment recommendations, clinical explanations, and medical reasoning for factual accuracy, completeness, and clinical safety.
- Identified hallucinations, unsupported claims, incomplete reasoning, unsafe medical recommendations, and inconsistencies with evidence-based medical practice.
- Evaluated whether AI responses correctly interpreted clinical information and reached medically defensible conclusions.
- Verified clinical information using PubMed, ADA, NICE, and internationally recognized medical references.
- Applied structured evaluation rubrics covering helpfulness, factual correctness, reasoning quality, completeness, clinical appropriateness, and evidence alignment.
- Performed evidence-based fact-checking of AI-generated medical claims and identified unsupported or misleading statements.
- Assessed differential diagnosis reasoning and whether proposed diagnoses were appropriately supported by the clinical information provided.
- Evaluated medical AI responses for potential safety issues and clinically inappropriate recommendations.
- Assessed whether models recognized uncertainty, insufficient data, urgent symptoms, escalation needs, and indications for specialist referral.
- Contributed subject-matter expertise as a Doctor of Dental Medicine to healthcare-focused AI evaluation and validation workflows.
Industry experience
See where this freelancer has spent most of their professional time.
Experienced in Healthcare.
Summary
Doctor of Dental Medicine and Medical AI Evaluation Specialist with a clinical background spanning oral and maxillofacial surgery, general dentistry, diagnostic imaging, and healthcare AI evaluation. Approximately four years of hospital-based exposure to oral and maxillofacial surgery, including trauma, reconstructive cases, perioperative care, and critically ill patients, followed by surgical-focused postgraduate clinical training and private dental practice. Experienced in evaluating AI-generated clinical content for factual accuracy, hallucinations, clinical reasoning, patient-safety risks, guideline alignment, escalation logic, and treatment appropriateness. Additional experience includes multilingual QA, machine-translation validation, ASR evaluation, comparative LLM ranking, and secondary reviewer workflows across English, Ukrainian, and Russian.
Skills
Ai Evaluation & Quality Assurance
- Ai Output Evaluation
- Ai Quality Assurance
- Secondary Quality Review
- Annotation Auditing
- Evaluator Performance & Consistency Review
- Evaluation Rubric Application
- Rubric-Based Quality Assessment
- Error Taxonomy & Classification
- Severity Assessment
- Guideline & Sop Compliance
- Edge-Case Analysis
- Quality Control
- Human-In-The-Loop Ai Evaluation
- Clinical Ai Safety & Adversarial Evaluation
- Safety-Critical Failure Analysis
- Escalation & Red-Flag Assessment
Machine Translation Evaluation
- Multilingual Machine Translation Evaluation
- Translation Quality Assessment
- Acceptability Scoring (0-100)
- Source-Target Semantic Comparison
- Translation Error Detection
- Translation Evaluation Validation
- Evaluation Justification Validation
- Finding & Error Severity Validation
- Accuracy / Fluency / Terminology Assessment
- Omission & Addition Detection
- Repetition Detection
- Style & Locale Evaluation
- Multilingual Linguistic Qa
Asr & Linguistic Evaluation
- Asr Evaluation
- Transcription Auditing
- Audio Segmentation
- Semantic Consistency Analysis
- Homophone & Pronunciation Error Detection
- Proper Noun Evaluation
- Hot Word Evaluation
- Function Word Evaluation
- Affix Error Identification
- Trigger Phrase Detection
- Speech Boundary Identification
- Overlapping Speech Analysis
- Background Audio Identification
- Multilingual Annotation
Llm & Medical Ai
- Llm Response Evaluation
- Prompt Evaluation
- Instruction-Following Assessment
- Rlhf Evaluation
- Medical Ai Validation
- Clinical Reasoning Evaluation
- Hallucination Detection
- Medical Response Evaluation
- Differential Diagnosis Assessment
- Treatment Recommendation Evaluation
- Evidence-Based Fact-Checking
- Pubmed Verification
- Ada & Nice Guideline Validation
- Comparative Llm Ranking & Preference Evaluation
- Clinical Safety Evaluation
- Unsupported Claim Detection
- Patient-Safety Risk Assessment
Technical & Clinical
- Dicom
- Pacs
- Cbct Interpretation
- 2d Radiographic Analysis
- Enterprise Annotation Platforms
- Microsoft Office
- Google Workspace
Additional Expertise
- Multilingual Ai Data Evaluation
- Linguistic Quality Assurance
- Machine Translation Quality
- Speech Recognition Evaluation
- Llm Evaluation
- Medical Ai & Healthcare Ai
- Human Evaluation Of Ai-Generated Content
- Large-Scale Annotation Workflows
- Quality Control & Secondary Review
- Evidence-Based Evaluation
- Clinical & Linguistic Error Analysis
Languages
Education
Ivano-Frankivsk National Medical University
Doctor of Dental Medicine (DMD) · Dental Medicine · Ivano-Frankivsk, Ukraine
Statistics
Experience
Expertise
Qualifications
Profile
Frequently asked questions
Have questions? Find more information here.
Ilya speaks the following languages: Russian (Native), Ukrainian (Native), English (Advanced).
Ilya has at least 3 years of experience. During this time, Ilya has worked in at least 1 role and for 1 company. The average length of individual experience is 3 years and 8 months. Note that Ilya may not have shared all experience and actually has more experience.
Based on recent experience, Ilya would be well-suited for roles such as: General Dentist, Oral & Maxillofacial Surgery Assistant, LLM Response Evaluation & Comparative Ranking.
Ilya's most recent position is General Dentist at Private Practice.
In recent years, Ilya has worked for Private Practice.
Ilya is most experienced in industries like Healthcare.
Ilya holds a Doctorate in Dental Medicine from Ivano-Frankivsk National Medical University.
Ilya is immediately available for suitable projects.
Similar freelancers
Discover other experts with similar qualifications and experience
Experts recently working on similar projects
Freelancers with hands-on experience in comparable project as a General Dentist
