Whitney L.-AI Evaluation and Annotation Specialist
Check rate
Experience
AI Evaluation and Annotation Specialist
- LLM Benchmarking: Evaluated and benchmarked 1,000+ LLM-generated responses for technical correctness, code efficiency, safety, and strict adherence to complex technical prompt constraints.
- Mathematical Auditing: Executed highly granular structural audits of model-generated code and multi-step mathematical derivations, documenting root-cause logic gaps, syntax errors, and hallucinations.
- Prompt Architecture: Authored and stress-tested high-complexity mathematical, logical, and coding prompts designed to uncover edge-case failures and assess advanced reasoning capabilities.
- Model Alignment: Led RLHF and SFT workflows, utilizing Google Collab and interactive Python scripts to parse massive datasets of model outputs and deliver highly detailed preference rankings.
Junior Software Engineer
Neurosurgical Research Unit
- Data Workflows and Pipelines: Develop, test, and maintain robust software tools and Python script pipelines to process, clean, and validate complex clinical research datasets.
- Requirement Translation: Collaborate closely with cross-functional research teams to translate high-level data collection needs into structured algorithmic logic, mathematical validation checks, and consistency filters.
- Quality Assurance: Conduct comprehensive edge-case testing, system debugging, and code reviews in Google Collab environments to ensure absolute data reliability and alignment with stringent guidelines.
- Technical Documentation: Document end-to-end software workflows, API integrations, and system specifications to streamline operational handoffs for large research teams.
Mathematics and Analytical Data Tutor, Evaluator
- Quantitative Logic: Analyzed and deconstructed advanced mathematical concepts (college algebra, quantitative reasoning, applied math) into pristine, step-by-step logical derivations.
- Error Analysis: Audited student and system-generated technical derivations to catch, classify, and correct structural errors and logical oversights.
- Performance Optimization: Designed structured problem-solving templates that accelerate accuracy rates in multi-step quantitative problem solving by over 25%.
Industry experience
See where this freelancer has spent most of their professional time.
Experienced in Information Technology, Healthcare, and Education.
Business area experience
See which departments and functions this freelancer has contributed to most.
Experienced in Information Technology, Quality Assurance, Research and Development, and Product Development.
Summary
Analytical Software Engineer and AI Evaluation Specialist with a strong foundation in algorithmic reasoning, mathematical logic, and advanced data workflows. Expert in designing rigorous prompt edge cases, executing Reinforcement Learning from Human Feedback (RLHF) frameworks, and conducting multi-step mathematical error analysis to optimize Large Language Model (LLM) alignment. Proven track record of translating complex STEM research requirements into scalable data pipelines and managing high-quality data collection for Machine Learning applications. Extensive experience utilizing Python and Google Colab for automated code auditing, data validation, and consistency filtering.
Skills
- Stem Data And Ai Evaluation: Rlhf, Supervised Fine-Tuning (Sft), Model Response Grading, Prompt Engineering, Edge-Case Identification, Preference Ranking, Red Teaming.
- Software Engineering And Mathematics: Software Architecture, Algorithmic Logic, Advanced Data Validation, Applied Mathematics, Quantitative Reasoning, System Debugging.
- Tools And Frameworks: Python, Google Collab, Git, Json, Markdown, Rest Apis, Advanced Data Analysis (Ms Excel / Google Sheets), Agile Principles.
- Workflow Operations: Cross-Functional Collaboration, Technical Writing, Data Annotation, Workflow Design, Quality Assurance & Data Integrity Enforcement.
Languages
Education
Humboldt University of Berlin
Master of Science · Software Engineering · Berlin, Germany
Humboldt University of Berlin
Bachelor of Science · Software Engineering · Berlin, Germany
Statistics
Experience
Expertise
Qualifications
Profile
Frequently asked questions
Have questions? Find more information here.
Whitney speaks the following languages: German (Native), English (Advanced).
Whitney has at least 9 years of experience. During this time, Whitney has worked in at least 3 different roles and for 1 company. The average length of individual experience is 3 years and 11 months. Note that Whitney may not have shared all experience and actually has more experience.
Based on recent experience, Whitney would be well-suited for roles such as: AI Evaluation and Annotation Specialist, Junior Software Engineer, Mathematics and Analytical Data Tutor, Evaluator.
Whitney's most recent position is AI Evaluation and Annotation Specialist.
In recent years, Whitney has worked for Neurosurgical Research Unit.
Whitney is most experienced in industries like Education, Information Technology, and Healthcare.
Whitney is most experienced in business areas like Information Technology, Quality Assurance, and Research and Development. Whitney also has some experience in Product Development.
Whitney has recently worked in industries like Education, Information Technology, and Healthcare.
Whitney has recently worked in business areas like Information Technology, Quality Assurance, and Research and Development.
Whitney holds a Master in Software Engineering from Humboldt University of Berlin and a Bachelor in Software Engineering from Humboldt University of Berlin.
Whitney is immediately available full-time for suitable projects.
Similar freelancers
Discover other experts with similar qualifications and experience
Experts recently working on similar projects
Freelancers with hands-on experience in comparable project as a AI Evaluation and Annotation Specialist
