Krisztian Simon-AI Trainer & Evaluation Specialist · Adversarial VQA · Human Data Review
Check rate
Experience
Freelance Evaluator & Human Data Reviewer
Independent AI Training & Evaluation Projects
- Evaluate AI-generated outputs against detailed rubrics covering relevance, accuracy, groundedness, safety, naturalness and task completion.
- Review simulated storefront and support-agent conversations for groundedness, policy compliance, naturalness and task completion.
- Compare AI-generated image and video outputs for prompt alignment, visual consistency, artifacts, realism and temporal coherence.
- Apply changing calibration guidance and edge-case policies consistently across high-volume evaluation workflows.
- Write concise, evidence-based scoring rationales that separate observed facts from assumptions.
- Handle e-commerce product classification, catalog QA, search quality evaluation, category tagging and attribute extraction for structured datasets.
Sole Developer & Project Lead
ZeroBounce
- Self-initiated and building solo: a native Salesforce integration approved by the executive team to replace an unreliable third-party connector.
- Writing the full solution in Apex and JavaScript, deploying and testing through Salesforce CLI, and validating behavior manually against production traffic.
- Building it alongside a full-time lead role by driving coding agents (Claude Code, Codex) through design, implementation and test, reviewing every change and verifying against live API behavior.
- Presented architecture, packaging and AppExchange distribution plans to engineering and executive leadership.
Adversarial VQA Benchmark Creation
micro1
- Create image-question-answer benchmark tasks (211 submitted to date) designed to make current multimodal models fail on dense charts, diagrams and detailed photographs.
- Write questions that require several reasoning steps, with exact ground-truth answers a reviewer can verify directly against the source image.
- Design benchmark items that target multimodal reasoning weaknesses while keeping answers objective, reproducible and verifiable.
- Apply reviewer feedback and calibration guidance to raise task difficulty and remove ambiguity. Every task is individually reviewed before acceptance.
- Source public-domain and CC0 imagery so every benchmark item is clean and reproducible.
Technical Enablement Lead (Technical Support Manager)
ZeroBounce
- Designed the support domain for an internal AI agent assessment system: simulated customer conversations, technical scenarios and scoring rubrics, with human final approval on every result.
- Administer Zendesk and drive AI Copilot tuning through workflow review, agent feedback, response-quality analysis and recurring case patterns.
- Lead technical enablement for a 30+ person Support and Sales team and directly manage one enablement specialist.
- Tier-3 escalation owner for REST API, CRM integration, SMTP, authentication and DNS failures. Cut escalation volume 28% by converting the recurring cases into SOPs and targeted training.
- Built and maintain a 210+ page knowledge base covering email validation, REST API behavior, status and substatus logic, and CRM integrations, alongside 15+ SOPs and decision-tree workflows.
- Rebuilt new-hire onboarding as a scenario-based curriculum around API authentication, JSON payloads and live technical cases, cutting new-hire ramp from six weeks to three.
- Monitor API traffic in Cloudflare, investigate abusive usage patterns, coordinate customer outreach, and refine traffic-control rules.
Technical Support Engineer
SendGrid (via IntelligentBee)
- Supported SMB through enterprise customers on SMTP, REST APIs, JSON payloads, authentication, SPF/DKIM/DMARC, suppression lists, bounce classification and IP warmup.
- Traced delivery behavior in Splunk, Snowflake and Postman, then explained root cause directly to customers.
- Handled 13+ tickets and 25+ live chats daily within SLA. Promoted to TSE+ for technical performance and for mentoring newer engineers.
Technical Support Engineer
DIGI (RCS-RDS)
- Handled B2B and residential network escalations across MPLS, BGP, broadband, mobile, GPON/PON and VLAN environments under SLA-driven conditions.
- Monitored network status across a three-county area on overnight operations and coordinated field teams to restore connectivity.
Customer Support Agent
DIGI (RCS-RDS)
- Handled 70+ daily support inquiries covering connectivity, router configuration and account issues, and helped onboard new agents.
Industry Experience
See where this freelancer has spent most of their professional time.
Experienced in Telecommunication, Information Technology, and Retail.
Business Area Experience
See which departments and functions this freelancer has contributed to most.
Experienced in Customer Service, Information Technology, Operations, Product Development, Quality Assurance, and Project Management.
Summary
AI trainer and evaluation specialist with a 9-year technical background in SaaS support, Tier-3 escalations, API operations and enablement. Builds adversarial VQA benchmark tasks for multimodal model testing and evaluates LLM outputs against detailed rubrics across text, image, audio and video. At ZeroBounce, owns the support-domain design of an internal AI agent assessment system: simulated conversations, technical scenarios, scoring guidance and guardrails, with human final approval. Years of root-cause analysis and technical writing underpin the evaluation work: precise rationales, consistent calibration, and answers a reviewer can verify independently.
Skills
- Ai Evaluation & Human-In-The-Loop Systems: Llm Output Review, Adversarial Vqa, Rlhf/Sft Tasks, Response Ranking, Rubric-Based Scoring, Hallucination Detection, Groundedness, Factuality, Safety, Scenario Design, Guardrails, Human Final Approval
- Human Data & Annotation: Visual Question Answering, Data Labeling, Product Classification, Catalog Qa, Category Tagging, Attribute Extraction, Search Quality Evaluation, Tts And Audio Review, Image And Video Prompt Alignment
- Agent Assessment & Quality: Simulated Support Conversations, Technical Knowledge Evaluation, Scoring Guidance, Calibration, Qa Feedback, Scenario-Based Assessment
- Saas & Api Systems: Rest Apis, Json, Webhooks, Authentication Flows, Postman, Cloudflare, Salesforce (Apex, Salesforce Cli), Hubspot, Zendesk, Jira, Splunk, Snowflake
- Email Infrastructure: Smtp, Dns, Spf, Dkim, Dmarc, Bimi, Ip And Domain Warmup, Bounce Analysis, Inbox Placement, Suppression Lists
Languages
Education
Simion Bărnuțiu National College
Romanian Baccalaureate · Philology · Șimleu Silvaniei, Romania
Certifications & licenses
HubSpot Sales Hub Software Certified
HubSpot Service Hub Software Certified
Statistics
Experience
Global Experience
Expertise
Qualifications
Profile
Frequently asked questions
Have questions? Find more information here.
Daily Rate Distribution
The rates shown represent the typical market range for freelancers in this position based on recent contracts on our platform.
Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.
Average rates for similar positions
Rates are based on recent contracts and do not include FRATCH margin.
The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.
The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.
Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.
Similar Freelancers
Discover other experts with similar qualifications and experience
Experts recently working on similar projects
Freelancers with hands-on experience in comparable project as a Freelance Evaluator & Human Data Reviewer
Nearby freelancers
Professionals working in or nearby Oradea, Romania
