Skip to main content
Top expert badge
Recommended expert
Profile header background

Majid Khan-Senior Software Engineer

Majid Khan - Senior Software Engineer - profile avatar
Profile header overlay
Available
Stockport, United Kingdom

Check rate

Experience

Jan 2024 - Present
United Kingdom

Senior AI / LLM Evaluation Engineer

OPENCAST

Position summary
Senior AI / LLM Evaluation Engineer at OPENCAST
Industries
Information Technology
Business areas
Business Intelligence
Information Technology
Product Development
Quality Assurance
Research and Development
Explore AI Engineer
  • Design and maintain evaluation frameworks for large language models across reasoning, coding, instruction following, factuality, tool use, agent behaviour, and safety.
  • Develop structured datasets and evaluation tasks used for supervised fine-tuning (SFT), RLHF, preference optimization, post-training evaluation, and model benchmarking.
  • Create prompts, reference solutions, scoring criteria, metadata schemas, evaluation rubrics, and reviewer guidelines for model-training datasets.
  • Review and rank competing model responses for correctness, completeness, reasoning quality, relevance, instruction adherence, code quality, and safety.
  • Build Python evaluation pipelines that execute test cases, aggregate grader results, categorize failure modes, and compare model performance across releases.
  • Develop deterministic and model-assisted graders, using rule-based validation where objective checks are possible and rubric-driven grading for open-ended tasks.
  • Design repository-level coding-agent tasks covering debugging, feature implementation, API integration, refactoring, dependency changes, and test repair.
  • Build isolated Docker execution environments where agents can modify repositories, run terminal commands, execute tests, and interact with development tools safely.
  • Implement hidden tests, integration tests, filesystem validation, timeout controls, and modification checks to detect reward hacking, test tampering, hard-coded outputs, and false-positive completion.
  • Investigate recurring model failures including hallucination, tool misuse, invalid assumptions, incomplete execution, brittle reasoning, and instruction violations; convert findings into new benchmark and training examples.
  • Contribute to pairwise preference ranking and evaluator calibration, analysing reviewer disagreement and refining annotation guidelines to improve consistency.
  • Produce model-quality reporting covering pass rates, regressions, grader disagreement, failure categories, and performance across evaluation dimensions.
Jan 2021 - Dec 2024
United Kingdom

Software Engineer / Machine Learning Engineer

SCOTT LOGIC

Position summary
Software Engineer / Machine Learning Engineer at SCOTT LOGIC
Industries
Information Technology
Professional Services
Business areas
Information Technology
Product Development
Research and Development
  • Developed Python backend services and REST APIs for enterprise applications using FastAPI, Flask, and Django, with emphasis on maintainability, testing, and clear service boundaries.
  • Built data-processing pipelines for analytics and machine-learning applications, integrating structured and unstructured information from internal services, databases, and third-party APIs.
  • Designed backend data models and worked with PostgreSQL, SQL, Redis, caching, asynchronous processing, and application performance optimization.
  • Developed applied NLP features including text classification, information extraction, semantic search, similarity matching, and document retrieval.
  • Experimented with BERT and Sentence Transformers and integrated selected transformer models into application workflows through Python-based inference services.
  • Built embedding-based document retrieval and semantic-search systems, later extending the work into early Retrieval-Augmented Generation prototypes.
  • Integrated LLM APIs into internal tools and client prototypes for document Q&A, summarisation, structured extraction, and controlled text generation.
  • Created prompt templates, structured-output validation, retry logic, model-response comparison utilities, and automated tests measuring correctness, relevance, consistency, and formatting.
  • Collaborated with engineers, product managers, and data teams to move ML and generative-AI prototypes into testable, maintainable production services.
Jan 2018 - Dec 2021
United Kingdom

Software Engineer

SOFTWARE

Position summary
Software Engineer at SOFTWARE
Industries
Information Technology
Business areas
Information Technology
Product Development
Quality Assurance
  • Developed backend applications, web services, and internal platforms using Python across requirements, implementation, testing, deployment, and production support.
  • Designed REST APIs and integrations connecting internal services with external platforms and third-party systems.
  • Built and maintained PostgreSQL schemas and SQL queries, including migrations, indexing, transactional workflows, and performance tuning.
  • Developed frontend functionality using JavaScript, TypeScript, and React while collaborating with backend engineers, designers, and QA teams.
  • Implemented authentication, authorization, request validation, error handling, logging, and application security controls.
  • Wrote unit and integration tests, investigated production defects, containerized applications with Docker, and contributed to Git-based CI/CD workflows.
  • Built Python automation and data-processing scripts and gained early exposure to data analysis, machine-learning experiments, and NLP applications.

Coding Agent Evaluation Environment

Position summary
Coding Agent Evaluation Environment
Industries
Information Technology
Business areas
Information Technology
Quality Assurance

Built a containerized environment where coding agents receive repository-level tasks, modify code, execute commands, and are evaluated against deterministic and hidden tests. Added filesystem validation, timeouts, dependency isolation, and checks for reward hacking and unintended modifications.

LLM Evaluation & Benchmarking Framework

Position summary
LLM Evaluation & Benchmarking Framework
Industries
Information Technology
Business areas
Information Technology
Product Development
Quality Assurance
Research and Development

Developed a Python framework for evaluating multiple LLMs against structured datasets with deterministic graders, rubric-based evaluation, LLM-assisted grading, pairwise comparison, regression tracking, and failure categorization.

RAG & Tool-Use Evaluation

Position summary
RAG & Tool-Use Evaluation
Industries
Information Technology
Business areas
Information Technology
Research and Development

Built embedding/vector-search RAG workflows and evaluation scenarios for retrieval quality, groundedness, factual consistency, tool selection, API execution, multi-step reasoning, and error recovery.

Industry experience

See where this freelancer has spent most of their professional time.

Experienced in Information Technology and Professional Services.

Information Technology
Professional Services
Profile match chart

Business area experience

See which departments and functions this freelancer has contributed to most.

Experienced in Information Technology, Product Development, Quality Assurance, Research and Development, and Business Intelligence.

Information Technology
Product Development
Quality Assurance
Research and Development
Business Intelligence
Profile match chart

Summary

Software Engineer and AI Engineer with 8+ years of experience building backend systems, APIs, data platforms, machine-learning applications, and AI evaluation infrastructure. Strong Python engineering background with hands-on work across NLP, transformers, semantic search, RAG, and generative AI. Recent specialization in LLM evaluation, RLHF, preference data, post-training, coding-agent evaluation, and model-quality assessment. Experienced in designing evaluation datasets, benchmarks, graders, hidden-test environments, and containerized execution systems for large language models and AI agents.

Skills

  • Programming: Python, Javascript, Typescript, Sql, Bash
  • Backend & Apis: Fastapi, Flask, Django, Rest Apis, Microservices
  • Llm / Genai: Llm Evaluation, Rlhf, Rlaif, Sft, Preference Data, Preference Optimization, Prompt Engineering, Response Ranking, Llm-As-A-Judge, Tool / Function Calling, Agent Evaluation
  • Evaluation: Benchmark Development, Rubric Design, Pairwise Ranking, Human Evaluation, Automated Grading, Model Regression Testing, Failure Analysis, Safety Evaluation, Coding-Agent Evaluation
  • Ml / Nlp: Pytorch, Hugging Face Transformers, Bert, Sentence Transformers, Embeddings, Semantic Search, Classification, Information Extraction
  • Rag & Agents: Retrieval-Augmented Generation, Faiss, Vector Databases, Langchain, Langgraph, Agentic Workflows, Retrieval Evaluation
  • Infrastructure: Docker, Git, Github Actions, Ci/Cd, Linux, Aws, Postgresql, Mysql, Redis, Mongodb, Pytest

Languages

Urdu
Native
English
Advanced
Punjabi
Advanced

Education

Oct 2016 - Jun 2018

University of Leeds

MSc Advanced Computer Science · Advanced Computer Science · United Kingdom

Oct 2013 - Jun 2016

University of Birmingham

BSc Computer Science · Computer Science · United Kingdom

Statistics

Experience

Total positions 6
Experience in Information Technology 8.5 y
Avg length 1 y 9 m
Longest experience 3 y 11 m

Global experience

Countries worked in 1 (United Kingdom)
Primary country United Kingdom

Expertise

Recent roles Senior AI / LLM Evaluation Engineer, Software Engineer / Machine Learning Engineer, Software Engineer
Main industries Information Technology, Professional Services
Main business areas Information Technology, Product Development, Quality Assurance

Qualifications

Highest degree Master

Profile

Member since
Need a freelancer? Find your match in seconds.
Try FRATCH GPT
More actions

Frequently asked questions

Have questions? Find more information here.

Majid is based in Stockport, United Kingdom.

Majid speaks the following languages: Urdu (Native), English (Advanced), Punjabi (Advanced).

Majid has at least 9 years of experience. During this time, Majid has worked in at least 3 different roles and for 3 different companies. The average length of individual experience is 3 years and 10 months. Note that Majid may not have shared all experience and actually has more experience.

Based on recent experience, Majid would be well-suited for roles such as: Senior AI / LLM Evaluation Engineer, Software Engineer / Machine Learning Engineer, Software Engineer.

Majid's most recent position is Senior AI / LLM Evaluation Engineer at OPENCAST.

In recent years, Majid has worked for OPENCAST, SCOTT LOGIC, and SOFTWARE.

Majid is most experienced in industries like Information Technology and Professional Services.

Majid is most experienced in business areas like Information Technology, Product Development, and Quality Assurance. Majid also has some experience in Research and Development and Business Intelligence.

Majid has recently worked in industries like Information Technology and Professional Services.

Majid has recently worked in business areas like Information Technology, Product Development, and Quality Assurance.

Majid holds a Master in Advanced Computer Science from University of Leeds and a Bachelor in Computer Science from University of Birmingham.

Majid is immediately available full-time for suitable projects.

Daily rate distribution

0 1 2 3 4
<€320 €480-​640 €640-​800 €800-​960 €960+

The rates shown represent the typical market range for freelancers in this position based on recent contracts on our platform.

Average rates for similar positions

Rates are based on recent contracts and do not include FRATCH margin.

1000
750
500
250
Rate comparison chart
Daily rate avg. 751 €

The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.

1000
750
500
250
Rate comparison chart
Median rate 800 €

The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.

Calculated based on our freelancers’ daily rates as of 23 Aug 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.