Skip to main content
🇩🇪GDPR-compliant
Find experienced

PySpark Experts in Munich

matched in minutes from over 15,000 CVs with the power of AI.

Hire experts who build Spark jobs, streaming pipelines, and scalable data processing with PySpark, Apache Spark, and Python. Get vetted, available freelancers for clean delivery, fast handover, and reliable support.

Meet FRATCH Experts in Munich, who have recently used PySpark

Verified expert

Ajay Kumar Deekonda

View profile

Senior BI and Analytics Engineer

Munich
Ajay Kumar Deekonda

Last position:

Senior BI and Analytics Engineer at Novartis

  • Led enterprise reporting modernization by migrating legacy SSRS reporting solutions to Power BI, supporting 500+ business users while ensuring full GDPR/DSGVO compliance.
  • Designed and optimized Power BI and Microsoft Fabric semantic models using star schema, dimensional modeling, advanced DAX, and performance optimization techniques, reducing query latency by 25%.
  • Delivered 20+ executive and operational dashboards featuring KPI scorecards, drill-through, bookmarks, and row-level security, improving reporting efficiency by 20%.
  • Enabled self-service analytics through governed Power BI datasets, dataflows, and gateway architecture, increasing business-led reporting adoption by 35%.
  • Configured an incremental refresh policy and query folding for a 50+ million row sales dataset, reducing daily report refresh times by 85%.
  • Deployed automated ETL/ELT pipelines using Azure Data Factory, Microsoft Fabric, and Snowflake, reducing reporting delivery timelines by 40% through workflow automation.
  • Spearheaded Microsoft Fabric analytics modernization initiatives including lakehouse architecture, OneLake integration, and centralized data platform development, reducing data latency from 2 hours to 20 minutes.
  • Translated business requirements from 15+ stakeholders into scalable Power BI semantic models and dashboards, improving reporting consistency and reducing ad-hoc reporting requests by 25%.
  • Applied Microsoft Copilot and generative AI tools to accelerate SQL development, DAX authoring, technical documentation, and testing activities, reducing development effort by approximately 15 hours per week.
Verified expert

Philipp Grunert

View profile

Machine Learning & Data Engineer

München
Philipp Grunert

Last position:

Data Scientist & ML Engineer at Data-Science Factory GmbH

  • Building, implementing and selling automated Data Science solutions such as Scorecard Factory and Forecast Factory
  • Implementation of automated end-to-end cloud processes
  • Development of LLM and NLP models
  • Creation of interactive reports
  • Support for national and international large corporations as well as medium-sized companies in implementing ML projects
Verified expert

Thomas Hoefkens

View profile

Senior MLOps, DevOps Engineer

Munich
Thomas Hoefkens

Last position:

Senior MLOps, DevOps Engineer at Trianel Energy

  • Build and operate an end-to-end MLOps platform on Azure ML and Kubernetes (Kubeflow) for the automated deployment, monitoring, and scaling of forecasting models (including Temporal Fusion Transformer, Informer, Autoformer).
  • Implement CI/CD pipelines in Azure DevOps for the full ML lifecycle – from resource provisioning (Terraform), data transformation (Hugging Face Datasets, Pandas, PyTorch, CUDA cluster) through training and evaluation to model registry and endpoint deployment.
  • Integrate MLflow for experiment tracking, model versioning, performance monitoring, and automated registration in the Azure Model Registry.
  • Develop and containerize PyTorch training jobs (Azure Notebook, Jupyter Notebooks) for price and time series forecasting (PFC models) with automatic rollout via Azure ML Endpoints and REST/gRPC interfaces, Docker containerization, secured with OAuth 2.0.
  • Set up monitoring and alerting mechanisms (Prometheus, MLflow Metrics), log centralization, and cost monitoring.
  • Automate infrastructure provisioning and model deployment using Terraform, Helm, and Azure CLI; connect to existing market data systems and event pipelines.
  • Migrate existing workloads and databases (IONOS → Azure, MongoDB) with integration into central MLOps workflows and internal networks.
  • Extend the platform with LLM-based tools (LangChain, LangServe) to integrate GPT-based analysis modules into existing Spring Boot services for market anomaly detection and automated reports.
  • Analyze and architect a software solution to process large volumes of data efficiently (>3000 messages/sec.) (market data store).
  • Spring Boot / Java 21 container development with RabbitMQ for distributing stock market data via MongoDB (Kubernetes) with fast storage of data in Redis RMaps, deduplication, forwarding messages to Read Model queues, and building Read Models for UI display in MongoDB.
  • Integration of RESTHeart to create a REST API for MongoDB.
  • Build an Angular frontend to simplify data queries and master data maintenance.
  • Agentic coding with remote and local LLMs (Claude Sonnet, Ollama Qwen) and MCP servers.
  • Develop Python scripts for transforming and cleaning incoming stock market data (Pandas, scikit-learn).
Verified expert

Krithika Chand

View profile

Professional Reorientation

Garching
Krithika Chand

Last position:

Professional Reorientation at Von Rundstedt

  • Engaged in a structured career development program while strengthening German language proficiency (B1 level) and evaluating opportunities in ADAS/AD systems and requirements engineering.
Verified expert

Serge Kalinin

View profile

MLOps (machine learning operations)

Munich
Serge Kalinin

Last position:

MLOps (machine learning operations) at REWE Digital GmbH

  • It is like a startup within REWE, where we have to build a new forecasting system on Google Cloud Platform from the scratch. Although, officially my role is called MLOps, my actual tasks also include development of data processing pipelines (data engineering) and data scientists tasks such as feature engineering and model trainings.
  • GCP: Terraform (tofu), Vertex AI (Kubeflow), Cloud Run, IAM, Google Cloud Storage, BigQuery, Artifact Registry
  • Data engineering: Snowflake as the main data warehouse, Terraform, DBT for data model implementations
  • CI/CD: GitLab. We have built a CI/CD pipeline that automates deployments of new releases up to production environment
Verified expert

Michael Ternes

View profile

Senior DWH Developer

Munich
Michael Ternes

Last position:

ETL Developer at Insurance service provider

DWH for customer and financial data

  • Extension of the DWH with new data sources
  • Report development
  • Data quality management

Methodology: Scrum

Tools: Atlassian Confluence & Jira

Databases: Microsoft SQL Server

Programming languages: SQL, T-SQL

ETL: Microsoft SQL Server Integration Services (SSIS)

Frontend platform: PowerBI, Microsoft Reporting Services

Verified expert

Alexandre Savio

View profile

Cloud Engineer

München
Alexandre Savio

Last position:

Cloud Engineer at Dectris AG

  • Build a scalable multi-region backend service in AWS to serve remote desktop virtual machines for scientific analysis
  • Stack: AWS, GitHub, Terraform, Python, Rust
  • Built and defined the core infrastructure of the backend system
  • Defined and coded the virtual machines provisioning supporting Ubuntu and Rocky Linux desktop setups
  • Programmed the API service running in ECS to manage virtual machines and build custom Docker images for users
Verified expert

Axel Kraus

View profile

Data Engineer & Business Analyst

Munich
Axel Kraus

Last position:

Data Engineer & Business Analyst at Metafinanz

  • Migration of existing data jobs from Cognos Data Manager to Tibco/IBI Datamigrator
  • Migration data jobs parametrisation for dynamic runs
  • Optimisation and cutting-back
  • Regression tests
  • Knowledge transfer and documentation
Verified expert

Christian Schulz

View profile

Data-Scientist/AI Engineer

Ismaning
Christian Schulz

Last position:

Data-Scientist/AI Engineer at The Marcom Engine GmbH & Co. KG

  • Concept creation and implementing AI Agents in AWS Cloud
  • Continuously alignment with stakeholders
  • Collaborate with DevOps
  • Technologies: Git, CI/CD (GitHub Actions), Python/ML, Streamlit, Deno/typescript, AWS SAM, AWS Bedrock, AWS Lambda, AWS Dynamo DB, AWS S3, AWS Event Bridge etc.
Verified expert

Stephan Sahm

View profile

Senior Data/ML Consultant & Technical Lead

München
Stephan Sahm

Last position:

Senior Data/ML Consultant & Technical Lead at Jolin.io

  • Role: Software Engineer & Applied Mathematician (Mathematical optimization for scheduling; duration: 1 months; team setting: Team of 2, remote; technologies: JuMP, Julia, Pluto, Svelte, JavaScript, TypeScript, JetBrains Space, Terraform, Nomad)

  • Role: Software & Cloud & Web Engineer (Building scalable data science compute cluster from scratch; duration: 11 months; team setting: Team of 1, on-site; technologies: Terraform, Kubernetes, k8s ingress, k8s services, k8s RBAC, k8s networking, k3s, etcd, S3, DNS, certificates, Julia, Pluto, JavaScript, Tailwind, Astro, npm, Parcel, Preact, MUI, JWT, AWS SQS, AWS RDS, Python, GitLab, GitHub)

  • Role: AI & Web Engineer (Custom ChatGPT service; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Poetry, LangChain, Tailwind, ChatGPT API, Flask, FastAPI)

  • Role: Architect & Data Engineer (Central datalake setup and ingestion; duration: 9 months; team setting: Team of 5, remote; technologies: Infrastructure-as-code, AWS CDK, Python, Boto3, PySpark, AWS Glue, IAM, S3, ECS, Fargate, Lambda, Apache Hudi, DeltaLake, Databricks, GitHub, Jira, Miro)

  • Role: Software Engineer (PoC Julia migration of scikit-decide; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Julia, GitHub)

Verified expert

Stephan Baier

View profile

Freelance Data Scientist

Munich
Stephan Baier

Last position:

Freelance Data Scientist at Baier Data & AI Consulting

Verified expert

Maziyar Khorrami

View profile

Senior Data Engineer

Taufkirchen
Maziyar Khorrami

Last position:

Data Engineer at MSD Germany

  • Lead Architect to design and implement the data lake and ETL Pipeline using AWS Stack
  • Performance Optimization of Data Ingestion of ETL Pipeline
  • Development of Data Validation using Great Expectations
  • Leading of the data migration for two sources exchanges
  • Data Modeling in AWS Redshift

MLOps

  • Model inference implementation by mlflow and AWS SageMaker
  • Feature Engineering for the running ML Models ( Recommender Engineer, Clustering )
  • Implementatino of Model Registry and artifactory using mlflow
  • Historization an Profiling of the Input Data Using AWS Glue Crawler and AWS Data Catalog
  • Feature importance using mlflow

Tech. Stack: Python 3, AWS Glue, AWS Step Fucntion, AWS Lambda, AWS EventBridge, AWS IAM Role, AWS SageMaker, AWS EC2, AWS Glue Crawler, AWS CloudWatch, MLFlow, ETL, Data lake, GitHub Action, Terraform, Jenkins, Ansible playbooks (Infrastructure as Code), CI/CD, GitLab, SQL, PySparkSCRUM, Agile, Jira, BigData, VSCode, DBeaver, MSSQL, MySQL, grafana, Docker, Linux, Bash, MapReduce, Data Modeling (ORM), Pandas, YAML, SQL-Alchemy

Verified expert

Stefan Corsten

View profile

SQL, ETL, Reporting, DWH Development

Munich
Stefan Corsten

Last position:

SSIS Development at Stadtsparkasse München

  • Replacement of a Java application and the Oracle DB for loading the internal WerWasWo system using SSIS.
  • Development of SSIS packages to load text files into the database (SQL Server)
  • Development of a database project for deployment on various servers
  • Creation of queries to monitor the loading runs
  • Development of a PowerShell script to automate the deployment of the SSDT projects.
  • Oracle, SQL Developer, Microsoft SQL Server 2022 on-premises, SQL Server Management Studio v21, Visual Studio 2022, SSIS, SSDT, PowerShell.
Verified expert

Borui Li

View profile

Spectral Analysis of Neural Network Kernels

Munich
Borui Li

Last position:

Spectral Analysis of Neural Network Kernels at Borui Li Projects

  • Explored the impact of neural network structure on network-inspired kernels, such as Neural Tangent Kernel (NTK).
  • Demonstrated through theoretical analysis and empirical studies that the RKHS of NNGP is a subspace of NTK.
  • Explored the connections between these kernels and the Matérn family.

Discover over 15,000 top freelancers

Statistics of experts using PySpark

Aggregated from the professional profiles of matched freelancers.

Experience

16 years (Germany: 13 years)

Position duration

1.9 years (Germany: 2.8 years)

Positions per freelancer

13 (Germany: 10)

Top business areas

Information Technology, Business Intelligence, Product Development

Top industries

Information Technology, Professional Services, Healthcare

Certification focus areas

Information Technology, Business Intelligence, Project Management

Bachelor's degree or higher

100% (Germany: 96%)

Master's degree or higher

85% (Germany: 71%)

Doctorate

30% (Germany: 13%)

Certifications per freelancer

3 (Germany: 4)

Most common languages

German, English, Spanish

Speak two or more languages

100% (Germany: 96%)

Based on our profile pool as of 30 Aug 2026.

Daily rate distribution

0 3 6 9 12
<€800 €800-​1200 €1200-​1600 €1600+

The chart shows how the daily rates of freelancers in this technology in Munich are distributed, based on recent contracts on our platform. Each bar covers a rate range — its height shows how many freelancers charge within that range.

Average rates of experts in Munich using PySpark

Rates are based on recent contracts and do not include FRATCH margin.

1000
750
500
250
Rate comparison chart
Daily rate avg. 772 €
Germany avg. 756 €

The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.

1000
750
500
250
Rate comparison chart
Median rate 800 €
Germany median 800 €

The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.

Calculated based on our freelancers’ daily rates as of 30 Aug 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.

About the technology

PySpark in practice

PySpark is the Python API for Apache Spark. Teams use it to process large data sets, transform raw feeds, and build batch or streaming pipelines that run across clusters. It is a fit when Python skills matter, but the work must scale beyond a single machine.

What specialists deliver

  • ETL and ELT jobs for data lakes and warehouses
  • Streaming pipelines for near real-time feeds
  • Data quality checks, joins, and large-scale transformations
  • Notebook work that moves into production code

Stack around it

Strong professionals know Spark DataFrames, Spark SQL, joins, partitions, caching, and job tuning. They also work with Py4J, Parquet, Delta Lake, Kafka, cloud storage, and orchestration tools such as Airflow. Good code keeps Python logic clear and Spark execution efficient.

When to bring help

Companies bring in freelance experts when a pipeline is slow, unstable, or hard to maintain. They also need support for migrations from pandas or old Spark code, new data products, and handovers to internal teams. In Munich, this often comes up in data-heavy sectors that need clean collaboration and clear documentation.

What strong experts do

They write code that is easy to test, keep data types explicit, and avoid unnecessary shuffles and wide transformations. They understand executor memory, partition sizing, and how Spark behaves on a cluster. They also explain trade-offs clearly, which makes reviews and production support much easier.

Munich projects

PySpark work in Munich often sits close to analytics, reporting, platform engineering, and cloud data stacks. Freelance specialists can work remote or on-site, depending on access needs and team setup. Clear English is common in technical teams, while some projects also expect German for coordination with business stakeholders.

Published on:
FRATCH GPT

FRATCH GPT delivers freelancer proposals with clear reasoning and transparent pricing in minutes, helping your hiring department quickly and compliantly find the best talent.

Give it a try:

Try FRATCH GPT

Frequently asked questions

What clients ask us most about PySpark — answered in short.

PySpark is used to process large data sets with Python on Apache Spark. Companies use it for ETL, data cleansing, feature preparation, batch pipelines, and streaming jobs. It is a practical choice when Python teams need distributed processing without switching languages.

PySpark gives Python users access to Spark’s distributed engine, while pandas stays in memory on one machine. Compared with Scala Spark, it is often easier for Python teams to adopt, though some low-level Spark work still benefits from Scala knowledge. The right choice depends on scale, team skills, and how much Spark tuning the project needs.

A strong PySpark specialist usually knows Spark SQL, DataFrames, partitioning, joins, caching, and performance tuning. Adjacent skills often include SQL, data modeling, Kafka, Airflow, Delta Lake, and cloud storage such as S3 or ADLS. Testing and clean Git-based delivery matter as much as code speed.

PySpark projects differ a lot, but anything that touches production data usually needs someone who has worked with Spark internals, not just notebooks. Simple transformations can be straightforward, while streaming, cost control, and cluster tuning need deeper knowledge. Ask for examples that match your pipeline shape, data size, and reliability goals.

Bring in PySpark specialists when a pipeline is too slow, a migration is blocked, or a data product must go live under time pressure. Freelancers are also useful for short focused work such as code reviews, optimization, or fixing broken jobs. This fits well when you need hands-on delivery without adding long-term headcount.

Yes, most PySpark work can be done remotely because the core tasks are code, data logic, and cluster configuration. On-site time in Munich can help when access, security, or stakeholder workshops are part of the project. Many teams use a hybrid setup with remote delivery and local kickoff meetings.

Look for clear examples of production Spark work, not just notebooks or tutorials. A good PySpark expert can explain why a job is slow, how to reduce shuffles, and how they handle schema changes, retries, and monitoring. Clean reviews, practical testing habits, and readable Spark SQL are strong signs.

PySpark is the Python interface for Apache Spark, not a separate engine. If a project mentions Spark, Spark SQL, or Spark jobs, the work may still be PySpark-based when the team codes in Python. Search for both names when you want the best match for a data engineering task.

The average hourly rate of freelancers in Munich, Germany who have used PySpark in their recent projects is 97 €, which corresponds to a daily rate of about 772 € based on an 8-hour working day.

Of the freelancers in Munich, Germany who have used PySpark in their recent projects, 100% hold at least a Bachelor's degree, 85% hold at least a Master's degree, and 30% hold a doctorate.

On average, freelancers in Munich, Germany who have used PySpark in their recent projects have 16 years of professional experience, with a single engagement typically lasting around 1.9 years.

The most common languages among freelancers in Munich, Germany who have used PySpark in their recent projects are German (100%), English (100%), and Spanish (32%).

The most common industries among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (95%), Professional Services (59%), and Healthcare (55%).

The most common business areas among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (95%), Business Intelligence (86%), and Product Development (77%).

Main locations of FRATCH Experts, who have recently used PySpark

Our freelancers and interim experts are at home across the DACH region — available on-site in the major business hubs or fully remote. Choose a location to discover matched specialists, local market insights and up-to-date availability.

Berlin Hamburg Munich Cologne Frankfurt Stuttgart Dusseldorf Leipzig Dortmund Essen Bremen Dresden Hanover Nuremberg

Request a free demo

Get in touch with the FRATCH team and we will get back to you within 4 hours.

Contact form

Would you rather directly get in touch?
We always have the time for a call or email!

FRATCH CEO avatar

Philipp Thomaschewski

FRATCH CEO

LinkedInFRATCH