PySpark Experts in Munich
matched in minutes from over 15,000 CVs with the power of AI.Hire experts who build Spark jobs, streaming pipelines, and scalable data processing with PySpark, Apache Spark, and Python. Get vetted, available freelancers for clean delivery, fast handover, and reliable support.
Meet FRATCH Experts in Munich, who have recently used PySpark
Michael Nelz
Last position:
Senior ML Engineer, AI Engineer at Lanxess AG
- Deployment and scaling of existing ML initiatives, including demand and cash flow forecasts.
- Building robust monitoring with mlflow for data stability, model performance, and drift detection, as well as implementing additional ML use cases.
- Further development of an Agentic AI chatbot for transparent and easy-to-understand model explanations.
Ajay Kumar Deekonda
Last position:
Senior BI and Analytics Engineer at Novartis
- Led enterprise reporting modernization by migrating legacy SSRS reporting solutions to Power BI, supporting 500+ business users while ensuring full GDPR/DSGVO compliance.
- Designed and optimized Power BI and Microsoft Fabric semantic models using star schema, dimensional modeling, advanced DAX, and performance optimization techniques, reducing query latency by 25%.
- Delivered 20+ executive and operational dashboards featuring KPI scorecards, drill-through, bookmarks, and row-level security, improving reporting efficiency by 20%.
- Enabled self-service analytics through governed Power BI datasets, dataflows, and gateway architecture, increasing business-led reporting adoption by 35%.
- Configured an incremental refresh policy and query folding for a 50+ million row sales dataset, reducing daily report refresh times by 85%.
- Deployed automated ETL/ELT pipelines using Azure Data Factory, Microsoft Fabric, and Snowflake, reducing reporting delivery timelines by 40% through workflow automation.
- Spearheaded Microsoft Fabric analytics modernization initiatives including lakehouse architecture, OneLake integration, and centralized data platform development, reducing data latency from 2 hours to 20 minutes.
- Translated business requirements from 15+ stakeholders into scalable Power BI semantic models and dashboards, improving reporting consistency and reducing ad-hoc reporting requests by 25%.
- Applied Microsoft Copilot and generative AI tools to accelerate SQL development, DAX authoring, technical documentation, and testing activities, reducing development effort by approximately 15 hours per week.
Philipp Grunert
Last position:
Data Scientist & ML Engineer at Data-Science Factory GmbH
- Building, implementing and selling automated Data Science solutions such as Scorecard Factory and Forecast Factory
- Implementation of automated end-to-end cloud processes
- Development of LLM and NLP models
- Creation of interactive reports
- Support for national and international large corporations as well as medium-sized companies in implementing ML projects
Thomas Hoefkens
Last position:
Senior MLOps, DevOps Engineer at Trianel Energy
- Build and operate an end-to-end MLOps platform on Azure ML and Kubernetes (Kubeflow) for the automated deployment, monitoring, and scaling of forecasting models (including Temporal Fusion Transformer, Informer, Autoformer).
- Implement CI/CD pipelines in Azure DevOps for the full ML lifecycle – from resource provisioning (Terraform), data transformation (Hugging Face Datasets, Pandas, PyTorch, CUDA cluster) through training and evaluation to model registry and endpoint deployment.
- Integrate MLflow for experiment tracking, model versioning, performance monitoring, and automated registration in the Azure Model Registry.
- Develop and containerize PyTorch training jobs (Azure Notebook, Jupyter Notebooks) for price and time series forecasting (PFC models) with automatic rollout via Azure ML Endpoints and REST/gRPC interfaces, Docker containerization, secured with OAuth 2.0.
- Set up monitoring and alerting mechanisms (Prometheus, MLflow Metrics), log centralization, and cost monitoring.
- Automate infrastructure provisioning and model deployment using Terraform, Helm, and Azure CLI; connect to existing market data systems and event pipelines.
- Migrate existing workloads and databases (IONOS → Azure, MongoDB) with integration into central MLOps workflows and internal networks.
- Extend the platform with LLM-based tools (LangChain, LangServe) to integrate GPT-based analysis modules into existing Spring Boot services for market anomaly detection and automated reports.
- Analyze and architect a software solution to process large volumes of data efficiently (>3000 messages/sec.) (market data store).
- Spring Boot / Java 21 container development with RabbitMQ for distributing stock market data via MongoDB (Kubernetes) with fast storage of data in Redis RMaps, deduplication, forwarding messages to Read Model queues, and building Read Models for UI display in MongoDB.
- Integration of RESTHeart to create a REST API for MongoDB.
- Build an Angular frontend to simplify data queries and master data maintenance.
- Agentic coding with remote and local LLMs (Claude Sonnet, Ollama Qwen) and MCP servers.
- Develop Python scripts for transforming and cleaning incoming stock market data (Pandas, scikit-learn).
Krithika Chand
Last position:
Professional Reorientation at Von Rundstedt
- Engaged in a structured career development program while strengthening German language proficiency (B1 level) and evaluating opportunities in ADAS/AD systems and requirements engineering.
Serge Kalinin
Last position:
MLOps (machine learning operations) at REWE Digital GmbH
- It is like a startup within REWE, where we have to build a new forecasting system on Google Cloud Platform from the scratch. Although, officially my role is called MLOps, my actual tasks also include development of data processing pipelines (data engineering) and data scientists tasks such as feature engineering and model trainings.
- GCP: Terraform (tofu), Vertex AI (Kubeflow), Cloud Run, IAM, Google Cloud Storage, BigQuery, Artifact Registry
- Data engineering: Snowflake as the main data warehouse, Terraform, DBT for data model implementations
- CI/CD: GitLab. We have built a CI/CD pipeline that automates deployments of new releases up to production environment
Michael Ternes
Last position:
ETL Developer at Insurance service provider
DWH for customer and financial data
- Extension of the DWH with new data sources
- Report development
- Data quality management
Methodology: Scrum
Tools: Atlassian Confluence & Jira
Databases: Microsoft SQL Server
Programming languages: SQL, T-SQL
ETL: Microsoft SQL Server Integration Services (SSIS)
Frontend platform: PowerBI, Microsoft Reporting Services
Alexandre Savio
Last position:
Cloud Engineer at Dectris AG
- Build a scalable multi-region backend service in AWS to serve remote desktop virtual machines for scientific analysis
- Stack: AWS, GitHub, Terraform, Python, Rust
- Built and defined the core infrastructure of the backend system
- Defined and coded the virtual machines provisioning supporting Ubuntu and Rocky Linux desktop setups
- Programmed the API service running in ECS to manage virtual machines and build custom Docker images for users
Axel Kraus
Last position:
Data Engineer & Business Analyst at Metafinanz
- Migration of existing data jobs from Cognos Data Manager to Tibco/IBI Datamigrator
- Migration data jobs parametrisation for dynamic runs
- Optimisation and cutting-back
- Regression tests
- Knowledge transfer and documentation
Christian Schulz
Last position:
Data-Scientist/AI Engineer at The Marcom Engine GmbH & Co. KG
- Concept creation and implementing AI Agents in AWS Cloud
- Continuously alignment with stakeholders
- Collaborate with DevOps
- Technologies: Git, CI/CD (GitHub Actions), Python/ML, Streamlit, Deno/typescript, AWS SAM, AWS Bedrock, AWS Lambda, AWS Dynamo DB, AWS S3, AWS Event Bridge etc.
Stephan Sahm
Last position:
Senior Data/ML Consultant & Technical Lead at Jolin.io
Role: Software Engineer & Applied Mathematician (Mathematical optimization for scheduling; duration: 1 months; team setting: Team of 2, remote; technologies: JuMP, Julia, Pluto, Svelte, JavaScript, TypeScript, JetBrains Space, Terraform, Nomad)
Role: Software & Cloud & Web Engineer (Building scalable data science compute cluster from scratch; duration: 11 months; team setting: Team of 1, on-site; technologies: Terraform, Kubernetes, k8s ingress, k8s services, k8s RBAC, k8s networking, k3s, etcd, S3, DNS, certificates, Julia, Pluto, JavaScript, Tailwind, Astro, npm, Parcel, Preact, MUI, JWT, AWS SQS, AWS RDS, Python, GitLab, GitHub)
Role: AI & Web Engineer (Custom ChatGPT service; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Poetry, LangChain, Tailwind, ChatGPT API, Flask, FastAPI)
Role: Architect & Data Engineer (Central datalake setup and ingestion; duration: 9 months; team setting: Team of 5, remote; technologies: Infrastructure-as-code, AWS CDK, Python, Boto3, PySpark, AWS Glue, IAM, S3, ECS, Fargate, Lambda, Apache Hudi, DeltaLake, Databricks, GitHub, Jira, Miro)
Role: Software Engineer (PoC Julia migration of scikit-decide; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Julia, GitHub)
Stephan Baier
Last position:
Freelance Data Scientist at Baier Data & AI Consulting
Maziyar Khorrami
Last position:
Data Engineer at MSD Germany
- Lead Architect to design and implement the data lake and ETL Pipeline using AWS Stack
- Performance Optimization of Data Ingestion of ETL Pipeline
- Development of Data Validation using Great Expectations
- Leading of the data migration for two sources exchanges
- Data Modeling in AWS Redshift
MLOps
- Model inference implementation by mlflow and AWS SageMaker
- Feature Engineering for the running ML Models ( Recommender Engineer, Clustering )
- Implementatino of Model Registry and artifactory using mlflow
- Historization an Profiling of the Input Data Using AWS Glue Crawler and AWS Data Catalog
- Feature importance using mlflow
Tech. Stack: Python 3, AWS Glue, AWS Step Fucntion, AWS Lambda, AWS EventBridge, AWS IAM Role, AWS SageMaker, AWS EC2, AWS Glue Crawler, AWS CloudWatch, MLFlow, ETL, Data lake, GitHub Action, Terraform, Jenkins, Ansible playbooks (Infrastructure as Code), CI/CD, GitLab, SQL, PySparkSCRUM, Agile, Jira, BigData, VSCode, DBeaver, MSSQL, MySQL, grafana, Docker, Linux, Bash, MapReduce, Data Modeling (ORM), Pandas, YAML, SQL-Alchemy
Stefan Corsten
Last position:
SSIS Development at Stadtsparkasse München
- Replacement of a Java application and the Oracle DB for loading the internal WerWasWo system using SSIS.
- Development of SSIS packages to load text files into the database (SQL Server)
- Development of a database project for deployment on various servers
- Creation of queries to monitor the loading runs
- Development of a PowerShell script to automate the deployment of the SSDT projects.
- Oracle, SQL Developer, Microsoft SQL Server 2022 on-premises, SQL Server Management Studio v21, Visual Studio 2022, SSIS, SSDT, PowerShell.
Borui Li
Last position:
Spectral Analysis of Neural Network Kernels at Borui Li Projects
- Explored the impact of neural network structure on network-inspired kernels, such as Neural Tangent Kernel (NTK).
- Demonstrated through theoretical analysis and empirical studies that the RKHS of NNGP is a subspace of NTK.
- Explored the connections between these kernels and the Matérn family.
Discover over 15,000 top freelancers
Statistics of experts using PySpark
Aggregated from the professional profiles of matched freelancers.
Experience
16 years (Germany: 13 years)
Position duration
1.9 years (Germany: 2.8 years)
Positions per freelancer
13 (Germany: 10)
Top business areas
Information Technology, Business Intelligence, Product Development
Top industries
Information Technology, Professional Services, Healthcare
Certification focus areas
Information Technology, Business Intelligence, Project Management
Bachelor's degree or higher
100% (Germany: 96%)
Master's degree or higher
85% (Germany: 71%)
Doctorate
30% (Germany: 13%)
Certifications per freelancer
3 (Germany: 4)
Most common languages
German, English, Spanish
Speak two or more languages
100% (Germany: 96%)
Based on our profile pool as of 30 Aug 2026.
Daily rate distribution
The chart shows how the daily rates of freelancers in this technology in Munich are distributed, based on recent contracts on our platform. Each bar covers a rate range — its height shows how many freelancers charge within that range.
Average rates of experts in Munich using PySpark
Rates are based on recent contracts and do not include FRATCH margin.
The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.
The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.
Calculated based on our freelancers’ daily rates as of 30 Aug 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.
About the technology
PySpark in practice
PySpark is the Python API for Apache Spark. Teams use it to process large data sets, transform raw feeds, and build batch or streaming pipelines that run across clusters. It is a fit when Python skills matter, but the work must scale beyond a single machine.
What specialists deliver
- ETL and ELT jobs for data lakes and warehouses
- Streaming pipelines for near real-time feeds
- Data quality checks, joins, and large-scale transformations
- Notebook work that moves into production code
Stack around it
Strong professionals know Spark DataFrames, Spark SQL, joins, partitions, caching, and job tuning. They also work with Py4J, Parquet, Delta Lake, Kafka, cloud storage, and orchestration tools such as Airflow. Good code keeps Python logic clear and Spark execution efficient.
When to bring help
Companies bring in freelance experts when a pipeline is slow, unstable, or hard to maintain. They also need support for migrations from pandas or old Spark code, new data products, and handovers to internal teams. In Munich, this often comes up in data-heavy sectors that need clean collaboration and clear documentation.
What strong experts do
They write code that is easy to test, keep data types explicit, and avoid unnecessary shuffles and wide transformations. They understand executor memory, partition sizing, and how Spark behaves on a cluster. They also explain trade-offs clearly, which makes reviews and production support much easier.
Munich projects
PySpark work in Munich often sits close to analytics, reporting, platform engineering, and cloud data stacks. Freelance specialists can work remote or on-site, depending on access needs and team setup. Clear English is common in technical teams, while some projects also expect German for coordination with business stakeholders.
Frequently asked questions
What clients ask us most about PySpark — answered in short.
PySpark is used to process large data sets with Python on Apache Spark. Companies use it for ETL, data cleansing, feature preparation, batch pipelines, and streaming jobs. It is a practical choice when Python teams need distributed processing without switching languages.
PySpark gives Python users access to Spark’s distributed engine, while pandas stays in memory on one machine. Compared with Scala Spark, it is often easier for Python teams to adopt, though some low-level Spark work still benefits from Scala knowledge. The right choice depends on scale, team skills, and how much Spark tuning the project needs.
A strong PySpark specialist usually knows Spark SQL, DataFrames, partitioning, joins, caching, and performance tuning. Adjacent skills often include SQL, data modeling, Kafka, Airflow, Delta Lake, and cloud storage such as S3 or ADLS. Testing and clean Git-based delivery matter as much as code speed.
PySpark projects differ a lot, but anything that touches production data usually needs someone who has worked with Spark internals, not just notebooks. Simple transformations can be straightforward, while streaming, cost control, and cluster tuning need deeper knowledge. Ask for examples that match your pipeline shape, data size, and reliability goals.
Bring in PySpark specialists when a pipeline is too slow, a migration is blocked, or a data product must go live under time pressure. Freelancers are also useful for short focused work such as code reviews, optimization, or fixing broken jobs. This fits well when you need hands-on delivery without adding long-term headcount.
Yes, most PySpark work can be done remotely because the core tasks are code, data logic, and cluster configuration. On-site time in Munich can help when access, security, or stakeholder workshops are part of the project. Many teams use a hybrid setup with remote delivery and local kickoff meetings.
Look for clear examples of production Spark work, not just notebooks or tutorials. A good PySpark expert can explain why a job is slow, how to reduce shuffles, and how they handle schema changes, retries, and monitoring. Clean reviews, practical testing habits, and readable Spark SQL are strong signs.
PySpark is the Python interface for Apache Spark, not a separate engine. If a project mentions Spark, Spark SQL, or Spark jobs, the work may still be PySpark-based when the team codes in Python. Search for both names when you want the best match for a data engineering task.
The average hourly rate of freelancers in Munich, Germany who have used PySpark in their recent projects is 97 €, which corresponds to a daily rate of about 772 € based on an 8-hour working day.
Of the freelancers in Munich, Germany who have used PySpark in their recent projects, 100% hold at least a Bachelor's degree, 85% hold at least a Master's degree, and 30% hold a doctorate.
On average, freelancers in Munich, Germany who have used PySpark in their recent projects have 16 years of professional experience, with a single engagement typically lasting around 1.9 years.
The most common languages among freelancers in Munich, Germany who have used PySpark in their recent projects are German (100%), English (100%), and Spanish (32%).
The most common industries among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (95%), Professional Services (59%), and Healthcare (55%).
The most common business areas among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (95%), Business Intelligence (86%), and Product Development (77%).
Main locations of FRATCH Experts, who have recently used PySpark
Our freelancers and interim experts are at home across the DACH region — available on-site in the major business hubs or fully remote. Choose a location to discover matched specialists, local market insights and up-to-date availability.
Request a free demo
Get in touch with the FRATCH team and we will get back to you within 4 hours.
Would you rather directly get in touch?
We always have the time for a call or email!

Berlin
Frankfurt