
PySpark Experts in Munich
matched in minutes from over 15,000 CVsHire experts who process large datasets, design reliable Spark pipelines and connect Python workloads to cloud data platforms. FRATCH matches you quickly and precisely with vetted, available freelancers who fit your project.
Meet FRATCH Experts in Munich, who have recently used PySpark
Michael N.
Last position:
Senior AI Engineer | Forward Deployed Engineer at Tiefbau
- Development of an AI-powered project organization tool for a civil engineering company that intelligently links project, task, tender, schedule, and document data through a knowledge graph.
- Implementation of AI features for document analysis, information extraction, context-based assistance, and voice-based data capture based on Microsoft Azure AI, reducing administrative effort, making information available faster, and supporting project teams in decision-making.
- Tech stack: Python, React, TypeScript, FastAPI, Claude Code, Codex, Graphify, PostgreSQL, Microsoft Azure AI Foundry, Azure OpenAI, Azure AI Speech, Azure AI Document Intelligence, Microsoft Graph, Microsoft Entra ID, Docker, Git, CI/CD.
Mirza K.
Last position:
Agentic Automation and a RAG system
- This project involved extraction of intelligence data to support report writing for a company that provides geopolitical, global, commercial intelligence. The data have been gathered from a number of resources (interview transcripts, online data, internal documents), and then a knowledge base has been build from it. This was the basis of a complex RAG system, that was evaluated against a golden dataset. Agents have been used to find out the contradicting intelligence, the statements supporting each other, and to store back the generated knowledge.
Used: Python, RAG, LangGraph, LangChain, deepeval, MCP
Ajay Kumar D.
Last position:
Senior BI and Analytics Engineer at Novartis
- Led enterprise reporting modernization by migrating legacy SSRS reporting solutions to Power BI, supporting 500+ business users while ensuring full GDPR/DSGVO compliance.
- Designed and optimized Power BI and Microsoft Fabric semantic models using star schema, dimensional modeling, advanced DAX, and performance optimization techniques, reducing query latency by 25%.
- Delivered 20+ executive and operational dashboards featuring KPI scorecards, drill-through, bookmarks, and row-level security, improving reporting efficiency by 20%.
- Enabled self-service analytics through governed Power BI datasets, dataflows, and gateway architecture, increasing business-led reporting adoption by 35%.
- Configured an incremental refresh policy and query folding for a 50+ million row sales dataset, reducing daily report refresh times by 85%.
- Deployed automated ETL/ELT pipelines using Azure Data Factory, Microsoft Fabric, and Snowflake, reducing reporting delivery timelines by 40% through workflow automation.
- Spearheaded Microsoft Fabric analytics modernization initiatives including lakehouse architecture, OneLake integration, and centralized data platform development, reducing data latency from 2 hours to 20 minutes.
- Translated business requirements from 15+ stakeholders into scalable Power BI semantic models and dashboards, improving reporting consistency and reducing ad-hoc reporting requests by 25%.
- Applied Microsoft Copilot and generative AI tools to accelerate SQL development, DAX authoring, technical documentation, and testing activities, reducing development effort by approximately 15 hours per week.
Philipp G.
Last position:
Data Scientist & ML Engineer at Data-Science Factory GmbH
- Building, implementing and selling automated Data Science solutions such as Scorecard Factory and Forecast Factory
- Implementation of automated end-to-end cloud processes
- Development of LLM and NLP models
- Creation of interactive reports
- Support for national and international large corporations as well as medium-sized companies in implementing ML projects
Thomas H.
Last position:
Senior MLOps, DevOps Engineer at Trianel Energy
- Build and operate an end-to-end MLOps platform on Azure ML and Kubernetes (Kubeflow) for the automated deployment, monitoring, and scaling of forecasting models (including Temporal Fusion Transformer, Informer, Autoformer).
- Implement CI/CD pipelines in Azure DevOps for the full ML lifecycle – from resource provisioning (Terraform), data transformation (Hugging Face Datasets, Pandas, PyTorch, CUDA cluster) through training and evaluation to model registry and endpoint deployment.
- Integrate MLflow for experiment tracking, model versioning, performance monitoring, and automated registration in the Azure Model Registry.
- Develop and containerize PyTorch training jobs (Azure Notebook, Jupyter Notebooks) for price and time series forecasting (PFC models) with automatic rollout via Azure ML Endpoints and REST/gRPC interfaces, Docker containerization, secured with OAuth 2.0.
- Set up monitoring and alerting mechanisms (Prometheus, MLflow Metrics), log centralization, and cost monitoring.
- Automate infrastructure provisioning and model deployment using Terraform, Helm, and Azure CLI; connect to existing market data systems and event pipelines.
- Migrate existing workloads and databases (IONOS → Azure, MongoDB) with integration into central MLOps workflows and internal networks.
- Extend the platform with LLM-based tools (LangChain, LangServe) to integrate GPT-based analysis modules into existing Spring Boot services for market anomaly detection and automated reports.
- Analyze and architect a software solution to process large volumes of data efficiently (>3000 messages/sec.) (market data store).
- Spring Boot / Java 21 container development with RabbitMQ for distributing stock market data via MongoDB (Kubernetes) with fast storage of data in Redis RMaps, deduplication, forwarding messages to Read Model queues, and building Read Models for UI display in MongoDB.
- Integration of RESTHeart to create a REST API for MongoDB.
- Build an Angular frontend to simplify data queries and master data maintenance.
- Agentic coding with remote and local LLMs (Claude Sonnet, Ollama Qwen) and MCP servers.
- Develop Python scripts for transforming and cleaning incoming stock market data (Pandas, scikit-learn).
Krithika C.
Last position:
Professional Reorientation at Von Rundstedt
- Engaged in a structured career development program while strengthening German language proficiency (B1 level) and evaluating opportunities in ADAS/AD systems and requirements engineering.
Serge K.
Last position:
MLOps (machine learning operations) at REWE Digital GmbH
- It is like a startup within REWE, where we have to build a new forecasting system on Google Cloud Platform from the scratch. Although, officially my role is called MLOps, my actual tasks also include development of data processing pipelines (data engineering) and data scientists tasks such as feature engineering and model trainings.
- GCP: Terraform (tofu), Vertex AI (Kubeflow), Cloud Run, IAM, Google Cloud Storage, BigQuery, Artifact Registry
- Data engineering: Snowflake as the main data warehouse, Terraform, DBT for data model implementations
- CI/CD: GitLab. We have built a CI/CD pipeline that automates deployments of new releases up to production environment
Michael T.
Last position:
ETL Developer at Insurance service provider
DWH for customer and financial data
- Extension of the DWH with new data sources
- Report development
- Data quality management
Methodology: Scrum
Tools: Atlassian Confluence & Jira
Databases: Microsoft SQL Server
Programming languages: SQL, T-SQL
ETL: Microsoft SQL Server Integration Services (SSIS)
Frontend platform: PowerBI, Microsoft Reporting Services
Axel K.
Last position:
Data Engineer & Business Analyst at Metafinanz
- Migration of existing data jobs from Cognos Data Manager to Tibco/IBI Datamigrator
- Migration data jobs parametrisation for dynamic runs
- Optimisation and cutting-back
- Regression tests
- Knowledge transfer and documentation
Stephan B.
Last position:
Freelance Data Scientist at Baier Data & AI Consulting
Alexandre S.
Last position:
Cloud Engineer at Dectris AG
- Build a scalable multi-region backend service in AWS to serve remote desktop virtual machines for scientific analysis
- Stack: AWS, GitHub, Terraform, Python, Rust
- Built and defined the core infrastructure of the backend system
- Defined and coded the virtual machines provisioning supporting Ubuntu and Rocky Linux desktop setups
- Programmed the API service running in ECS to manage virtual machines and build custom Docker images for users
Christian S.
Last position:
Data-Scientist/AI Engineer at The Marcom Engine GmbH & Co. KG
- Concept creation and implementing AI Agents in AWS Cloud
- Continuously alignment with stakeholders
- Collaborate with DevOps
- Technologies: Git, CI/CD (GitHub Actions), Python/ML, Streamlit, Deno/typescript, AWS SAM, AWS Bedrock, AWS Lambda, AWS Dynamo DB, AWS S3, AWS Event Bridge etc.
Stephan S.
Last position:
Senior Data/ML Consultant & Technical Lead at Jolin.io
Role: Software Engineer & Applied Mathematician (Mathematical optimization for scheduling; duration: 1 months; team setting: Team of 2, remote; technologies: JuMP, Julia, Pluto, Svelte, JavaScript, TypeScript, JetBrains Space, Terraform, Nomad)
Role: Software & Cloud & Web Engineer (Building scalable data science compute cluster from scratch; duration: 11 months; team setting: Team of 1, on-site; technologies: Terraform, Kubernetes, k8s ingress, k8s services, k8s RBAC, k8s networking, k3s, etcd, S3, DNS, certificates, Julia, Pluto, JavaScript, Tailwind, Astro, npm, Parcel, Preact, MUI, JWT, AWS SQS, AWS RDS, Python, GitLab, GitHub)
Role: AI & Web Engineer (Custom ChatGPT service; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Poetry, LangChain, Tailwind, ChatGPT API, Flask, FastAPI)
Role: Architect & Data Engineer (Central datalake setup and ingestion; duration: 9 months; team setting: Team of 5, remote; technologies: Infrastructure-as-code, AWS CDK, Python, Boto3, PySpark, AWS Glue, IAM, S3, ECS, Fargate, Lambda, Apache Hudi, DeltaLake, Databricks, GitHub, Jira, Miro)
Role: Software Engineer (PoC Julia migration of scikit-decide; duration: 1 months; team setting: Team of 2, remote; technologies: Python, Julia, GitHub)
Maziyar K.
Last position:
Data Engineer at MSD Germany
- Lead Architect to design and implement the data lake and ETL Pipeline using AWS Stack
- Performance Optimization of Data Ingestion of ETL Pipeline
- Development of Data Validation using Great Expectations
- Leading of the data migration for two sources exchanges
- Data Modeling in AWS Redshift
MLOps
- Model inference implementation by mlflow and AWS SageMaker
- Feature Engineering for the running ML Models ( Recommender Engineer, Clustering )
- Implementatino of Model Registry and artifactory using mlflow
- Historization an Profiling of the Input Data Using AWS Glue Crawler and AWS Data Catalog
- Feature importance using mlflow
Tech. Stack: Python 3, AWS Glue, AWS Step Fucntion, AWS Lambda, AWS EventBridge, AWS IAM Role, AWS SageMaker, AWS EC2, AWS Glue Crawler, AWS CloudWatch, MLFlow, ETL, Data lake, GitHub Action, Terraform, Jenkins, Ansible playbooks (Infrastructure as Code), CI/CD, GitLab, SQL, PySparkSCRUM, Agile, Jira, BigData, VSCode, DBeaver, MSSQL, MySQL, grafana, Docker, Linux, Bash, MapReduce, Data Modeling (ORM), Pandas, YAML, SQL-Alchemy
Stefan C.
Last position:
SSIS Development at Stadtsparkasse München
- Replacement of a Java application and the Oracle DB for loading the internal WerWasWo system using SSIS.
- Development of SSIS packages to load text files into the database (SQL Server)
- Development of a database project for deployment on various servers
- Creation of queries to monitor the loading runs
- Development of a PowerShell script to automate the deployment of the SSDT projects.
- Oracle, SQL Developer, Microsoft SQL Server 2022 on-premises, SQL Server Management Studio v21, Visual Studio 2022, SSIS, SSDT, PowerShell.
Discover over 15,000 top freelancers
Statistics of experts using PySpark
Aggregated from the professional profiles of matched freelancers.
Experience
16 years (Germany: 13 years)

Position duration
1.8 years (Germany: 2.8 years)

Positions per freelancer
13 (Germany: 10)

Top business areas
Information Technology, Business Intelligence, Product Development

Top industries
Information Technology, Professional Services, Education

Certification focus areas
Information Technology, Business Intelligence, Project Management
Bachelor's degree or higher
100% (Germany: 96%)
Master's degree or higher
86% (Germany: 71%)
Doctorate
33% (Germany: 15%)

Certifications per freelancer
3 (Germany: 4)

Most common languages
German, English, Spanish

Speak two or more languages
100% (Germany: 96%)
Based on our profile pool as of 19 Sep 2026.
Daily rate distribution
The chart shows how the daily rates of freelancers in this technology in Munich are distributed, based on recent contracts on our platform. Each bar covers a rate range — its height shows how many freelancers charge within that range.
Average rates of experts in Munich using PySpark
Rates are based on recent contracts and do not include FRATCH margin.
The average daily rate is the mean of all daily rates from recent contracts of comparable freelancers on our platform.
The median daily rate is the middle value of all daily rates — half of comparable freelancers charge less, half charge more. Unlike the average, it is barely affected by outliers.
Calculated based on our freelancers’ daily rates as of 19 Sep 2026. Actual rates may vary depending on seniority level, experience, skill specialization, project complexity, and engagement length.
PySpark experts industry focus
See which industries our matched freelancers work in most often — every figure is calculated live from the freelancers on FRATCH.
- Information Technology (96%)
- Professional Services (61%)
- Education (52%)
- Banking and Finance (52%)
- Healthcare (52%)
- Manufacturing (52%)
- Automotive (43%)
- Insurance (43%)
Please note that freelancers can work across multiple industries, so percentages overlap.
About the technology
Distributed data processing
PySpark is the Python API for Apache Spark, a framework for processing data across distributed computing clusters. It helps teams transform large datasets, prepare features for machine learning and run repeatable batch or streaming workloads. Python syntax makes Spark accessible while Spark supplies parallel execution and fault tolerance.
Core capabilities
Strong PySpark work covers data ingestion, transformation, aggregation and validation from raw sources to trusted outputs.
- Build DataFrame and Spark SQL transformations
- Create batch and structured streaming pipelines
- Tune joins, partitions, caching and file layouts
- Prepare datasets for analytics and machine learning
- Test, monitor and document production workflows
Ecosystem and tooling
PySpark specialists usually work with Apache Spark, Python, Spark SQL and Delta Lake or other lakehouse formats. They may also use Kafka for event streams, Airflow or managed orchestration for scheduling, and Parquet on cloud object storage. Familiarity with notebooks, Git, Docker and cloud services helps connect experimentation with production delivery.
Where it runs
Companies use PySpark for data lakes, reporting platforms, recommendation pipelines, fraud analysis, log processing and feature engineering. It fits cloud environments such as Databricks, AWS, Microsoft Azure and Google Cloud, as well as self-managed Spark clusters. In Munich, specialists may support automotive, manufacturing, finance, insurance and research workloads, with remote collaboration often combined with occasional on-site sessions.
When to bring in an expert
Freelance expertise is useful when a prototype must become a stable pipeline, a warehouse migration creates performance issues or data volume has outgrown single-machine tools.
- Spark jobs run slowly or fail unpredictably
- Streaming and batch data need one reliable design
- A team needs a lakehouse or Delta Lake migration
- Data quality, lineage or operational ownership is unclear
- Internal Python skills do not cover distributed execution
What good work looks like
A strong professional explains why Spark is needed and when pandas, SQL or another tool would be simpler. They produce readable transformations, controlled schemas, efficient partitioning and tests for correctness. They also understand cluster configuration, serialization, observability and data security, then document decisions so the team can operate the result after handover.
Frequently asked questions
What clients ask us most about PySpark — answered in short.
PySpark is used to process and transform large datasets across Spark clusters with Python. Companies use it for batch ETL, streaming, analytics, feature engineering and data lake workloads.
PySpark distributes work across a cluster, making it suitable when data or processing demands exceed a single machine. pandas is often simpler for smaller in-memory datasets, while SQL can be preferable for transformations already handled efficiently inside a warehouse.
A capable PySpark specialist often brings Spark SQL, Python testing and data modeling skills. Experience with Kafka, Airflow, Delta Lake, Parquet, cloud storage and platforms such as Databricks can be important depending on the delivery environment.
The right level depends on the work rather than a fixed career duration. A straightforward transformation may need solid DataFrame and SQL knowledge, while production streaming, cluster tuning or a lakehouse migration calls for a professional who has operated comparable systems.
Yes. PySpark projects are commonly developed through shared repositories, cloud workspaces, tickets and data-platform documentation. A Munich-based team can collaborate remotely, with on-site workshops useful for architecture, access setup and stakeholder alignment; German or English expectations should be agreed early.
PySpark may add unnecessary overhead for small datasets, simple scripts or low-latency transactional applications. A database query, pandas workflow or a specialized stream processor can be a better fit when distributed execution is not needed.
Review whether PySpark pipelines have clear schemas, reliable tests, sensible partitioning and measurable monitoring. Ask the professional to explain failure handling, data quality checks, cost control and why the chosen design is better than a simpler alternative.
Before starting, a PySpark professional should clarify data sources, volume patterns, batch or streaming requirements, Spark version, cluster access and deployment ownership. They should also confirm privacy constraints, release processes and how success will be assessed.
The average hourly rate of freelancers in Munich, Germany who have used PySpark in their recent projects is 100 €, which corresponds to a daily rate of about 796 € based on an 8-hour working day.
Of the freelancers in Munich, Germany who have used PySpark in their recent projects, 100% hold at least a Bachelor's degree, 86% hold at least a Master's degree, and 33% hold a doctorate.
On average, freelancers in Munich, Germany who have used PySpark in their recent projects have 16 years of professional experience, with a single engagement typically lasting around 1.8 years.
The most common languages among freelancers in Munich, Germany who have used PySpark in their recent projects are German (100%), English (100%), and Spanish (30%).
The most common industries among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (96%), Professional Services (61%), and Education (52%).
The most common business areas among freelancers in Munich, Germany who have used PySpark in their recent projects are Information Technology (96%), Business Intelligence (87%), and Product Development (78%).
Main locations of FRATCH Experts, who have recently used PySpark
Our freelancers and interim experts are at home across the DACH region — available on-site in the major business hubs or fully remote. Choose a location to discover matched specialists, local market insights and up-to-date availability.
Request a free demo
Get in touch with the FRATCH team and we will get back to you within 4 hours.
Would you rather directly get in touch?
We always have the time for a call or email!

Berlin
Frankfurt