Skip to main content
🇩🇪GDPR-compliant
Find the perfect

Observability Experts in Berlin

in minutes from over 15,000 CVs with the power of AI

Hire experts who design logging, metrics, and tracing setups, connect OpenTelemetry with Grafana, Prometheus, or Datadog, and turn noisy systems into clear signals. Fast, precise matching with vetted, available freelancers.

Meet FRATCH Experts in Berlin, who have recently used Observability

Verified expert

Jorge Nuricumbo

View profile

Senior AI Engineer | Backend Developer C#/.NET | RAG, LLM Integration, Semantic Kernel | Azure, GCP, AWS

Berlin
Jorge Nuricumbo

Last position:

Senior Developer at SafeXSmart KI Solutions UG

AI Platform Backend – Senior Developer

Brought in to design and build a backend for an AI platform from scratch, including multi-provider LLM orchestration and real-time infrastructure for AI influencer personas at scale.

Tasks and responsibilities

  • Architected and implemented a multi-LLM orchestration layer with Semantic Kernel to integrate GPT-4 and other providers for core platform logic and AI influencer personas, reducing model-switching overhead by abstracting provider APIs behind a single interface.
  • Designed and developed a backend from scratch in C# / .NET 10, including domain modeling with DDD, a versioned RESTful API layer, and cloud infrastructure setup on Azure.
  • Built a real-time chat infrastructure with Server-Sent Events (SSE), message persistence, and delivery guarantees for live operation of AI influencer personas at scale.
  • Developed a media management service with integration of cloud object storage for upload and retrieval of influencer-generated content.
  • Created an integration and unit test suite with data seeding for reliable regression testing across all core platform flows, significantly reducing the production error rate.

Tools and technologies: C#, .NET, ASP.NET Core, Python, TypeScript, MySQL, Semantic Kernel, EF Core, Minimal APIs, LLM Orchestration, Prompt Engineering, Agentic AI, Generative AI, AI-Assisted Engineering, Claude Code, GitHub Copilot, Google Gemini, OpenAI API, Ollama, Redis, Azure, Azure Container Apps, Azure Database for MySQL, Docker, GitHub Actions, Clean Architecture, Vertical Slice Architecture, CQRS, Domain-Driven Design, REST API, xUnit, Integration Testing, Unit Testing, Jira, Confluence, Scrum

Verified expert

Julius Herrera Glomm

View profile

Freelancer

Berlin
Julius Herrera Glomm

Last position:

Freelancer at Freelancer — Pharma Industry

  • Led migration to GCP using Terraform, GKE, and GitOps, improving deployment consistency and scalability
  • Implemented Datadog observability stack via Terraform and datadog-operator
  • Established automated end-to-end tests and on-call processes, improving incident response and service reliability
  • Migrated from NGINX Ingress Controller to Kubernetes Gateway API (NGINX Gateway Fabric)
  • Migrated stateful services (PostgreSQL and Redis) to GCP, improving scalability and operational reliability
Verified expert

Santhosh Kannan

View profile

Freelance Software Engineer

Berlin
Santhosh Kannan

Last position:

Freelance Software Engineer at Zalando SE

  • Support Authorization as a Service initiative for enterprise-scale authorization platform
  • Incorporate comprehensive observability solutions into authorization infrastructure
  • Provision and manage AWS infrastructure for authorization services
  • Mentor development team on AWS and Kubernetes best practices
  • Tech Stack: Java/Kotlin, Golang, Python, OPA, Spring Boot, AWS, Kubernetes, Terraform, ELK Stack, Prometheus, Grafana
Verified expert

Sejal Vaidya

View profile

Data & ML Engineering

Berlin
Sejal Vaidya

Last position:

Data & ML Engineering at Consulting

  • Fractional leadership; consulting growth-stage startups and scale-ups on data strategy, ML products, and platform foundations
  • Building decisioning systems for growth, personalization, & product experimentation, across e-Commerce, Digital Health, Energy, and Logistics
  • Exploring Agentic AI & LLM-based tooling for production readiness patterns
Verified expert

Imran Ali

View profile

Software Engineer II

Berlin
Imran Ali

Last position:

Software Engineer II at LivePerson Germany GmbH

  • Led development of 15+ microservices (Java 17, Spring Boot) driving customer interactions; migrated from on-prem to GCP Kubernetes, improving scalability and reducing infra cost by 20%.
  • Optimized user services with CouchDB caching and API refactoring, cutting response times by 35% and enhancing customer experience.
  • Implemented canary deployments, FluxCD GitOps, and CI/CD optimizations in GitLab, reducing release lead time by 25% and enabling zero-downtime rollouts.
  • Set up Grafana health checks and Anodot alerts for latency, error, and throughput monitoring, reducing MTTR by 40%.
  • Built secure APIs using OAuth2, DPoP, and Gatekeeper, integrated REST and GraphQL, and achieved 90%+ test coverage with unit and E2E tests.
  • Mentored junior developers, promoted Agile best practices, and collaborated cross-functionally to deliver high-impact, reliable customer-facing features.
Verified expert

Viktor Shcherban

View profile

AI Engineer & Full-Stack Developer

Berlin
Viktor Shcherban

Last position:

AI Engineer (Freelance) at Empion

Enterprise AI content categorization and AI-powered web research.

  • Built multi-LLM evaluation framework with annotated data
  • Iterated LLM error rates based on annotated datasets
  • Implemented AI-powered web research pipeline Stack: LLM, evals, OpenRouter, Python, Node.js, TypeScript, React
Verified expert

Wolfram Knan

View profile

Certified AI & Machine Learning Engineer · Senior Consultant

Berlin
Wolfram Knan

Last position:

AI / Machine Learning Engineer (Projects & Applied AI) at UNIVERSITÉ PARIS 1 PANTHEON-SORBONNE & LIORA

  • Designed and implemented a hybrid recommendation system (content-based + collaborative filtering)
  • Built end-to-end ML pipelines including data processing, feature engineering, model training, and evaluation
  • Developed RAG-based LLM systems using LangChain and vector databases for semantic search and knowledge retrieval
  • Established MLOps workflows with MLflow for experiment tracking, versioning, and deployment readiness
  • Implemented deep learning models (computer vision & classification) using PyTorch and TensorFlow
Verified expert

Abhiroop Basu

View profile

Software Engineer III

Berlin
Abhiroop Basu

Last position:

Software Engineer III at Foundry Digital

  • Developed and deployed microservices in Kotlin and Spring Boot, integrated AWS Secrets Manager to secure credentials and decreased network calls using Spring cache.
  • Refactored Kafka consumer using Spring Kafka with semaphore-based backpressure to cap records and keep heap memory stable under spikes; switched to batch upserts to cut down on database invocations; added Testcontainers integration tests for Kafka and database to pave the way for future changes.
  • Automated the financial reconciliation workflow in Spring Boot (Kotlin) using Spring Scheduler, transactional boundaries, JPA/Hibernate on MySQL, and Flyway migrations, saving the accounts team 16+ hours per week.
  • Designed and dockerized payments end-to-end test framework in Robot (Python) with reusable keyword libraries and profiles; integrated with GitLab CI (JaCoCo XML and HTML reports) to accelerate releases and lift code coverage to 80%.
  • Implemented end-to-end observability on Datadog by instrumenting services with Datadog APM, correlating metrics and logs, provisioning dashboards, and creating monitors with burn-rate alerts and anomalies to harden reliability and give stakeholders clear visibility.
Verified expert

Nune Isabekyan

View profile

Engineering Leader · Fractional CTO of OpsWorker

Berlin
Nune Isabekyan

Last position:

Fractional CTO at OpsWorker

OpsWorker turns Kubernetes alerts into root-cause analyses, on top of the monitoring a team already runs. I lead the technical side: the agent architecture, the AWS infrastructure it runs on (fully inside EU regions), and the engineering decisions behind it, read-only in the cluster by default, human in the loop for judgment. The stack underneath: Amazon Bedrock and Bedrock AgentCore, agents built with the Strands Agents SDK, the Claude and OpenAI APIs, and the Kubernetes API.

Verified expert

Can Savastürk

View profile

Software Development for People

Berlin
Can Savastürk

Last position:

Platform Engineer at ClimateChoice

In a lean, execution-focused environment, I took ownership beyond a narrow engineering lane, shaping and implementing systems across backend, data, and infrastructure. Partnered directly with the three founders in a fast-moving, high-stakes environment, turning strategic priorities into concrete technical decisions and production outcomes.

  • Owned core platform development across backend (Django/Rest Framework/Postgres), ETL (Python/Dagster), infrastructure (Terraform/Kubernetes/AWS), and frontend (typescript/react) for a climate-tech SaaS product, driving continuous cross-stack development across five repositories from October 2021 to this day.
  • Architected and owned a standalone internal Python scoring framework for CRC assessments, using YAML-driven rules and metaprogramming to enable non-technical users to define complex evaluation logic without hardcoded implementations.
  • Built and stabilized ETL and scraping pipelines using Dagster and Scrapfly, improving document ingestion, tagging, retry behavior, deployment flow, and operational resilience.
  • Contributed to platform modernization and reliability through Django/Python upgrades, Postgres/RDS and EKS changes, CDN/TLS updates, test and performance improvements, and observability hardening.
  • Drove backend engineering for product features, translating requirements into technical specifications, API contracts, data structures, and scalable implementation plans.
Verified expert

Sebastian Striebig

View profile

Group Product Manager – Digital Platform Discovery

Berlin
Sebastian Striebig

Last position:

Group Product Manager – Digital Platform Discovery at SPREAD.AI

  • Developed and implemented organization-wide discovery framework based on Ulwick’s Outcome-Driven Innovation; enabled 7 Product Owners to systematically identify and quantify unrealized value through shared outcome language and opportunity scoring methodology
  • Transformed Product Owner role from backlog clerks to strategic experimenters; established dedicated time budget for autonomous hypothesis testing and discovery activities
  • Rebuilt customer journey maps to start at actual user need (tool selection phase) instead of platform entry point; eliminated manual data aggregation work previously done by project teams
  • Implemented OKR framework across 4 product teams; defined quarterly objectives with measurable key results (e.g., 40% reduction in manual integration effort, self-service adoption increase)
  • Unified 3 separate platform roadmaps through cross-team dependency mapping and shared service agreements
  • Supported enterprise sales cycle with ROI modeling and technical due diligence for automotive and defense customers
Verified expert

Chiemela Ogu

View profile

Product Management Consultant

Berlin
Chiemela Ogu

Last position:

AI Enthusiast – Independent Projects at Chiemela Ogu Consulting

  • Too Good To Throw (AI powered social impact webapp focused on reducing food waste in Nigeria):

  • Integrated Paystack Split Payments to automatically route payments between the platform and partner vendors.

  • Configured automated subaccount creation workflows so new businesses get a settlement account instantly.

  • Setup a scalable cloud backend using Supabase.

  • Implemented role-based access control (RBAC) for Users, Partners, and Admin.

  • FaithFlow (AI powered webapp supporting Christian teens on their spiritual journey):

  • Designed and implemented an AI-driven scripture search engine that interprets natural language questions and maps them to relevant Bible texts, commentary, devotionals, and cross-references.

  • Designed a spiritual growth dashboard enabling users to track reading progress, prayer streaks, and devotional completion milestones.

Verified expert

Ashwin Parthasarathy

View profile

Data Scientist

Berlin
Ashwin Parthasarathy

Last position:

Data Scientist at Mercor Intelligence

  • Elevated LLM output reliability by engineering domain-specific prompts and evaluation logic, improving reasoning consistency across production language model workflows.
  • Designed advanced coding benchmarks and validated solutions to strengthen training and evaluation datasets, improving model performance on technical problem-solving tasks.
  • Designed and implemented automated evaluation frameworks for technical reasoning tasks; optimized LLM output reliability by 15% through rigorous prompt engineering and rubric-based benchmarking.
Verified expert

Gautam Dhameja

View profile

Enterprise AI and Engineering Leader

Berlin
Gautam Dhameja

Last position:

Founder at Proferent

  • Shipped Memorable, a production iOS app using on-device CLIP-based semantic photo search. Owned the full stack: Core ML conversion, local inference pipeline, App Store release, and post-launch iteration.
  • Built a practical AI deployment framework that covers workflow redesign, use-case prioritization, system integration, eval planning, and human-in-the-loop controls.
  • Conducting AI use-case discovery and advisory conversations with professionals in legal, tax, and real estate sectors.

Discover over 15,000 top freelancers

Statistics of experts using Observability

Aggregated from the professional profiles of matched freelancers.

Experience

15 years

Position duration

2 years

Positions per freelancer

7

Top business areas

Information Technology, Product Development, Project Management

Top industries

Information Technology, Banking and Finance, Retail

Certification focus areas

Information Technology, Product Development, Project Management

Bachelor's degree or higher

96%

Master's degree or higher

61%

Certifications per freelancer

2

Most common languages

English, German, Spanish

Speak two or more languages

97%

Based on our profile pool as of 30 Aug 2026.

About the technology

What it covers

Observability helps teams see how systems behave in production, not just whether they are up or down. It combines logs, metrics, traces, and alerts so experts can explain latency, failures, and user-impacting issues across services, APIs, and infrastructure.

Common stack

  • OpenTelemetry for standard instrumentation
  • Prometheus for metrics and alerting
  • Grafana for dashboards and exploration
  • Elasticsearch, Loki, or the ELK stack for log search
  • Jaeger or Tempo for distributed tracing

Where it helps

Companies bring in observability specialists when incidents are hard to trace, dashboards are inconsistent, or teams need one shared view of service health. It is common in cloud platforms, SaaS products, fintech, e-commerce, and data-heavy systems that change often and need clear operational evidence.

What strong experts do

Good professionals instrument code and infrastructure with care, choose useful signals, and avoid noisy dashboards that nobody trusts. They know how to connect traces, logs, and metrics into one investigation flow, and they can tune retention, alert thresholds, and naming conventions so teams can act fast.

Delivery work

A freelance observability expert may set up service dashboards, define alerts for critical paths, standardize OpenTelemetry instrumentation, or clean up an overloaded logging strategy. In Berlin, this often supports product teams working with distributed systems, remote platform groups, or hybrid teams that need clear handover notes and English-first collaboration.

How to choose

Look for experience with real production incidents, not just tool names. Strong experts can explain why a signal matters, how they reduce alert fatigue, and how they would observe a new service from the first release onward. They should also understand developer workflows, Kubernetes, cloud services, and incident response.

Published on:
FRATCH GPT

FRATCH GPT delivers freelancer proposals with clear reasoning and transparent pricing in minutes, helping your hiring department quickly and compliantly find the best talent.

Give it a try:

Try FRATCH GPT

Frequently asked questions

Need clarity? These are the questions we hear most often about Observability.

Observability helps teams understand why a system behaves the way it does. It connects logs, metrics, and traces so specialists can follow a request across services, spot bottlenecks, and see the impact of failures on users. That makes it much more useful than simple uptime checks.

Observability is broader than monitoring. Monitoring tells you when something is wrong; observability helps you investigate what happened and why. In practice, teams usually need both, but observability is the layer that supports deeper debugging and system understanding.

A Observability project often includes OpenTelemetry, Prometheus, Grafana, and one log or trace backend such as Loki, Elasticsearch, Jaeger, or Tempo. The exact stack depends on the environment and on whether the team needs dashboards, alerting, tracing, or centralized log search. Strong specialists can work across more than one tool if the setup is mixed.

A good Observability specialist usually understands Kubernetes, cloud infrastructure, application architecture, and incident response. They should also be comfortable with instrumentation in the main language stack and with CI/CD, because good visibility starts in the delivery pipeline as well as in production.

Observability work can start with a small assessment, but deeper setups need access to service maps, deployment patterns, and recent incident history. If the system is distributed or heavily regulated, the expert should review current alerting, retention, and access controls before changing anything. That context prevents wasted effort and noisy results.

Yes, Observability specialists often work remotely, especially for tooling, instrumentation, and dashboard cleanup. For Berlin teams, on-site time is mainly useful for incident workshops, architecture sessions, or quick alignment with platform and product stakeholders. English is usually enough, though German can help in some internal settings.

A strong Observability freelancer can explain trade-offs clearly and show how they reduce alert fatigue, improve traceability, and make incidents easier to debug. Ask for examples of dashboards, alert rules, and instrumentation decisions, not just screenshots. Good experts focus on signal quality and team usability, not tool decoration.

Observability matters most when systems are distributed, change quickly, or support critical user flows. It is especially valuable for microservices, cloud platforms, data pipelines, and APIs where failures can be partial and hard to reproduce. In those cases, clear visibility saves time during every release and incident.

Of the freelancers in Berlin, Germany who have used Observability in their recent projects, 96% hold at least a Bachelor's degree and 61% hold at least a Master's degree.

On average, freelancers in Berlin, Germany who have used Observability in their recent projects have 15 years of professional experience, with a single engagement typically lasting around 2 years.

The most common languages among freelancers in Berlin, Germany who have used Observability in their recent projects are English (100%), German (94%), and Spanish (13%).

The most common industries among freelancers in Berlin, Germany who have used Observability in their recent projects are Information Technology (100%), Banking and Finance (61%), and Retail (32%).

The most common business areas among freelancers in Berlin, Germany who have used Observability in their recent projects are Information Technology (100%), Product Development (97%), and Project Management (55%).

Main locations of FRATCH Experts, who have recently used Observability

Our freelancers and interim experts are at home across the DACH region — available on-site in the major business hubs or fully remote. Choose a location to discover matched specialists, local market insights and up-to-date availability.

Berlin Hamburg Munich Cologne Frankfurt Stuttgart Dusseldorf Leipzig Dortmund Essen Bremen Dresden Hanover Nuremberg

Request a free demo

Get in touch with the FRATCH team and we will get back to you within 4 hours.

Contact form

Would you rather directly get in touch?
We always have the time for a call or email!

FRATCH CEO avatar

Philipp Thomaschewski

FRATCH CEO

LinkedInFRATCH