Blog Post

How SAP achieved world-class uptime through modern observability

Updated
Published
July 22, 2025
#
 mins read

in this blog post

SAP Customer Experience (CX) runs one of the most demanding e-commerce platforms in the enterprise world. Thousands of global customers, including Alphabet, Shell, Cigna, and Mercedes-Benz Group, depend on SAP Commerce for always-on digital storefronts. Over the past several years, the SAP CX observability team, led by Martin Norato Auer, VP of Observability, has transformed its operations from fragmented, reactive monitoring into a scalable, automated system that sets a new standard for enterprise reliability.

SAP Commerce (formerly SAP Hybris) is an enterprise-grade e-commerce platform designed to unify and manage digital commerce across B2B, B2C, and B2B2C business models. It empowers businesses to deliver consistent, personalized, and seamless customer experiences across web, mobile, social, and physical channels.

This case study breaks down how the team cut SLA violations from 16% to 0.1%, reduced incident tickets by two-thirds, and compressed customer notification times from 180 minutes to 2 minutes.

Raising the Bar: SLA and Uptime Breakthroughs

  • SLA Violation Reduction: SAP CX cut SLA violations from 16% to just 0.1%, bringing them close to their long-term goal of zero downtime for customers.
  • Dramatic Ticket Reduction: The team slashed incident tickets from approximately 1,500 per year to 500, a two-thirds drop reflecting improved stability and proactive issue prevention.
  • Lightning-Fast Customer Notifications: Average time to inform customers about incidents dropped from 180 minutes to just 2 minutes after detection. Each notification is qualified with specific context, such as "there's an issue on your storefront, it's most likely your CDN, we're working on it." That level of detail lets customers act immediately rather than wait for diagnosis.

__wf_reserved_inherit

For SAP Commerce customers, uptime and performance are critical because even brief disruptions or slowdowns can lead to lost revenue, diminished customer trust, and missed opportunities in highly competitive, always-on global markets.

Strategies for Faster Problem Identification

  • Unified Observability Stack: Before consolidation, Black Friday operations looked like controlled chaos. The team ran three war rooms across the globe, each packed with screens showing seven or eight different tools. Dedicated operators watched individual dashboards, shouting across the room when something looked off in one tool while someone else tried to cross-reference a different one. It didn't scale.The team consolidated to two core platforms: Dynatrace for application performance and Catchpoint, a LogicMonitor company, for visibility across the Internet path. The reasoning was straightforward: "You can only run a business on scale if you consolidate. If I need three experts for each tool because I need to do 24/7 support, it doesn't scale."The move to Catchpoint also changed how the team thought about external monitoring. Initially, the value wasn't obvious: "Pinging websites from outside. What's the real benefit? What can I see from the Internet that I can't see internally?" But IPM extended visibility beyond the firewall into CDNs, DNS, and ISPs, areas where Dynatrace had no reach. The team could now tell customers, "We've identified an issue with your CDN. We can't fix it on our side, but you should be aware and take action." That kind of proactive, specific communication was impossible with internal monitoring alone.Internet Sonar started as an exploratory project during its early release. Today, it's part of daily operations and core business. Internet Stack Map added another layer of context, giving the team real-time visibility into the Internet infrastructure their customers depend on.
  • Automated Alert Relevance Filtering: Rather than simply multiplying alert volumes, SAP CX built an internal platform called Marvin that consolidates alerts from Dynatrace and LM IPM into a single visualization. Getting the algorithm right took multiple iterations. The first approach was volume-based (more alerts equals a bigger problem), but it didn't work. The second attempt used weighted scoring, which also fell short. The team eventually developed a custom algorithm that combines alert data with business-impact logic to surface what actually matters.The goal was clarity for non-experts: "Someone who has no deep knowledge about the product can see right away whether this is something where they need help, or something they can just watch." That shift from raw alert volume to business-contextualized relevance is what made the tool operationally useful.
  • Workflow Automation: The team clocked every step of the notification chain: alert detection, relevance decision, expert validation, message drafting, customer lookup, and message send. They found bottlenecks at each stage and introduced automation and APIs to eliminate manual handoffs wherever possible. Trust was built incrementally. Early on, they kept verification steps in place because they didn't trust the automation. As confidence grew through consistent results, they automated verification too, collapsing the entire chain into a process that delivers qualified customer notifications within two minutes.
  • Proactive and Predictive Monitoring: The shift from reactive to predictive monitoring enabled SAP CX to spot potential issues and intervene before they impacted customers, using dashboards that flag availability risk in advance.

Best Practices That Drove Transformation

Large IT operations teams face significant barriers when implementing transformative changes. These challenges span organizational, cultural, and technical dimensions, and they compound as teams and systems grow. Here's how SAP CX addressed them.

  • Continuous Data Analysis: Every incident, whether missed or delayed, triggered a root cause analysis and process iteration. The team improved detection logic and reduced blind spots systematically, treating each failure as input for the next version of their process.
  • Process Transparency: Clear mapping of all steps needed for detection, triage, customer communication, and escalation allowed for targeted automation and efficiency gains.
  • Global Team Collaboration: A dispersed, cross-continental team structure enabled 24/7 coverage and quick mobilization for major events like Black Friday.
  • Leadership Engagement and Mobile Insights: The team built an internal mobile app that gives leadership real-time, high-level incident summaries: situation status, who's working on it, a contact person, and an AI-generated summary. The app started as a student project through SAP's "New Horizons" innovation program (the 10% allocation for exploratory work). The first cohort's version didn't make it, but a second cohort rebuilt it successfully. The result: "Within thirty seconds, I can give a qualified answer to every customer about a certain situation." That speed comes from AI summarization of bridge call transcripts and communication channels, condensed for leadership consumption.
  • AI for Summarization and Analysis: The team uses generative AI heavily for data analysis at scale (correlating tickets with customer sale events across thousands of customers), summarizing bridge calls, and condensing information for different audiences. But they're deliberate about where AI ends. For incident management, the team avoids agentic AI in favor of consistent, repeatable automation: "When you're in operations, you want a consistent way of doing it because every deviation from the standard creates edge cases." Simple, governed automation beats agents for reliability in production operations.
  • Structured Innovation Time: SAP CX operates on a 60/30/10 resource split: 60% on core business (mission-critical, non-negotiable), 30% on emerging business (known upcoming work), and 10% on "New Horizons" (exploratory innovation). The philosophy is simple: "If you lose innovation, you will lose business." This structure produced both Internet Sonar adoption and the leadership mobile app.

Addressing these barriers requires strong, visible leadership, clear communication, and aligning strategy to daily work. Only with an integrated approach that accounts for people, process, and platform can large organizations drive lasting operational change.

Lessons for the Enterprise

SAP CX's transformation offers a clear playbook for enterprise teams looking to move from reactive monitoring to proactive, automated operations. Several patterns stand out.

  • Tool consolidation is essential for reducing complexity and noise, enabling true end-to-end visibility. Running seven or eight tools with dedicated experts for each one doesn't scale. Two integrated platforms covering application performance and Internet path visibility gave SAP CX the signal clarity they needed.
  • Automation alongside monitoring is key to shrinking response times and scaling operations. But automation earns trust through incremental adoption: keep verification steps until the system proves itself, then automate those too.
  • Customer-centric communication transforms incident management from technical firefighting into a trust-building opportunity. Qualified, contextual notifications within minutes show customers you understand the problem and are acting on it.
  • Combining APM insights with Internet performance visibility enabled SAP Commerce to achieve comprehensive, end-to-end observability. Integrating deep application diagnostics (system logs, infrastructure metrics, code traces) with Internet Performance Monitoring gave the team the ability to identify, prevent, and resolve incidents across the full digital path.
  • Continuous improvement and data-driven iteration are the foundation of durable operational excellence. Every incident is an input, every process has a metric, and every bottleneck is a candidate for automation.

By adopting these practices, SAP CX drove measurable, world-class improvements in SLA adherence, availability, and incident response, and built a system that earns the trust of organizations demanding the highest standards of reliability, performance, and customer transparency.

SAP's journey illustrates what becomes possible when application diagnostics, Internet visibility, and intelligent automation work as one system. That's the unified, user-to-code visibility LogicMonitor delivers through a single connected platform, bringing together LM Envision, LM Internet Performance Monitoring, and Edwin AI so teams can see the full picture and act on it faster.

Summary

SAP Customer Experience (CX) runs one of the most demanding e-commerce platforms in the enterprise world. Thousands of global customers, including Alphabet, Shell, Cigna, and Mercedes-Benz Group, depend on SAP Commerce for always-on digital storefronts. Over the past several years, the SAP CX observability team, led by Martin Norato Auer, VP of Observability, has transformed its operations from fragmented, reactive monitoring into a scalable, automated system that sets a new standard for enterprise reliability.

SAP Commerce (formerly SAP Hybris) is an enterprise-grade e-commerce platform designed to unify and manage digital commerce across B2B, B2C, and B2B2C business models. It empowers businesses to deliver consistent, personalized, and seamless customer experiences across web, mobile, social, and physical channels.

This case study breaks down how the team cut SLA violations from 16% to 0.1%, reduced incident tickets by two-thirds, and compressed customer notification times from 180 minutes to 2 minutes.

Raising the Bar: SLA and Uptime Breakthroughs

  • SLA Violation Reduction: SAP CX cut SLA violations from 16% to just 0.1%, bringing them close to their long-term goal of zero downtime for customers.
  • Dramatic Ticket Reduction: The team slashed incident tickets from approximately 1,500 per year to 500, a two-thirds drop reflecting improved stability and proactive issue prevention.
  • Lightning-Fast Customer Notifications: Average time to inform customers about incidents dropped from 180 minutes to just 2 minutes after detection. Each notification is qualified with specific context, such as "there's an issue on your storefront, it's most likely your CDN, we're working on it." That level of detail lets customers act immediately rather than wait for diagnosis.

__wf_reserved_inherit

For SAP Commerce customers, uptime and performance are critical because even brief disruptions or slowdowns can lead to lost revenue, diminished customer trust, and missed opportunities in highly competitive, always-on global markets.

Strategies for Faster Problem Identification

  • Unified Observability Stack: Before consolidation, Black Friday operations looked like controlled chaos. The team ran three war rooms across the globe, each packed with screens showing seven or eight different tools. Dedicated operators watched individual dashboards, shouting across the room when something looked off in one tool while someone else tried to cross-reference a different one. It didn't scale.The team consolidated to two core platforms: Dynatrace for application performance and Catchpoint, a LogicMonitor company, for visibility across the Internet path. The reasoning was straightforward: "You can only run a business on scale if you consolidate. If I need three experts for each tool because I need to do 24/7 support, it doesn't scale."The move to Catchpoint also changed how the team thought about external monitoring. Initially, the value wasn't obvious: "Pinging websites from outside. What's the real benefit? What can I see from the Internet that I can't see internally?" But IPM extended visibility beyond the firewall into CDNs, DNS, and ISPs, areas where Dynatrace had no reach. The team could now tell customers, "We've identified an issue with your CDN. We can't fix it on our side, but you should be aware and take action." That kind of proactive, specific communication was impossible with internal monitoring alone.Internet Sonar started as an exploratory project during its early release. Today, it's part of daily operations and core business. Internet Stack Map added another layer of context, giving the team real-time visibility into the Internet infrastructure their customers depend on.
  • Automated Alert Relevance Filtering: Rather than simply multiplying alert volumes, SAP CX built an internal platform called Marvin that consolidates alerts from Dynatrace and LM IPM into a single visualization. Getting the algorithm right took multiple iterations. The first approach was volume-based (more alerts equals a bigger problem), but it didn't work. The second attempt used weighted scoring, which also fell short. The team eventually developed a custom algorithm that combines alert data with business-impact logic to surface what actually matters.The goal was clarity for non-experts: "Someone who has no deep knowledge about the product can see right away whether this is something where they need help, or something they can just watch." That shift from raw alert volume to business-contextualized relevance is what made the tool operationally useful.
  • Workflow Automation: The team clocked every step of the notification chain: alert detection, relevance decision, expert validation, message drafting, customer lookup, and message send. They found bottlenecks at each stage and introduced automation and APIs to eliminate manual handoffs wherever possible. Trust was built incrementally. Early on, they kept verification steps in place because they didn't trust the automation. As confidence grew through consistent results, they automated verification too, collapsing the entire chain into a process that delivers qualified customer notifications within two minutes.
  • Proactive and Predictive Monitoring: The shift from reactive to predictive monitoring enabled SAP CX to spot potential issues and intervene before they impacted customers, using dashboards that flag availability risk in advance.

Best Practices That Drove Transformation

Large IT operations teams face significant barriers when implementing transformative changes. These challenges span organizational, cultural, and technical dimensions, and they compound as teams and systems grow. Here's how SAP CX addressed them.

  • Continuous Data Analysis: Every incident, whether missed or delayed, triggered a root cause analysis and process iteration. The team improved detection logic and reduced blind spots systematically, treating each failure as input for the next version of their process.
  • Process Transparency: Clear mapping of all steps needed for detection, triage, customer communication, and escalation allowed for targeted automation and efficiency gains.
  • Global Team Collaboration: A dispersed, cross-continental team structure enabled 24/7 coverage and quick mobilization for major events like Black Friday.
  • Leadership Engagement and Mobile Insights: The team built an internal mobile app that gives leadership real-time, high-level incident summaries: situation status, who's working on it, a contact person, and an AI-generated summary. The app started as a student project through SAP's "New Horizons" innovation program (the 10% allocation for exploratory work). The first cohort's version didn't make it, but a second cohort rebuilt it successfully. The result: "Within thirty seconds, I can give a qualified answer to every customer about a certain situation." That speed comes from AI summarization of bridge call transcripts and communication channels, condensed for leadership consumption.
  • AI for Summarization and Analysis: The team uses generative AI heavily for data analysis at scale (correlating tickets with customer sale events across thousands of customers), summarizing bridge calls, and condensing information for different audiences. But they're deliberate about where AI ends. For incident management, the team avoids agentic AI in favor of consistent, repeatable automation: "When you're in operations, you want a consistent way of doing it because every deviation from the standard creates edge cases." Simple, governed automation beats agents for reliability in production operations.
  • Structured Innovation Time: SAP CX operates on a 60/30/10 resource split: 60% on core business (mission-critical, non-negotiable), 30% on emerging business (known upcoming work), and 10% on "New Horizons" (exploratory innovation). The philosophy is simple: "If you lose innovation, you will lose business." This structure produced both Internet Sonar adoption and the leadership mobile app.

Addressing these barriers requires strong, visible leadership, clear communication, and aligning strategy to daily work. Only with an integrated approach that accounts for people, process, and platform can large organizations drive lasting operational change.

Lessons for the Enterprise

SAP CX's transformation offers a clear playbook for enterprise teams looking to move from reactive monitoring to proactive, automated operations. Several patterns stand out.

  • Tool consolidation is essential for reducing complexity and noise, enabling true end-to-end visibility. Running seven or eight tools with dedicated experts for each one doesn't scale. Two integrated platforms covering application performance and Internet path visibility gave SAP CX the signal clarity they needed.
  • Automation alongside monitoring is key to shrinking response times and scaling operations. But automation earns trust through incremental adoption: keep verification steps until the system proves itself, then automate those too.
  • Customer-centric communication transforms incident management from technical firefighting into a trust-building opportunity. Qualified, contextual notifications within minutes show customers you understand the problem and are acting on it.
  • Combining APM insights with Internet performance visibility enabled SAP Commerce to achieve comprehensive, end-to-end observability. Integrating deep application diagnostics (system logs, infrastructure metrics, code traces) with Internet Performance Monitoring gave the team the ability to identify, prevent, and resolve incidents across the full digital path.
  • Continuous improvement and data-driven iteration are the foundation of durable operational excellence. Every incident is an input, every process has a metric, and every bottleneck is a candidate for automation.

By adopting these practices, SAP CX drove measurable, world-class improvements in SLA adherence, availability, and incident response, and built a system that earns the trust of organizations demanding the highest standards of reliability, performance, and customer transparency.

SAP's journey illustrates what becomes possible when application diagnostics, Internet visibility, and intelligent automation work as one system. That's the unified, user-to-code visibility LogicMonitor delivers through a single connected platform, bringing together LM Envision, LM Internet Performance Monitoring, and Edwin AI so teams can see the full picture and act on it faster.

This is some text inside of a div block.

You might also like

Blog post

SRE Report: Why fast is what users trust

Blog post

Why Synthetic Tracing Delivers Better Data, Not Just More Data

Blog post

Creating the IPM Category: Catchpoint’s Journey to Leadership and the LogicMonitor Era