AI workloads are reshaping SRE and platform engineering. New Dynatrace research shows how AI is changing SRE and platform engineering in two ways: teams are using AI to streamline and scale IT and DevOps while taking on greater responsibility for the reliability, performance, and cost of AI workloads.
Site reliability engineers (SREs) and platform engineers pursue different missions with shared inputs, and AI is transforming how both operate. AI applications are becoming mission-critical, and observability must now span traditional reliability signals in addition to AI health indicators. But AI workloads fail differently than conventional software, and as projects move from pilot to production, SREs and platform engineers are under pressure to deliver:
- AI workloads that behave as expected in production
- AI-driven automation that facilitates reliable AI workloads
Another challenge is how teams bridge the gap between where AI is built and where it runs. The tools teams use to evaluate and iterate on AI during development are typically separate from the observability platforms that monitor it in production. As more AI workloads reach production, teams need greater continuity between how they evaluate AI during development and how they observe it in production.
The State of SRE and Platform Engineering 2026 research explores how adopting AI is affecting SRE and platform engineering practices, where teams are struggling with automation and scale, and how AI-powered observability provides crucial shared context, trust, and coordination across domains.

Key insights for leaders
- AI workloads are changing development operations. SRE and platform engineering practices are widespread and mature, but the adoption of AI workloads is shifting the challenge to scale, integration, and decision complexity.
- Monitoring AI systems is now SREs’ top use case (58%), while 67% rank AI-powered features as their most important observability platform capability, reflecting the growing role of AI in both what SRE teams monitor and how they operate.
- Platform engineers’ top priorities are AI-powered developer support (55%) and delivering self-service observability dashboards (74%) through internal developer platforms (IDPs), pointing to the importance of integrating systems, enforcing standards, and automating workflows.
- Teams embed security and compliance into platforms-as-code (nearly 70%) and CI/CD workflows (60%+), indicating the emphasis teams place on centralized governance facilitated by unified observability for consistent, scalable control.
- Observability is implemented broadly but not deeply. Many teams don’t fully embed it across deployments (40%), leaving environments unoptimized as AI workloads come online (source: The State of SRE and Platform Engineering 2026 research).
AI is changing the landscape for SRE and platform engineering
AI is shifting SRE and platform engineering toward monitoring AI systems, enabling AI-driven automation, and requiring observability for model behavior, cost, and performance.
Monitoring AI systems is SRE’s #1 use case (58%) ahead of automation, service-level objectives, and quality gates. Likewise, providing developers with access to AI-powered tools (chatbots and copilots) is platform engineers’ top AI focus (55%).
In parallel, 67% of SREs and 63% of platform engineers rank AI-powered features as their most important observability platform capability. Yet evaluation and production monitoring often remain separate. Teams may evaluate AI during experimentation, testing, and model iteration in one set of tools, then monitor those workloads in another once they reach production. That separation makes it harder to connect how an AI system was evaluated with how it actually behaves in production.
What are the top SRE trends?
SLOs and quality gates drive SRE automation, but fragmented telemetry is slowing progress.
| Trend | Key data | Implication |
| AI monitoring | 58% – top use case | Expand observability to AI |
| SLO adoption | 89% usage | Automation depends on signal clarity |
| Fragmented telemetry | 50% experience challenges | Tool consolidation needed |
- SRE adoption and automation look strong: 92% report executive leadership support for SRE to some degree, and 64% say most or all IT operations processes are fully automated.
- SLOs anchor that automation: 89% use SLOs in at least some teams or systems, and quality gates are also widely automated to test thresholds, performance metrics, and SLO compliance.
- But teams hit a ceiling when telemetry fragments: ~50% say too many data sources hinder defining and creating SLOs. Top challenges include too many metrics, complex monitoring tools, and uncertainty about what to monitor. AI workloads exacerbate this trend.
The report answers these challenges directly: to keep pace with cloud- and AI-native workloads, teams should use observability as the central source of operational data and reliability insights.
What are the top platform engineering trends?
Platform engineering scales self-service, but teams struggle most with integrating tools and systems.
| Trend | Key data | Implication |
| Self-service observability | 74% adoption | Scale self-service access |
| Reusable assets | 63% preconfigured | Expand CI/CD pipeline preconfiguration |
| Deployment consistency | 41% ad-hoc | Broaden mandatory platform controls |
Internal developer platforms (IDPs) advance self-service goals when teams treat the platform like a product and bake in observability and guardrails by default.
- Platform engineering has scaled quickly: 89% of organizations with platform engineering have an internal developer platform, and 60% report broad adoption across teams.
- Platform engineering teams lean hard into self-service: 74% currently provide self-service access to observability/monitoring dashboards—the most common self-service capability reported.
- Integration at scale is a challenge: 37% report integrating existing tools and systems as their top pain point. AI tools require their own provisioning and governance methods, intensifying this challenge.
Unified observability and AI-driven analysis are increasingly important for integrating systems, advancing standards, and automating workflows.
Centralized governance is critical to successful SRE and platform engineering practices
SREs and platform engineering teams are also embedding governance directly into the platforms and workflows they use to automate operations.
The report shows a strong coupling between SRE and platform engineering:
- 92% of organizations with SRE programs also have or are initiating platform engineering programs
- 82% of platform engineering programs also have or are initiating SRE
Collaboration follows: 73% say SRE and platform engineering teams collaborate and share responsibilities.
Teams also embed governance into automation:
- Nearly 70% embed security into platforms as code
- 60%+ build regulatory compliance into infrastructure and CI/CD workflows
As AI-driven operations expand, these embedded guardrails can help teams manage risk, maintain compliance, and establish clearer controls over automated actions.
How SRE and platform engineering teams use observability to advance AI initiatives
Observability adoption is widespread, but implementation remains uneven.
| Trend | Key data | Implication |
| Observability in production | 55% post-deploytment | Embed observability for greater production visibility |
| Observability gaps | 37% limited coverage | AI workloads risk under-optimization |
Most SRE and platform engineering teams have adopted observability broadly, but unevenly. For platform engineers, 77% embed observability into at least some services, but only 40% embed it into all deployments. On the SRE side, 97% use observability either through minimal logging/monitoring or full observability tools.
Uneven coverage also makes it harder to connect pre-production evaluation with production behavior. When AI workloads reach parts of the stack that observability doesn’t cover, teams can lose visibility into whether systems continue to behave as they did during testing.
Observability for AI workloads is essential. For conventional workloads, observability detects system degradation and failures. For AI, it must also identify model drift, unexpected agent behavior, and rising inference costs.
Platform engineering and SRE best practices for AI reliability
Reliable AI workloads require SRE and platform engineering teams to connect observability, automation, and AI systems more closely.
- Design for resilience first by treating reliability, security, and observability as one requirement.
- Enable humans in the loop by pairing AI-powered operations with guardrails and oversight.
- Streamline SLO and quality gate design with strong signals, smart templates, and automated evaluation that avoids dashboard overload.
- Operationalize governance-as-code (config/infra/observability-as-code) to standardize controls with auditability and consistency.
- Use observability as the control plane that converts telemetry into insight and supports governed action across services, pipelines, and AI workloads.
- AI workloads will continue increasing operational complexity and scale requirements for SRE and platform engineering teams.
The next evolution isn’t just broader observability coverage, but deeper connectivity between how AI is built and how it runs. Teams that can carry evaluation signals from development into production, and feed production behavior back into the engineering cycle, will be the ones who scale AI reliably.
FAQs
What are the key SRE best practices in 2026?
SRE best practices in 2026 focus on using observability to monitor AI workloads, defining service-level objectives (SLOs), automating incident response, and reducing telemetry complexity to improve reliability at scale.
Why is observability important for AI workloads?
Observability is critical for AI workloads because it helps teams monitor model performance, detect model drift, identify unexpected agent behavior, and manage inference costs in production environments.
How are AI workloads changing SRE and platform engineering?
AI workloads are shifting SRE and platform engineering toward predictive operations, AI-driven automation, and new observability requirements that go beyond traditional monitoring to include model behavior, cost, and accuracy.
What challenges do SRE teams face with AI workloads?
SRE teams face challenges such as fragmented telemetry, too many metrics, difficulty defining SLOs, and limited visibility into AI system behavior, which makes it harder to ensure reliability and performance.
What are the top platform engineering trends in 2026?
Platform engineering trends in 2026 include scaling internal developer platforms (IDPs), enabling self-service observability, integrating AI-powered tools, and embedding governance and security into platforms-as-code.
How does platform engineering support AI adoption?
Platform engineering supports AI adoption by providing standardized platforms, self-service tools, and built-in observability and governance, enabling developers to build and operate AI-powered applications at scale.
How do SRE and platform engineering teams use observability together?
SRE and platform engineering teams use observability as a shared system of record to unify telemetry, automate workflows, enforce governance, and improve collaboration across development and operations.
What role does governance play in AI-driven operations?
Governance further supports security, compliance, and explainability for AI-driven operations by embedding guardrails into platforms, CI/CD pipelines, and observability workflows.
What does “observability for AI” include?
Observability for AI includes monitoring model accuracy, detecting drift, tracking inference costs, analyzing agent behavior, and correlating AI performance with system-level telemetry.
Looking for answers?
Start a new discussion or ask for help in our Q&A forum.
Go to forum