tl;dr: Platform engineering is making real progress, but fragmented tooling, uneven automation gains, and new AI governance demands are keeping teams from realizing their full potential. The data shows where teams are getting stuck and what they can do to move forward.
The idea behind platform engineering is simple: Give developers consistent tools, paved paths, and fewer reasons to fight their way through one-off workflows. The field is maturing quickly as organizations adopt internal developer platforms (IDPs) to accelerate software delivery: 89% of organizations now have an IDP, and 60% report broad adoption. Observability and monitoring dashboards are the leading IDP capability, available in 74% of organizations. But as teams bolt on new tools, cloud services, and AI assistants, the strain is starting to show. AI systems do not fail like conventional software. They drift, behave probabilistically, and introduce new questions around accuracy, cost, security, and governance.
The State of SRE and Platform Engineering 2026 report from Dynatrace details platform engineering challenges related to integration, automation outcomes, compliance, and AI governance. Despite progress, several real-world friction points are keeping platform engineering from reaching its potential. This post digs into the key issues the data exposes.
Integration and standardization are the biggest headwinds
Integration is still where many platforms struggle. More than a third of organizations—37%—say their biggest challenge is simply getting existing tools and systems to play nicely with their platforms, while another 32% struggle to keep standards consistent across teams.
The pressure is only growing as AI enters the picture. Connecting internal systems to LLMs, model monitoring tools, and workflows powered by AI introduces yet another layer of complexity. Without shared standards and predictable integration points, the gap between what teams want from AI and what they can realistically wire together continues to widen. The good news is that early interoperability efforts like Model Context Protocol (MCP), OpenLLMetry, and Agent-to-Agent protocol (A2A) hint at a better standards-based path. Leading organizations are already using MCP to link disparate systems into unified coding assistants.
What to do about it
- Standardize your instrumentation layer to consolidate fragmented observability tooling across teams and systems, then enrich it with domain context — ownership, baselines, criticality — so everyone is working from meaningful, connected signals rather than noise.
- Create standard golden-path templates that embed telemetry, Service Level Objectives (SLOs), and security controls by default, and use a consistent protocol like MCP to give teams unified access to infrastructure and application context rather than stitching together ad-hoc data pulls across tools.
- Use observability as the shared context layer across platform services, pipelines, and AI tools, so automation is grounded in correlated telemetry, topology, ownership, and business context rather than fragmented data pulls.
Automation benefits aren’t matching expectations
For all the talk about automation reducing toil, the numbers show a more nuanced story. Most teams walked into AI adoption expecting a noticeable drop in routine work—53% anticipated increased automation and reduced toil—but only 45% have actually seen those gains materialize. The gap is similar on the deployment side: Nearly half of respondents expected more consistent releases, but only 42% observed that improvement in practice. It’s the kind of mismatch that suggests automation is still hitting the limits of integration, context, or both.
A bright spot is that 55% observed improved developer productivity from AI, close to the 58% who expected it. However, 55% expected lower operational costs and only 45% observed lower costs, and 53% expected AI to improve MTTD/MTTR but only 46% saw improvement.
Automation is improving, but it’s not compounding the way teams hoped. Developer productivity is rising, but broader gains in cost reduction, incident response, and toil reduction depend on something harder: connecting automation to reliable operational context. For workflows powered by AI, that means pairing automation with observability, guardrails, and human oversight before handing systems more autonomy.
What to do about it
- Design automation for resilience first by treating reliability, security, and observability as requirements for every automated workflow, not just the services those workflows support.
- Treat model drift as an operational risk. Monitor AI behavior continuously in production so teams can catch degraded outputs before they trigger unreliable automation.
- Keep humans in the loop with clear guardrails, escalation paths, and explainable workflows until automations powered by AI are trustworthy enough to expand.
Proving value and reducing toil remain pain points
Proving the value of platform engineering is still harder than it should be. 30% of organizations say they struggle to make a clear case. Another 30% still deal with excessive manual work. Additionally, compliance adds its own drag—nearly half of teams handle at least some compliance steps by hand, and 31% cite security and compliance complexity as an ongoing pain point.
AI was supposed to chip away at some of this, but it’s become a double-edged responsibility. More than half of platform engineers—52%—now spend time monitoring AI tools to ensure data security, model performance, and accuracy. That’s work no one was planning for a year or two ago.
Platform engineering is making developers more productive, but the platform teams themselves are still carrying the weight of manual work, compliance noise, and AI governance.
What to do about it
- Use configuration-as-code to version and automate observability and platform settings across environments with full auditability, giving you a clear, demonstrable record of what your platform is doing and the value it’s delivering.
- Embed observability into platform pipelines to validate releases and enforce reliability standards continuously, reducing manual compliance.
- Use observability as a unified intelligence layer to correlate telemetry, topology, ownership, and business context, so AI observability becomes a governed capability for tracking accuracy, drift, latency, cost, and behavioral consistency.
Putting it together
The numbers point to a discipline that is maturing fast but still fighting the drag of fragmented systems, uneven automation, and AI workloads that do not behave like traditional software. Even the most advanced teams are learning that progress now depends on understanding how their platforms behave under real‑world pressure. The organizations that can bring that picture into sharper focus will be the ones that turn today’s growing pains into tomorrow’s advantage.
Looking for answers?
Start a new discussion or ask for help in our Q&A forum.
Go to forum