Overview
High customer expectations and increasingly distributed systems mean disruptions to digital service can have catastrophic effects on sales, brand loyalty, and operating costs. The PagerDuty Operations Cloud deflects unnecessary work from teams and subject matter experts so they can focus on delivering business value. Urgent work is escalated to the right teams and routine work is made self-service. Teams can automate and accelerate issue resolutions with minimal human interruption-and improve system resilience and team capacity while reducing the strain of operational complexity and the unexpected.
With more than 700 integrations, APIs, and apps for customer service, the PagerDuty Operations Cloud empowers rapid responses in any environment. And thanks to more than 10 years of data ingestion, its machine learning-powered AIOps functionality can reduce alert noise by up to 98% and drive down MTTR with critical context for faster triage and effective automation.
PagerDuty integrates with various AWS services, including AWS CloudWatch, Amazon GuardDuty, AWS CloudTrail, AWS Personal Health Dashboard, Amazon EventBridge, AWS Security Hub, Amazon DevOps Guru, AWS Control Tower, AWS Outposts, and AWS S3 Storage Lens.
AIOps PagerDuty AIOps helps teams reduce noise, triage efficiently to drive the right actions towards resolution, and remove manual, repetitive work from the incident response process. Noise reduction baked in with an ML model that learns and adapts based on user behavior means teams see fewer incidents overall. And automating toil from manual event processing results in greater efficiency, saving teams valuable time for innovating.
Process Automation PagerDuty Runbook Automation is a managed cloud service that enables DevOps teams and SREs to create and delegate operational tasks in automated runbooks to other stakeholders such as developers, NOC personnel, and incident responders. Runbook Automation provides automated workflows and task automation focused on IT and developer process automation. Examples include service provisioning, CI/CD, configuration management, incident diagnosis and remediation, and more. With PagerDuty Runbook Automation, you can resolve requests in minutes, rather than days, optimize security and compliance, and give your engineers more time to spend on innovation rather than firefighting.
Incident Response PagerDuty helps you save time and money by bringing together the right teams with the right information to resolve incidents faster. Replace manual processes with automation to streamline incident response, freeing up time and resources for more innovation. Orchestrate end-to-end incident response with a service ownership model that only brings in the teams you need. Over 21K organizations trust PagerDuty to help them adopt DevOps best practices and build more resilient operational practices to minimize costly downtime and protect the customer experience.
Custom Private Offer All PagerDuty platforms are available through a custom private offer. Please use the contact form at <www.pagerduty.co.jp/contact-us/ > for help to create a custom offer tailored to your needs.
Highlights
- Incident Response - Manage incidents end-to-end
- Process Automation - Automate and delegate business and IT processes
- AIOps - Maximize IT capacity with fewer incidents and faster resolution
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Buyer guide

Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/12 months |
|---|---|---|
PagerDuty Subscription | PagerDuty Operations Cloud on AWS | $50,000.00 |
The following dimensions are not included in the contract terms, which will be charged based on your usage.
Dimension | Cost/unit |
|---|---|
Additional events over contracted value | $0.0805 |
Vendor refund policy
All fees are non-cancellable and non-refundable except as required by law.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Software as a Service (SaaS)
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
Support
Vendor support
Our team provides multiple resources for customers to find answers to questions and get help with our product. Users may browse our integration guides (pagerduty.com/integrations) to integrate with partner tools, our knowledge base (support.pagerduty.com) to learn more about using PagerDuty, and our developer docs (developer.pagerduty.com) to use our APIs. Additionally, anyone can interact with other PagerDuty users and PagerDuty employees via the PagerDuty Community (community.pagerduty.com). Our Support team is available during regular business hours around the globe, Monday through Friday, and can be contacted at: Email: support@pagerduty.com or via a ticket submitted at tickets.pagerduty.com
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products

Customer reviews
Unified incident response has reduced alert noise and improves on-call focus and coordination
What is our primary use case?
I have used PagerDuty Operations Cloud for two years, so I have over two years of experience with it.
My main duties in PagerDuty Operations Cloud are incident routing, escalation policies, on-call management, incident response coordination, and integrations.
For incident routing, I connected AWS CloudWatch alarms to PagerDuty using the Event API. When our EC2 CPU crosses 85%, CloudWatch pushes the alert to PagerDuty, which then routes it to the backend services team. I also used event rules to auto-tag alerts and suppress duplicates. Another use case is on-call management. I used PagerDuty's mobile app to acknowledge alerts during off hours. For example, on weekends, such as Sunday morning, when there is an incident I acknowledged a high-priority alert from my phone and joined the Slack room directly from the PagerDuty incident link. This is helpful for those who are on call during off shifts. Regarding integrations, since we use many integrations on PagerDuty, we integrated PagerDuty with Slack so incidents automatically created a dedicated channel. I also connected it to Jira so incidents could generate tickets for follow-up actions, and also to ServiceNow as well.
What is most valuable?
PagerDuty Operations Cloud excels at routing alerts, managing on-call schedules, coordinating incidents, and automating remediation. These features make it helpful for the on-call team and the people who are working with PagerDuty Operations Cloud.
The feature I use the most is incident routing. We use incident routing more than anything else because it is the main foundation of PagerDuty Operations Cloud. If routing is incorrect, nothing else is going to work, no escalations, no on-call, and no automation at all. I use it the most because it is the first step in every incident; it ensures that alerts from different cloud platforms or any other third-party providers such as AWS , Azure , and other platforms including Datadog reach the right service. It prevents misrouted alerts and missed incidents and it directly impacts the MTTA, because responders get the alert instantly. Every other PagerDuty Operations Cloud feature depends on routing being accurate. I use incident routing daily in this way. This is the main foundation of PagerDuty Operations Cloud, to work with it and to make it a very easy platform for us to work with.
PagerDuty Operations Cloud excels at routing alerts, managing on-call schedules, coordinating incidents, and automating remediation. It ensures critical alerts reach the right teams instantly, reducing the MTTA and MTTR. It strengthened the on-call process with dependable escalations and real-time incident coordination, preventing missed alerts and speeding up resolution tasks. Automation and noise reduction features help cut down repetitive manual work and alert fatigue, allowing engineers to focus on real issues.
PagerDuty Operations Cloud significantly reduced alert noise and improved MTTA. Our acknowledgement time dropped almost by half, and MTTR improved by around 15–20% due to improved automation and routing.
The alert noise dropped by 20–25% after tuning event rules. After implementing PagerDuty Operations Cloud, we saw a 20–25% reduction in alert noise and a noticeable improvement in MTTA. Our acknowledgement time dropped by almost half and the MTTR improved by around 15–20%. MTTR improved by 15–20% due to major incident automation and better routing, which also reduced unnecessary on-call interruptions. It made a good impact on our current schedules and in my current organization.
PagerDuty Operations Cloud's alert reduction features helped lower operational cost by cutting unnecessary alerts and reducing on-call interruptions by grouping and suppressing noisy alerts. The team spent less time triaging low-value incidents, which directly reduced overtime and burnout. The biggest impact for us was fewer false alarms, meaning engineers could focus on real issues instead of constant alert fatigue.
What needs improvement?
PagerDuty Operations Cloud could improve its noise reduction by making deduplication and suppression more automated instead of manually tuned, and the service dependency graph could be more intuitive with cleaner visuals and easier-to-understand root cause tracing during major incidents. Automation could go further with smarter runbook triggers and AI-driven suggestions to help find root causes, which could save a lot of time for engineers who are struggling to understand what is actually happening. These AI capabilities could lower the time by maybe 50–60%. Analytics and reporting could be more flexible, allowing custom dashboards and filters and team-level MTTA and MTTR breakdowns, so that it is segregated based on teams and it is much easier to have custom dashboards for the teams to understand more.
Alert storms were a recurring frustration for the on-call team, and escalation overrides and service dependency graphs can get a bit confusing in larger environments. More customizable analytics and smarter automation would make the platform even easier, more flexible and more powerful for the team to understand.
PagerDuty Operations Cloud could improve its analytics flexibility with customized dashboards and AI capabilities to be more trained and more reliable. Service dependency mapping can also feel cluttered in big environments, making it harder to trace upstream and downstream impacts during bigger incidents.
We implemented PagerDuty Operations Cloud's AI to help with alert grouping and early incident insights, but accuracy was not consistent enough to rely on during critical events. It occasionally grouped unrelated alerts or missed correlations, which limited the operational efficiency gains we expected. The on-call team still depended heavily on manual triage because AI suggestions were not always aligned with the real root cause. Overall, AI added some value but it has not yet reached the reliability needed to significantly improve the incident response efficiency. It needs more training or more work.
What do I think about the stability of the solution?
PagerDuty Operations Cloud is stable in day-to-day use with no major outages or reliability issues affecting our on-call workflows. Alert delivery, incident routing and escalation chains have been consistently working without delays or missed notifications. Its cloud-native architecture also means updates and patches roll out smoothly without impacting uptime. Overall, the stability has been one of the strongest features for our team.
What do I think about the scalability of the solution?
PagerDuty Operations Cloud scales well for growing teams. Adding services, integrations and new on-call groups does not impact performance. Its native cloud architecture handles higher alert volumes smoothly, especially when paired with strong incident routing. We have seen it manage increased workloads without slowing down or causing notification delays. On the whole, scalability has been reliable even as our environment and service footprint has expanded.
How are customer service and support?
The customer support has been responsive and generally helpful when we have raised any issue. They provided clear troubleshooting steps and follow-ups. Resolution speed can vary depending on the severity of the issue, but support has been good and reliable for day-to-day operational needs.
Which solution did I use previously and why did I switch?
Before PagerDuty Operations Cloud, we used basic native alerting tools such as AWS CloudWatch since AWS is our primary cloud platform. We switched because those tools lacked valuable incident routing and consistent on-call escalation workflows. PagerDuty Operations Cloud offered stronger coordination, better noise reduction, and more mature incident response features. The move was mainly driven by the need for a unified and dependable on-call and incident management platform.
What was our ROI?
We have actually seen a moderate return on investment mainly through reduced alert noise and more efficient incident routing. Alert storms dropped by roughly 20–25%, saving the team several hours per week that used to be spent on low-value triage. We estimate a small cost reduction from fewer overtime hours and less on-call fatigue, though not enough to reduce headcount. Overall, the return on investment is positive but incremental — more time saved than money saved.
What's my experience with pricing, setup cost, and licensing?
PagerDuty Operations Cloud's pricing felt reasonable but it definitely is not the cheapest option. The value comes more from reliability than cost savings. Setup costs were minimal since it is a SaaS platform and onboarding did not require any infrastructure or hidden implementation costs or fees. Licensing is straightforward but scaling seats for larger teams can get expensive, especially when adding advanced features. Overall the experience was smooth but the price point could be more flexible for growing teams. Our team is small now but in the future it might become bigger.
Which other solutions did I evaluate?
Before adopting PagerDuty Operations Cloud, we evaluated options such as Opsgenie and Splunk On-Call for incident management. We also looked at other native cloud tools such as AWS CloudWatch, but they lacked the escalation workflows. Opsgenie had good features but did not match PagerDuty Operations Cloud's reliability and integrations for bigger teams. Overall, PagerDuty Operations Cloud offered stronger coordination and more consistent on-call performance which made us choose it.
What other advice do I have?
Escalation policies are a very helpful feature for managing the team, such as multi-step escalation logic which is extremely dependable and prevents missed alerts, and also the automation runbook. I rely on incident routing daily because it ensures the alerts reach the right team. The other features, such as integrations, are really good to integrate PagerDuty Operations Cloud with different platforms, such as Jira and ServiceNow platforms to automatically open an incident. These are really helpful top features.
I would give PagerDuty Operations Cloud eight out of ten, considering all the features I have been using and also the improvements I have suggested. I chose eight out of ten because PagerDuty Operations Cloud is generally strong in the areas that matter the most, especially incident routing and escalation management which are extremely reliable and directly reduce MTTA and MTTR. It stands out for dependable on-call orchestration and smooth incident response coordination, especially with Slack and Jira integrations. But it does not reach a ten because noise reduction still needs more automation, analytics lack deeper customization such as custom dashboards to filter among the teams, and the service dependency graph gets cluttered in large environments. With smarter automation and improved visualization of the analytics, it could easily move closer to a perfect ten.
Which deployment model are you using for this solution?
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Automation has reduced manual support effort and has improved incident response accuracy
What is our primary use case?
My main use case for PagerDuty Operations Cloud is for doing the automation part and minimizing the efforts. It decreases the count of the L1 support team because most of the tasks have been handled by the operation of PagerDuty Operations Cloud and the AI.
What is most valuable?
PagerDuty Operations Cloud offers several valuable features that have significantly benefited our operations. The solution keeps the person available and only rings the alert while there is any outage or any issue generated in the live environment. Additionally, it has an automated section where it handles similar issues if they happen again and is able to troubleshoot them by itself by running those jobs or actions that have been put in PagerDuty Operations Cloud.
PagerDuty Operations Cloud positively impacts our organization by helping us increase the overall SLA delivery. Earlier, we were having a delivery count of 87% or 88%, but after setting up PagerDuty Operations Cloud, we were able to achieve the SLA up to 99.5% post-installation. The improvement in SLA and SLI is mostly because of the faster incident response and fixing the issues that are most generic and reoccurring, which has been minimized and helped us increase the SLA and SLI.
What needs improvement?
PagerDuty Operations Cloud is already at its best, but AI can be integrated just to set it up and continue learning and provide a list of fixes that can be applied as suggestions, as that might help improve PagerDuty Operations Cloud operations.
For how long have I used the solution?
I have been using PagerDuty Operations Cloud for around three years or more.
What do I think about the stability of the solution?
PagerDuty Operations Cloud is stable and much more stable than previous solutions.
What do I think about the scalability of the solution?
PagerDuty Operations Cloud's scalability is much more stable than expected. We can scale it up whatever the requirements are, so scalability is good.
How are customer service and support?
Regarding customer support, we never found any reason to get support from the customer. We were already supported while setting up the infrastructure and it was good. I would rate the customer support on a scale of one to ten as ten out of ten.
Which solution did I use previously and why did I switch?
Previously, we were using Teams for generating the alerts and Slack, but we prefer PagerDuty Operations Cloud, which offers many more options and setups that can be used.
How was the initial setup?
To deploy PagerDuty Operations Cloud in our organization, we have used public and private cloud. Earlier it was on-premises but we switched to private cloud.
What about the implementation team?
We have implemented the AI and automation through PagerDuty Operations Cloud for incident response. As mentioned earlier, it increased the SLA from 87% to 99% and our operations have been improved, and there is a very low count of clear or false alerts.
What was our ROI?
We have seen a return on investment. We have cost-cutting on the employees as we have decreased the headcount since there is less man labor required while PagerDuty Operations Cloud is able to handle all the alerts first as an L1. There is a lot of money saved compared to the pricing or the license that we have spent.
What's my experience with pricing, setup cost, and licensing?
My experience with pricing, setup cost, and licensing for PagerDuty Operations Cloud is that pricing and license are good to go. Everything is perfect as per the licensing and pricing setup cost. There is no more suggestion I can provide based on that. We completely achieved whatever we are spending.
Which other solutions did I evaluate?
Before choosing PagerDuty Operations Cloud, we evaluated other options and chose Teams.
What other advice do I have?
We are using PagerDuty Operations Cloud to our maximum capacity and we appreciate the support that it has for continuing to learn from the issues that we have and to predict the outages or any alerts or whatever escalation is required. It can fulfill that.
PagerDuty Operations Cloud's embedded AI has significantly influenced revenue protection in terms of reducing alert fatigue and incident costs. We have improved a lot, and we have improved 25 to 27% of our overall incident management.
The alert reduction feature of PagerDuty Operations Cloud has a significant impact on preventing costly incidents in our organization. Alert reduction has been decreased as a result of removing false alerts. Earlier, we were getting 30 to 35 alerts per day. Now, there are only four or five general alerts that are genuine. That is a significant amount of alert reduction.
For others looking into using PagerDuty Operations Cloud, I would completely suggest that PagerDuty Operations Cloud is a good option to install or set up in infrastructure. It handles the infrastructure very well. I would rate this review nine out of ten overall.
Reliable incident paging has kept outages under control but separate client notes are still missing
What is our primary use case?
My usual use cases with PagerDuty Operations Cloud involve handling incidents through a full flow. When there is an outage, an incident is created that can be either severity one or severity two. As the person on call that day, I receive a page from PagerDuty on the app and three calls on my cell phone. When I pick up the call, PagerDuty IVR asks me to acknowledge the incident. Once I acknowledge the incident from the call, I go to PagerDuty through the website, which is much easier to navigate than the mobile app. I then page other teams responsible for the incident, as well as the stakeholders and product owners.
PagerDuty is integrated with Microsoft Teams, so I open a Teams bridge call to resolve the issue and update all incident details in PagerDuty notes, which automatically integrates with ServiceNow incident and sends messages to stakeholders' phone numbers.
One time while hanging out in the mountains, there was no internet signal but there was cell reception. An incident happened while I was on call that day. Normally, without internet, I would not be able to know about it, but because of PagerDuty, I was paged three times on my cell phone as well as through text message. I managed to call another colleague from my work and told him to take care of the incident. This helped me avoid breaching the SLAs on incident acknowledgment and allowed me to access remote incidents without relying solely on the internet.
What is most valuable?
The features of PagerDuty Operations Cloud that I find most valuable involve being automatically paged whenever an incident is triggered. PagerDuty has group names embedded into it, and when we set up PagerDuty in our organization, we embedded the group name, allowing me to page other respondents without having to go separately into Microsoft Teams, add everybody's name, and then ping and call them. I can do this directly from PagerDuty itself.
The notes update feature allows me to put the details of the incident in the notes and click post, and it is integrated everywhere. Everything is centralized.
PagerDuty Operations Cloud has improved my team's ability to focus on core tasks rather than routine issues primarily due to its availability and reliability. We do not worry about whether PagerDuty will call us when an incident triggers, allowing us to focus on that incident. The notes update process ensures that everyone gets informed with just one update.
What needs improvement?
I think PagerDuty Operations Cloud could be improved by having two fields for incident updates. In my work, I handle incidents that have two fields: work notes visible only to developers working on the incident and additional comments visible to the client. When I update through PagerDuty, everything gets updated into the additional comments. There should be two fields, perhaps based on how it is integrated with ServiceNow.
I have not used PagerDuty's autonomous AI agents or generative AI, so I am not sure whether it is integrated.
I can evaluate the effectiveness of PagerDuty Operations Cloud in providing insights for decision-making, and I rate it around seven out of ten. PagerDuty Operations Cloud is really helpful. With generative AI integrated and a chatbot, I think it would bump the rating up to nine, but I have not used it yet, so I cannot say for certain.
For how long have I used the solution?
I have been using PagerDuty Operations Cloud for around one and a half years.
What do I think about the stability of the solution?
Regarding the reliability and stability of this product, I find it really stable. Based on my personal experience in the mountains, I can vouch for it. My organization measures the acknowledgment rate by checking how many times a call rings and how quickly the responder acknowledges the incident. The acknowledgment rate is really good, around ninety to ninety-five percent.
What do I think about the scalability of the solution?
I think the scalability of PagerDuty Operations Cloud is superb. It can handle many requests according to demand. As other members and teams are added to our organization, it has not impacted the latency, call failure, or anything else. The performance is really good.
How are customer service and support?
I do not often communicate with the technical support of PagerDuty. The manuals are created by our team, so we use those.
Which solution did I use previously and why did I switch?
I did not use a different solution for the same use case before PagerDuty Operations Cloud. PagerDuty Operations Cloud was the first solution I used.
How was the initial setup?
I did not participate in the initial setup of PagerDuty Operations Cloud. When I joined the organization after one and a half years of use, it was already set up for me by the IT team.
What about the implementation team?
I have not personally implemented automation through PagerDuty for incident response, but I think some of my colleagues may have done so.
What was our ROI?
I am not aware of the impact of PagerDuty's alert reduction feature on preventing costly incidents in our organization, as that metric is not shared with me.
What's my experience with pricing, setup cost, and licensing?
I am not sure about the influence of PagerDuty Operations Cloud on revenue protection in terms of reducing alert fatigue and incident costs, as these are organizational-level decisions, and employees are not involved in those discussions.
Which other solutions did I evaluate?
Before PagerDuty Operations Cloud was chosen, I did not evaluate other options, as the decision was made by the board members.
What other advice do I have?
My review rating for PagerDuty Operations Cloud is seven point five out of ten.
Which deployment model are you using for this solution?
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Alert de-duplication has reduced noise and now improves response time and root cause analysis
What is our primary use case?
I have integrated other monitoring tools like LogicMonitor , and alerts come to PagerDuty Operations Cloud where we acknowledge them and work upon the issues. I have created multiple services that send alerts to out-of-hours groups for the on-call engineers. We detect issues faster and this helps in root cause analysis. Alert noise reduction is a major use case for us, as it groups duplicate alerts, which is very useful. The mobile application is also excellent.
What is most valuable?
I appreciate the event de-duplication feature in PagerDuty Operations Cloud because my company has many alerts for similar devices or servers, and it groups them together. This helps us see when a particular server's CPU and memory are both spiking, which aids significantly in root cause analysis.
Another feature I value is push notifications. We receive calls, SMS messages, and emails for the same alert, so we do not miss any notifications.
My organization has reduced noise by approximately 20% because of the de-duplication feature in PagerDuty Operations Cloud and the report feature. The report feature sends us a weekly report showing how many similar alerts occurred that week, and we work on reducing those alerts. By following this policy for three months, we reduced noise by 20%, which is a huge achievement for us.
PagerDuty Operations Cloud has improved our response time and mean time to resolution in my organization. We have integrated many monitoring tools through PagerDuty Operations Cloud, and the integration feature is excellent. It integrates very well with other monitoring tools via API and through email. I recommend other organizations use this integration feature.
The platform generates weekly reports showing how many alerts we received and the response time for each service and alert. I now pull daily reports via API. Since my company operates from 7:30 AM to 4:30 PM, with on-calls after hours, I need to know how many alerts occur outside business hours. Using a report scheduled through PagerDuty Operations Cloud API, the system sends me the alerts. I then analyze how many alerts came that night and work with the application team to reduce noise and resolve incidents. I value the report feature completely.
As a technical engineer, I observe that noise is being reduced and platform stability is increasing. My company is product-based with many products, and they are becoming more stable because we receive alert notifications faster. PagerDuty Operations Cloud is helping my organization tremendously.
What needs improvement?
Overall, I have positive feedback about PagerDuty Operations Cloud, but as an enhancement, I would suggest the reporting feature could be improved. I generate reports based on the service, but it has a limitation where it cannot send all alerts. The limitation is that it can only send 1,000 incidents using the API. If that capacity could be enhanced to send 2,000 alerts in one report, that would be beneficial.
Currently, we have not applied any automation through PagerDuty Operations Cloud. However, it does help with automation in that when we receive more alerts for a similar issue or for only one server, we know that server's health is not good. We then find the root cause and apply automation directly on the server, not through PagerDuty Operations Cloud. The feature would be useful, but my company does not have the automation feature enabled. It shows as a request trial, so I think we need to try that.
For how long have I used the solution?
I have been using PagerDuty Operations Cloud for one year and two months.
What do I think about the stability of the solution?
We do not experience downtime. However, I have observed one issue: we integrated with LogicMonitor , which is a monitoring tool, and alerts come from there to PagerDuty Operations Cloud. When alerts are resolved in LogicMonitor, they should also resolve in PagerDuty Operations Cloud, but sometimes they do not resolve. This should happen, and I think this is an API issue that needs to be addressed. I am not certain whether other customers of PagerDuty Operations Cloud are experiencing the same issue.
How are customer service and support?
I have no complaints about customer service because PagerDuty Operations Cloud is an incident management tool and it performs that function very well.
Which solution did I use previously and why did I switch?
We preferred PagerDuty Operations Cloud over ServiceNow , which we used previously for the same purpose. When an alert came, we would call engineers, and ServiceNow has that feature as well. However, PagerDuty Operations Cloud is much more advanced in terms of notifying users and reducing the time to respond. We are satisfied with it and are not planning to move to other tools currently.
How was the initial setup?
I joined this organization one year and two months ago, and the initial setup was already done. I only enhanced that setup and created new integrations and new event orchestrations. I cannot comment on the initial setup itself, but I am confident it would have been easy.
Which other solutions did I evaluate?
Overall, I can say PagerDuty Operations Cloud is a critical part of our incident management process. Reliability and alert delivery are strong compared to other tools such as ServiceNow. The area where we see the biggest opportunity is AI-driven event correlation, richer alert context, and improved analytics. I do not think any other tool is near that level. We tried ServiceNow because we have it as well, but it does not match PagerDuty Operations Cloud. The overall feedback is positive.
What other advice do I have?
I would recommend that organizations with high alert noise, whether similar to my company or larger companies, should try PagerDuty Operations Cloud. They should use its event and alert de-duplication features and integration with other tools, which are excellent. The calling notification feature is also very good. Overall, it is a strong solution. I rate PagerDuty Operations Cloud as nine out of ten because I do not see any gaps in what I use on a daily basis.
Integration workflows have become seamless and now power AI-driven incident management
What is our primary use case?
I started as a user working in an operations team where we handled the AWS infrastructure deployed for a particular company. Whenever any issue occurred, we received pages using PagerDuty Operations Cloud . I gradually learned about PagerDuty Operations Cloud and started integrating it into different workflows. I began integrating it to make Slack bots, and right now I am using PagerDuty Operations Cloud API endpoints to make AI agents as well. I can think of myself as an integration engineer who works extensively with integrating different services, one of which is PagerDuty Operations Cloud.
I also use xMatters alongside PagerDuty Operations Cloud. Speaking from an integrations engineer's perspective, I have not integrated xMatters heavily, but I have been a user of xMatters more lately. The major difference I observed was the workflow management. xMatters has better workflow management than PagerDuty Operations Cloud. Let me explain what I mean by workflow management. If you have a company with ten teams working on a particular product, every member of those teams may or may not receive a page. Every member should have an orchestrator where they can define custom rules such as a service should be paged directly to X team, or if a page comes from ABC issue, it should directly go to Y team without having me to manually put the team name or team details regarding where to page. This capability is lacking in PagerDuty Operations Cloud while in xMatters, it is flawlessly integrated where you can add custom rules and custom rule sets. In PagerDuty Operations Cloud, we have to create separate pages for that functionality. However, when I talk about integration, the xMatters API toolkit is confusing and disorganized. The tree structure is not present in xMatters, but I appreciate that about PagerDuty Operations Cloud. Integration-wise, PagerDuty Operations Cloud is flawless. I love PagerDuty Operations Cloud from an integration perspective, but it makes my life difficult if someone wants me to integrate xMatters.
What is most valuable?
I appreciate the overall API toolkit very much. It is one of the simplest API toolkits I have seen that lets me do literally anything via API calls, which I can essentially do via the browser. I do not even need to log into my browser to do anything if I have a CLI or any tools integrated with it.
I have built an AI agent which detects if any page comes into PagerDuty Operations Cloud. PagerDuty Operations Cloud has webhooks, which is great. If anything comes into PagerDuty Operations Cloud, I basically poll every detail of that page, perform some incident resolution, and do something on the infrastructure according to whatever page I receive. I add comments in the pages via PagerDuty Operations Cloud API and then do the whole incident life cycle using all the APIs. There is also a very good Python library called PDPYRAS, which I use extensively, which uses PagerDuty Operations Cloud APIs to build SDK. I have developed my own CLI toolkit using PagerDuty Operations Cloud APIs itself, which is on my GitHub .
The UI is good and looks good, but sometimes when pages come very frequently, such as receiving ten to fifteen pages per five to ten seconds, it works flawlessly. However, when you tie your PagerDuty Operations Cloud instance to very large-scale infrastructure where you have millions of instances and get at least five to ten pages per second, the UI starts to hang. The API works flawlessly even then, but if someone does not know how to use all the integrations that PagerDuty Operations Cloud provides, they have only one choice but to log into the UI and check for the pages. Then they will have to face the lag in the UI.
Integration is very easy. I have completed entire integrations, deployments, and testing within six hours. It is just so easy.
One person can do everything end to end on their own. I have done it multiple times, and I have seen other people doing it multiple times as well. Everything is very seamless. I should also appreciate the official documentation that you have. Usually official documentation is not that good, but yours feels like someone has taken time to write those documents. The commands which you have written for back-end integration are straightforward. I literally have to just copy and paste after setting environment variables.
What needs improvement?
I believe you really need to work on your UI. The UI is good and looks good, but when pages come very frequently, such as receiving ten to fifteen pages per five to ten seconds, it works flawlessly. However, when you tie your PagerDuty Operations Cloud instance to very large-scale infrastructure where you have millions of instances and get at least five to ten pages per second, the UI starts to hang. The API works flawlessly even then, but if someone does not know how to use all the integrations that PagerDuty Operations Cloud provides, they have only one choice but to log into the UI and check for the pages. Then they will have to face the lag in the UI. This is the area where I feel improvements are needed.
PagerDuty Operations Cloud does require a fair amount of maintenance. Many incidents go into triage and need to be regularly cleaned up. If an incident comes for which we have not set rules to auto acknowledge and close, it basically stays in the triage and we have to manually clean it up. I think an auto-detection mechanism can be implemented out of the box. Because it is not there currently, we have to develop modules around that. An alternative would be webhooks that you can come up with which we can utilize to implement this out of the box.
The only lag I have experienced is the UI lag that I have already described.
For how long have I used the solution?
I have been using PagerDuty Operations Cloud for more than five years now.
What do I think about the scalability of the solution?
PagerDuty Operations Cloud is scalable. I have not faced any issues with scalability. It is pretty good when it comes to scalability.
How are customer service and support?
I have contacted support multiple times.
There are ticket levels which we can create. I have not called the support team on the phone, but I have mailed and raised tickets with them. There have been instances where I had to integrate a very old server, AIX server framework to PagerDuty Operations Cloud for which the modules were not present in PagerDuty Operations Cloud. I am not blaming them for this; I am blaming the company who are using AIX servers. However, they did not have the module, so I had to raise a ticket. It was around Severity 3. The response was not super fast, which is expected. However, as soon as it escalated to Severity 2, I received an immediate email from PagerDuty Operations Cloud team. I think the support is fine. I have faced just one time some inconvenience. I do not know what was the reason, but there was one time where I did not receive any response for two days. Thank goodness it was not a Severity 1 incident for us, but two days was unacceptable at that point. Hence we started looking at other products including xMatters. There has been just one instance, but the company which I work with, even one instance sometimes causes a lot of friction. People start looking at other options. However, there has been just once.
How was the initial setup?
Integration is very easy. I have completed entire integrations, deployments, and testing within six hours. It is just so easy.
What about the implementation team?
One person can do everything end to end on their own. I have done it multiple times, and I have seen other people doing it multiple times as well. Everything is very seamless. I should also appreciate the official documentation that you have. Usually official documentation is not that good, but yours feels like someone has taken time to write those documents. The commands which you have written for back-end integration are straightforward. I literally have to just copy and paste after setting environment variables. Everything is very easy.
Which other solutions did I evaluate?
I use xMatters alongside PagerDuty Operations Cloud. Speaking from an integrations engineer's perspective, I have not integrated xMatters heavily, but I have been a user of xMatters more lately. The major difference I observed was the workflow management. The workflow management in xMatters is better than PagerDuty Operations Cloud. Let me explain what I mean by workflow management. If you have a company with ten teams working on a particular product, every member of those teams may or may not receive a page. Every member should have an orchestrator where they can define custom rules such as a service should be paged directly to X team, or if a page comes from ABC issue, it should directly go to Y team without having me to manually put the team name or team details regarding where to page. This capability is lacking in PagerDuty Operations Cloud while in xMatters, it is flawlessly integrated where you can add custom rules and custom rule sets. In PagerDuty Operations Cloud, we have to create separate pages for that functionality. However, when I talk about integration, the xMatters API toolkit is confusing and disorganized. The tree structure is not present in xMatters, but I appreciate that about PagerDuty Operations Cloud. Integration-wise, PagerDuty Operations Cloud is flawless. I love PagerDuty Operations Cloud from an integration perspective, but it makes my life difficult if someone wants me to integrate xMatters.
What other advice do I have?
Regarding pricing, I do not remember the current prices, but I used the first tier about two years back for one of the startup failures which I was working on. That startup did not work out, but I integrated PagerDuty Operations Cloud with a lot of things there.
For the enterprise and for large-scale enterprises, the pricing is good. I will not even say it is fine; it is good for large-scale enterprises. However, for small-scale startups and small businesses, because they already are in a very nascent stage, the pricing is a little on the higher side. There is no custom module which I can just add to my cart which gets me custom pricing. It is just one bucket. For small-scale operations, I think the pricing is a bit pricey.
I do not use PagerDuty Operations Cloud's AI assistant, but I integrate the back-end to create agents. I do not use their default ones. I have never used it, and I do not know how good or bad that is. I am integrating PagerDuty Operations Cloud modules, APIs, and SDKs to develop AI agents, not using anything which comes out of the box.
My overall review rating for PagerDuty Operations Cloud is eight out of ten.