AWS Cloud Operations Blog

Using Amazon Bedrock and Amazon Nova for AI-Powered Incident Response

In today’s cloud-native world, incident response teams face overwhelming challenges. When critical applications fail, engineers must sift through mountains of observability data across multiple services; all while under intense pressure to restore service quickly. This manual correlation process is time-consuming, error-prone, and often delays resolution, resulting in extended outages and frustrated customers. Traditional monitoring tools […]

Launching Amazon CloudWatch generative AI observability  (Preview)

Launching Amazon CloudWatch generative AI observability  (Preview)

As organizations rapidly deploy large language models (LLMs) and generative AI agents to power increasingly intelligent workloads, they struggle to monitor and troubleshoot the complex interactions within their AI applications. Traditional monitoring tools fall short in providing the visibility across components, leading to developers and AI/ML engineers to manually correlate interaction logs or building custom […]

SAP on AWS – Streamlined Operations and Monitoring

SAP ERP (Enterprise Resource Planning) systems are at the core of many enterprises, supporting a wide range of mission-critical processes, including Procure to Pay, Order to Cash, Production Planning, Financial Accounting, Supply Chain Management (SCM), and Human Capital Management. Given the critical role of SAP ERP, maintaining the stability, security, and efficiency of these ERP […]

Automate installing AWS Systems Manager agent on unmanaged Amazon EC2 nodes

Automate installing AWS Systems Manager agent on unmanaged Amazon EC2 nodes

Managing a fleet of AWS resources at scale can be challenging. Organizations rely on multiple solutions to automate tasks, collect inventory, patch instances, and maintain security compliance. Organizations need to access instances without opening inbound ports or managing SSH keys. AWS Systems Manager (SSM) simplifies this by serving as a centralized management solution that supports […]

Blog Post title image

Simulating partial failures with AWS Fault Injection Service

Modern distributed systems must be resilient to unexpected disruptions to maintain availability, performance, and stability. Chaos engineering helps teams uncover hidden weaknesses by deliberately injecting faults into a system and observing how it recovers. While traditional testing validates expected behavior, chaos engineering tests system resilience during failures. AWS Fault Injection Service (AWS FIS) is a […]

Observing Agentic AI workloads using Amazon CloudWatch agent

Introduction As the adoption of agentic AI applications continues to grow, ensuring the reliability, performance, and overall observability of these systems becomes increasingly critical. Agentic AI applications, powered by large language models (LLM) and integrated with various data sources and APIs, can quickly become complex, making it challenging to gain visibility into their inner workings […]

Best practices for utilizing AWS Systems Manager with AWS Fault Injection Service

Introduction In today’s cloud-centric world, ensuring the resilience of mission-critical applications is paramount. The ability to withstand and recover from unexpected failures, including degradation of cloud provider services, can mean the difference between seamless operation and costly downtime. This is where the powerful combination of AWS Systems Manager (SSM) and AWS Fault Injection Service (AWS […]

Optimizing Queries with Amazon Managed Service for Prometheus

Optimizing Queries with Amazon Managed Prometheus

Introduction In today’s cloud-native environments, organizations rely on metrics monitoring to maintain application reliability and performance. Amazon Managed Service for Prometheus serves as a tool for storing and analyzing application and infrastructure metrics. As applications and platforms evolve, teams often discover opportunities to optimize their metrics querying patterns. Common scenarios like expanding service deployments, growing […]

How Indegene Optimizes User Experience with Amazon CloudWatch

In today’s digital healthcare landscape, optimal application performance and user experience are crucial for business success. Indegene, a digital-first life sciences commercialization company, combines deep medical expertise with domain-contextualized technology to help clients accelerate innovation, modernize operations, and improve customer experience. With the world’s top 20 pharma companies among its clientele, Indegene brings an AI-first […]

Exporting a subset of AWS CloudTrail Lake events to Amazon S3

Introduction Monitoring and managing your AWS environment is critical to maintaining security and operational excellence. With the availability of AWS CloudTrail Lake data for zero-ETL analysis in Amazon Athena, you can use Athena to query your activity logs in CloudTrail Lake without the operational complexity of moving data or building data processing pipelines. CloudTrail Lake […]