View a markdown version of this page

Alerts from baseline monitoring in AMS - AMS Advanced User Guide

End of support notice: On June 30, 2027, AWS will end support for AMS Advanced. After June 30, 2027, you will no longer be able to access the AMS Advanced console or AMS Advanced resources. For more information, see AMS Advanced end of support.

Alerts from baseline monitoring in AMS

Learn about AMS monitoring defaults. For more information, see Monitoring and event management in AMS.

The following table shows what is monitored, and the default alerting thresholds. You can change the alerting thresholds with a Management | Other | Other | Update (ct-0xdawir96cy7k) RFC after determining what changes you want and subscribing to the relevant CloudWatch Amazon SNS topic. For information about creating and subscribing to topics, see Subscribe to a Topic. For general information, see Amazon SNS FAQs. To be notified directly when alarms cross their threshold, in addition to AMS's standard alerting process, follow these instructions about how to overwrite alarm configurations, Receiving alerts generated by AMS.

Amazon CloudWatch provides extended retention of metrics. For more information, see CloudWatch Limits.

Note

AMS calibrates its baseline monitoring on a periodic basis. New accounts are always onboarded with the latest baseline monitoring and the table describes the baseline monitoring for an account that is newly onboarded. AMS updates the baseline monitoring in existing accounts on a periodic basis and you may experience a time lag before the updates are in place. For more information, see Viewing the monitoring configuration for an AMS account.

Note

The EC2 instance alert Non-root volume usage is DISABLED by default. If you require alert generation based on this alarm, then you must enable it using the RFC Change Type ct-0erkoad6uyvvg

Alerts from baseline monitoring

Service

Security alert

Alert name and trigger condition

Notes

For starred (*) alerts, AMS proactively assesses impact and remediates when possible; if remediation is not possible, AMS creates an incident. Where automation fails to correct the issue, AMS informs you of the incident case and an AMS engineer is engaged. In addition, these alerts can be sent directly to your email (if you have opted in to the Direct-Customer-Alerts SNS topic).

Amazon EC2 instance – Windows

No

SecureChannelFailure

> 0.0 for 10 out of the last 15 data points.

CloudWatch alarm on Windows instances to alert when Secure a Channel connection has failed.

Aurora instance

No

CPUUtilization

> 85% for 5 mins, 2 consecutive times.

CloudWatch alarm.

AWS Backup

Yes

DeleteRecoveryPoint

An unexpected IAM role principal or IAM user principal has deleted an AWS Backup recovery point.

CloudWatch event. Emitted when a backup recovery point is deleted.

AWS Outposts

Yes

AMSOutpostsInstanceFamilyCapacityAvailability InstanceFamilyCapacityAvailability

= 80% for 5 minutes, 12 consecutive times.

CloudWatch alarm on instance family capacity availability of the AWS Outposts resource.

AMSOutpostsInstanceTypeCapacityAvailability TypeCapacityAvailability

= 80% for 5 minutes, 12 consecutive times.

CloudWatch alarm on instance type capacity availability of the AWS Outposts resource.

AMSOutpostsConnectedStatusConnectedStatus

< 1 for 5 minutes, 1 consecutive time.

CloudWatch alarm on AWS Outposts service link connection, less than 1 count is impaired.

AMSOutpostsCapacityExceptionCapacityExceptions

0 for 5 minutes, 1 consecutive time.

CloudWatch alarm on insufficient capacity errors for instance launches for AWS Outpostss resource

.

EC2 instance - all OSs

No

CPUUtilization*

>= 95% for 5 mins, 6 consecutive times, if the instance is unresponsive to a polling Systems Manager command.

CloudWatch alarm. High CPU utilization is an indicator of a change in application state such as dead locks, infinite loops, malicious attacks, and other anomalies.

StatusCheckFailed

> 0 for 5 minutes, 3 consecutive times.

CloudWatch alarm.

Root Volume Usage

>= 95% for 5 mins, 6 consecutive times.

Non-root Volume Usage

> 85% for 5 mins, 2 consecutive times.

Disabled by default; for additional information, see https://docs.aws.amazon.com/managedservices/latest/ctref/management-monitoring-cloudwatch-enable-non-root-volumes-monitoring.html#management-monitoring-cloudwatch-enable-non-root-volumes-monitoring-info.

Memory Free*

MemoryFree < 5% for 5 minutes, 6 consecutive times.

Yes

EPS Malware

Malware found on instance.

CloudWatch event.

Amazon EC2 instance - Linux

No

Root Volume Inode Usage

Average >= 95% for 5 mins, 6 consecutive times.

CloudWatch alarm. Applied to Linux instances only.

Swap Free*

Memory Swap < 5% for 5 minutes, 6 consecutive times.

ElastiCache Cluster

No

CurrConnections = 65000

This alarm notifies AMS of the maximum connection limit of an ElastiCache Host.

CloudWatch Alarm. If you would like to update this threshold, contact AMS support.

ElastiCache Node

No

CPUUtilization

Average > predefined value for 15 mins, 2 consecutive times.

CloudWatch alarm. Default is 90. If Redis, use one the following values based on instance type:

  • cache.t1.micro: 90%

  • cache.m1.small: 90%

  • cache.m1.medium: 90%

  • cache.m1.large: 45%

  • cache.m1.xlarge: 22.5%

  • cache.m2.xlarge: 45%

  • cache.m2.4xlarge: 11.25%

  • cache.c1.xlarge: 11.25%

  • cache.t2.micro: 90%

  • cache.t2.small: 90%

  • cache.t2.medium: 45%

  • cache.m3.medium: 90%

  • cache.m3.large: 45%

  • cache.m3.xlarge: 22.5%

  • cache.m3.2xlarge: 11.25%

  • cache.r3.large: 45%

  • cache.r3.xlarge: 22.5%

  • cache.r3.2xlarge: 11.25%

  • cache.r3.4xlarge: 5.625%

  • cache.r3.8xlarge: 2.8125%

ElastiCache Node - memcached

No

SwapUsage

maximum > 50,000,000 bytes for 5 mins, 5 consecutive times.

CloudWatch alarm. Applied to memcached only.

OpenSearch cluster

No

ClusterStatus.red

maximum is >= 1 for 1 minute, 1 consecutive time.

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

CloudWatch alarm. At least one primary shard and its replicas are not allocated to a node. To learn more, see Red Cluster Status.

OpenSearch domain

No

KMSKeyError

>= 1 for 1 minute, 1 consecutive time.

CloudWatch alarm. The KMS encryption key that is used to encrypt data at rest in your domain is disabled. Re-enable it to restore normal operations. To learn more, see Encryption of Data at Rest for OpenSearch Service Service.

ClusterStatus.yellow

maximum is >= 1 for 1 minute, 1 consecutive time

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

At least one replica shard is not allocated to a node. To learn more, see Yellow Cluster Status.

FreeStorageSpace

minimum is <= 20480 for 1 minute, 1 consecutive time

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

A node in your cluster is down to 20 GiB of free storage space. To learn more, see Lack of Available Storage Space.

ClusterIndexWritesBlocked

>= 1 for 5 minutes, 1 consecutive time

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

The cluster is blocking write requests. To learn more, see ClusterBlockException.

Nodes

minimum is < x for 1 day, 1 consecutive time

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

x is the number of nodes in your cluster. This alarm indicates that at least one node in your cluster has been unreachable for one day. To learn more, see Failed Cluster Nodes.

CPUUtilization

average is >= 80% for 15 minutes, 3 consecutive times

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

100% CPU utilization is common, but sustained high averages are problematic. Consider using larger instance types or adding instances.

JVMMemoryPressure

maximum is >= 80% for 5 minutes, 3 consecutive times

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

The cluster could encounter out of memory errors if usage increases. Consider scaling vertically. Amazon ES uses half of an instance's RAM for the Java heap, up to a heap size of 32 GiB. You can scale instances vertically up to 64 GiB of RAM, at which point you can scale horizontally by adding instances.

MasterCPUUtilization

average is >= 50% for 15 minutes, 3 consecutive times

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

Consider using larger instance types for your dedicated master nodes. Because of their role in cluster stability and blue/green deployments, dedicated master nodes should have lower average CPU usage than data nodes.

MasterJVMMemoryPressure

maximum is >= 80% for 15 minutes, 1 consecutive time

AMS takes pro-active actions to reduce operational impact, when this alert is triggered.

Consider using larger instance types for your dedicated master nodes. Because of their role in cluster stability and blue/green deployments, dedicated master nodes should have lower average CPU usage than data nodes.

OpenSearch instance

No

AutomatedSnapshotFailure

maximum is >= 1 for 1 minute, 1 consecutive time.

CloudWatch alarm. An automated snapshot failed. This failure is often the result of a red cluster health status. See Red Cluster Status.

Elastic Load Balancing instance

No

SurgeQueueLength

> 100 for 1 minute, 15 consecutive times.

CloudWatch alarm if an excess number of requests are pending routing.

HTTPCode_ELB_5XX_Count

sum > 0 for 5 min, 3 consecutive times.

CloudWatch alarm on excess number of HTTP 5XX response codes that originate from the load balancer.

SpilloverCount

> 1 for 1 minute, 15 consecutive times.

CloudWatch alarm if an excess number of requests that were rejected because the surge queue is full.

GuardDuty service

Yes

Not applicable; all findings (threat purposes) are monitored. Each finding corresponds to an alert.

Changes in the GuardDuty findings. These changes include newly generated findings or subsequent occurrences of existing findings.

List of supported GuardDuty finding types are on GuardDuty Active Finding Types.

Health

Varies

AWS Health Dashboard

Notifications are sent when there are changes in the status of AWS Health Dashboard (AWS Health) events in relation to baseline services supported by AMS requiring action from AMS Operations. For more information, see Supported services.

AWS Managed Microsoft AD

No

Active Directory Status

AWS Managed Microsoft AD instance sends an active status event.

Service event. Emitted when the directory is operating normally after an event.

Impaired Directory Status

AWS Managed Microsoft AD instance sends an impaired directory status event.

Service event. Emitted when the directory is running in a degraded state. One or more issues have been detected, and not all directory operations may be working at full operational capacity.

Inoperable Directory Status

AWS Managed Microsoft AD instance sends an inoperable status event.

Service event. Emitted when the directory is not functional. All directory endpoints have reported issues.

Deleting Directory Status

AWS Managed Microsoft AD instance sends a deleting directory status event.

Service event. Emitted when the directory is currently being deleted.

Failed Directory Status

AWS Managed Microsoft AD instance sends a failed status event.

Service event. Emitted when the directory could not be created.

RestoreFailed Directory Status

AWS Managed Microsoft AD instance sends a restore failed directory status event.

Service event. Emitted when restoring the directory from a snapshot failed.

Amazon RDS instance

No

Low Storage alert triggers when the allocated storage for the DB instance has been exhausted.

RDS-EVENT-0007, see details at Using Amazon RDS event notification.

DB instance fail

The DB instance has failed due to an incompatible configuration or an underlying storage issue. Begin a point-in-time-restore for the DB instance.