Blazalek.com

Monitoring and Postmaster Signals

On this page

Why is Google Postmaster Tools data missing, stale, or not updating?

Severity: MediumPause sending: Pause conditionally

Treat missing or stale Postmaster Tools data as an observability gap, not proof that Gmail delivery failed. The dashboards are not real time, low-volume days can be blank, and some views depend on DKIM-authenticated traffic. Do not pause solely for an expected reporting delay or low-volume omission.

First 15 minutes

  1. Confirm that the affected traffic is authenticated with DKIM for dashboards that only expose DKIM-authenticated messages.
  2. Confirm that the domain is verified, follows sender guidelines, and has a sample message passing SPF and DKIM before preparing a Postmaster Tools delivery-issue report.

Now

  • Restore DKIM authentication on the affected stream if the dashboard requires it.

Next 24 hours

  • Check whether DKIM-authenticated traffic begins populating the affected dashboard.

Next 7 days

  • Keep the affected stream DKIM-authenticated so dashboard visibility is not lost for that reason.

Technical checks

Deliverability

  • Identify whether the missing view is one that only shows DKIM-authenticated traffic.

Verification criteria

  • For the Compliance status dashboard, check again after seven days and confirm whether the rolling multi-day view reflects the configuration change.

Escalation criteria

  • Submit a Postmaster Tools delivery-issue report only after the domain is verified and compliant and the sample passes SPF and DKIM with a matching From domain.

Prevention

  • Keep relevant Gmail traffic DKIM-authenticated for dashboards that depend on DKIM.
  • Maintain the verified-domain and authentication prerequisites needed for a Postmaster Tools delivery-issue report.

Business impact

  • Delayed or omitted dashboard data can slow a sending decision, but it does not itself establish a customer-facing delivery failure.

Provider notes

  • Postmaster Tools data is typically updated within 24 hours but can take longer, and low-volume days may be omitted.
  • The Compliance status dashboard uses a rolling multi-day average; Google recommends checking again after seven days.

Open questions

  • Google does not publish one universal numeric daily-volume threshold that guarantees dashboard visibility.
  • A dashboard gap alone does not reveal current acceptance, rejection, or inbox placement; sender logs and received samples are still required.
  • Different Postmaster Tools dashboards use different filters and aggregation windows, so their visible dates can differ.
Sources (5)
  1. Postmaster Tools dashboardsDashboard data, lines 31-36Google Gmail
  2. Postmaster Tools dashboardsDashboard data, line 35Google Gmail
  3. Postmaster Tools dashboardsDashboard data, line 34Google Gmail
  4. Postmaster Tools dashboardsTroubleshoot compliance status issues, lines 117-125Google Gmail
  5. Report delivery issues in Postmaster ToolsEligibility and reporting steps, lines 25-46Google Gmail

Related runbooks

Why is Google Postmaster Tools setup failing to produce usable data?

Severity: LowPause sending: Keep sending

An empty Postmaster Tools dashboard alone does not establish a Gmail delivery outage. Confirm that the setup targets the SPF or DKIM authentication domain and personal gmail.com or googlemail.com traffic, not all Google Workspace recipients. Even a correct setup can omit low-volume days, and Google publishes no universal message count that guarantees visibility.

First 15 minutes

  1. Confirm that the expected traffic is addressed to personal gmail.com or googlemail.com recipients rather than only Google Workspace recipients.
  2. Check that Postmaster Tools contains the domain used by SPF, DKIM, or both; add subdomains separately when independent views are required.
  3. Confirm DNS verification and allow up to 10 minutes for its status to update.

Now

  • Add the live SPF or DKIM authentication domain to Postmaster Tools.

Next 24 hours

  • Complete DNS verification and allow up to 10 minutes for verification status to update.

Next 7 days

  • Operate real-time SMTP and ESP telemetry alongside Postmaster Tools because dashboard data usually updates within 24 hours and can take longer.

Technical checks

IT / DNS

  • Compare the configured Postmaster domain with the live SPF and DKIM authentication domains.
  • Verify the domain and check status again after up to 10 minutes.

Deliverability

  • Confirm the stream sends to personal Gmail and assess whether low volume could trigger privacy-based omission.
  • Use real-time SMTP and ESP telemetry for the incident window because Postmaster Tools usually updates within 24 hours and can take longer.

Verification criteria

  • The configured authentication domain shows verified after allowing up to 10 minutes.
  • Dashboard data appears after the usual update interval, or the remaining absence is consistent with documented low-volume privacy omission and is not treated as a proven outage.

Escalation criteria

  • Escalate the setup only after the correct authentication domain is verified and the normal update delay has been allowed, documenting that data usually appears within 24 hours but can take longer.

Prevention

  • Document the personal-Gmail scope and the exact SPF or DKIM domains that must be added and verified, and never use an empty dashboard alone as outage proof.

Business impact

  • The empty dashboard removes a useful diagnostic signal but, without independent SMTP, bounce, or recipient evidence, does not establish a delivery outage.

Provider notes

  • Google can omit Postmaster Tools data on low-volume days for privacy and publishes no universal visibility threshold.

Open questions

  • The configured Postmaster domain, verification status, live SPF and DKIM domains, personal-Gmail daily volume, and time since setup are not supplied.
  • Google does not publish a universal daily-volume threshold that guarantees dashboard visibility.
Sources (6)
  1. Set up Postmaster ToolsMonitor outgoing email to Gmail accountsGoogle Gmail
  2. Set up Postmaster ToolsStep 1: Add your sending domainGoogle Gmail
  3. Set up Postmaster ToolsStep 2: Verify your sending domainsGoogle Gmail
  4. Set up Postmaster ToolsTroubleshoot Postmaster Tools setup: expected data missingGoogle Gmail
  5. Postmaster Tools dashboardsDashboard dataGoogle Gmail
  6. Set up Postmaster ToolsSetup steps and troubleshootingGoogle Gmail

Related runbooks

Can DMARC reports help us detect spoofing and phishing attempts?

Severity: MediumPause sending: Keep sending

Use DMARC aggregate reports to identify source IPs using the domain and inspect authentication alignment, but treat unexpected sources as investigation leads rather than proof of phishing or account compromise. Reconcile report rows with the approved sender inventory before taking action. Reporting coverage is receiver-dependent, so a missing row does not prove that spoofing did not occur.

First 15 minutes

  1. Group report rows by connecting source IP and RFC5322.From domain, retaining message counts, disposition, and DKIM and SPF alignment results.
  2. Reconcile those groups against the approved sender inventory.
  3. Prioritize unexpected high-volume sources, repeated alignment failures, and unexpected policy dispositions.

Now

  • Group aggregate rows by source IP and header-from domain and compare them with the approved sending inventory.

Next 24 hours

  • Investigate unexpected high-volume sources, repeated alignment failures, and unexpected policy dispositions in priority order.

Next 7 days

  • Maintain an authoritative sender inventory so later aggregate periods can be reconciled consistently.

Technical checks

IT / DNS

  • Confirm that the valid DMARC record publishes the intended aggregate-report destination in rua.

Security

  • Review source IP, message count, applied disposition, header-from domain, and DKIM and SPF alignment in each relevant aggregate row.

Verification criteria

  • In later aggregate periods, confirm that expected legitimate sources align and that the suspect source disappears or receives the intended disposition.
  • Corroborate the aggregate trend with security logs, recognizing that receivers may omit reports and are only encouraged to report at least once every 24 hours.

Escalation criteria

  • Escalate to security when a source remains unexplained after inventory reconciliation; do not treat missing receiver reports or an alignment failure alone as proof of compromise.

Prevention

  • Publish a valid DMARC record with the intended rua reporting destination.
  • Keep the legitimate-sender inventory current and reconcile it against aggregate reports over time.

Business impact

  • Unreconciled sources can obscure possible unauthorized use of the domain, while a false malicious label can disrupt legitimate senders.

Open questions

  • The observed source IPs, counts, alignment results, approved-sender inventory, and corroborating security logs are not provided, so no source can be labeled malicious from this pack alone.
  • Receiver reporting coverage and timing vary, and some receivers do not send aggregate reports for policy, resource, or privacy reasons.
Sources (6)
  1. RFC 9989: Domain-Based Message Authentication, Reporting, and Conformance (DMARC)Section 4.8, rua tagIETF RFC 9989
  2. RFC 9990: DMARC Aggregate ReportingSections 3.1.1.8 through 3.1.1.10IETF RFC 9990
  3. RFC 9990: DMARC Aggregate ReportingSection 1, IntroductionIETF RFC 9990
  4. RFC 9989: Domain-Based Message Authentication, Reporting, and Conformance (DMARC)Section 5.3.8, Send Aggregate ReportsIETF RFC 9989
  5. RFC 9990: DMARC Aggregate ReportingSections 1 and 3.1.1.8 through 3.1.1.10IETF RFC 9990
  6. RFC 9990: DMARC Aggregate ReportingReport schema and Section 9.3, Report StorageIETF RFC 9990

Related runbooks

How should we investigate a deliverability incident when monitoring systems provide no usable visibility?

Severity: HighPause sending: Pause conditionally

Treat missing visibility as an incident constraint, not evidence that delivery is healthy: a provider's accepted or delivered event does not prove inbox placement, and an empty Gmail dashboard does not prove no errors occurred. Reconstruct the narrowest evidence chain from enqueue and provider timestamps, message identifiers, SMTP responses or DSNs, webhook receipts, and privacy-safe controlled headers. Keep missing telemetry explicit so business owners understand that delivery outcome may remain unknown.

First 15 minutes

  1. Identify which event destinations and sending streams were configured and selected for affected sends.
  2. Separate provider request success from recipient-server delivery and from inbox placement.
  3. Reconstruct a privacy-safe timeline from retained timestamps, message identifiers, SMTP responses or DSNs, webhooks, and controlled sample headers.

Now

  • Reconstruct the incident from retained enqueue and provider timestamps, message identifiers, SMTP responses or DSNs, webhooks, and controlled headers.

Next 24 hours

  • Enable or repair event publication for every affected sending stream and configured destination.

Next 7 days

  • Alert on gaps between submitted messages and terminal outcomes after validating event publication with a controlled send.

Technical checks

Engineering

  • Check whether configuration sets and event destinations captured send, delivery, bounce, complaint, rejection, rendering-failure, and delivery-delay events for the stream.

Deliverability

  • Group retained evidence by stream, recipient domain, and mailbox provider while preserving missing telemetry as an explicit gap.

Verification criteria

  • Verify that a controlled send produces the expected submission and terminal outcome events at the configured destination.
  • Confirm that monitoring alerts on gaps between submitted messages and terminal outcomes, while recognizing that the test does not prove inbox placement across providers.

Escalation criteria

  • Escalate to engineering and the ESP when event publication remains broken, required identifiers or outcomes cannot be reconstructed, or DeliveryDelay payloads show unresolved temporary recipient-server failures.

Prevention

  • Configure event publication for every sending stream and select the intended event destination on each send.
  • Continuously test controlled event flow and alert on missing terminal outcomes.

Business impact

  • Messages can remain temporarily unaccepted by recipient servers, while apparent provider success can still leave inbox placement and user receipt unknown.

Provider notes

  • In Amazon SES, Send means the request succeeded and SES will attempt delivery, while Delivery means the recipient mail server accepted it; neither proves inbox placement.
  • Gmail Postmaster Tools is provider-scoped, not real time, and can omit low-volume data for privacy.

Open questions

  • The provider, retained message identifiers, SMTP responses, DSNs, event destination health, stream, recipient domains, and incident time window are unavailable; the actual delivery outcome is therefore not knowable from the prompt.
  • Event names, retention, dashboard sampling, and low-volume visibility differ by ESP and mailbox provider; SES and Gmail behaviors cannot be generalized.
Sources (6)
  1. Monitor email sending using Amazon SES event publishingOverviewAmazon SES
  2. Monitor email sending using Amazon SES event publishingEvent publishing terminology > Email sending eventAmazon SES
  3. Monitor email sending using Amazon SES event publishingEvent publishing terminology > DeliveryDelayAmazon SES
  4. Postmaster Tools dashboardsDashboard dataGoogle Gmail
  5. Monitor email sending using Amazon SES event publishingHow event publishing works with configuration sets and message tagsAmazon SES
  6. Monitor email sending using Amazon SES event publishingHow to use event publishing and Event publishing terminologyAmazon SES

Related runbooks

How should we interpret Google Postmaster domain-reputation and spam-rate dashboards?

Severity: HighPause sending: Pause conditionally

Use Google Postmaster as delayed, Gmail-specific telemetry rather than a real-time or complete complaint ledger. Missing or low data can reflect privacy thresholds, DKIM-only coverage, or automatic spam placement. Interpret Spam Rate against Gmail guidance and account for the legacy reputation-dashboard transition before changing traffic.

First 15 minutes

  1. Confirm the dashboard date range and allow for the usual 24-hour update lag, longer delays, and low-volume privacy omissions.
  2. Interpret Spam Rate using its DKIM-authenticated, engaged-inbox denominator and check for automatic spam placement effects.
  3. Compare the reported spam rate with Gmail guidance to keep it below 0.10% and avoid 0.30% or higher.

Now

  • Reduce the implicated Gmail-bound stream when Postmaster spam reaches 0.30% or higher.

Next 24 hours

  • Recheck Compliance status with its rolling multi-day behavior rather than expecting an immediate status change after a fix.

Next 7 days

  • Continue checking the rolling Compliance status because a fix can take time to appear in its multi-day averages.

Technical checks

Deliverability

  • Record dashboard dates, missing days, and update timing before comparing trends.
  • Compare Spam Rate, rolling Compliance status, and authenticated Delivery Errors for the same Gmail-bound domain and period.
  • Plan for the eventual retirement of legacy Domain and IP Reputation dashboards without assuming an unpublished deadline.

Marketing / CRM

  • Identify the Gmail-bound stream associated with spam at or above 0.30% and keep the operating target below 0.10%.

Verification criteria

  • Postmaster data for the relevant date has updated, with any missing low-volume days explicitly treated as privacy-filtered rather than zero complaints.
  • The rolling Compliance status reflects the correction and authenticated Delivery Errors decline for the same traffic scope.

Escalation criteria

  • Escalate to a deliverability specialist when spam remains at or above 0.30% or authenticated Delivery Errors persist after the implicated stream is reduced.

Prevention

  • Monitor Gmail Postmaster spam against the below-0.10% recommendation and the avoid-at-or-above-0.30% boundary, alongside rolling compliance data.

Business impact

  • A rising Gmail spam-rate signal indicates engaged inbox recipients are manually rejecting the mail, while automatic spam placement can hide additional impact from the displayed percentage.

Provider notes

  • Legacy Domain and IP Reputation dashboards are scheduled to retire rather than move unchanged into Postmaster Tools v2, but Google publishes no final retirement date.

Open questions

  • No dashboard export, sending volume, date range, domain set, or provider-side SMTP error sample is supplied, so current severity cannot be classified from the question alone.
  • Google publishes no final retirement date for the legacy reputation dashboards and says the deprecation was postponed.
  • Missing or unusually low spam-rate data can reflect privacy thresholds, DKIM-only coverage, or automatic spam placement; it is not equivalent to zero complaints.
Sources (6)
  1. Postmaster Tools dashboardsDashboard dataGoogle Gmail
  2. Postmaster Tools dashboardsSpam RateGoogle Gmail
  3. Email sender guidelinesMonitoring and troubleshooting > Postmaster Tools > Spam rateGoogle Gmail
  4. Learn about the deprecation of the old Postmaster Tools interfaceWhat about the dashboards in the old Postmaster Tools?Google Gmail
  5. Postmaster Tools dashboardsCompliance status and troubleshoot compliance status issuesGoogle Gmail
  6. Postmaster Tools dashboardsPostmaster Tools dashboards overview > Delivery errorsGoogle Gmail

Related runbooks

Which Amazon SES deliverability alert should trigger incident response?

Severity: CriticalPause sending: Pause conditionally

Open an incident when the SES account-level bounce alarm reaches 5%, the complaint alarm reaches 0.1%, or the Reputation page shows Under review, Pending sending pause, or Sending paused. The account-state alert is the most urgent because unresolved review can progress to a pause and a paused account cannot send through SES. Confirm the affected AWS account and Region before applying the response.

First 15 minutes

  1. Confirm the AWS account and Region, then record whether SES shows Under review, Pending sending pause, or Sending paused.
  2. Open the AWS Support case and identify the trigger and the corrective action AWS expects before replying.
  3. Correlate the SES complaint metric with event notifications, mailbox-provider data, campaigns, and list sources because its feedback coverage is not universal.

Now

  • If the account is under review and the situation allows, stop mail and correct the trigger identified in the AWS case.

Next 24 hours

  • Correlate complaints with events, providers, campaigns, and list sources, then reply to AWS with corrective and preventive changes already implemented.

Next 7 days

  • Review complaint-feedback coverage and retain the supporting cohort evidence used to explain the corrected cause.

Technical checks

Deliverability

  • Check the account-level Reputation.BounceRate against the documented 0.05 alarm and Reputation.ComplaintRate against the documented 0.001 alarm.
  • Interpret complaint rate with SES feedback-domain and representative-volume coverage, then correlate it with other event sources.

ESP support

  • Verify the exact SES Reputation page state in the affected account and Region.

Verification criteria

  • The SES bounce and complaint reputation metrics are below the incident alarm conditions and continue to be observed at account level.
  • The affected account and Region no longer show Under review, Pending sending pause, or Sending paused on the Reputation page.

Escalation criteria

  • Escalate immediately through the AWS Support case for Under review, Pending sending pause, or Sending paused, and describe completed corrective and preventive changes.

Prevention

  • Maintain CloudWatch alarms at the documented SES account-level 5% bounce and 0.1% complaint conditions, with lower internal warnings where earlier detection is needed.
  • Correlate SES complaint-rate alerts with event notifications, providers, campaigns, and list sources rather than treating the metric as complete complaint coverage.

Business impact

  • If SES reaches Sending paused, the affected account and Region cannot send through SES, interrupting every dependent email stream there.

Provider notes

  • SES complaint rate covers complaints from domains that provide feedback to SES and uses representative volume rather than a fixed time window.

Open questions

  • The AWS account, Region, current SES account status, alarm state, and measured bounce and complaint rates were not provided.
  • The affected sending streams and the cause of any reputation event are unresolved until SES metrics, notifications, and recipient cohorts are reviewed.
  • Campaign criticality, list provenance, complaint-feedback coverage, and business impact are not available.
Sources (5)
  1. Creating reputation monitoring alarms using CloudWatchCreate alarm > Conditions > Reputation.BounceRateAmazon SES
  2. Creating reputation monitoring alarms using CloudWatchCreate alarm > Conditions > Reputation.ComplaintRateAmazon SES
  3. Using reputation metrics to track bounce and complaint ratesAccount status; Bounce Rate and Complaint Rate status messagesAmazon SES
  4. Amazon SES Sending review process FAQsAccount under review FAQ > What should I do if my account is under review?Amazon SES
  5. Amazon SES Sending review process FAQsSES complaints through feedback loops FAQ > Q2, Q6, Q7Amazon SES

Related runbooks

Which Google Postmaster dashboard should we use to diagnose the current deliverability change?

Severity: HighPause sending: Pause conditionally

Choose the Postmaster dashboard from the symptom: Delivery Errors for rejects or deferrals, Spam Rate for user complaints, Reputation for inbox-versus-spam context, and Authentication for identity failures. Use live sender logs and message artifacts first because Postmaster data is delayed. Do not wait for a dashboard update before acting on current delivery evidence.

First 15 minutes

  1. Capture live SMTP responses and open Delivery Errors for the matching authenticated domain and time window when the symptom is rejection or deferral.
  2. Open Authentication for the exact affected identity when DNS or message authentication may have changed.

Now

  • Route the investigation to Delivery Errors, Spam Rate, Reputation, or Authentication according to the observed symptom and affected identity.

Next 24 hours

  • Correct the exact Delivery Errors reason or authentication failure and compare the same authenticated cohort after the change.

Next 7 days

  • Maintain a symptom-to-dashboard runbook covering delivery errors, complaints, reputation, and authentication rather than relying on one chart.

Technical checks

Deliverability

  • Match live SMTP outcomes to Delivery Errors reasons for authenticated traffic to personal Gmail accounts.
  • Use Spam Rate for manual user complaints, not as direct proof of automatic inbox placement.
  • Check IP and Domain Reputation for the exact represented IP and authenticated domain, noting shared-IP effects.
  • Account for reporting lag and low-volume omissions before interpreting a blank or unchanged chart.

IT / DNS

  • Use Authentication pass percentages for the matching From, SPF, or DKIM identity, then confirm the result in a received message.

Verification criteria

  • Live sender logs and the updated Delivery Errors view show the incident reject or deferral reason declining for the affected authenticated cohort.
  • Authentication reports the expected pass percentages for the selected identity and a received message confirms its message-level result.

Escalation criteria

  • Escalate to the ESP or a deliverability specialist when live Gmail errors persist but Postmaster is blank because of volume limits, or when shared-IP or domain reputation cannot be isolated internally.

Prevention

  • Keep live SMTP telemetry beside Postmaster monitoring, select dashboards by symptom and identity, and document reporting lag and low-volume coverage limits.

Business impact

  • Rejects and temporary failures reduce or delay Gmail reach, while reputation and complaint signals can indicate broader inbox-versus-spam risk.

Provider notes

  • Postmaster Tools is typically updated within 24 hours but can take longer, so it is not a real-time incident feed.
  • Postmaster coverage is for personal Gmail traffic, and Google can omit low-volume data for privacy.

Open questions

  • The current symptom, selected Postmaster domain, sender volume, incident time window, and affected Gmail recipient population were not provided.
  • The correct first dashboard is unresolved until the symptom is classified as temporary deferral, permanent rejection, spam placement, reputation, or authentication.
  • Live Postmaster Tools data is not available in this research environment, and low-volume or reporting-lag gaps can limit diagnosis.
Sources (6)
  1. Postmaster Tools dashboardsPostmaster Tools dashboards overview; Delivery ErrorsGmail Postmaster Tools
  2. Postmaster Tools dashboardsSpam RateGmail Postmaster Tools
  3. Postmaster Tools dashboardsIP Reputation & Domain ReputationGmail Postmaster Tools
  4. Postmaster Tools dashboardsAuthenticationGmail Postmaster Tools
  5. Postmaster Tools dashboardsDashboard dataGmail Postmaster Tools
  6. Postmaster Tools dashboardsDashboard dataGmail Postmaster Tools

Related runbooks

Why does Google Postmaster report a problem when LearnDMARC does not?

Severity: MediumPause sending: Pause conditionally

Treat both results as valid within their different scopes rather than choosing one as the truth. LearnDMARC proves one tested authentication path, while Google Postmaster aggregates production mail to personal Gmail accounts and may lag or omit low-volume data. Align the domain, message path, and observation time before acting on the difference.

First 15 minutes

  1. Record the exact Postmaster dashboard, verified domain, date range, and visible warning.
  2. Repeat the point test with the same visible From, envelope, and DKIM domains used by the production path.
  3. Align the selected Postmaster domain and production view with the From, DKIM, and SPF domains used by the point test.

Now

  • Use the point test and Postmaster production aggregates together to isolate the affected path.

Next 24 hours

  • Separate From-domain, DKIM-domain, and SPF-domain views and investigate forwarding or mailing-list traffic that changes the aggregate.

Next 7 days

  • Keep both point authentication tests and Postmaster production monitoring in the operating checklist.

Technical checks

IT / DNS

  • Align the visible From, envelope, DKIM, and SPF domains between the production cohort and LearnDMARC test.

Deliverability

  • Compare the selected Postmaster production dashboards and time range only after allowing for delayed, filtered aggregate data.

Verification criteria

  • The point test and production comparison use the same visible From, envelope, DKIM, and SPF domain path.
  • The point test and Postmaster production evidence are used together to identify and resolve the scoped sending issue.
  • The selected production dashboard shows an understood or improved authentication, reputation, spam, encryption, or delivery-error signal for the verified domain.

Escalation criteria

  • Use the Postmaster reporting route for an unresolved Gmail classification, rejection, or temporary-failure issue only when the sample meets Google's SPF, DKIM, and matching From-domain requirements.

Prevention

  • Monitor production Gmail aggregates alongside point authentication tests so neither is treated as a substitute for the other.
  • Account for Postmaster's delayed and filtered production data in every routine comparison with a point test.

Business impact

  • Dismissing the aggregate production signal after one clean test can leave a Gmail authentication, reputation, spam, encryption, or delivery-error cohort uninvestigated.

Provider notes

  • Postmaster covers aggregate production mail to personal Gmail accounts and is delayed and privacy-filtered; LearnDMARC observes one message at its own receiving server.

Open questions

  • The exact Postmaster warning, selected domain and date range, dashboard view, LearnDMARC test message, and production-message authentication samples are not provided, so the affected production cohort remains unknown.
  • A one-message tester and Gmail's delayed, filtered production aggregates have different scopes, inputs, and visibility; a clean test does not invalidate a Postmaster warning.
Sources (6)
  1. Learn and Test DMARCWelcome and DMARC ResultsLearnDMARC
  2. Postmaster Tools dashboardsDashboard data and dashboards overviewGoogle Gmail
  3. Postmaster Tools dashboardsDashboard dataGoogle Gmail
  4. Postmaster Tools dashboardsSpam rate doesn't match third-party spam reportsGoogle Gmail
  5. Postmaster Tools dashboardsAuthentication dashboardGoogle Gmail
  6. Report delivery issues in Postmaster ToolsEligibility and Report a delivery issueGoogle Gmail

Related runbooks

How should we process DMARC aggregate reports at scale without ignoring them?

Severity: Medium

Process DMARC aggregate reports as RFC 9990 XML. Use their sending-IP, volume, policy, disposition, SPF, DKIM, alignment, and domain signals to distinguish legitimate and unrecognized mail streams.

First 15 minutes

  1. Validate a sample of incoming XML and quarantine malformed reports instead of mixing them into trusted aggregates.
  2. Measure the raw-report backlog and identify which accepted data is not yet indexed for dashboards or monitoring.

Now

  • Sideline malformed XML and keep any salvaged data explicitly untrusted until it validates.

Next 24 hours

  • Extract and index accepted report data so operators can query it instead of reading raw XML manually.

Next 7 days

  • Expose the indexed data through reporting, dashboards, and monitoring appropriate to the report volume.

Technical checks

Engineering

  • Validate each document against the RFC 9990 format before trusting its data.
  • Normalize by policy domain and configuration, then retain source IP, count, evaluated disposition, identifiers, and SPF and DKIM results for each record.

Verification criteria

  • Every report admitted to trusted aggregates validates against the RFC 9990 format.
  • Normalized rows preserve one policy domain and configuration plus the source IP, message count, disposition, identifiers, and authentication results.

Escalation criteria

  • Contact the report generator when recurring format defects prevent its reports from validating.
  • Escalate to the monitoring owner when decisions assume daily complete coverage, because receivers may opt out despite the 24-hour reporting recommendation.

Prevention

  • Keep accepted report data indexed for dashboards and monitoring instead of depending on manual raw-XML review.
  • Design coverage indicators to tolerate receivers that do not generate reports because reporting is recommended rather than guaranteed.

Business impact

  • Unprocessed reports hide sending-IP, volume, policy, disposition, authentication, alignment, and domain signals needed to separate legitimate from unrecognized mail.

Open questions

  • DMARC aggregate-report coverage is incomplete and provider-variable because receivers may choose not to generate reports.
Sources (6)
  1. RFC 9990: Domain-Based Message Authentication, Reporting, and Conformance (DMARC) Aggregate ReportingAbstract and Status of This MemoIETF RFC 9990
  2. RFC 9990: Domain-Based Message Authentication, Reporting, and Conformance (DMARC) Aggregate ReportingSections 1 and 3.1, report dataIETF RFC 9990
  3. RFC 9990: Domain-Based Message Authentication, Reporting, and Conformance (DMARC) Aggregate ReportingSections 3.1 and 3.1.1.7-3.1.1.11IETF RFC 9990
  4. RFC 9990: Domain-Based Message Authentication, Reporting, and Conformance (DMARC) Aggregate ReportingSection 3.1.1 and Section 9.2 Report EvaluationIETF RFC 9990
  5. RFC 9990: Domain-Based Message Authentication, Reporting, and Conformance (DMARC) Aggregate ReportingSections 9.2 Report Evaluation and 9.3 Report StorageIETF RFC 9990
  6. RFC 9989: Domain-Based Message Authentication, Reporting, and Conformance (DMARC)Section 5.3.8, Send Aggregate ReportsIETF RFC 9989

Related runbooks

Which Mimecast Service Monitor notifications should trigger an email incident?

Severity: High

Open an incident when a Mimecast Service Monitor alert reports a configured queue-threshold breach or monitored-service failure. Identify whether the sampled service is outbound delivery, inbound delivery, or directory synchronization.

First 15 minutes

  1. Identify whether the alert concerns outbound delivery, inbound delivery, or directory synchronization and inspect the latest Service Monitor snapshot.
  2. Confirm whether notifications are repeating every 15 minutes while the condition remains active.
  3. Compare the configured threshold with the account-specific Recommended Threshold and local business criticality.

Now

  • Verify the notification configuration and establish an independent alert channel for a business-critical condition.

Next 24 hours

  • Tune the queue threshold from the account-specific recommended starting point and local business-criticality requirements.

Next 7 days

  • Test both email and configured SMS notification paths for critical Service Monitor alerts.

Technical checks

Deliverability

  • Determine whether the alert is a queue-threshold breach or service failure and name the affected monitored service.
  • Review the configured threshold, recent account history, repeated-notification sequence, and current service snapshots.

Verification criteria

  • Subsequent 15-minute service snapshots no longer show the affected condition and repeated notifications stop.
  • Configured email and independent SMS notification paths are both exercised successfully for critical alerts.

Escalation criteria

  • If the default escalation level remains five and the condition persists, verify that escalation subscribers receive the sixth and later notifications at the documented 1.5-hour point.

Prevention

  • Review queue thresholds against recent account history and local criticality instead of treating the recommended value as a universal threshold.
  • Use and test an independent notification channel because Service Monitor email alerts may sometimes not be sent.

Business impact

  • The alert can concern outbound delivery, inbound delivery, or directory synchronization, so the possible impact depends on which monitored service is affected.

Provider notes

  • Service Monitor alerts on a configured queue threshold breach or monitored-service failure.
  • With the documented default escalation level of five, escalation subscribers receive the sixth and later notifications after 1.5 hours of 15-minute alerts.

Open questions

  • The tenant's configured queue thresholds, critical services, and incident-severity mapping are unknown and must be supplied by the local operating policy.
Sources (6)
  1. Service Monitor - Managing Alert NotificationsOverviewMimecast
  2. Service Monitor - Monitoring ServicesMonitoring Services overviewMimecast
  3. Service Monitor - Managing Alert NotificationsEmail NotificationsMimecast
  4. Service Monitor - Managing Alert NotificationsCreating an Alert Notification, Recommended ThresholdMimecast
  5. Service Monitor - Managing Alert NotificationsCreating an Alert Notification, Escalation LevelMimecast
  6. Service Monitor - Managing Alert NotificationsExample Alert Notifications and Troubleshooting Notification IssuesMimecast

Related runbooks

Wojtek BlazalekWojtek BlazalekEmail deliverability expert

Stuck in an email incident? I help teams get delivery, reputation and auth back on track.

Hands-on deliverability work for teams that send at scale

Book a free diagnostic callSee services
Need help? Contact us!